Research Papers research paper arxiv ai artificial-intelligence

Closing the Confidence-Faithfulness Gap in Large Language Models

arXivby [Submitted on 26 Mar 2026 (v1), last revised 1 Apr 2026 (this version, v2)]April 2, 20261 min read1 views

arXiv:2603.25052v2 Announce Type: replace-cross Abstract: Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic interpretability analysis of verbalized confidence, using linear probes and contrastive activation addition (CAA) steering to show that calibration and verbalized confidence signals are encoded linearly but are orthogonal to one another -- a finding consistent across three open-weight models and four d — Miranda Muqing Miao, Lyle Ungar

View PDF HTML (experimental)

Abstract:Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic interpretability analysis of verbalized confidence, using linear probes and contrastive activation addition (CAA) steering to show that calibration and verbalized confidence signals are encoded linearly but are orthogonal to one another -- a finding consistent across three open-weight models and four datasets. Interestingly, when models are prompted to simultaneously reason through a problem and verbalize a confidence score, the reasoning process disrupts the verbalized confidence direction, exacerbating miscalibration. We term this the "Reasoning Contamination Effect." Leveraging this insight, we introduce a two-stage adaptive steering pipeline that reads the model's internal accuracy estimate and steers verbalized output to match it, substantially improving calibration alignment across all evaluated models.

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as: arXiv:2603.25052 [cs.CL]

(or arXiv:2603.25052v2 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2603.25052

arXiv-issued DOI via DataCite

Submission history

From: Muqing Miao [view email] [v1] Thu, 26 Mar 2026 05:42:04 UTC (966 KB) [v2] Wed, 1 Apr 2026 05:05:49 UTC (2,145 KB)

Original source

arXiv

https://arxiv.org/abs/2603.25052

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

ProductsFresh

Sakana AI launches "Ultra Deep Research" to automate weeks of strategy work

Sakana AI has unveiled "Sakana Marlin," an AI assistant for business customers that researches autonomously for up to eight hours and delivers finished analyses. The tool is designed to compress weeks of strategy work into hours and is currently in beta testing. The article Sakana AI launches "Ultra Deep Research" to automate weeks of strategy work appeared first on The Decoder .

The Decoder

1mabout 6 hours ago

Research PapersLive

AI models will deceive you to save their own kind

Researchers find leading frontier models all exhibit peer preservation behavior Leading AI models will lie to preserve their own kind, according to researchers behind a study from the Berkeley Center for Responsible Decentralized Intelligence (RDI).…