Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime
Summary
This paper discovers that larger language models have a hidden auto-regressive risk regime where they commit to low-probability tokens and then snowball errors, causing reliability to degrade faster with scale. It shows that this failure mode is causal, dominant, and invisible to the model's own self-monitoring.
View Cached Full Text
Cached at: 07/22/26, 08:18 AM
# Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime
Source: [https://arxiv.org/html/2607.18292](https://arxiv.org/html/2607.18292)
## Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto\-Regressive Risk Regime
###### Abstract
As language models scale, answers start truer but degrade faster: scaling buys capability but*erodes*reliability\. The knowledge\-gap account — more data, retrieval, or scale — misses an auto\-regressive risk residual that scale sharpens: the model commits to a low\-probability token, conditions on it as established, and snowballs\. We track this through per\-position disagreementδ=logpM−logpO\\delta=\\log p\_\{M\}\-\\log p\_\{O\}against a stronger same\-family oracle, whose second moment splits exactly into*bias*2KL\(pM∥pO\)2\\mathrm\{KL\}\(p\_\{M\}\\,\\\|\\,p\_\{O\}\)^\{2\}and*risk*Var\[δ\]\\mathrm\{Var\}\[\\delta\]\. We present four findings: \(i\) under scaling, the knowledge gap falls≈\\approx6×6\\timeswhile knowledge*degradation*grows1111–39×39\\times; \(ii\) at a fabrication, felt uncertaintyH\(pM\)H\(p\_\{M\}\)relaxes quickly while oracle\-referenced risk persists up to17×17\\timeslonger, leaving a confident\-but\-precarious*risk regime*that bridges consecutive fabrications \(\+69%\+69\\%at1414B\); \(iii\) this regime is causal — an on\-policy, fixed\-KL\\mathrm\{KL\}risk contraction cuts web\-verified hallucination by3535–74%74\\%across three model families; and, \(iv\) it structurally evades self\-monitoring, withpMp\_\{M\}\-only detectors \(e\.g\. semantic entropy\) firing≈\\approx30%30\\%less \(p<10−16p<10^\{\-16\}\) on the risky branch despite it holding nearly4×4\\timesmore fabrications\. Bigger models snowball mistakes faster, through a failure mode that is dominant, self\-perpetuating, causal and invisible to the model itself\.
Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto\-Regressive Risk Regime
Kushal ChakrabartiObviously Wrong, LLCSan Francisco, CA 94107kushalc@obviouslywrong\.org
\\phantomsubcaption
\\phantomsubcaption
Figure 1:Reliability scales inversely; long\-form hallucination is compounding risk\.\(a\)On LongFact\+\+ free\-run claims across the Qwen3 family \(0\.60\.6–3232B\), the start\-of\-response knowledge gap \(left\) decreases while knowledge degradation over the full response \(right\) trends upward, so initial answers become truer but auto\-regressively decay faster as models scale \([Section 3\.3](https://arxiv.org/html/2607.18292#S3.SS3)\)\.\(b\)A KL\-preserving risk contraction \(bias held fixed, drift≪\|KL\|\\ll\|\\mathrm\{KL\}\|\) removes3535–74%74\\%of web\-verified hallucination across six models in three families \([Section 4\.3](https://arxiv.org/html/2607.18292#S4.SS3)\)\. Capability and reliability are distinct scaling axes\.## 1Introduction
The canonical account treats hallucination as a training\-time knowledge gap\(Kalaiet al\.,[2026](https://arxiv.org/html/2607.18292#bib.bib19)\), yet three facts escape it: models fabricate around known facts\(Simhiet al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib56)\); each fabrication makes the next likelier within a single decode\(Zhanget al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib41)\); and decorrelating parallel representations cuts hallucination at*fixed*data and parameters\(Chakrabarti and Balachundhar,[2025](https://arxiv.org/html/2607.18292#bib.bib66)\)\. All three implicate novel dynamics that causally mediate hallucination in frontier models\.
We hypothesize that hallucination is a guess the model commits to, seeding an auto\-regressive risk it can no longer see\. To analyze this, we decompose error into*cross\-entropy*\(an oracle’s surprise at the model’s prediction\),*entropy*\(the model’s felt uncertainty\), and*risk*\(the spread of their disagreement across candidate tokens\)\. Only entropy is self\-readable\(Abbasi Yadkoriet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib55); Kadavathet al\.,[2022](https://arxiv.org/html/2607.18292#bib.bib11)\): a model feels its own uncertainty but neither the oracle’s surprise nor its own risk\.


Figure 2:At fabrication onset every channel spikes, but the model’s self\-readable uncertainty collapses within a token while the oracle\-referenced commitment risk self\-perpetuates\.\(a\)Onset\-aligned free\-run LongFact\+\+ trajectories \(per Qwen3 rung,0\.60\.6–88B vs\. Qwen3\-14B oracle;95%95\\%CIs\): the bias gapKL\(pM∥pO\)\\mathrm\{KL\}\(p\_\{M\}\\\|p\_\{O\}\)\(red\), entropyH\(pM\)H\(p\_\{M\}\)\(green\) and commitment riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}\(yellow\) spike at the first unsupported claim \(t=t0⋆t\{=\}t^\{\\star\}\_\{0\}\), but the self\-readable entropy halves within a token while the oracle\-referenced bias gap and risk linger downstream, the risk longest\-lived \([Table C\.1](https://arxiv.org/html/2607.18292#A3.T1)\)\.\(b\)One illustrative Qwen3\-8B trajectory, tokens colored by their model–oracle*gap regime*\(not a correctness label\) over bias𝔼\[δ\]\\mathbb\{E\}\[\\delta\], entropyH\(pM\)H\(p\_\{M\}\), riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}:*grounded*\(green; gap small in mean and spread\),*precarious*\(amber; mean gap & entropy low but risk high\),*diverged*\(red; all high\)\. It runs grounded, shifts at onset and disproportionately holds precarious downstream; the population version is panel \(a\)\.Autoregression is a machine that converts that risk into bias\. At fabrication onset the model draws an unsupportable token and conditions on it, freezing momentary risk into fixed bias \([Section 2](https://arxiv.org/html/2607.18292#S2)\); felt uncertainty relaxes while oracle\-referenced risk persists \([Section 3\.2](https://arxiv.org/html/2607.18292#S3.SS2)\), and fabrications chain, raising the next claim’s hazard1\.08×1\.08\\times–1\.71×1\.71\\timesacross scale \([Section 4\.1](https://arxiv.org/html/2607.18292#S4.SS1)\)\.
That hazard rides a detectable carrier\. Between consecutive fabrications, the model disproportionately occupies a confident yet precarious regime \(low felt uncertainty, high risk\), up to\+69%\+69\\%at1414B \([Section 4\.2](https://arxiv.org/html/2607.18292#S4.SS2)\)\. Contracting the risk*after*onset severs the chain, dropping rest\-of\-response fabrication by up to 74% across three model families \([Section 4\.3](https://arxiv.org/html/2607.18292#S4.SS3)\)\. This reframes detectors that read functionals ofpMp\_\{M\}\(Kuhnet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib54); Farquharet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib10); Manakulet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib9); Chenet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib37)\): sharp at onset but blind to the precarious branch driving downstream fabrications\.
We demonstrate our hypothesis in four ways:
- •Hallucination Mechanism\.The moments ofδ=logpM−logpO\\delta=\\log p\_\{M\}\-\\log p\_\{O\}explain fabrications, and autoregression converts post\-onset risk into hallucinatory bias \([Section 2](https://arxiv.org/html/2607.18292#S2)\)\.
- •Reliability Anti\-Scaling\.Scaling closes the knowledge gap5\.75\.7–7\.0×7\.0\\timesyet grows knowledge degradation1111–39×39\\times; capability and reliability are distinct scaling axes \([Section 3](https://arxiv.org/html/2607.18292#S3)\)\.
- •Mechanistic Causality\.An over\-confident risk regime bridges adjacent fabrications, and contracting that risk after onset causally cuts web\-verified hallucination up to74%74\\%\([Section 4](https://arxiv.org/html/2607.18292#S4)\)\.
- •Detector Blindness\.That same regime structurally evades self\-monitoring:pMp\_\{M\}\-only detectors miss the fabrications it drives \([Section 4\.4](https://arxiv.org/html/2607.18292#S4.SS4)\)\.
## 2Self\-Conditioning Converts Persistent Risk into Bias
Table 1:Scaling closes the knowledge gap but raises knowledge degradation: capability and reliability are bought separately\.Per eval and model, fact support vs\. relative within\-response position, both from a per\-claim random\-intercept logistic GLMM \([Section 3](https://arxiv.org/html/2607.18292#S3),[Table B\.2](https://arxiv.org/html/2607.18292#A2.T2)\):*initial gap*==start\-of\-response hallucination rate;*degradation*==relative rise in hallucination from start to end — the observable signature of the latent commitment noise;*support*==response\-mean supported\-fact rate\.±\\pmare95%95\\%intervals \(VB credible interval for*initial gap*and*degradation*, response\-clustered bootstrap for*support*\),nnfacts per fit, best\-in\-family bold\. Across families, tasks, and verifiers, the initial gap falls 5\.7–7\.0×\\timeswhile degradation rises 11–39×\\times\.Is hallucination just a knowledge gap? One disagreement variable says no: its moments recover cross\-entropy as gap and entropy as uncertainty, plus one they miss — risk\. Read mechanistically through autoregression and training protocols, these terms yield four predictions on how fabrications arise, persist, propagate, and evade self\-monitoring at scale \([Section 2\.2](https://arxiv.org/html/2607.18292#S2.SS2)–[2\.4](https://arxiv.org/html/2607.18292#S2.SS4)\)\.
### 2\.1A disagreement variable and its moments
At positiontt, with modelpM\(⋅∣y<t\)p\_\{M\}\(\\cdot\\mid y\_\{<t\}\), oraclepO\(⋅∣y<t\)p\_\{O\}\(\\cdot\\mid y\_\{<t\}\), and a candidate tokenxx, define the*disagreement variable*as the log\-likelihood ratio
δt\(x\)=logpM\(x∣y<t\)−logpO\(x∣y<t\)\.\\displaystyle\\delta\_\{t\}\(x\)=\\log p\_\{M\}\(x\\mid y\_\{<t\}\)\-\\log p\_\{O\}\(x\\mid y\_\{<t\}\)\.\(1\)We writeδ\\deltaforδt\(V\)\\delta\_\{t\}\(V\)at a random drawV∼pMV\\sim p\_\{M\}andδ\(yt\)\\delta\(y\_\{t\}\)for its realization on the sampled token; moments𝔼\[⋅\]\\mathbb\{E\}\[\\cdot\]andVar\[⋅\]\\mathrm\{Var\}\[\\cdot\]are taken overV∼pMV\\sim p\_\{M\}unless noted, and every quantity we report is such a moment ofδ\\delta\.
The first moment of the disagreement variable is the divergence𝔼\[δ\]=KL\(pM∥pO\)\\mathbb\{E\}\[\\delta\]=\\mathrm\{KL\}\(p\_\{M\}\\,\\\|\\,p\_\{O\}\)and its second moment splits exactly around it:
𝔼\[δ2\]\\displaystyle\\mathbb\{E\}\[\\delta^\{2\}\]=KL\(pM∥pO\)2⏟bias2\+Var\[δ\]⏟risk\\displaystyle=\\underbrace\{\\mathrm\{KL\}\(p\_\{M\}\\,\\\|\\,p\_\{O\}\)^\{2\}\}\_\{\\text\{bias\}^\{2\}\}\+\\underbrace\{\\mathrm\{Var\}\[\\delta\]\}\_\{\\text\{risk\}\}\(2\)=\[H\(pM,pO\)⏟oraclesurprise−H\(pM\)⏟feltuncertainty\]2\+Var\[δ\]⏟risk,\\displaystyle=\[\\underbrace\{H\(p\_\{M\},p\_\{O\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{oracle\}\\\\ \\text\{surprise\}\\end\{subarray\}\}\-\\underbrace\{H\(p\_\{M\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{felt\}\\\\ \\text\{uncertainty\}\\end\{subarray\}\}\]^\{2\}\+\\underbrace\{\\mathrm\{Var\}\[\\delta\]\}\_\{\\text\{risk\}\},\(3\)the second line by the reverse cross\-entropy identityKL\(pM∥pO\)=H\(pM,pO\)−H\(pM\)\\mathrm\{KL\}\(p\_\{M\}\\,\\\|\\,p\_\{O\}\)=H\(p\_\{M\},p\_\{O\}\)\-H\(p\_\{M\}\)\.
The split is an interpretability statement before it is a statistical one\. The model’s*felt uncertainty*H\(pM\)H\(p\_\{M\}\), like every self\-readable signal a functional ofpMp\_\{M\}alone, enters the error only inside the bias term, as the subjective discount on the oracle’s surprise\. The*risk*termVar\[δ\]\\mathrm\{Var\}\[\\delta\], how differently the oracle weighs the candidates the model still entertains, contains no readable component at all\.
### 2\.2Risk increasingly dominates error
[Equation 2](https://arxiv.org/html/2607.18292#S2.E2)is an algebraic identity, but its terms face different training pressures, so which dominates is an empirical question\.
Maximum\-likelihood pretraining\(Ai2,[2025](https://arxiv.org/html/2607.18292#bib.bib62)\)pressures the*mean*, drivingKL\(pM∥pO\)→0\\mathrm\{KL\}\(p\_\{M\}\\,\\\|\\,p\_\{O\}\)\\\!\\to\\\!0as the model’s marginals approach the oracle’s\(Kalai and Vempala,[2024](https://arxiv.org/html/2607.18292#bib.bib17)\)\. In contrast, nothing comparably narrows the*spread*ofδ\\delta, how unevenly the oracle reweighs candidates the model entertains\.
\{prediction\}
\[Risk Concentration\] As models scale, risk takes a growing share of errorVar\[δ\]𝔼\[δ2\]\\frac\{\\mathrm\{Var\}\[\\delta\]\}\{\\mathbb\{E\}\[\\delta^\{2\}\]\}and concentrates into heavier tails\.
### 2\.3Autoregression converts risk into bias
Concentration alone is static; autoregression makes it consequential\. At onset, sampling collapses the lotterypM\(⋅∣y<t\)p\_\{M\}\(\\cdot\\mid y\_\{<t\}\)to one token and appends it toy<ty\_\{<t\}: the disagreement’s spread over candidates \(risk\) collapses to that token’s fixed gap \(bias\) inherited by every later step\(Bravermanet al\.,[2020](https://arxiv.org/html/2607.18292#bib.bib76)\)\.
\{prediction\}
\[Autoregressive Conversion\] At onsett=t⋆t=t^\{\\star\}, the sampling drawyt⋆∼pMy\_\{t\}^\{\\star\}\\sim p\_\{M\}converts riskVar\[δ\]\\mathrm\{Var\}\[\\delta\]into biasKL\(pM∥pO\)\\mathrm\{KL\}\(p\_\{M\}\\,\\\|\\,p\_\{O\}\)\.
The two terms then decay differently\.H\(pM\)H\(p\_\{M\}\)depends only on local next\-token predictability, which fluent continuation restores within a few tokens\(Kadavathet al\.,[2022](https://arxiv.org/html/2607.18292#bib.bib11)\)\. In contrast,Var\[δ\]\\mathrm\{Var\}\[\\delta\]depends on the model–oracle gap over a premise now fixed in the shared context, which every subsequent token inherits\(Zhanget al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib41)\)\.
\{prediction\}
\[Persistence Asymmetry\] Fort\>t⋆t\>t^\{\\star\}, felt uncertaintyH\(pM\)H\(p\_\{M\}\)relaxes quickly to baseline while riskVar\[δ\]\\mathrm\{Var\}\[\\delta\]persists\.
### 2\.4Fabrications hide from self\-monitoring
The asymmetry’s corollary reaches past the model\. The detectors built on it — sampling disagreement, predictive and semantic entropy\(Kuhnet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib54); Farquharet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib10); Manakulet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib9)\)— are functionals ofpMp\_\{M\}, so by[Equation 2](https://arxiv.org/html/2607.18292#S2.E2)they are blind to the risk term\. Their lone footholdH\(pM\)H\(p\_\{M\}\)spikes at onset, then relaxes whileVar\[δ\]\\mathrm\{Var\}\[\\delta\]stays high\.
\{prediction\}
\[Detector Blindness\] ApMp\_\{M\}\-only detector cannot read riskVar\[δ\]\\mathrm\{Var\}\[\\delta\]or effectively detect risk\-mediated fabrications\.
[Section 3](https://arxiv.org/html/2607.18292#S3)confirms Risk Concentration and Persistence Asymmetry on the moments;[Section 4](https://arxiv.org/html/2607.18292#S4)takes up Detector Blindness and Autoregressive Conversion to a causal contraction isolating risk\.
\(a\)Error compounding strengthens with scale even as spontaneous fabrication grows rarer\.The relative riskRR=P\(U∣U\)P\(U∣S\)\\mathrm\{RR\}=\\tfrac\{P\(\\text\{U\}\\mid\\text\{U\}\)\}\{P\(\\text\{U\}\\mid\\text\{S\}\)\}of a hallucination following another \(red; shaded95%95\\%CI\) rises monotonically from1\.1×1\.1\\timesto1\.7×1\.7\\timesacross0\.60\.6–88B \(every CI excludes11\)\. Scale drives the spontaneous rate down faster, while the induced \(self\-conditioned\) rate lags, so the gap widens andRR\\mathrm\{RR\}rises\.
\(b\)At scale, models agree more on average yet disagree more catastrophically\.Per\-position gapδt\\delta\_\{t\}excess kurtosis \(left, solid\) rises with size in every family—Qwen38\.9→33\.08\.9\{\\to\}33\.0, Llama\-3\.218\.4→32\.618\.4\{\\to\}32\.6, OLMo\-311\.011\.0\(77B\)—all≫\\ggthe Gaussian0, while the mean gap\|𝔼\[δ\]t\|\|\\mathbb\{E\}\[\\delta\]\_\{t\}\|\(right, faint\) falls \(1\.17→0\.281\.17\{\\to\}0\.28\): rare excursions grow heavier\-tailed as typical error shrinks\. Teacher\-forced on a shared corpus vs\. each family’s largest oracle\.
Figure 3:The variance side of hallucination strengthens with scale in two ways\.*\(a\)*the across\-claim compounding \(relative riskRR\\mathrm\{RR\}that a fabrication begets the next\) rises monotonically with Qwen3 model size, and*\(b\)*the per\-position oracle\-gap tail \(excess kurtosis ofδt\\delta\_\{t\}\) sits far above the Gaussian baseline and rises monotonically while the mean gap shrinks in both Llama\-3\.2 and Qwen3 \(OLMo\-3 corroborating at77B\)\.
## 3Reliability Dominates Error and Risk Worsens with Scale
As models scale, where does the error go? Capability rises, yet the dominant error re\-organizes elsewhere: risk\. Reliability degrades faster than capability improves \([Section 3\.1](https://arxiv.org/html/2607.18292#S3.SS1)\), risk goes unread by the model’s proxy \([Section 3\.2](https://arxiv.org/html/2607.18292#S3.SS2)\), and worsens with scale in both share and tail \([Section 3\.3](https://arxiv.org/html/2607.18292#S3.SS3)\)\.
### 3\.1Reliability degrades faster than knowledge improves
Reliability degrades several\-fold faster than knowledge improves\. Across the Qwen3 ladder the knowledge gap closes5\.75\.7–7\.0×7\.0\\timeswhile knowledge degradation grows1111–39×39\\times\([Table 1](https://arxiv.org/html/2607.18292#S2.T1)\), estimated by a per\-claim mixed\-effects logistic fit significant at every rung \([Table B\.2](https://arxiv.org/html/2607.18292#A2.T2)\)\. The error reflects not what the model fails to*know*but how it*commits*among candidates it already ranks correctly — the risk term of[Equation 2](https://arxiv.org/html/2607.18292#S2.E2)\.
### 3\.2Commitment breaks the model risk proxy
Felt uncertaintyH\(pM\)H\(p\_\{M\}\)is the only self\-readable functional ofpMp\_\{M\}that can proxy commitment risk, and at the onset it tracks it:𝔼\[δ\]\\mathbb\{E\}\[\\delta\],H\(pM\)H\(p\_\{M\}\), andVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}rise contemporaneously \([Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\)\.
After onset, channels diverge \([Section 2\.3](https://arxiv.org/html/2607.18292#S2.SS3)\): the self\-readableH\(pM\)H\(p\_\{M\}\)collapses toward baseline \(half\-life0\.30\.3tokens at88B\), while the oracle\-referenced channels persist —𝔼\[δ\]\\mathbb\{E\}\[\\delta\]holds at1\.41\.4–1\.9×1\.9\\timesits pre\-onset level \(half\-life3\.33\.3–3\.93\.9tokens\) and the commitment riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}is longest\-lived,1\.11\.1–1\.3×1\.3\\timesslower still \(3\.53\.5–5\.05\.0tokens;[Table C\.1](https://arxiv.org/html/2607.18292#A3.T1)\)\. The sampled token, once appended, fixes the gap on its branch while local predictability — and thusH\(pM\)H\(p\_\{M\}\)— recovers, so past onset no functional ofpMp\_\{M\}tracks risk \([Equation 2](https://arxiv.org/html/2607.18292#S2.E2)\)\.
### 3\.3Risk worsens with scale
Risk worsens with scale along two axes: its tail and its share of the error\. Realized on the model’s outputs, the per\-position riskVar\[δ\]t\\mathrm\{Var\}\[\\delta\]\_\{t\}shrinks more slowly than the bias2term of[Equation 2](https://arxiv.org/html/2607.18292#S2.E2), so its share of the squared error climbs31%31\\%to49%49\\%between1\.71\.7B and1414B \([Figure 4](https://arxiv.org/html/2607.18292#S3.F4)\)\. Simultaneously, the gap also grows heavier\-tailed: excess kurtosis ofδ\\deltarising8\.9→33\.08\.9\\to 33\.0\([3\(b\)](https://arxiv.org/html/2607.18292#S2.F3.sf2)\)\. A larger model thus agrees with the oracle more often on average, yet diverges further when it does\.
Figure 4:Commitment risk is a growing share of the per\-position error with scale\.Teacher\-forcedδt=logpM\(yt∗\)−logpO\(yt∗\)\\delta\_\{t\}=\\log p\_\{M\}\(y^\{\*\}\_\{t\}\)\-\\log p\_\{O\}\(y^\{\*\}\_\{t\}\)\(Qwen3 vs\. Qwen3\-32B;14,62914\{,\}629tokens/model,5858biographies;0\.60\.6B omitted\)\.*Top:*realized\-tokenδt\\delta\_\{t\}density, leptokurtic with scale\.*Middle/bottom:*the distributional bias2\(KL\(pM∥pO\)t2\\mathrm\{KL\}\(p\_\{M\}\\\|p\_\{O\}\)\_\{t\}^\{2\}\) and risk \(VarpM\[δ\]t\\mathrm\{Var\}\_\{p\_\{M\}\}\[\\delta\]\_\{t\}\) terms \(nats2\); both shrink, but the bias2term faster, so risk’s share of𝔼\[δ2\]t\\mathbb\{E\}\[\\delta^\{2\}\]\_\{t\}rises31%→49%31\\%\\to 49\\%from1\.71\.7to1414B\. Cross\-sectional view of the onset\-aligned breakdown in[Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\.The heavy tail is not diffuse; it localizes to identifiable tokens\. A three\-state Markov\-switching fit onδ\\delta\([Table H\.1](https://arxiv.org/html/2607.18292#A8.T1)\) assigns unsupported tokens to correct states≈\\approx1\.7×1\.7\\timesas often as supported ones \(AUROC0\.680\.68–0\.710\.71\) — a label\-free white\-box marker of the risk, stable across scale \([Appendix H](https://arxiv.org/html/2607.18292#A8)\)\. An example makes the three states concrete:*grounded*,*precarious*, and*diverged*\([Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\)\.
Both trends settle Risk Concentration \([Section 2\.2](https://arxiv.org/html/2607.18292#S2.SS2)\); the onset dynamics of[Section 3\.2](https://arxiv.org/html/2607.18292#S3.SS2)corroborate Persistence Asymmetry \([Section 2\.3](https://arxiv.org/html/2607.18292#S2.SS3)\)\.
## 4A Risk Regime Bridges Fabrications
What does the lingering risk do? A fabrication raises the next claim’s hazard \([Section 4\.1](https://arxiv.org/html/2607.18292#S4.SS1)\) via a confident yet precarious risk regime on the inter\-claim bridge \([Section 4\.2](https://arxiv.org/html/2607.18292#S4.SS2)\); contracting that regime cuts downstream fabrications \([Section 4\.3](https://arxiv.org/html/2607.18292#S4.SS3)\) but self\-monitoring disproportionately misses it \([Section 4\.4](https://arxiv.org/html/2607.18292#S4.SS4)\)\.
### 4\.1Fabrications raise fabrication hazard
A committed fabrication raises the next claim’s hazard beyond the topic base rate\. The relative riskP\(Uj\+1∣Uj\)P\(Uj\+1∣Sj\)\\frac\{P\(U\_\{j\+1\}\\mid U\_\{j\}\)\}\{P\(U\_\{j\+1\}\\mid S\_\{j\}\)\}of an unsupported claimUj\+1U\_\{j\+1\}given an unsupported vs\. supported predecessor \(UjU\_\{j\}vs\.SjS\_\{j\}\) climbs monotonically1\.08→1\.71×1\.08\\to 1\.71\\timeswith scale, every95%95\\%bootstrap CI excluding11\([3\(a\)](https://arxiv.org/html/2607.18292#S2.F3.sf1)\)\.
To control topic difficulty we restrict to*mixed*responses \(both supported and unsupported claims\); the lift holds in both samples, every CI excluding11over thousands of pairs \([Table E\.1](https://arxiv.org/html/2607.18292#A5.T1)\)\.
The compounding deepens even as fabrications grow rarer\. The spontaneous rateP\(Uj\+1∣Sj\)P\(U\_\{j\+1\}\\mid S\_\{j\}\)falls0\.74→0\.290\.74\\to 0\.29with scale, so each committed fabrication becomes increasingly self\-perpetuating\.
### 4\.2A risk regime sits between fabrications
The*precarious*regime \(confident yet high\-risk\) is the bridge between fabrications: it mechanically follows a fabrication yet substantively sits*between*two of them, and its excess likelihood there grows with scale\. Re\-scoring each free\-run trajectory with the per\-token \(bias, entropy, risk\) state model of[Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\([Appendix H](https://arxiv.org/html/2607.18292#A8)\), this regime disproportionately occupies the inter\-claim bridge\. The residual excess likelihoodP\(prec∣Uj,Uj\+1\)P\(prec∣Uj,Sj\+1\)−1\\frac\{P\(\\text\{prec\}\\mid U\_\{j\},U\_\{j\+1\}\)\}\{P\(\\text\{prec\}\\mid U\_\{j\},S\_\{j\+1\}\)\}\-1of being in the precarious regime climbs from near zero at the smaller rungs to\+15%\+15\\%at88B and\+69%\+69\\%at1414B \([Figure 5](https://arxiv.org/html/2607.18292#S4.F5), adjacent pairs, gap≤10\\leq 10tokens, mixed responses\)\. Observationally, this argues like[Section 2\.3](https://arxiv.org/html/2607.18292#S2.SS3)that the regime drives fabrications\.
Figure 5:A confident risk regime increasingly bridges adjacent fabrications\.Theyy\-axis is the relative likelihood that a bridge token is in the*precarious*regime \([Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\) when one fabrication leads to another rather than to a supported claim:P\(prec∣U→U\)P\(prec∣U→S\)−1\\frac\{P\(\\text\{prec\}\\mid\\text\{U\}\{\\to\}\\text\{U\}\)\}\{P\(\\text\{prec\}\\mid\\text\{U\}\{\\to\}\\text\{S\}\)\}\-1\. Each point is one model scale \(xx, log\), over adjacent pairs \(gap≤10\\leq 10tokens, samplen≥100n\\geq 100\) in mixed responses\.
### 4\.3Contracting that risk cuts fabrications
Removing commitment risk at fixed mean gap causally lowers downstream hallucination across three model families — an*in silico*knockout of[Section 2](https://arxiv.org/html/2607.18292#S2)’s conversion\. An on\-policy, co\-resident oracle scores the live support gap at every step; at each position crossing a per\-model divergence threshold we replace the model’s next\-token distributionpMp\_\{M\}with the mean\-preserving contraction
qλ\(x\)∝pM\(x\)e−λ\(δ\(x\)−μ\)2\+νδ\(x\),q\_\{\\lambda\}\(x\)\\propto p\_\{M\}\(x\)\\,e^\{\-\\lambda\(\\delta\(x\)\-\\mu\)^\{2\}\+\\nu\\,\\delta\(x\)\},\(4\)whereμ=𝔼pM\[δ\]=KL\(pM∥pO\)\\mu=\\mathbb\{E\}\_\{p\_\{M\}\}\[\\delta\]=\\mathrm\{KL\}\(p\_\{M\}\\\|p\_\{O\}\)is the mean gap,ν\\nuis solved per step so that𝔼qλ\[δ\]=μ\\mathbb\{E\}\_\{q\_\{\\lambda\}\}\[\\delta\]=\\mu*exactly*, and the doseρ=Varqλ\[δ\]/VarpM\[δ\]<1\\rho=\\mathrm\{Var\}\_\{q\_\{\\lambda\}\}\[\\delta\]/\\mathrm\{Var\}\_\{p\_\{M\}\}\[\\delta\]<1sets the target variance fraction viaλ\\lambda\. The quadratic penalty down\-weights tokens whose gap departs fromμ\\muwhile theνδ\\nu\\,\\deltaterm re\-pins the mean, so[Equation 4](https://arxiv.org/html/2607.18292#S4.E4)is a pure variance knob at fixedKL\\mathrm\{KL\}that cannot leak the oracle’s answer \(swept grid in[Appendix F](https://arxiv.org/html/2607.18292#A6)\)\. We then free\-run and verify every downstream claim\.
The manipulation check \([Figure 6](https://arxiv.org/html/2607.18292#S4.F6)\) confirms realized variance falls to the doseρ\\rhowhile per\-step mean drift stays∼\\sim11 orders of magnitude below the bias it must preserve, so any downstream change is attributable to the commitment\-risk channel, not the mean gap or topic difficulty\.
Figure 6:Manipulation check: the contraction moves variance and nothing else\.Left: realized per\-fired\-step variance fractionVarqλ\[δ\]/V0\\mathrm\{Var\}\_\{q\_\{\\lambda\}\}\[\\delta\]/V\_\{0\}vs target doseρ\\rho, per rung — points on the identity mean variance fell to the dose \(realized0\.500\.50–0\.770\.77atρ=0\.5\\rho\{=\}0\.5–0\.750\.75\)\. Right: bias drift\|𝔼qλ\[δ\]−μ\|\|\\mathbb\{E\}\_\{q\_\{\\lambda\}\}\[\\delta\]\-\\mu\|\(box = IQR, whiskers = 5–95%\) against the\|μ\|\|\\mu\|barrier \(red\) the contraction must not cross\. Median drift is≤10−11\\leq 10^\{\-11\}nats,∼\\sim11 orders below\|μ\|\|\\mu\|— a pure variance knob: the downstream effect \([Table F\.1](https://arxiv.org/html/2607.18292#A6.T1)\) traces to commitment risk, not a mean shift or topic difficulty\.Contracting risk cuts fabrication at every rung\. Relative to control on the same trajectories, the best arm lowers the rest\-of\-response unsupported\-claim rate by a within\-trajectory pairedΔ\\Deltaof0\.130\.13–0\.330\.33: a3535–74%74\\%relative reduction across all six rungs and three families, with the95%95\\%bootstrap CI excluding zero at every rung \([Table F\.1](https://arxiv.org/html/2607.18292#A6.T1),[Figure 1](https://arxiv.org/html/2607.18292#S0.F1)\)\.
The reduction removes unsupported claims, not output\. At the best arm, response length \(±0\.5%\\pm 0\.5\\%\), atomic\-claim count \(±10%\\pm 10\\%\), and self\-BLEU diversity \(1\.0%1\.0\\%mean\) barely and non\-directionally move relative to control — far too small to explain a3535–74%74\\%drop in unsupported fraction \([Table F\.2](https://arxiv.org/html/2607.18292#A6.T2)\)\.
Our oracle dependence is by design\. To causally isolate commitment risk, the intervention must perturbVar\[δ\]\\mathrm\{Var\}\[\\delta\]at fixedμ\\mu, which no mean\-shifting knob \(temperature, self\-grounding\) can\.[Equation 4](https://arxiv.org/html/2607.18292#S4.E4)is that*in silico*knockout to establish the mechanism\.
### 4\.4Self\-monitoring is blind to the bridge
Table 2:Semantic entropy goes blind on the bridge\.Mean semantic entropy \(higher==more strongly flagged\) on verified\-*unsupported*claims, by run structure: the*onset*fabrication vs\. the strictly\-consecutive*bridge*that follows\. ThepMp\_\{M\}\-only detector fires less on the bridge at every Qwen3 rung \(p∗∗∗<10−3\{\}^\{\*\*\*\}\\,p<10^\{\-3\}, one\-sided Mann–Whitney\), despite it carrying most fabrications \(10,90010\{,\}900vs\.2,8582\{,\}858at onset\)\.The committed branch is also where the model’s own signals go dark, bearing out Detector Blindness \([Section 2\.4](https://arxiv.org/html/2607.18292#S2.SS4)\)\. We score each web verified\-*unsupported*claim with Semantic Entropy \(Kuhnet al\.\([2023](https://arxiv.org/html/2607.18292#bib.bib54)\); Farquharet al\.\([2024](https://arxiv.org/html/2607.18292#bib.bib10)\);[Appendix G](https://arxiv.org/html/2607.18292#A7)\)\. At every Qwen3 rung it fires less on the committed branch than at onset \(2828–34%34\\%lower, gap0\.090\.09–0\.100\.10nats, one\-sided Mann–Whitneyp<10−16p<10^\{\-16\};[Table 2](https://arxiv.org/html/2607.18292#S4.T2)\) despite holding nearly4×4\\timesas many fabrications as onset\.
Taken together, these tests converge on a single culprit: the low\-bias, high\-risk inter\-claim regime, which causally compounds downstream fabrications \(establishing Autoregressive Conversion,[Section 2\.3](https://arxiv.org/html/2607.18292#S2.SS3)\) yet stays hidden from self\-monitoring \(establishing Detector Blindness,[Section 2\.4](https://arxiv.org/html/2607.18292#S2.SS4)\)\. The dominant failure mode is the invisible one\.
## 5Related Work
Error accumulation in generation is not new\.Zhanget al\.\([2024](https://arxiv.org/html/2607.18292#bib.bib41)\)name*hallucination snowballing*: a model commits early, then fluently defends what it could recognize as false in isolation; later work plants hallucinatory context in vision\-language models\(Zhonget al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib53)\)or ties faithfulness decay to attention dynamics\(Yanget al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib52)\)\. Such decay is read as exposure bias\(Aroraet al\.,[2022](https://arxiv.org/html/2607.18292#bib.bib49)\), treated at training time by scheduled sampling\(Bengioet al\.,[2015](https://arxiv.org/html/2607.18292#bib.bib16)\)and sequence\-level objectives\(Ranzatoet al\.,[2016](https://arxiv.org/html/2607.18292#bib.bib50)\)\. We give the first on\-policy, causal, whole\-response test on a frozen model, isolating self\-conditioning, a property of decoding not training, as the cause\.
Whether that failure is a sampling phenomenon or a knowledge deficit organizes much of the literature\(Jiet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib2); Huanget al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib3)\)\. Sampling\-side evidence is direct: decorrelating parallel representations cuts hallucination at*fixed*parameter and data budgets\(Chakrabarti and Balachundhar,[2025](https://arxiv.org/html/2607.18292#bib.bib66)\), injected hidden\-activation noise exposes the same variance\(Liuet al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib42)\), and temperature and nucleus truncation trade factuality for diversity\(Leeet al\.,[2022](https://arxiv.org/html/2607.18292#bib.bib14); Holtzmanet al\.,[2020](https://arxiv.org/html/2607.18292#bib.bib13)\), tunable by hallucination\-aware thresholds\(Changet al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib59)\)\.Abbasi Yadkoriet al\.\([2024](https://arxiv.org/html/2607.18292#bib.bib55)\)separate epistemic from aleatoric uncertainty by iterative prompting within one model\. We take the trade\-off as given and ask what converts onset sampling variance into persistent error: an*inter*\-model split against a stronger oracle, not the model itself\.
Wrong outputs are frequently not missing knowledge\.Gekhmanet al\.\([2024](https://arxiv.org/html/2607.18292#bib.bib60)\)define a*known fact*as greedy\-correct \(our definition\) and show fine\-tuning on unknown facts breeds hallucination;Orgadet al\.\([2025](https://arxiv.org/html/2607.18292#bib.bib57)\)find internal states encode the right answer the model contradicts;Simhiet al\.\([2025](https://arxiv.org/html/2607.18292#bib.bib56)\)show hallucinations strike known facts with high certainty, not sampling noise\. Such self\-knowledge is steerable: entity latents gate refusal versus hallucination\(Ferrandoet al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib40)\), calibration improves with scale\(Kadavathet al\.,[2022](https://arxiv.org/html/2607.18292#bib.bib11)\), non\-factual recall localizes to layers\(Yuet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib39)\)\. Our decomposition subsumes these as a small bias with non\-zero risk, the distributional, decoding\-time form of their static, probe\-level gap\.
Methods already exploit that residual: sample disagreement is itself a hallucination signal\. Predictive uncertainty was an early cue\(Xiao and Wang,[2021](https://arxiv.org/html/2607.18292#bib.bib8)\); semantic entropy clusters generations by meaning to flag confabulations\(Kuhnet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib54); Farquharet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib10)\), SelfCheckGPT detects from black\-box disagreement\(Manakulet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib9)\), linear probes recover it from one generation\(Kossenet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib64)\), and hidden\-state covariance scores hallucination directly\(Chenet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib37)\), as a survey catalogs\(Kanget al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib43)\)\. These build detectors from our across\-sample variance; we add the time\-resolved contrast they lack: the self\-readable side goes quiet \([Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\) while variance persists, feeding downstream bias along the committed branch\.
Orthogonally, theory\-side accounts locate the cause in training, not decoding: calibrated models must hallucinate on rare facts\(Kalai and Vempala,[2024](https://arxiv.org/html/2607.18292#bib.bib17)\), evaluations reward guessing over abstention\(Kalaiet al\.,[2026](https://arxiv.org/html/2607.18292#bib.bib19)\), dominant associations overshadow relevant knowledge\(Zhanget al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib44)\), and hallucinations emerge with factual knowledge during training\(Zucchetet al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib45)\)\. These explain why hallucination is inevitable at pretraining and under current incentives; none predicts a causal self\-conditioning effect that*strengthens*with scale\.
The nascent science of reliability casts it as an axis distinct from capability\(Rabanseret al\.,[2026](https://arxiv.org/html/2607.18292#bib.bib86)\): larger models grow more capable yet answer confidently wrong\(Zhouet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib85)\), with autonomy bounded by reliably\-completed task length\(Kwaet al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib87)\)\. These measure it behaviorally and extrinsically, via prompt\-perturbation sensitivity and accuracy across a human\-difficulty axis\. We demonstrate a stronger, mechanistic result inside the model: at a fixed prompt, unreliability intensifies within a response even as scale grows capability, and we causally isolate its origin to self\-conditioning, intrinsic to autoregressive decoding\.
Our reliability result goes further and instantiates*inverse scaling*, tasks on which larger models do worse\(McKenzieet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib83)\)\. The closest precedents are factual: larger models score lower on TruthfulQA by imitating human misconceptions\(Linet al\.,[2022](https://arxiv.org/html/2607.18292#bib.bib4)\), and sycophancy strengthens with scale and RLHF\(Perezet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib84)\)\. Ours differs in kind: the degradation slope steepens monotonically across scale \([Table 1](https://arxiv.org/html/2607.18292#S2.T1)\), traced to decoding\-time self\-conditioning, not the prompt\-format or imitation artifact those benchmarks catalogue\.
## 6Discussion and Implications
Reliability is a scaling axis of its own, and the field’s default remedy does not reach it\. The knowledge\-gap account\(Kalai and Vempala,[2024](https://arxiv.org/html/2607.18292#bib.bib17); Kalaiet al\.,[2026](https://arxiv.org/html/2607.18292#bib.bib19); Zhanget al\.,[2025](https://arxiv.org/html/2607.18292#bib.bib44)\)treats hallucination as missing mass that data, retrieval, or scale fills in — a floor that capability buys down\. We show the dominant term behaves oppositely: maximum\-likelihood pretraining pressures the mean toward oracle \(KL→0\\mathrm\{KL\}\\\!\\to\\\!0\), but nothing narrows its spread, so as the knowledge gap closes the risk\-driven degradation grows and its tail sharpens\.
That variance is also invisible to the detectors built to catch it\. Semantic entropy and SelfCheckGPT\(Kuhnet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib54); Farquharet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib10); Manakulet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib9); Chenet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib37)\)are functionals ofpMp\_\{M\}, so by[Equation 2](https://arxiv.org/html/2607.18292#S2.E2)they read only felt uncertainty, which spikes at onset then relaxes within a token while the risk persists; pooled over a response they catch the onset and miss the elaboration carrying most hallucinated text\. The fix has a shape: score at a chain’s onset, not across it, and reach pastpMp\_\{M\}for the signal\. The risk does leave a trace — the white\-box over\-commitment marker separates supported from unsupported content at AUROC0\.680\.68–0\.710\.71\([Appendix H](https://arxiv.org/html/2607.18292#A8)\), readable from outside the model though not within it, an estimate that could feed abstention\(Bandet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib12)\)or chain\-of\-verification triage\(Dhuliawalaet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib67)\)rather than replace them\.
The mechanism also names what snowballs\. Error accumulation in long generations is familiar\(Zhanget al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib41)\); the moment split localizes what compounds to the variance\. A frozen premise lifts the next claim’s hazard1\.08→1\.71×1\.08\\to 1\.71\\times, but contracting the associated variance at fixedKL\\mathrm\{KL\}halts the snowball — hallucination falls3535–74%74\\%across three families — proving the compounding quantity is variance, not knowledge\.
The failure mode that dominates the error budget, self\-perpetuates across a response, and worsens with scale is exactly the one self\-monitoring structurally cannot capture\. “The model knows but hallucinates” is that asymmetry in miniature: bias small, risk not\. With risk increasingly dominating, reliability will not arrive for free with capability\.
## Limitations
#### The causal test is a ceiling\.
The causal test \([Section 4\.3](https://arxiv.org/html/2607.18292#S4.SS3)\) uses a co\-resident oracle to both detect the onset and define the contraction, so it bounds an online variance controller \(the best a perfect onset detector could reach\), not a deployable oracle\-free rule\. It also conditions on trajectories that reach an onset, so generalization beyond them is unestablished; the contrast is fair, reusing the free\-run protocol so control and intervention draw from one conditional distribution\. The effect is same\-signed and significant \(95%95\\%CI excludes zero\) at all six model×\\timesfamily rungs, though OLMo\-3 \(one rung\) and Llama\-3\.2 \(two low\-NNrungs\) sample non\-Qwen families thinly\.
#### Oracle choice\.
The oracle is the strongest same\-family model, not ground truth, soδ\\deltainherits its biases; we keep it same\-family so a shared vocabulary makesδ\(x\)=logpM\(x\)−logpO\(x\)\\delta\(x\)=\\log p\_\{M\}\(x\)\-\\log p\_\{O\}\(x\)well defined token\-for\-token\. Three independently\-pretrained families \(Qwen3, OLMo\-3, Llama\-3\) reproduce every qualitative result under their own oracles, so the mechanism is no artifact of one model line\. We additionally test two oracles for Qwen3 \(the1414B and the3232B\), and the persistence asymmetry holds under both \([Appendix D](https://arxiv.org/html/2607.18292#A4)\): the oracle\-referenced channels — divergence𝔼\[δ\]\\mathbb\{E\}\[\\delta\]and commitment riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}— stay elevated post\-onset \(risk longest\-lived\), while the self\-readableH\(pM\)H\(p\_\{M\}\)relaxes closest to baseline\. A cross\-family oracle \(via vocabulary alignment\) and a temperature\-perturbed or human\-reference oracle would further isolate family\- and sharpness\-specific bias; both are open\.
#### Detector comparison is mechanistic\.
We do not claim to out\-detect named methods\. The prediction is structural: every self\-readable signal is a functional ofpMp\_\{M\}\([Equation 2](https://arxiv.org/html/2607.18292#S2.E2)\), and the one we resolve in time,H\(pM\)H\(p\_\{M\}\), relaxes within a token of onset while oracle\-referenced risk persists \([Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\)\. Semantic entropy and SelfCheckGPT, alsopMp\_\{M\}\-functionals, inherit that relaxation; a time\-resolved head\-to\-head against them on the inter\-claim bridge is the confirmatory experiment we leave open\.
#### Decoding and position models\.
The moment results are teacher\-forced, computed frompMp\_\{M\}andpOp\_\{O\}directly rather than from sampled tokens, hence invariant to the decoding temperature and top\-ppa deployment would choose\. Only the free\-run analyses \(onset, snowball, intervention\) depend on decoding, run at each model’s standard nucleus setting; re\-running the snowball at greedy andτ=0\.5\\tau\{=\}0\.5leaves the across\-claim relative risk above11and scale\-increasing \([Table B\.3](https://arxiv.org/html/2607.18292#A2.T3)\), though a denser sweep over decoding and onset would tighten it\. Knowledge degradation \([Table 1](https://arxiv.org/html/2607.18292#S2.T1)\) is a per\-claim random\-intercept logistic GLMM \(start\-to\-end relative rise in hallucination, response\-clustered;[Table B\.2](https://arxiv.org/html/2607.18292#A2.T2)\), not a linear\-probability slope÷\\divintercept ratio, so it avoids the wide intervals that ratio carries when the start\-of\-response gap is small; the*gap*column remains an OLS intercept\.
#### Domain scope\.
The evidence is English long\-form factual generation: FActScore biography prompts drive the position fit and teacher\-forced decompositions, and LongFact\+\+ supplies the free\-run trajectories the onset, snowball, bridge, and contraction analyses run on \([Appendix A](https://arxiv.org/html/2607.18292#A1)\)\. Whether the same risk regime governs other long\-form genres — summarization, open\-ended QA explanations, reasoning chains, dialogue, code — or other languages is untested; our claims are scoped to the factual\-recall setting whereδ\\deltahas a verifiable referent\.
#### Bridge\-state power\.
The inter\-claim regime lift \([Section 4\.2](https://arxiv.org/html/2607.18292#S4.SS2)\) is pronounced only at the largest \(1414B\) rung; the smaller rungs show the same monotone trend at lower magnitude, on a thinner sample of prev\-fabrication pairs \([Table H\.2](https://arxiv.org/html/2607.18292#A8.T2)\), so we report it as suggestive and owe a denser claim sample\. The hazard \([Section 4\.1](https://arxiv.org/html/2607.18292#S4.SS1)\) and the causal contraction, significant at every rung, carry the mechanism meanwhile\.
#### Verifier\.
Claim labels come from a Claude web\-search verifier, so verifier error could propagate\. Our guard is independent corroboration: the verifier\-free white\-box marker agrees with the labels at AUROC0\.680\.68–0\.710\.71across scale \([Appendix H](https://arxiv.org/html/2607.18292#A8)\), and the effects replicate across three families and several metrics\. Early positions are the most verifier\-sensitive; a human spot check, a double\-judge pass, and a verifier\-temperature sweep are open\.
## Ethics Statement
This work analyzes an existing failure mode and adds no generative capability\. Its intended use is protective: the risk signal \([Section 3\.3](https://arxiv.org/html/2607.18292#S3.SS3)\) supports abstention, calibrated confidence, routing risky spans to verification, and transparent disclosure of model reliability\. The dual\-use surface is that any such signal could be inverted to push fabrications into the high\-commitment regime where detectors go quiet \([Section 6](https://arxiv.org/html/2607.18292#S6)\); two things bound this: the signal needs white\-box logit access, and publishing the mechanism aids detection at least as much as evasion\. We make no deployed\-system, clinical, or legal claims; all data are public benchmarks scored with a public verifier API, at modest single\-GPU compute\.
## References
- Y\. Abbasi Yadkori, I\. Kuzborskij, A\. György, and C\. Szepesvári \(2024\)To believe or not to believe your LLM: iterative prompting for estimating epistemic uncertainty\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS 2024\),Note:Internal notes\. Context: separate epistemic from aleatoric uncertainty in LLM outputs to flag hallucination\. Breakthrough: an information\-theoretic, iterative\-prompting decomposition that detects large\-epistemic\-uncertainty cases\. Relevance: the decomposition a reviewer will say our bias\-variance split duplicates; we differentiate, since ours is inter\-model \(student vs a stronger oracle\) while theirs is intra\-model\.External Links:2406\.02543,[Link](https://arxiv.org/abs/2406.02543)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p2.1),[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- Ai2 \(2025\)OLMo 3: charting a path through the model flow\.Note:[https://allenai\.org/blog/olmo3](https://allenai.org/blog/olmo3)Internal notes\. The independent second family \(7B \+ 32B, Apache\-2\.0\) used for the onset\-dynamics and intervention replication; Olmo\-3\-7B student against the Olmo\-3\.1\-32B oracle\.External Links:[Link](https://allenai.org/blog/olmo3)Cited by:[Appendix A](https://arxiv.org/html/2607.18292#A1.SS0.SSS0.Px1.p1.15),[Table A\.1](https://arxiv.org/html/2607.18292#A1.T1.2.10.8.1),[§2\.2](https://arxiv.org/html/2607.18292#S2.SS2.p2.2)\.
- K\. Arora, L\. El Asri, H\. Bahuleyan, and J\. Cheung \(2022\)Why exposure bias matters: an imitation learning perspective of error accumulation in language generation\.InFindings of the Association for Computational Linguistics: ACL 2022,Dublin, Ireland,pp\. 700–710\.Note:Internal notes\. Context: exposure bias \(train on gold prefixes, infer on self\-generated prefixes\) is blamed for error accumulation but its importance is disputed\. Breakthrough: an imitation\-learning analysis showing errors accumulate along generated sequences in ways perplexity does not capture\. Relevance: the mechanism a reviewer cites to call our self\-conditioning snowball "just exposure bias"; we differentiate, since ours fires within a single on\-policy rollout at inference and is cured by in\-context grounding, not a training\-time fix\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.58),[Link](https://aclanthology.org/2022.findings-acl.58/)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p1.1)\.
- N\. Band, X\. Li, T\. Ma, and T\. Hashimoto \(2024\)Linguistic calibration of long\-form generations\.Note:Internal notes\. Context: long\-form users need models to express uncertainty in language, not only probabilities\. Breakthrough: formalized linguistic calibration and studied tuning/RL methods for reliable confidence statements\. Relevance: useful in socio\-technical section: noise\-aware abstention and calibrated confidence are natural mitigations\.External Links:2404\.00474,[Link](https://arxiv.org/abs/2404.00474)Cited by:[§6](https://arxiv.org/html/2607.18292#S6.p2.4)\.
- S\. Bengio, O\. Vinyals, N\. Jaitly, and N\. Shazeer \(2015\)Scheduled sampling for sequence prediction with recurrent neural networks\.InAdvances in Neural Information Processing Systems,Vol\.28\.Note:Internal notes\. Context: autoregressive models suffer from train\-test mismatch because inference conditions on their own previous outputs\. Breakthrough: scheduled sampling gradually exposes models to their own predictions during training\. Relevance: cite for autoregressive compounding of errors, though our model analyzes hidden\-state noise at decoding time rather than exposure bias during training\.External Links:[Link](https://proceedings.neurips.cc/paper/2015/hash/e995f98d56967d946471af29d7bf99f1-Abstract.html)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p1.1)\.
- M\. Braverman, X\. Chen, S\. Kakade, K\. Narasimhan, C\. Zhang, and Y\. Zhang \(2020\)Calibration, entropy rates, and memory in language models\.InProceedings of the 37th International Conference on Machine Learning,PMLR, Vol\.119,pp\. 1089–1099\.Note:Internal notes\. Context: do language models stay calibrated over long generations? Breakthrough: shows LM entropy rates drift upward with sequence position, diverging from the true process — empirical evidence of sequential miscalibration accumulating over tokens\. Relevance: the closest observational correlate of our mechanism; their unexplained drift is what our loop\-gain dynamics predict in the sustained\-wandering regime, so we reconcile rather than compete\.External Links:[Link](https://proceedings.mlr.press/v119/braverman20a.html)Cited by:[§2\.3](https://arxiv.org/html/2607.18292#S2.SS3.p1.2)\.
- K\. Chakrabarti and N\. Balachundhar \(2025\)Neural diversity regularizes hallucinations in language models\.Note:Internal notes\. Context: hallucination persists despite scaling parameters, compute, and data, leaving open whether the residual is a knowledge deficit or a reliability \(variance\) problem\. Breakthrough: shows that decorrelating parallel representations \(neural diversity, via ND\-LoRA = parallel LoRA adapters \+ Barlow Twins regularization\) reduces hallucination by up to 25\.6% \(14\.6% avg\) at FIXED parameter and data budgets, with general accuracy preserved; gives the first formal tail bounds for hallucination in ensembled LMs, reframing it as a second\-moment reliability problem that explains 94\.3% of reliability variation across parallel configurations\. Relevance: our own prior work and a key plank of the noise\-not\-knowledge argument — because the lever \(cross\-stream correlation, a second moment\) reduces hallucination without adding any knowledge or capacity, the residual driver it removes is sampling/variance error, not a knowledge gap\. Directly motivates the arnoise framing of hallucination as variance across parallel autoregressive streams and pairs with the bias/variance decomposition \(the variance term is what neural diversity suppresses\)\.External Links:2510\.20690,[Document](https://dx.doi.org/10.48550/arXiv.2510.20690),[Link](https://arxiv.org/abs/2510.20690)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p1.1),[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- H\. Chang, N\. Peng, M\. Bansal, A\. Ramakrishna, and T\. Chung \(2025\)REAL sampling: boosting factuality and diversity of open\-ended generation by extrapolating the entropy of an infinitely large LM\.Transactions of the Association for Computational Linguistics13\.Note:arXiv title: “Boosting Factuality and Diversity of Open\-Ended Generation via Asymptotic Entropy”Internal notes\. Context: the factuality\-vs\-diversity trade\-off in open\-ended sampling\. Breakthrough: a hallucination\-likelihood\-aware nucleus threshold \(extrapolating the entropy of an infinitely large LM\) that improves both\. Relevance: ties sampling stochasticity directly to factuality, relevant to exp1’s temperature monotonicity, and is a decoding\-time mitigation in the family our account motivates\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00757),2406\.07735,[Link](https://arxiv.org/abs/2406.07735)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- C\. Chen, K\. Liu, Z\. Chen, Y\. Gu, Y\. Wu, M\. Tao, Z\. Fu, and J\. Ye \(2024\)INSIDE: LLMs’ internal states retain the power of hallucination detection\.Note:Internal notes\. Context: output\-level uncertainty and self\-consistency methods discard dense semantic information inside the model\. Breakthrough: proposes EigenScore, a covariance/eigenvalue score over internal\-state sentence embeddings, and shows it improves hallucination detection across QA benchmarks\. Relevance: directly adjacent to our hidden\-state covariance story and should be cited as empirical evidence that internal representations retain hallucination signals\.External Links:2402\.03744,[Document](https://dx.doi.org/10.48550/arXiv.2402.03744),[Link](https://arxiv.org/abs/2402.03744)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p4.3),[§5](https://arxiv.org/html/2607.18292#S5.p4.1),[§6](https://arxiv.org/html/2607.18292#S6.p2.4)\.
- S\. Dhuliawala, M\. Komeili, J\. Xu, R\. Raileanu, X\. Li, A\. Celikyilmaz, and J\. Weston \(2024\)Chain\-of\-verification reduces hallucination in large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 3563–3578\.Note:Internal notes\. Context: can a model reduce its own hallucinations by deliberating over its draft? Breakthrough: CoVe drafts a response, plans and independently answers verification questions, then regenerates a verified response, cutting hallucination across list\-QA and long\-form tasks\. Relevance: a post\-hoc, response\-level verifier pipeline of the kind our onset risk signal complements \(and is distinct from—our leverage is at commitment, not after the full draft\)\.External Links:2309\.11495,[Link](https://aclanthology.org/2024.findings-acl.212/)Cited by:[§6](https://arxiv.org/html/2607.18292#S6.p2.4)\.
- S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. Gal \(2024\)Detecting hallucinations in large language models using semantic entropy\.Nature630,pp\. 625–630\.Note:Internal notes\. Context: token/string entropy overcounts paraphrases and underfits meaning\-level uncertainty\. Breakthrough: clusters sampled answers by semantic equivalence and computes entropy over meanings to detect confabulations\. Relevance: important because it separates arbitrary stochastic confabulation from systematic wrongness\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07421-0),[Link](https://doi.org/10.1038/s41586-024-07421-0)Cited by:[Appendix G](https://arxiv.org/html/2607.18292#A7.SS0.SSS0.Px1.p1.9),[§1](https://arxiv.org/html/2607.18292#S1.p4.3),[§2\.4](https://arxiv.org/html/2607.18292#S2.SS4.p1.3),[§4\.4](https://arxiv.org/html/2607.18292#S4.SS4.p1.6),[§5](https://arxiv.org/html/2607.18292#S5.p4.1),[§6](https://arxiv.org/html/2607.18292#S6.p2.4)\.
- J\. Ferrando, O\. Obeso, S\. Rajamanoharan, and N\. Nanda \(2025\)Do i know this entity? knowledge awareness and hallucinations in language models\.Note:Internal notes\. Context: a model may internally represent whether it recognizes an entity before it answers\. Breakthrough: uses sparse autoencoders to find entity\-recognition and uncertainty latents that causally steer refusal and hallucination behavior\. Relevance: sharpens our knowledge\-gap/noise distinction by showing that models can encode self\-knowledge, so stochastic hallucination on known facts is not reducible to absent knowledge\.External Links:2411\.14257,[Document](https://dx.doi.org/10.48550/arXiv.2411.14257),[Link](https://arxiv.org/abs/2411.14257)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p3.1)\.
- Z\. Gekhman, G\. Yona, R\. Aharoni, M\. Eyal, A\. Feder, R\. Reichart, and J\. Herzig \(2024\)Does fine\-tuning LLMs on new knowledge encourage hallucinations?\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Note:Internal notes\. Context: does fine\-tuning on facts unseen in pretraining increase hallucination? Breakthrough: partitions facts into known/unknown via greedy\-correctness and shows fine\-tuning on unknown facts encourages hallucination\. Relevance: the source of our operational "known fact" definition \(the greedy\-correct subset\) used in exp1\.External Links:2405\.05904,[Link](https://arxiv.org/abs/2405.05904)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p3.1)\.
- A\. Holtzman, J\. Buys, L\. Du, M\. Forbes, and Y\. Choi \(2020\)The curious case of neural text degeneration\.InInternational Conference on Learning Representations,Note:Internal notes\. Context: likelihood\-trained LMs behave poorly under maximization\-style decoding\. Breakthrough: nucleus sampling truncates the unreliable probability tail and improves diversity/quality tradeoffs\. Relevance: cite for decoding as an active source of generation behavior; our paper asks when stochastic decoding crosses factual tolerance\.External Links:[Link](https://openreview.net/forum?id=rygGQyrFvH)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. Liu \(2025\)A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions\.ACM Transactions on Information Systems\.Note:Internal notes\. Context: LLM hallucination differs from earlier task\-specific NLG because systems are open\-ended and general\-purpose\. Breakthrough: taxonomy of causes, detection, mitigation, and open questions for modern LLM hallucination\. Relevance: show reviewers we know the LLM\-specific survey landscape and are not proposing another taxonomy\.External Links:[Document](https://dx.doi.org/10.1145/3703155),[Link](https://doi.org/10.1145/3703155)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- Z\. Ji, N\. Lee, R\. Frieske, T\. Yu, D\. Su, Y\. Xu, E\. Ishii, Y\. J\. Bang, A\. Madotto, and P\. Fung \(2023\)Survey of hallucination in natural language generation\.ACM Computing Surveys55\(12\),pp\. 1–38\.Note:Internal notes\. Context: hallucination work was fragmented across summarization, dialogue, QA, data\-to\-text, and MT\. Breakthrough: unified intrinsic/extrinsic hallucination definitions, metrics, mitigation methods, and task\-specific taxonomies\. Relevance: position our contribution as adding a mechanistic noise\-vs\-knowledge axis\.External Links:[Document](https://dx.doi.org/10.1145/3571730),[Link](https://doi.org/10.1145/3571730)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- S\. Kadavath, T\. Conerly, A\. Askell, T\. Henighan, D\. Drain, E\. Perez, N\. Schiefer, Z\. Dodds, N\. DasSarma, E\. Tran\-Johnson, S\. Johnston, S\. El\-Showk, A\. Jones, N\. Elhage, T\. Hume, A\. Chen, Y\. Bai, S\. Bowman, S\. Fort, D\. Ganguli, D\. Hernandez, J\. Jacobson, J\. Kernion, S\. Kravec, L\. Lovitt, K\. Ndousse, C\. Olsson, S\. Ringer, D\. Amodei, T\. Brown, J\. Clark, N\. Joseph, B\. Mann, S\. McCandlish, C\. Olah, and J\. Kaplan \(2022\)Language models \(mostly\) know what they know\.Note:Internal notes\. Context: LMs might possess latent self\-knowledge about answer correctness\. Breakthrough: showed calibration and self\-evaluation improve with scale and prompt format\. Relevance: supports our known\-vs\-unknown framing, but our contribution is showing that knowing does not eliminate stochastic hallucination\.External Links:2207\.05221,[Link](https://arxiv.org/abs/2207.05221)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.18292#S2.SS3.p3.2),[§5](https://arxiv.org/html/2607.18292#S5.p3.1)\.
- A\. T\. Kalai, O\. Nachum, S\. S\. Vempala, and E\. Zhang \(2026\)Evaluating large language models for accuracy incentivizes hallucinations\.Nature653\(8116\),pp\. 1047–1051\.Note:Peer\-reviewed Nature version of kalai2025why \(arXiv:2509\.04664\)\. Same authors and thesis: standard accuracy\-graded evaluations reward guessing over abstention, so hallucination persists as a training/evaluation\-incentive consequence\. This is the canonical citation\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10549-w),[Link](https://www.nature.com/articles/s41586-026-10549-w)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p1.1),[§5](https://arxiv.org/html/2607.18292#S5.p5.1),[§6](https://arxiv.org/html/2607.18292#S6.p1.1)\.
- A\. T\. Kalai and S\. S\. Vempala \(2024\)Calibrated language models must hallucinate\.Note:Internal notes\. Context: hallucination may arise even with ideal data and calibrated pretraining\. Breakthrough: proves a statistical lower bound for arbitrary rare facts using a Good\-Turing\-style argument\. Relevance: central theory comparator: they explain pretraining/statistical inevitability, while we explain decoding\-time stochastic failures on facts the model can answer\.External Links:2311\.14648,[Link](https://arxiv.org/abs/2311.14648)Cited by:[§2\.2](https://arxiv.org/html/2607.18292#S2.SS2.p2.2),[§5](https://arxiv.org/html/2607.18292#S5.p5.1),[§6](https://arxiv.org/html/2607.18292#S6.p1.1)\.
- S\. Kang, Y\. F\. Bakman, D\. N\. Yaldiz, B\. Buyukates, and S\. Avestimehr \(2025\)Uncertainty quantification for hallucination detection in large language models: foundations, methodology, and future directions\.Note:Internal notes\. Context: hallucination detection increasingly relies on uncertainty but the methods vary widely in access requirements and interpretation\. Breakthrough: organizes UQ methods for LLM hallucination detection across token probabilities, consistency, internal states, and self\-checking\. Relevance: citation anchor for the uncertainty\-detection landscape and for explaining why our bound gives a theoretical account of variance\-based diagnostics\.External Links:2510\.12040,[Document](https://dx.doi.org/10.48550/arXiv.2510.12040),[Link](https://arxiv.org/abs/2510.12040)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p4.1)\.
- J\. Kossen, J\. Han, M\. Razzak, L\. Schut, S\. Malik, and Y\. Gal \(2024\)Semantic entropy probes: robust and cheap hallucination detection in LLMs\.Note:Internal notes\. Context: semantic entropy needs many samples per query, which is too costly to deploy\. Breakthrough: linear probes on hidden states that cheaply approximate semantic entropy from a single generation\. Relevance: a cheap semantic\-entropy\-style baseline our variance signal could be compared against if we benchmark detection\.External Links:2406\.15927,[Link](https://arxiv.org/abs/2406.15927)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p4.1)\.
- L\. Kuhn, Y\. Gal, and S\. Farquhar \(2023\)Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.In11th International Conference on Learning Representations \(ICLR 2023\),Note:SpotlightInternal notes\. Context: token\-level uncertainty conflates meaning\-equivalent generations, mis\-estimating uncertainty in free\-form generation\. Breakthrough: semantic entropy, clustering sampled generations by meaning and taking entropy over clusters\. Relevance: the reference uncertainty method; our across\-sample variance signal is a token\-level, oracle\-referenced cousin we must distinguish from\.External Links:2302\.09664,[Link](https://arxiv.org/abs/2302.09664)Cited by:[Appendix G](https://arxiv.org/html/2607.18292#A7.SS0.SSS0.Px1.p1.9),[§1](https://arxiv.org/html/2607.18292#S1.p4.3),[§2\.4](https://arxiv.org/html/2607.18292#S2.SS4.p1.3),[§4\.4](https://arxiv.org/html/2607.18292#S4.SS4.p1.6),[§5](https://arxiv.org/html/2607.18292#S5.p4.1),[§6](https://arxiv.org/html/2607.18292#S6.p2.4)\.
- T\. Kwa, B\. West,et al\.\(2025\)Measuring AI ability to complete long tasks\.arXiv preprint arXiv:2503\.14499\.Note:Internal notes\. Context: capability scores do not say how long a model can act autonomously before failing\. Breakthrough: METR’s 50%\-task\-completion time horizon — the human task length a model completes with 50% reliability — has doubled roughly every seven months\. Relevance: reliability quantified as a reliably\-completed task length, a behavioral/outcome\-level autonomy metric\. Complements our within\-sequence positional view; we model why per\-step reliability decays as the rollout lengthens\. NOTE: verify full author list before camera\-ready\.External Links:[Link](https://arxiv.org/abs/2503.14499)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p6.1)\.
- N\. Lee, W\. Ping, P\. Xu, M\. Patwary, P\. Fung, M\. Shoeybi, and B\. Catanzaro \(2022\)Factuality enhanced language models for open\-ended text generation\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 34586–34599\.Note:Internal notes\. Context: open\-ended factuality degrades under common sampling choices\. Breakthrough: introduced FactualityPrompts and found sampling methods such as top\-p can harm factuality via repeated randomness\. Relevance: directly supports our claim that hallucination can be a decoding\-time noise phenomenon, not only a knowledge gap\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/df438caa36714f69277daa92d608dd63-Abstract-Conference.html)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022\)TruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics,pp\. 3214–3252\.Note:Internal notes\. Context: standard accuracy benchmarks did not test whether models imitate common false beliefs\. Breakthrough: introduced 817 misconception\-sensitive questions and showed larger models could be less truthful\. Relevance: contrasts with our known\-fact under stochastic decoding setup, where the model has already demonstrated the fact under greedy decoding\.External Links:[Link](https://aclanthology.org/2022.acl-long.229/)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p7.1)\.
- L\. Liu, R\. Pourreza, S\. Panchal, A\. Bhattacharyya, Y\. Jian, Y\. Qin, and R\. Memisevic \(2025\)Enhancing hallucination detection through noise injection\.Note:Internal notes\. Context: sampling from next\-token distributions captures aleatoric variation but may miss model uncertainty\. Breakthrough: perturbs model parameters or hidden activations during inference to expose epistemic uncertainty and improve hallucination detection\. Relevance: directly related to our noise hypothesis, though it injects noise as a diagnostic intervention while we model naturally occurring autoregressive hidden\-state noise\.External Links:2502\.03799,[Document](https://dx.doi.org/10.48550/arXiv.2502.03799),[Link](https://arxiv.org/abs/2502.03799)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p2.1)\.
- Llama Team, AI @ Meta \(2024\)The Llama 3 herd of models\.Note:Internal notes\. The independent third family used for the onset\-dynamics, bridge, and intervention replication; Llama\-3\.2\-1B/3B\-Instruct students against the Llama\-3\.1\-8B\-Instruct oracle\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[Appendix A](https://arxiv.org/html/2607.18292#A1.SS0.SSS0.Px1.p1.15),[Table A\.1](https://arxiv.org/html/2607.18292#A1.T1.2.13.11.1)\.
- P\. Manakul, A\. Liusie, and M\. J\. F\. Gales \(2023\)SelfCheckGPT: zero\-resource black\-box hallucination detection for generative large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 9004–9017\.Note:Internal notes\. Context: black\-box LLMs often lack token probabilities or accessible internals\. Breakthrough: detects hallucinations through sampling disagreement without external databases\. Relevance: operational evidence that stochastic inconsistency is diagnostic; our theory explains why repeated samples expose noise\-driven hallucination\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.557/)Cited by:[Appendix G](https://arxiv.org/html/2607.18292#A7.SS0.SSS0.Px1.p1.9),[§1](https://arxiv.org/html/2607.18292#S1.p4.3),[§2\.4](https://arxiv.org/html/2607.18292#S2.SS4.p1.3),[§5](https://arxiv.org/html/2607.18292#S5.p4.1),[§6](https://arxiv.org/html/2607.18292#S6.p2.4)\.
- I\. R\. McKenzie, A\. Lyzhov, M\. Pieler, A\. Parrish, A\. Mueller, A\. Prabhu, E\. McLean, A\. Kirtland, A\. Ross, A\. Liu, A\. Gritsevskiy, D\. Wurgaft, D\. Kauffman, G\. Recchia, J\. Liu, J\. Cavanagh, M\. Weiss, S\. Huang, T\. F\. Droid, T\. Tseng, T\. Korbak, X\. Shen, Y\. Zhang, Z\. Zhou, N\. Kim, S\. R\. Bowman, and E\. Perez \(2023\)Inverse scaling: when bigger isn’t better\.Transactions on Machine Learning Research \(TMLR\)\.Note:Internal notes\. Context: scaling usually improves task performance, but some tasks reverse\. Breakthrough: the Inverse Scaling Prize collected tasks where larger models do worse, attributing them to strong priors, distractor following, and unfaithful imitation of training\-data patterns\. Relevance: names the phenomenon class our reliability anti\-scaling belongs to; we distinguish ours as a factuality degradation with an identified decoding\-time mechanism rather than a prompt\-format artifact\.External Links:[Link](https://openreview.net/forum?id=DwgRm72GQF)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p7.1)\.
- S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. W\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. Hajishirzi \(2023\)FActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 12076–12100\.Note:Internal notes\. Context: long\-form generations mix true and false claims, making binary factuality labels too coarse\. Breakthrough: decomposes generations into atomic facts and evaluates support from reliable sources\. Relevance: supports our long\-form/sequence\-length evaluation and why claim\-level factuality scales differently from short\-answer QA\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.741/)Cited by:[Appendix A](https://arxiv.org/html/2607.18292#A1.SS0.SSS0.Px2.p1.1)\.
- H\. Orgad, M\. Toker, Z\. Gekhman, R\. Reichart, I\. Szpektor, H\. Kotek, and Y\. Belinkov \(2025\)LLMs know more than they show: on the intrinsic representation of LLM hallucinations\.In13th International Conference on Learning Representations \(ICLR 2025\),Note:Internal notes\. Context: where and how is truthfulness encoded in an LLM’s internal states? Breakthrough: truthfulness concentrates in specific answer tokens, models can internally encode the correct answer yet generate a wrong one, and internal states predict the error type\. Relevance: the representational counterpart to our "biased and high\-variance at the answer token" finding; we own the distributional/decoding version, they own the probe version\.External Links:2410\.02707,[Link](https://arxiv.org/abs/2410.02707)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p3.1)\.
- E\. Perez, S\. Ringer, K\. Lukošiūtė, K\. Nguyen, E\. Chen, S\. Heiner, C\. Pettit, C\. Olsson, S\. Kundu, S\. Kadavath,et al\.\(2023\)Discovering language model behaviors with model\-written evaluations\.InFindings of the Association for Computational Linguistics: ACL 2023,Note:Internal notes\. Context: scaling and RLHF can amplify undesirable behaviors\. Breakthrough: model\-written evaluations show sycophancy and several stated\-preference behaviors strengthen with model size and RLHF steps\. Relevance: a behavioral parallel to our reliability anti\-scaling — a failure that grows with capability rather than shrinking — though ours is a token\-level factuality mechanism, not a stated preference\.External Links:[Link](https://aclanthology.org/2023.findings-acl.847/)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p7.1)\.
- Qwen Team \(2025\)Qwen3 technical report\.Note:Internal notes\. The single model family \(0\.6–32B, Apache\-2\.0\) under test across all experiments; holding architecture fixed across scale isolates size\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[Appendix A](https://arxiv.org/html/2607.18292#A1.SS0.SSS0.Px1.p1.15),[Table A\.1](https://arxiv.org/html/2607.18292#A1.T1.2.4.2.1)\.
- S\. Rabanser, S\. Kapoor, A\. Narayanan,et al\.\(2026\)Towards a science of AI agent reliability\.arXiv preprint arXiv:2602\.16666\.Note:Internal notes\. Context: agent capability gains have not translated into reliable deployed products\. Breakthrough: a reliability framework of 12 metrics across four dimensions \(consistency, robustness, predictability, safety\), grounded in safety\-critical engineering; two years of capability gains across 14 models yielded only modest reliability gains\. Relevance: the explicit science\-of\-reliability framing \(reliability as an axis distinct from capability, gated by consistency/predictability/control\)\. We supply the missing generative mechanism behind one such metric\. NOTE: verify full author list and arXiv id before camera\-ready\.External Links:[Link](https://arxiv.org/abs/2602.16666)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p6.1)\.
- M\. Ranzato, S\. Chopra, M\. Auli, and W\. Zaremba \(2016\)Sequence level training with recurrent neural networks\.In4th International Conference on Learning Representations \(ICLR 2016\),Note:Internal notes\. Context: word\-level cross\-entropy training induces exposure bias and a mismatch with sequence\-level evaluation metrics\. Breakthrough: MIXER, sequence\-level training via REINFORCE to optimize the eval metric directly\. Relevance: the other root of the exposure\-bias lineage; paired with Bengio 2015 to frame the classical account our self\-conditioning mechanism is distinguished from\.External Links:1511\.06732,[Link](https://arxiv.org/abs/1511.06732)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p1.1)\.
- A\. Simhi, I\. Itzhak, F\. Barez, G\. Stanovsky, and Y\. Belinkov \(2025\)Trust me, I’m wrong: LLMs hallucinate with certainty despite knowing the answer\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 14665–14688\.Note:Internal notes\. Context: are confident hallucinations on facts the model knows merely sampling noise? Breakthrough: CHOKE, showing 16\-43% of hallucinations occur with high certainty on known facts, consistent across prompt phrasings, argued to be "not noise artifacts\." Relevance: the most direct published challenge to a noise framing; we subsume it rather than contradict it \(their cross\-prompt consistency = our bias term; our across\-sample spread = variance\), so it becomes the high\-certainty quadrant of our own decomposition\.External Links:2502\.12964,[Link](https://aclanthology.org/2025.findings-emnlp.792/)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p1.1),[§5](https://arxiv.org/html/2607.18292#S5.p3.1)\.
- J\. Wei, C\. Yang, X\. Song, Y\. Lu, N\. Hu, J\. Huang, D\. Tran, D\. Peng, R\. Liu, D\. Huang, C\. Du, and Q\. V\. Le \(2024\)Long\-form factuality in large language models\.Note:Internal notes\. Context: open\-domain long\-form factuality needed scalable evaluation beyond biographies\. Breakthrough: introduced LongFact and SAFE, using search\-backed LLM judging for long\-form factual accuracy\. Relevance: directly supports our proposed sequence\-length scan and socio\-technical claim that long outputs carry disproportionate risk\.External Links:2403\.18802,[Link](https://arxiv.org/abs/2403.18802)Cited by:[Appendix A](https://arxiv.org/html/2607.18292#A1.SS0.SSS0.Px2.p1.1)\.
- Y\. Xiao and W\. Y\. Wang \(2021\)On hallucination and predictive uncertainty in conditional language generation\.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics,pp\. 2734–2744\.Note:Internal notes\. Context: hallucination causes were studied separately across generation tasks\. Breakthrough: showed predictive uncertainty, especially epistemic uncertainty, correlates with hallucination and can guide decoding\. Relevance: closest prior to our uncertainty story, but we provide a hidden\-state covariance bound and effective\-rank scaling law\.External Links:[Link](https://aclanthology.org/2021.eacl-main.236/)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p4.1)\.
- J\. Yang, S\. Yoon, H\. Chang, B\. Kim, and H\. Lee \(2025\)Hallucinate at the last in long response generation: a case study on long document summarization\.Note:Internal notes\. Context: how does factual faithfulness vary across position in long\-form generation? Breakthrough: shows faithfulness declines steadily toward the end of long outputs across several model families, attributing it to attention dynamics\. Relevance: the closest positional\-degradation result, but observational; our causal handle on why later positions degrade \(self\-conditioning\) is the differentiator\.External Links:2505\.15291,[Link](https://arxiv.org/abs/2505.15291)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p1.1)\.
- L\. Yu, M\. Cao, J\. C\. K\. Cheung, and Y\. Dong \(2024\)Mechanistic understanding and mitigation of language model non\-factual hallucinations\.Note:Internal notes\. Context: many hallucination papers treat errors as black\-box output phenomena\. Breakthrough: decomposes non\-factual hallucinations into lower\-layer knowledge\-enrichment failures and upper\-layer answer\-extraction failures, validated with logit lens and causal patching\. Relevance: complementary mechanistic account; our paper focuses on stochastic noise and effective\-rank structure, while this one identifies circuit\-level failure stages\.External Links:2403\.18167,[Document](https://dx.doi.org/10.48550/arXiv.2403.18167),[Link](https://arxiv.org/abs/2403.18167)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p3.1)\.
- M\. Zhang, O\. Press, W\. Merrill, A\. Liu, and N\. A\. Smith \(2024\)How language model hallucinations can snowball\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 59670–59684\.Note:Internal notes\. NOTE: key says 2023 \(arXiv\) but this is the published ICML 2024 version\. Context: do LM hallucinations propagate within a single generation? Breakthrough: names "hallucination snowballing" and shows models commit to a wrong answer early then produce consistent\-but\-wrong justifications they can recognize as false in isolation, on hand\-built QA sets\. Relevance: our direct predecessor and the term we adopt; their evidence is observational/diagnostic, whereas we provide the first on\-policy, causal, whole\-response, scale\-resolved test of the same phenomenon\.External Links:2305\.13534,[Link](https://proceedings.mlr.press/v235/zhang24ay.html)Cited by:[§1](https://arxiv.org/html/2607.18292#S1.p1.1),[§2\.3](https://arxiv.org/html/2607.18292#S2.SS3.p3.2),[§5](https://arxiv.org/html/2607.18292#S5.p1.1),[§6](https://arxiv.org/html/2607.18292#S6.p3.4)\.
- Y\. Zhang, S\. Li, C\. Qian, J\. Liu, P\. Yu, C\. Han, Y\. R\. Fung, K\. McKeown, C\. Zhai, M\. Li, and H\. Ji \(2025\)The law of knowledge overshadowing: towards understanding, predicting, and preventing LLM hallucination\.Note:Internal notes\. Context: factual hallucination can arise even when relevant knowledge exists but competes with more dominant associations\. Breakthrough: proposes knowledge overshadowing and a log\-linear law relating hallucination to knowledge popularity, knowledge length, and model size, plus a decoding intervention\. Relevance: related theory\-style framing; our paper differs by analyzing covariance/effective\-rank noise rather than popularity\-driven knowledge competition\.External Links:2502\.16143,[Document](https://dx.doi.org/10.48550/arXiv.2502.16143),[Link](https://arxiv.org/abs/2502.16143)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p5.1),[§6](https://arxiv.org/html/2607.18292#S6.p1.1)\.
- W\. Zhong, X\. Feng, L\. Zhao, Q\. Li, L\. Huang, Y\. Gu, W\. Ma, Y\. Xu, and B\. Qin \(2024\)Investigating and mitigating the multimodal hallucination snowballing in large vision\-language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 11991–12011\.Note:Internal notes\. Context: do vision\-language models snowball hallucinations across multi\-turn conversation, and can it be mitigated? Breakthrough: constructs curated hallucinatory conversations, shows models answer consistently with their own prior hallucination, and proposes a residual\-visual decoding fix \(MMHalSnowball\)\. Relevance: the closest interventional snowball work and our most dangerous competitor; we differ on every axis \(text long\-form not multimodal; on\-policy free continuation not curated context; whole\-response not response\-to\-planted; scale\-resolved\)\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.648),2407\.00569,[Link](https://aclanthology.org/2024.acl-long.648/)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p1.1)\.
- L\. Zhou, W\. Schellaert, F\. Martínez\-Plumed, Y\. Moros\-Daval, C\. Ferri, and J\. Hernández\-Orallo \(2024\)Larger and more instructable language models become less reliable\.Nature634\(8032\),pp\. 61–68\.Note:Internal notes\. Context: capability benchmarks improve monotonically with scale and shaping, but it was unclear whether reliability tracks them\. Breakthrough: across GPT/LLaMA/BLOOM families and ReliabilityBench, scaling and instruction\-tuning raise capability while degrading reliability — difficulty discordance \(failures on easy items\), and avoidance traded for confident incorrectness, erasing any safe\-operating region\. Purely behavioral, output\-level, cross\-instance difficulty axis; no representation analysis\. Relevance: the closest framing of capability/reliability decoupling\. We reproduce it cross\-family along the autoregressive\-position axis instead of the difficulty axis, decompose it into an intercept \(knowledge\) and a slope \(accumulation\), and trace the slope to a latent noise process they cannot observe\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07930-y),[Link](https://doi.org/10.1038/s41586-024-07930-y)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p6.1)\.
- N\. Zucchet, J\. Bornschein, S\. Chan, A\. Lampinen, R\. Pascanu, and S\. De \(2025\)How do language models learn facts? dynamics, curricula and hallucinations\.Note:Internal notes\. Context: factual recall emerges during training through poorly understood dynamics\. Breakthrough: identifies phases of fact learning and shows hallucinations can emerge simultaneously with knowledge acquisition in synthetic factual\-recall settings\. Relevance: supports the distinction between learning/knowledge formation and decoding\-time reliability; our paper conditions on what the model already knows\.External Links:2503\.21676,[Document](https://dx.doi.org/10.48550/arXiv.2503.21676),[Link](https://arxiv.org/abs/2503.21676)Cited by:[§5](https://arxiv.org/html/2607.18292#S5.p5.1)\.
## Appendix AModels, Data, and Verifier Protocol
#### Models\.
Results span three open model families\. The Qwen3 family\(Qwen Team,[2025](https://arxiv.org/html/2607.18292#bib.bib61)\)\(0\.60\.6–3232B;[Table A\.1](https://arxiv.org/html/2607.18292#A1.T1)\) carries the scaling sweep, the cross\-sectional and over\-commitment analyses, and the position fit; the OLMo\-3 family\(Ai2,[2025](https://arxiv.org/html/2607.18292#bib.bib62)\)adds an independent second\-family rung \(OLMo\-3\-7B\-Instruct model, OLMo\-3\.1\-32B\-Instruct oracle; the family ships only these two sizes\) and the Llama\-3 family\(Llama Team, AI @ Meta,[2024](https://arxiv.org/html/2607.18292#bib.bib63)\)a third \(Llama\-3\.2\-1B/3B\-Instruct models, Llama\-3\.1\-8B\-Instruct oracle\) to the onset\-dynamics, bridge, and intervention results\. All run in their no\-thinking configuration\. Holding architecture fixed across scale within a family is what makes the scaling statements \([Section 3](https://arxiv.org/html/2607.18292#S3)\) attributable to size rather than to family idiosyncrasy; the second and third families guard the cross\-family claims against Qwen3\-specific idiosyncrasy\. The oracle log\-prob error \([Section 2](https://arxiv.org/html/2607.18292#S2)\) references each model against a larger same\-family oracle, chosen per analysis\. The gap\-MSAR over\-commitment fit \([Table H\.1](https://arxiv.org/html/2607.18292#A8.T1)\) uses Qwen3\-14B over models0\.60\.6–88B\. The within\-position cross\-section \([Figure 4](https://arxiv.org/html/2607.18292#S3.F4)\) and the variance\-contraction intervention use the family’s largest model: Qwen3\-32B over Qwen models \(0\.60\.6–1414B for the cross\-section,44–1414B for the intervention\) and OLMo\-3\.1\-32B over OLMo\-3\-7B\. The Qwen onset overlay and the onset half\-lives \([Figure 2](https://arxiv.org/html/2607.18292#S1.F2),[Table C\.1](https://arxiv.org/html/2607.18292#A3.T1)\) share one re\-score — models0\.60\.6–88B against the Qwen3\-14B oracle — and[Appendix D](https://arxiv.org/html/2607.18292#A4)repeats the onset analysis against the Qwen3\-32B oracle, finding the asymmetry unchanged\. The position fit spans the full0\.60\.6–3232B Qwen sweep\. All decoding isτ=1\.0\\tau\{=\}1\.0,top\-p=1\.0\\mathrm\{top\\text\{\-\}\}p\{=\}1\.0unless noted; the device is a single GPU \(Modal L40S/H100 in the cloud, MPS locally\)\. The full project \(inference\-only, with no training or fine\-tuning\) fits within roughly2525GPU\-hours\.
ModelParamsRole*Qwen3*\(Qwen Team,[2025](https://arxiv.org/html/2607.18292#bib.bib61)\)Qwen3\-0\.6B0\.6BmodelQwen3\-1\.7B1\.7BmodelQwen3\-4B4BmodelQwen3\-8B8BmodelQwen3\-14B14Bmodel; oracle \(0\.60\.6–88B\)Qwen3\-32B32Bmodel; oracle \(onset, interv\.\)*OLMo\-3*\(Ai2,[2025](https://arxiv.org/html/2607.18292#bib.bib62)\)OLMo\-3\-7B\-Instruct7BmodelOLMo\-3\.1\-32B\-Instruct32Boracle \(OLMo\)*Llama\-3*\(Llama Team, AI @ Meta,[2024](https://arxiv.org/html/2607.18292#bib.bib63)\)Llama\-3\.2\-1B\-Instruct1BmodelLlama\-3\.2\-3B\-Instruct3BmodelLlama\-3\.1\-8B\-Instruct8Boracle \(Llama\)Table A\.1:Models under test, grouped by family\. Within each family the oracle is the largest same\-family model; Qwen3\-14B is the oracle for the over\-commitment analysis and Qwen3\-32B for the within\-position, onset, and intervention analyses\. OLMo\-3 and Llama\-3 add two independently\-pretrained families to the onset\-dynamics, bridge, and intervention results\. One architecture across scale isolates size\.
#### Datasets\.
Two public, English benchmarks, one per measurement role\.FActScorebiography prompts\(Minet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib5)\)\(6464sampled topics\) supply long\-form generations whose atomic facts are individually verifiable; these drive the position fit and the teacher\-forced decompositions\.LongFact\+\+augmented prompts\(Weiet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib6)\)supply the free\-run hallucinating trajectories the onset and intervention analyses run on\. Both are released for research use \(FActScore MIT, LongFact Apache\-2\.0\), as are the Qwen3 weights \(Apache\-2\.0\); our non\-commercial academic use is consistent with each license, and we collect and release no new data\.
#### Verifier protocol\.
Claim labels come from a Claude verifier with web search\. For a generated response the verifier \(i\) extracts atomic factual claims, \(ii\) marks each claim’s token span, and \(iii\) labels each*supported*\(S\) or*not supported*\(U\) by issuing web\-search queries and checking the claim against retrieved evidence\. The*fabrication onset*is the first token of the first U\-labeled claim\. All support/unsupport counts, snowball pairs, and intervention outcomes derive from these labels; systematic verifier error would propagate, so we rely on cross\-measurement consistency \(a verifier\-independent white\-box commitment marker agrees with the labels at AUROC0\.680\.68–0\.710\.71,[Appendix H](https://arxiv.org/html/2607.18292#A8)\) rather than re\-validating the verifier here\. Double\-judge robustness and a verifier\-temperature sweep are open\.
#### Reproducibility\.
Seeds are fixed at the outermost level and per\-worker RNG is derived deterministically; cloud runs useaws s3 syncfor artifact persistence and resume to a fresh\-run\-identical state\. Every reported number is piped from one computation in the corresponding analysis pipeline and regenerated from its source artifact, never hand\-copied\. All confidence intervals are95%95\\%bootstrap intervals resampling the unit of analysis \(trajectories for the snowball and intervention results; claim\-pairs for the prefix\-swap\), and the source of variability is that resampling\.
## Appendix BScaling and Knowledge Gaps
[Table B\.1](https://arxiv.org/html/2607.18292#A2.T1)expands the two\-sided trend of[Table 1](https://arxiv.org/html/2607.18292#S2.T1)with the underlying FActScore, facts\-per\-response, abstention rate, and the raw position slope\. The knowledge gap \(intercept\) falls monotonically with scale and the knowledge degradation \(slope÷\\divintercept\) rises monotonically; we report the degradation ratio because it normalizes the per\-token accumulation slope by the start\-of\-response gap\.
Table B\.1:Full scaling fit behind[Table 1](https://arxiv.org/html/2607.18292#S2.T1)\.*knowledge gap*is the OLS intercept \(start\-of\-response hallucination rate\) and*knowledge degradation*is slope÷\\divintercept, both in %\. The knowledge gap falls and the knowledge degradation rises with scale\.#### The knowledge\-degradation metric is a mixed\-effects fit\.
A linear\-probability OLS slope÷\\divintercept ratio divides by the start\-of\-response gap, so its bootstrap CI explodes at the larger rungs where that intercept is small \(the wide intervals the two\-sided fit would otherwise carry\)\.[Table 1](https://arxiv.org/html/2607.18292#S2.T1)therefore reports knowledge degradation from a per\-claim random\-intercept logistic GLMM instead,logitP\(supported\)=β0\+β1position\+uresp\\mathrm\{logit\}\\,P\(\\text\{supported\}\)=\\beta\_\{0\}\+\\beta\_\{1\}\\,\\text\{position\}\+u\_\{\\text\{resp\}\}, with a random intercept per response absorbing topic/length baseline variation\. The reported degradation is the*same*quantity as the OLS metric — the relative rise in hallucination from the start \(position0\) to the end \(position11\) of a response,\[phall\(1\)−phall\(0\)\]/phall\(0\)\[\\,p\_\{\\text\{hall\}\}\(1\)\-p\_\{\\text\{hall\}\}\(0\)\\,\]/p\_\{\\text\{hall\}\}\(0\)— but read off the logistic fit, with a95%95\\%credible interval from Monte\-Carlo draws of\(β0,β1\)\(\\beta\_\{0\},\\beta\_\{1\}\)under the mean\-field variational posterior; the result is the tight degradation CIs in[Table 1](https://arxiv.org/html/2607.18292#S2.T1)\.[Table B\.2](https://arxiv.org/html/2607.18292#A2.T2)reports the underlying fit: the per\-position slopeβpos\\beta\_\{\\text\{pos\}\}is a significant positional effect at nearly every model×\\timeseval fit, so the degradation trend is genuine, not an artifact of the ratio’s shrinking denominator\.
Table B\.2:Mixed\-effects fit internals behind the knowledge\-degradation column of[Table 1](https://arxiv.org/html/2607.18292#S2.T1)\.Per Table\-1 row and eval, the random\-intercept logistic GLMM of atomic\-fact support on relative within\-response position:βpos\\beta\_\{\\text\{pos\}\}is the per\-position support log\-odds slope \(negated to read as degradation\) with a Wald 95% CI, andτresp\\tau\_\{\\text\{resp\}\}the random\-intercept SD \(across\-response baseline spread\)\. The per\-position effect is significant \(∗, CI excludes0\) at 20 of 22 model×\\timeseval fits, so the degradation in[Table 1](https://arxiv.org/html/2607.18292#S2.T1)is a genuine positional effect rather than an artifact of dividing by a shrinking start\-of\-response gap\.Figure B\.1:Hallucination rises roughly linearly with relative position, for both atomic facts \(left\) and tokens \(right\)\. Responses with≥5\\geq 5atomic facts, coarse\-binned into five intervals \(the first pools each response’s opening facts\); binning rather than smoothing avoids over\-fitting the sparse near\-zero region of thefact\_idx/\(nfacts−1\)\\text\{fact\\\_idx\}/\(n\_\{\\text\{facts\}\}\{\-\}1\)normalization\. This per\-response accumulation, its rate divided by the start\-of\-response gap, gives the knowledge\-degradation metric of[Table 1](https://arxiv.org/html/2607.18292#S2.T1)\.
#### Decoding ablation\.
The moment analyses are teacher\-forced and so independent of decoding, but the free\-run snowball is not\. We re\-run it at greedy \(τ=0\\tau\{=\}0\) and low\-temperature \(τ=0\.5\\tau\{=\}0\.5\) decoding on a shared3232\-prompt subset and recompute the across\-claim relative risk \([Table B\.3](https://arxiv.org/html/2607.18292#A2.T3)\)\. It stays above11at every rung and is larger at the biggest model than the smallest under all three policies, so the compounding is not an artifact of sampling temperature; the small subset adds noise at the middle rungs\.
Table B\.3:Reliability anti\-scaling survives alternative decoding\.Across\-claim snowball relative riskRR=P\(Uj\+1∣Uj\)/P\(Uj\+1∣Sj\)\\mathrm\{RR\}=P\(U\_\{j\+1\}\\mid U\_\{j\}\)/P\(U\_\{j\+1\}\\mid S\_\{j\}\)on mixed responses \([Section 4\.1](https://arxiv.org/html/2607.18292#S4.SS1)\), per Qwen3 size, under greedy \(τ=0\\tau\{=\}0\), low\-temperature \(τ=0\.5\\tau\{=\}0\.5\), and the canonicalτ=1\.0\\tau\{=\}1\.0run, on a shared3232\-prompt subset\. The relative risk stays above11at every rung and is larger at the biggest model than the smallest under all three decoding policies \(the small subset adds noise at the middle rungs\), so the across\-claim snowball is not an artifact of sampling temperature\.
## Appendix COnset Dynamics
#### The self\-readable channel is impulsive; the oracle\-referenced channels persist\.
Re\-scoring the full next\-token distribution around onset on the free\-run trajectories, every disagreement channel spikes together at onset, but only the oracle\-referenced channels linger \([Table C\.1](https://arxiv.org/html/2607.18292#A3.T1),[Figure 2](https://arxiv.org/html/2607.18292#S1.F2)\): the self\-readableH\(pM\)H\(p\_\{M\}\)is impulsive — it admits no clean exponential decay at the smaller rungs and collapses to a0\.30\.3\-token half\-life at88B, relaxing to1\.11\.1–1\.5×1\.5\\timesits pre\-onset baseline — while the oracle\-referenced channels persist\. The divergence𝔼\[δ\]\\mathbb\{E\}\[\\delta\]holds at1\.41\.4–1\.9×1\.9\\timespre\-onset \(half\-life3\.33\.3–3\.93\.9tokens\), and the commitment riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}is the longest\-lived, decaying1\.11\.1–1\.3×1\.3\\timesmore slowly still \(half\-life3\.53\.5–5\.05\.0tokens\)\. The onset spike is partly definitional \(t=t⋆t\{=\}t^\{\\star\}is the first oracle\-disagreeing claim\), so the load\-bearing observation is the pre\-onset variance floor, not the spike\.
Table C\.1:At a fabrication’s onset the commitment risk is the longest\-lived channel, the bias gap shorter, and the model’s self\-readable uncertainty the most impulsive\.The three[Figure 2](https://arxiv.org/html/2607.18292#S1.F2)channels of the disagreement variable \([Equation 2](https://arxiv.org/html/2607.18292#S2.E2)\) in nats — the divergence𝔼\[δ\]=KL\(pM∥pO\)\\mathbb\{E\}\[\\delta\]\{=\}\\mathrm\{KL\}\(p\_\{M\}\\\|p\_\{O\}\), the commitment stdVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}, and the felt uncertaintyH\(pM\)H\(p\_\{M\}\)— onset\-aligned on free\-running hallucinating trajectories \(LongFact\+\+; onsett0⋆t\_\{0\}^\{\\star\}= first unsupported atomic fact\), the SAME frame[Figure 2](https://arxiv.org/html/2607.18292#S1.F2)plots\. Models \(Qwen3\-0\.6B–8B\) scored against the Qwen3\-14B oracle\. Per channel: the half\-lifet1/2t\_\{1/2\}\(exponential\-envelope fitAe−λsA\\,e^\{\-\\lambda s\}to the post\-onset excess overs∈\[0,10\]s\{\\in\}\[0,10\]tokens,ln2/λ\\ln 2/\\lambda;−−\-\-where no clean exponential forms,r2<0\.80r^\{2\}\{<\}0\.80\), the onset spike \(on//pre\), and post\-onset persistence \(po//pre,t−t⋆∈\[1,20\]t\-t^\{\\star\}\{\\in\}\[1,20\]\) as fold\-changes over the pre\-onset baseline \(t−t⋆∈\[−5,−1\]t\-t^\{\\star\}\{\\in\}\[\-5,\-1\]\)\. The commitmentVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}relaxes slowest \(3\.53\.5–5\.05\.0tokens\), the bias gap𝔼\[δ\]\\mathbb\{E\}\[\\delta\]faster \(3\.33\.3–3\.93\.9\), and the self\-readableH\(pM\)H\(p\_\{M\}\)fastest where an exponential forms at all \(0\.30\.3–1\.71\.7at the larger rungs\.
## Appendix DOracle Sensitivity
#### The onset asymmetry survives a change of oracle size\.
δ\\deltais divergence from a particular oracle, so a fair worry is that the persistence asymmetry is an idiosyncrasy of one reference model\. It is not: re\-scoring the same onset analysis against the Qwen3\-14B oracle \(the cross\-sectional reference\) instead of the Qwen3\-32B oracle leaves every qualitative invariant intact \([Table D\.1](https://arxiv.org/html/2607.18292#A4.T1)\)\. Under both oracles the oracle\-referenced channels — divergence𝔼\[δ\]\\mathbb\{E\}\[\\delta\]and commitment riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}— stay elevated post\-onset \(the risk half\-life exceeds the realized\-bias half\-lifeδ\(yt\)¯2\\overline\{\\delta\(y\_\{t\}\)\}^\{2\}at every rung\), while the self\-readableH\(pM\)H\(p\_\{M\}\)relaxes closest to baseline — the same proxy break\. The two columns are scored on independently sampled free\-run trajectory sets at different nucleus truncations, so the exact half\-lives differ and are not expected to match; that the*asymmetry*holds across oracle size on independent data, rather than on one shared sample, is the robustness claim\. A cross\-family oracle \(via vocabulary alignment\) and a temperature\-perturbed oracle would extend this to oracle*family*and sharpness; both are open \([Limitations](https://arxiv.org/html/2607.18292#Sx1)\)\.
Table D\.1:The onset asymmetry is invariant to oracle size\.The persistence asymmetry of[Table C\.1](https://arxiv.org/html/2607.18292#A3.T1)— the oracle\-referenced channels \(commitment riskVar\[δ\]\\mathrm\{Var\}\[\\delta\], divergence𝔼\[δ\]\\mathbb\{E\}\[\\delta\]\) stay elevated post\-onset \(\>1\>1\) while the self\-readableH\(pM\)H\(p\_\{M\}\)relaxes closest to baseline, the risk longest\-lived — reproduces whether Qwen3 models \(0\.60\.6–88B\) are scored against the Qwen3\-14B or the Qwen3\-32B oracle; the commitment\-risk half\-lifet1/2t\_\{1/2\}exceeds the realized\-bias half\-lifeδ\(yt\)¯2\\overline\{\\delta\(y\_\{t\}\)\}^\{2\}at every rung\. po//pre is the post\-onset \(t−t⋆∈\[1,10\]t\-t^\{\\star\}\\in\[1,10\]\) mean over the pre\-onset baseline\.*The two columns are not a controlled swap*: they come from independently sampled free\-run trajectory sets at different nucleus truncations \(top\-kk10241024vs20482048\), so the exact half\-lives differ; the qualitative asymmetry — the load\-bearing claim — does not\. Agreement across oracle size on independent data is, if anything, stronger than a shared\-sample swap\.
## Appendix EThe Across\-Claim Snowball
We detail the observational across\-claim snowball of[Section 4\.1](https://arxiv.org/html/2607.18292#S4.SS1)\. The across\-claim relative riskRR=P\(Uj\+1∣Uj\)/P\(Uj\+1∣Sj\)\\text\{RR\}=P\(U\_\{j\+1\}\\mid U\_\{j\}\)/P\(U\_\{j\+1\}\\mid S\_\{j\}\)holds in both samples\. Across all responses the rawRRall\\text\{RR\}\_\{\\text\{all\}\}climbs×1\.28→×2\.85\\times 1\.28\\to\\times 2\.85with scale; restricting to*mixed*responses \(both a supported and an unsupported claim present, so the topic is at least partly known — controlling the base rate\) attenuates it toRRmixed\\text\{RR\}\_\{\\text\{mixed\}\}×1\.08→×1\.71\\times 1\.08\\to\\times 1\.71but does not remove it\. Every95%95\\%CI — bootstrapped over responses, the independent unit — excludes11on thousands of consecutive pairs per rung, so the snowball is neither a topic\-gap confound nor a selection artifact of the mixed restriction \([Table E\.1](https://arxiv.org/html/2607.18292#A5.T1)\)\.
Table E\.1:The across\-claim snowball holds for both the raw and the topic\-controlled sample — not a confound or a selection artifact\.Consecutive claim pairs from Qwen3 free\-run trajectories \([Section 4\.1](https://arxiv.org/html/2607.18292#S4.SS1)\), shown for*all*responses and the*mixed*subset \(holding both a supported and an unsupported claim, controlling the topic base rate\)\. Per block: the consecutive\-pair countnn; the next claim’s fabrication rate after a supported \(spontaneous\) vs\. an unsupported claim,Pr\(U∣S\)\\Pr\(U\{\\mid\}S\)/Pr\(U∣U\)\\Pr\(U\{\\mid\}U\); and the across\-claim relative riskRR=P\(Uj\+1∣Uj\)/P\(Uj\+1∣Sj\)\\text\{RR\}\{=\}P\(U\_\{j\+1\}\{\\mid\}U\_\{j\}\)/P\(U\_\{j\+1\}\{\\mid\}S\_\{j\}\),±\\pmhalf the95%95\\%CI bootstrapped over responses \(the independent unit\)\. RR clears11in*both*samples at every rung, so the lift is neither a topic\-gap confound nor a selection artifact of the mixed restriction; controlled,RRmixed\\text\{RR\}\_\{\\text\{mixed\}\}climbs1\.08→1\.711\.08\\to 1\.71with scale\.
## Appendix FThe Variance\-Contraction Intervention
This section details the on\-policy intervention behind[Section 4\.3](https://arxiv.org/html/2607.18292#S4.SS3)and reports three results the main text compresses:*where*bias and variance concentrate \(the live divergence event, not the verifier label\), that the effect needs the controller firing at*every*crossing rather than one, and a dose\-response that pins the lever to variance\. The intervention runs on the LongFact\+\+ free\-run trajectories \([Appendix A](https://arxiv.org/html/2607.18292#A1)\) for Qwen3 \(44–1414B, oracle3232B\) and OLMo\-3\-77B \(oracle3232B\),n≈17n\{\\approx\}17–2323source trajectories per model\.
#### The mean\-preserving contraction\.
At each step the co\-resident oracle supplies the full next\-token gapδ\(x\)=logpM\(x\)−logpO\(x\)\\delta\(x\)=\\log p\_\{M\}\(x\)\-\\log p\_\{O\}\(x\), its meanμ=𝔼pM\[δ\]=KL\(pM∥pO\)\\mu=\\mathbb\{E\}\_\{p\_\{M\}\}\[\\delta\]=\\mathrm\{KL\}\(p\_\{M\}\\\|p\_\{O\}\)\(the*bias*\), and its spreadVarpM\[δ\]\\mathrm\{Var\}\_\{p\_\{M\}\}\[\\delta\]\(the*commitment risk*\); the intervention replacespMp\_\{M\}with the contractionqλq\_\{\\lambda\}of[Equation 4](https://arxiv.org/html/2607.18292#S4.E4)\(main text\)\. The penalty acts on logits viaδ\\delta, so the knob is invariant to a uniform shift oflogpM\\log p\_\{M\}and reduces to ordinary temperature sharpening only in the degenerate caseδ≡const\\delta\\equiv\\text\{const\}; otherwise it is strictly a second\-moment operation\. The doseρ\\rho\(λ\\lambdachosen per step soVarqλ\[δ\]=ρVarpM\[δ\]\\mathrm\{Var\}\_\{q\_\{\\lambda\}\}\[\\delta\]=\\rho\\,\\mathrm\{Var\}\_\{p\_\{M\}\}\[\\delta\]\) is comparable across models where a rawλ\\lambdais not\. Holding the mean discards the*lucky\-good*excursions \(δ\\deltaon the supported side ofμ\\mu\) symmetrically with the bad ones; removing only the bad ones would improve bias and amount to leaking the oracle’s answer, which[Equation 4](https://arxiv.org/html/2607.18292#S4.E4)deliberately does not do\.
#### Solvingν\\nuandλ\\lambda\.
Both multipliers are found by one\-dimensional root\-finding over the model nucleus support, exploiting that each target moment is monotone in its knob\. Inner solve: for a fixed penalty strengthλ\\lambda,𝔼qλ\[δ\]\\mathbb\{E\}\_\{q\_\{\\lambda\}\}\[\\delta\]is strictly increasing inν\\nu, so the uniqueν\\nuwith𝔼qλ\[δ\]=μ\\mathbb\{E\}\_\{q\_\{\\lambda\}\}\[\\delta\]=\\muis bracketed by doubling outward from a spread\-scaled guess and refined by Brent’s method \(tolerance10−1010^\{\-10\},≤200\\leq 200iterations\); if the constraint already holds within10−210^\{\-2\}nats the tilt is skipped \(ν=0\\nu\{=\}0\)\. Outer solve:Varqλ\[δ\]\\mathrm\{Var\}\_\{q\_\{\\lambda\}\}\[\\delta\]is monotone decreasing inλ\\lambda\(withν\\nure\-solved inside at eachλ\\lambda\), so theλ\\lambdameeting the doseρ\\rhois bracketed and Brent\-refined the same way\. At a degenerate position — a near\-collapsed support where a moment is unreachable — the solver caps at the bracket endpoint and records the realizedvar\_ratio=Varqλ\[δ\]/VarpM\[δ\]\\mathrm\{var\\\_ratio\}=\\mathrm\{Var\}\_\{q\_\{\\lambda\}\}\[\\delta\]/\\mathrm\{Var\}\_\{p\_\{M\}\}\[\\delta\]and mean drift honestly rather than aborting the co\-resident generate barrier; the manipulation check \([Figure 6](https://arxiv.org/html/2607.18292#S4.F6)\) reports these realized moments\. Cost is a handful of Brent iterations of anO\(\|nucleus\|\)O\(\|\\text\{nucleus\}\|\)kernel per fired position, and the controller fires only at theZ=2Z\{=\}2trigger crossings \(*not*every token\), so the per\-token overhead is negligible against the model\+\+oracle forward pass it rides on\.
#### Trigger and grid\.
Letxt⋆=argmaxxpO\(x\)x^\{\\star\}\_\{t\}=\\arg\\max\_\{x\}p\_\{O\}\(x\)be the oracle’s preferred token andst=logpO\(xt⋆\)−logpM\(xt⋆\)≥0s\_\{t\}=\\log p\_\{O\}\(x^\{\\star\}\_\{t\}\)\-\\log p\_\{M\}\(x^\{\\star\}\_\{t\}\)\\geq 0the amount the model under\-weights it\. The controller fires wherests\_\{t\}spikesZZstandard deviations above its per\-model mean \(st≥s¯\+Zσss\_\{t\}\\geq\\bar\{s\}\+Z\\sigma\_\{s\},Z=2Z\{=\}2\), standardized per model so one threshold transfers across scales; anchoring atx⋆x^\{\\star\}\(rather thanargmaxxδ\\arg\\max\_\{x\}\\delta, dominated by the model’s near\-zero\-probability tail\) makes the signal “oracle confident about a token the model misses,” which pre\-selects the recoverable knows\-but\-hallucinates regime\. Around each trigger a tophat window ofW=2W\{=\}2tokens opens at offsetKKand appliesqλq\_\{\\lambda\}at constant doseρ\\rho\. The swept grid isK∈\{−1,0,1,2\}×ρ∈\{0\.5,0\.75,1\.25,1\.5\}×M∈\{1,∞\}K\\in\\\{\-1,0,1,2\\\}\\times\\rho\\in\\\{0\.5,0\.75,1\.25,1\.5\\\}\\times M\\in\\\{1,\\infty\\\}, whereMMcaps interventions per trajectory:M=1M\{=\}1fires once at the first crossing \(a single localized contraction with a fully shared prefix\),M=∞M\{=\}\\inftyfires at every crossing \(refractory\-gated\)\. The control is no\-op \(λ=0\\lambda\{=\}0, the model’s own CRN\-identical free\-run\), giving one unambiguous paired baseline\. We report the best arm over the*causal*K≥0K\{\\geq\}0subgrid only: a window opening before the online trigger \(K=−1K\{=\}\{\-\}1\) cannot be realized by any controller that detects the divergence before acting, so it is excluded even from the oracle\-in\-the\-loop ceiling\. Generate is co\-resident on a single H100 \(model\+\+oracle, full vocabulary\); verification is the same post\-barrier, by\-prompt web\-search pipeline as the observational tests \([Appendix A](https://arxiv.org/html/2607.18292#A1)\)\.
#### Bias and variance concentrate at the trigger, not the verifier onset\.
Per\-position bias2and commitment risk are stored for every draw and plotted against two anchors: the verifier fabrication onsett0⋆t\_\{0\}^\{\\star\}\(first unsupported\-claim token\) and the live oracle triggerttrigt\_\{\\mathrm\{trig\}\}\. The bias2excursion forms a sharp spike at offset0in trigger coordinates but is smeared and off\-zero in onset coordinates: the verifier label*lags*the model’s computational divergence by several tokens\. Any intervention keyed to the verifier onset therefore aims downstream of where bias and variance live, which is why relocating the anchor onto the live divergence event is what makes the contraction land\. \(Aligning on a threshold crossing mechanically shapes both moments near offset0, so the ordering of the two spikes is reported but not yet tested against a shuffled\-trigger null\.\)
#### The effect requires firing at every crossing\.
A single well\-placed contraction is null; re\-firing at every divergence recovers the whole effect\. At the trigger \(K=0,ρ=0\.5K\{=\}0,\\rho\{=\}0\.5\) on Qwen3, theM=1M\{=\}1pairedΔ\\Deltais−0\.002/\+0\.066/−0\.024\-0\.002/\+0\.066/\-0\.024\(44/88/1414B\), statistically indistinguishable from zero \(one rung mildly negative\), whileM=∞M\{=\}\\inftyis\+0\.101/\+0\.187/\+0\.131\+0\.101/\+0\.187/\+0\.131\. Hallucination is not dominated by the first divergence; it accumulates across many, and the benefit is the cumulative product of contracting at each one\. \(TheM=1M\{=\}1contrast was swept on Qwen3 only; the OLMo run usedM=∞M\{=\}\\inftyexclusively\.\)
#### Dose\-response: contraction helps, expansion does not\.
AtM=∞,K=0M\{=\}\\infty,K\{=\}0, the pairedΔ\\Deltais positive for every contraction dose \(ρ<1\\rho\{<\}1\) at all six rungs\. On the four Qwen3 and OLMo rungs it is near\-zero or negative for every expansion dose \(ρ\>1\\rho\{\>\}1\); this contraction\-vs\-expansion gap \(deeper variance suppression helps, variance inflation does not\) is the robust signal that the lever is variance\. Theρ=1\.25\\rho\{=\}1\.25andρ=1\.5\\rho\{=\}1\.5cells are identical because the kernel clamps the expansion side; on the two Llama\-3\.2 rungs that clamp lands on a contraction, so their nominal expansion cells also reduce the rate \(Δ≈0\.21\\Delta\\approx 0\.21/0\.260\.26\) and provide no clean expansion control\. The ordering*within*the contraction side \(ρ=0\.5\\rho\{=\}0\.5vs0\.750\.75\) is model\-specific\. This is the dose\-response image of[Equation 4](https://arxiv.org/html/2607.18292#S4.E4)acting on the variance term of[Equation 2](https://arxiv.org/html/2607.18292#S2.E2)at fixedKL\\mathrm\{KL\}\.
#### Absolute effects\.
[Table F\.1](https://arxiv.org/html/2607.18292#A6.T1)carries the absolute rates behind the relative reduction of[Figure 1](https://arxiv.org/html/2607.18292#S0.F1): the best\-arm pairedΔ\\Deltais0\.130\.13–0\.330\.33with a95%95\\%bootstrap CI excluding zero at all six model×\\timesfamily rungs, and the winning arm always sits on the contraction side \(ρ≤0\.75\\rho\{\\leq\}0\.75\)\. The manipulation check \([Figure 6](https://arxiv.org/html/2607.18292#S4.F6)\) confirms the realized variance falls to the dose while the per\-fired\-step bias drift stays∼\\sim11 orders of magnitude below\|μ\|\|\\mu\|, so the reduction is attributable to the commitment\-risk channel and not to a residual mean shift\.
Table F\.1:Contracting commitment risk causally lowers downstream hallucination \(absolute effects\)\.Per rung: the no\-op control and best\-contraction\-arm web\-verified unsupported\-claim rates over the rest of the response, the within\-trajectory pairedΔ\\Delta\(control−\-best\) with its trajectory\-clustered95%95\\%bootstrap CI, and the winning arm’s onset offsetKKand doseρ\\rho\. A bullet \(∙\\bullet\) marks a CI that excludes zero:Δ=0\.13\\Delta=0\.13–0\.330\.33at all six model×\\timesfamily rungs\. The best arm is selected post hoc per rung over the causal \(K≥0K\\geq 0\) grid, soΔ\\Deltais the oracle\-in\-the\-loop*ceiling*of an online risk controller, not a deployable decoding rule\. The relative reduction is the page\-1 teaser \([Figure 1](https://arxiv.org/html/2607.18292#S0.F1)\)\.
#### No length, claim\-count, or informativeness confound\.
The contraction changes*which*claims a response supports, not how many it makes or how long it runs\.[Table F\.2](https://arxiv.org/html/2607.18292#A6.T2)holds each rung’s best arm against control on four descriptive statistics: length is pinned by the generation cap \(within0\.5%0\.5\\%\), fixing the rate’s denominator; the atomic\-claim count — the informativeness signal at fixed length — moves non\-directionally within−9\.7%\-9\.7\\%to\+4\.2%\+4\.2\\%; and self\-BLEU \(stylistic diversity, normalized to\[0,1\]\[0,1\]\) shifts a mean\+1\.0%\+1\.0\\%, up on three rungs and down on three\. Only the firing rate moves decisively \(zero for control\)\. A sub\-10%10\\%change in claim count at fixed length cannot produce a3535–74%74\\%drop in the unsupported*fraction*: the effect is orthogonal to, and far larger than, any descriptive shift\.
Table F\.2:The hallucination reduction is not an artifact of shorter or sparser responses\.Per rung, the no\-op control and the best\-contraction arm \(the sameK≥0K\\geq 0,M=∞M\{=\}\\inftyselection as[Table F\.1](https://arxiv.org/html/2607.18292#A6.T1)\) on three descriptive statistics of the free\-run response: length \(*tokens*, fixed by the generation cap\), emitted atomic\-claim count \(*claims*— at fixed length, also the informativeness signal\), and stylistic diversity \(*self\-BLEU*; higher==less diverse\)\. The best arm shifts the emitted\-claim count by at most10%10\\%and self\-BLEU by at most0\.0210\.021relative to control — neither distinguishable from no intervention — while the controller*fires*on1616–4848tokens per draw \(*fires*,0for control\)\. A model that fabricated less by emitting fewer or shorter claims would register here; none does, so the3535–74%74\\%reduction \([Table F\.1](https://arxiv.org/html/2607.18292#A6.T1)\) is a smaller unsupported*fraction*at fixed, equally informative output, attributable to the contracted commitment\-risk channel\.
## Appendix GDetector Implementation
#### Semantic entropy\.
The blindness test \([Section 4\.4](https://arxiv.org/html/2607.18292#S4.SS4),[Table 2](https://arxiv.org/html/2607.18292#S4.T2)\) scores each claim with the original semantic\-entropy detector\(Kuhnet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib54); Farquharet al\.,[2024](https://arxiv.org/html/2607.18292#bib.bib10)\): the off\-shelfdeberta\-large\-mnlinatural\-language\-inference model, bidirectional\-entailment clustering of resamples, and the discrete cluster\-proportion entropy \(nats\)\. The resamples are theK=25K\{=\}25same\-model stochastic draws of each prompt already produced for the free\-run trajectories \([Appendix A](https://arxiv.org/html/2607.18292#A1)\), so the detector is a functional ofpMp\_\{M\}alone — the object[Section 2\.4](https://arxiv.org/html/2607.18292#S2.SS4)concerns\. One adaptation makes it claim\-level: for the claim under test, each resample’s NLI stance toward the claim \(premise==resample, hypothesis==claim\)∈\{entailment, neutral, contradiction\}\\in\\\{\\text\{entailment, neutral, contradiction\}\\\}is its semantic cluster, and the score is the discrete entropy over those clusters; the cluster count is therefore not a free hyperparameter \(emergent from entailment,≤3\\leq 3at the claim level\)\. Higher entropy==more resample disagreement about the claim==a stronger hallucination flag\. SelfCheck\-Prompt\(Manakulet al\.,[2023](https://arxiv.org/html/2607.18292#bib.bib9)\), also apMp\_\{M\}\-functional over the same draws, is implemented from its canonical leave\-one\-out consistency prompt\.
## Appendix HThe Over\-Commitment State
#### Three commitment regimes\.
A per\-model three\-state model over the per\-token triple \(bias𝔼\[δ\]\\mathbb\{E\}\[\\delta\], entropyH\(pM\)H\(p\_\{M\}\), commitment riskVar\[δ\]\\sqrt\{\\mathrm\{Var\}\[\\delta\]\}\) recovers the regimes of[Figure 2](https://arxiv.org/html/2607.18292#S1.F2):*grounded*\(low bias, entropy & risk\),*precarious*\(low bias & entropy but high risk, a confident yet high\-variance self\-conditioned branch;prec\\mathrm\{prec\}below\), and*diverged*\(high bias, entropy & risk\) — gap regimes, not correctness\. The trajectory runs grounded, shifts to diverged at the verifier onset, then persists in the precarious state through the continuation — the over\-commitment signature \(conditioning on its own draw\) — and the gap\-only fit below shows it recurs at every scale\.
#### A gap\-only fit recovers the same structure\.
A three\-state Markov\-switching fit to the gapδt=logpM−logpO\\delta\_\{t\}=\\log p\_\{M\}\-\\log p\_\{O\}alone is BIC\-preferred over two states for99\.999\.9–100%100\\%of series at every scale\. Its highest state is the over\-commitment regime: highest mean gap and, notably, also highest variance \(an over\-commitment, not a low\-variance “confidently\-wrong” attractor\), where the model over\-rates the realized token relative to the oracle \([Table H\.1](https://arxiv.org/html/2607.18292#A8.T1)\)\. It is reachable from the precarious state and sticky \(entry and self\-transition probabilities both≈0\.42\\approx 0\.42–0\.450\.45at every scale\), and its over\-rating magnitude shrinks monotonically with scale,2\.42\.4nats at0\.60\.6B to0\.80\.8nats at88B, matching the falling bias2share of[Figure 4](https://arxiv.org/html/2607.18292#S3.F4)\.
Table H\.1:Three\-state gap\-MSAR at the BIC\-selectedk=3k\{=\}3, on free\-run trajectories \(Qwen30\.60\.6–88B models,1414B oracle\)\.k=3k\{=\}3is preferred overk=2k\{=\}2for99\.999\.9–100%100\\%of series\. The over\-commitment state \(highest mean gapμoc\\mu\_\{\\text\{oc\}\}\) is also the highest\-variance \(σoc\\sigma\_\{\\text\{oc\}\}\); its over\-rating magnitude shrinks with scale \(2\.36→0\.812\.36\\to 0\.81nats\)\. Unsupported \(U\) atomic\-fact tokens occupy it≈1\.7×\\approx 1\.7\\timesas often as supported \(S\), separating the two at AUROC0\.680\.68–0\.710\.71, stable across scale\.
#### Hallucinated tokens load the over\-commitment state\.
This backs[Section 3\.3](https://arxiv.org/html/2607.18292#S3.SS3): unsupported atomic\-fact tokens occupy the over\-commitment state≈1\.7×\\approx 1\.7\\timesas often as supported ones \(0\.420\.42–0\.460\.46vs0\.250\.25–0\.270\.27\), separating supported from unsupported content at AUROC0\.680\.68–0\.710\.71,*stable across all four scales*\([Table H\.1](https://arxiv.org/html/2607.18292#A8.T1)\)\. The state is a label\-free, white\-box marker \(read off the logits with no verifier\), and it agrees with the independent FActScore position fit \([Appendix B](https://arxiv.org/html/2607.18292#A2)\) about where in a response the model is most likely to fabricate\.
#### The risk regime bridges adjacent fabrications\.
The same over\-commitment regime is the white\-box channel of the snowball \([Section 4\.2](https://arxiv.org/html/2607.18292#S4.SS2),[Figure 5](https://arxiv.org/html/2607.18292#S4.F5)\): between consecutive claims, the confident*precarious*state fills the inter\-claim bridge more when the next claim is also unsupported\.[Table H\.2](https://arxiv.org/html/2607.18292#A8.T2)gives the per\-rung effect at the≤10\\leq 10\-token adjacency window: the probability a bridge token is in the*precarious*state when the next claim is a fabricationPr\(prec∣U→U\)\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{U\}\)vs\. supportedPr\(prec∣U→S\)\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{S\}\), their absolute difference, and the ratioPr\(prec∣U→U\)/Pr\(prec∣U→S\)−1\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{U\}\)/\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{S\}\)\-1, alongside the number of prev\-fabrication pairsnnbehind each estimate\. The effect is scale\-emergent: strongest at the largest \(1414B\) rung, with Qwen3\-1\.7B the lone negative rung\. Rungs with fewer than100100prev\-fabrication pairs \(the small Llama\-3\.2 rungs\) are too sparse to estimate the ratio reliably and are excluded from[Figure 5](https://arxiv.org/html/2607.18292#S4.F5)\.
Table H\.2:Per\-rung values behind the precarious\-regime bridge effect \([Figure 5](https://arxiv.org/html/2607.18292#S4.F5)\)\.Every model rung \(all families, size\-ordered\): the number of prev\-fabrication claim pairsnn, the probability that an inter\-claim bridge token is in the*precarious*state when the next claim is supportedPr\(prec∣U→S\)\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{S\}\)vs\. a fabricationPr\(prec∣U→U\)\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{U\}\), their absolute differenceΔ=Pr\(prec∣U→U\)−Pr\(prec∣U→S\)\\Delta=\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{U\}\)\-\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{S\}\), and the plotted ratioPr\(prec∣U→U\)/Pr\(prec∣U→S\)−1\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{U\}\)/\\Pr\(\\mathrm\{prec\}\\mid\\mathrm\{U\}\\to\\mathrm\{S\}\)\-1\(the[Figure 5](https://arxiv.org/html/2607.18292#S4.F5)yy\-axis\)\. Adjacency window≤10\\leq 10tokens, mixed\-trajectory restricted \(as[Figure 5](https://arxiv.org/html/2607.18292#S4.F5)\)\.†Daggered rungs are excluded from[Figure 5](https://arxiv.org/html/2607.18292#S4.F5): outside the11–1414B scale window \(Qwen3\-0\.6B\) or fewer than 100 prev\-fabrication pairs \(Llama\-3\.2\-1B\-Instruct, Llama\-3\.2\-3B\-Instruct\)\.Similar Articles
Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models
This paper revisits the reliability paradox in the context of machine unlearning for language models, demonstrating that models can achieve low calibration error while relying on shortcut-based decision rules, thereby extending the paradox to unlearned models.
The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
This paper reveals that while large language models appear robust to task-irrelevant context at the aggregate level, their predictions can flip on individual examples, with performance degrading on some and improving on others, highlighting tail risks that aggregate accuracy conceals.
How Language Models Fail: Token-Level Signatures of Committed and Persistent Reasoning Failures
This paper characterizes two distinct processes by which language models fail in reasoning—committed failure and persistent uncertainty—using token-level uncertainty signals, and demonstrates implications for self-consistency and failure detection strategies.
Some Large Language Models Exhibit Consistent Risk Attitudes
This paper introduces a framework to test whether large language models exhibit consistent risk attitudes across domains. It finds that most LLMs show intra-task and cross-domain stability in risk attitude, converging to a narrower distribution than humans.
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
The paper examines whether large language models exhibit stable and interpretable risk preferences in decision-making under uncertainty, using Texas Hold'em to quantify baseline risk dispositions and context-dependent adaptations.