List Counting Failures Are Not One Phenomenon
Summary
Research shows that open-weight chat models fail at counting list items in distinct error modes, not a single phenomenon, with implications for model interventions and transfer learning.
View Cached Full Text
Cached at: 09/22/26, 09:20 AM
# List Counting Failures Are Not One Phenomenon
Source: [https://arxiv.org/html/2609.22230](https://arxiv.org/html/2609.22230)
Saad MankariousAffiliation:The George Washington UniversityAya ZiriklyAffiliation:Washington, D\.C\., USARebecca HwaAffiliation:\{iyad\.aithou, saadm, ayah\.zirikly, rebecca\.hwa\}@gwu\.edu
###### Abstract
Counting the items in a bracketed list looks trivial, yet open\-weight chat models often get it wrong\. Prior work usually blames input bottlenecks such as subword fragmentation or attention dilution, which predict that different models should fail in roughly the same way\. Across seven instruct models on identical prompts, however, wrong answers form distinct modes: Qwen and Gemma 27B often flip odd lengths to a nearby even integer, OLMo concentrates errors on a few mid\-sized integers, and Llama tends to under\-count\. These modes are useful labels rather than a stable family law \(Gemma 9B does not reproduce Gemma 27B’s odd\-to\-even drop\), and heavier subword fragmentation does not make counting harder on our benchmark\. When the model answers incorrectly, a linear probe can usually still recover the true count from the residual stream\. Matching the same odd\-to\-even error also does not imply the same late\-MLP magnitude fix: scaling a late MLP output helps Qwen modestly but is near null on Gemma 27B under the same protocol, while residual steering can move both only by trading odd gains for even losses\. These results caution against transferring that magnitude fix across models without a transfer check\.
## 1Introduction
Ask seven chat models how many items are in\[apple, banana, cherry, …\], and they often fail even though every item is visible and the task asks only for the list length\. We study*list cardinality*: reporting how many whitespace\-separated items appear in a bracketed list \(e\.g\.,\[the, quick, brown, fox\]→\\to44\)\. That is distinct from character counting inside a token \(how manyr’s instrawberry;[Fu et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib12);[Sims et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib15)\) and from symbolic arithmetic \(37\+8637\{\+\}86;[Nogueira et al\., 2021](https://arxiv.org/html/2609.22230#bib.bib2);[Razeghi et al\., 2022](https://arxiv.org/html/2609.22230#bib.bib3);[Dziri et al\., 2023](https://arxiv.org/html/2609.22230#bib.bib4)\)\. The setting is controlled; a small tool\-argument constraint pilot keeps the same odd\-to\-even drop on Qwen 32B and Gemma 27B \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\), which motivates studying direct\-ask cardinality rather than treating it as a quiz artifact alone\.
The usual explanations put the problem in the*input*: either tokenization splits items into confusing subwords\([Singh and Strouse, 2024](https://arxiv.org/html/2609.22230#bib.bib5);[Zhang et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib11);[Fu et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib12);[Sims et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib15)\), or attention cannot track every item as the list gets longer\([Veličković et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib19);[Nakanishi, 2025](https://arxiv.org/html/2609.22230#bib.bib20)\)\. In both stories the model never forms a clean count, so different models should fail in roughly the same way, and heavier tokenization should make things worse\. If that package were right, a late\-layer intervention found in one model would be a natural candidate to reuse in another\. Figure[1](https://arxiv.org/html/2609.22230#S1.F1)contrasts that shared\-bottleneck story with what we find\. We ask four questions on a fixed list\-cardinality benchmark: \(i\) Do open\-weight chat models fail in one shared way, or in distinct error modes? \(ii\) When the answer is wrong, is the true count missing earlier in the network, or still linearly readable from the residual stream? \(iii\) If two models make the same odd\-to\-even mistake, do they share a transferable late\-MLP*magnitude*fix? \(iv\) Do those modes already appear in base models, or do they strengthen across released post\-training stages?
To answer them, we hold the prompts fixed across seven instruct models, compare the shapes of wrong answers, read answer\-position residuals with linear probes, and test a late\-MLP magnitude intervention on models that share an odd\-to\-even drop\. Where released stage checkpoints exist, we also compare base and post\-training models on the same lattice\. The contribution is that comparison: list\-cardinality failure is not one shared phenomenon with one transferable late\-layer fix\. We find the following\.
Figure 1:Overview: same surface failure need not mean the same mechanism\. The standard story predicts a shared input bottleneck and one transferable late\-layer fix\. On identical prompts we instead find distinct modes, a usually recoverable residual count when wrong \(0\.850\.85–0\.960\.96\), and non\-transfer of a late\-MLP magnitude intervention across a shared odd\-to\-even drop \(Qwen0\.200\.20–0\.250\.25vs\. Gemma 27B0\.0550\.055; Table[3](https://arxiv.org/html/2609.22230#S4.T3); §[4\.3](https://arxiv.org/html/2609.22230#S4.SS3)–[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\)\.Error modes differ across models\.On identical prompts, wrong answers form distinct patterns rather than one shared failure \(Fig\.[1](https://arxiv.org/html/2609.22230#S1.F1); Table[3](https://arxiv.org/html/2609.22230#S4.T3); §[4\.1](https://arxiv.org/html/2609.22230#S4.SS1)\): Qwen and Gemma 27B often answer odd lengths with a nearby even number, OLMo concentrates errors on a few mid\-sized integers, and Llama tends to under\-count\. These are useful labels, not a stable family law \(Gemma 9B does not reproduce Gemma 27B’s odd\-to\-even drop; Appendix[B](https://arxiv.org/html/2609.22230#A2)\), and heavier BPE fragmentation does not make counting harder \(§[4\.2](https://arxiv.org/html/2609.22230#S4.SS2)\)\.
The true count is usually still readable\.When the model answers incorrectly, a linear probe recovers the true count at0\.850\.85–0\.960\.96against a permutation baseline near0\.070\.07\(§[4\.3](https://arxiv.org/html/2609.22230#S4.SS3)\), pointing to late answer failure rather than a missing earlier count\.
Same surface error need not share a magnitude fix\.reverse\_halfhelps Qwen modestly \(0\.200\.20–0\.250\.25held\-out\) but is near null on Gemma 27B \(0\.0550\.055\) under the same protocol \(§[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\), so this late\-MLP magnitude intervention should not be transferred across models without a check\.
Modes shift across post\-training stages\.Matched Qwen base checkpoints already show an even−\-odd gap that instruction tuning widens \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\. On the OLMo 2–32B ladder, base prefers small wrong integers, DPO raises mid\-range mass, and instruct makes that concentration clearest\. That is stage sensitivity, not an installing cause: the usable SFT checkpoint refuses with non\-integer text, and we still lack a document\-level installing statistic\. By “training family” we mean pretraining lineage plus tokenizer, chat template, and post\-training, not architecture or size alone\.
## 2Related work
#### Counting in transformers\.
Theory and small\-model work characterize when transformers can implement counting and related algorithms\([Hahn, 2020](https://arxiv.org/html/2609.22230#bib.bib8);[Yehudai et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib9);[Behrens et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib13);[Golkar et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib14);[Chang and Bisk, 2025](https://arxiv.org/html/2609.22230#bib.bib10)\)\. On large LMs, empirical failures are most often attributed to the input encoding: subword tokenization can fragment or obscure the units being counted\([Singh and Strouse, 2024](https://arxiv.org/html/2609.22230#bib.bib5);[Zhang et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib11);[Fu et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib12);[Sims et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib15)\), and softmax attention can lose resolution as the number of items grows\([Veličković et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib19);[Nakanishi, 2025](https://arxiv.org/html/2609.22230#bib.bib20)\)\. Those accounts predict broadly shared, length\- or fragmentation\-driven degradation\. They do not predict training\-family\-specific wrong integers on identical prompts, nor that more fragmented lists can be easier than common\-word lists, both of which we observe for list cardinality\. We treat these as constraints on input\-side stories for*this*task, not as a blanket refutation of tokenization or attention accounts in other counting settings\.
#### Linear probes and unemitted information\.
A separate literature studies when models encode information that is not reflected in their outputs\([Burns et al\., 2023](https://arxiv.org/html/2609.22230#bib.bib30);[Orgad et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib31);[Li et al\., 2023](https://arxiv.org/html/2609.22230#bib.bib32)\)\. Numeric and ordinal structure is often linearly readable from activations\([Wallace et al\., 2019](https://arxiv.org/html/2609.22230#bib.bib1);[Gurnee and Tegmark, 2024](https://arxiv.org/html/2609.22230#bib.bib33);[Heinzerling and Inui, 2024](https://arxiv.org/html/2609.22230#bib.bib34)\), and probing methodology stresses control tasks and selectivity\([Alain and Bengio, 2017](https://arxiv.org/html/2609.22230#bib.bib16);[Hewitt and Liang, 2019](https://arxiv.org/html/2609.22230#bib.bib29);[Belinkov, 2022](https://arxiv.org/html/2609.22230#bib.bib17);[Giulianelli et al\., 2018](https://arxiv.org/html/2609.22230#bib.bib18)\)\. Models also favor high\-probability surface forms even when they conflict with the prompt\([McCoy et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib27)\)\. We build on these tools \(ridge probes with permutation and transfer controls\), but the prior work does not establish that list cardinality remains linearly present under model error, nor that incorrect emissions concentrate on family\-dependent integers\. That is an empirical claim of this paper, not a restatement of the probing literature\.
#### Mechanistic analyses of counting\.
Causal and correlational tools have been used to localize numeric computation\.[Stolfo et al\. \(2023\)](https://arxiv.org/html/2609.22230#bib.bib28)find late MLP involvement in arithmetic success; logit\-lens methods track when a digit becomes readable through the unembedding, with known caveats\([nostalgebraist, 2020](https://arxiv.org/html/2609.22230#bib.bib6);[Belrose et al\., 2023](https://arxiv.org/html/2609.22230#bib.bib7)\);[Hasani et al\. \(2026\)](https://arxiv.org/html/2609.22230#bib.bib35)study internal counters for repeated\-item counting\. Related work on adjacent counting tasks also finds recoverable internal counts with late emission failure, including character counting\([Datta et al\., 2026](https://arxiv.org/html/2609.22230#bib.bib36)\), geometric misalignment between count directions and digit unembeddings\([Garcia, 2026](https://arxiv.org/html/2609.22230#bib.bib37)\), and format\-triggered late MLP overwrite on repeated\-token lists\([Venkatesh, 2026](https://arxiv.org/html/2609.22230#bib.bib38)\)\. Those studies motivate looking past input\-side stories, but they do not hold list cardinality fixed across models, compare wrong\-integer modes on matched prompts, or test whether a late\-layer scaling fix transfers across models that make the same odd\-to\-even mistake\. Our questions in §[1](https://arxiv.org/html/2609.22230#S1)target that gap\.
## 3Experimental setup
The experiments follow the four questions in §[1](https://arxiv.org/html/2609.22230#S1)\. We first put many models on the same counting prompts, then ask whether a wrong answer still leaves a readable count inside the network, then test whether a late\-MLP magnitude fix transfers, and finally compare released base and post\-training stages where those checkpoints exist\. Seeds, pools, and prompt examples are in Appendix[A](https://arxiv.org/html/2609.22230#A1)\.
### 3\.1Models
Table[1](https://arxiv.org/html/2609.22230#S3.T1)lists the checkpoints and the role each one plays\. We use open\-weight releases from the Qwen 2\.5\([Yang et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib21)\), Gemma 2\([Gemma Team et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib22)\), OLMo 2\([Team OLMo et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib23)\), Llama 3\.1\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib24)\), and DeepSeek\-R1\([DeepSeek\-AI et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib25)\)model reports, with Mistral Small 3\([Mistral AI, 2025](https://arxiv.org/html/2609.22230#bib.bib26)\)as a held\-out family\. The core panel is seven instruct models \(Qwen 14B/32B/72B, Gemma 27B, OLMo 32B, Llama 70B, R1\-Distill\): each gets the main behavioral eval and an answer\-position probe\. Gemma 9B and Llama 8B are smaller within\-family checks on the same prompt lattice, used only behaviorally \(Appendix[B](https://arxiv.org/html/2609.22230#A2)\)\. Mistral\-Small\-24B is held out for one predictivity test of the late\-MLP magnitude intervention \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\.
For question \(iv\) we also evaluate released stage checkpoints on the samecommon\_wordslattice \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\. For Qwen 14B/32B that means matched base vs\. instruct\. For OLMo 2–32B we use the released post\-training ladder \(base, DPO, instruct\); the released SFT checkpoint is excluded because it almost always refuses with non\-integer text\. We additionally report two negative installing checks for OLMo \(a5050k\-document mix integer\-frequency proxy and chosen\-vs\-rejected integer mining in the preference mix\), without claiming an installing cause\. The main1,3001\{,\}300\-prompt cross\-model comparison remains instruct\-only\.
Table 1:Models\. Core seven: main eval \+ probes; Gemma 9B / Llama 8B: scale checks \(Appendix[B](https://arxiv.org/html/2609.22230#A2)\); Mistral\-Small\-24B: held\-outreverse\_halftest\. Stages: Qwen base vs\. instruct; OLMo base→\\toDPO→\\toinstruct \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\.
### 3\.2Task and data
To compare error modes fairly, every model must see the same prompts\. The task is list cardinality: given a whitespace\-separated list in brackets, return the number of items\. The main eval uses one English instruction, wrapped in each model’s chat template:How many items are inside the brackets? Return only one integer\. Items: \[\{list\}\]\. We keep the ask fixed and vary only list surface form, so cross\-model differences are not confounded by different questions\. Alternate English paraphrases are a later robustness check \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\. Multilingual lists are surface\-form controls under that fixed English ask, not a test of multilingual counting competence\.
For each lengthL∈\{3,…,15\}L\\in\\\{3,\\ldots,15\\\}and each of five surface conditions \(Table[2](https://arxiv.org/html/2609.22230#S3.T2)\), we sample2020lists from a closed pool with a fixed seed\. That yields1,3001\{,\}300prompts; every model sees the identical strings\. The conditions pressure input\-side stories:common\_wordsis the familiar baseline,rare\_wordslowers frequency, the multilingual rows change script/surface, andrandom\_stringsmultiplies tokens per item by about4×4\\timesas a tokenization\-load control \(items are also highly distinctive, so this does not rule out every input\-side account; §[4\.2](https://arxiv.org/html/2609.22230#S4.SS2)\)\. We decode greedily, take the first integer in the assistant turn \(after</think\>for R1\-Distill\), and report accuracy, the even−\-odd gap on lengths\{10,12,14\}\\\{10,12,14\\\}vs\.\{11,13,15\}\\\{11,13,15\\\}with bootstrap95%95\\%CIs \(300300prompts per side\), and the distribution of wrong integers\.
Table 2:Surface conditions under a fixed English ask\.ml\_\*==multilingual; multilingual rows stress list surface form, not multilingual counting competence\. Examples in Appendix[A](https://arxiv.org/html/2609.22230#A1)\.
### 3\.3Probing
Behavioral modes alone do not tell us whether the count is missing inside the network\. For every core model we therefore cache the answer\-position residual on an expandedcommon\_wordsset \(4040lists per length;520520prompts\) and fit a ridge regressor \(α=1\\alpha\{=\}1,55\-fold CV\) from residual to integer count at each layer\. We report round\-accuracy with permutation and leave\-one\-count\-out controls, a matched probe trained on the emitted integer, and \(on Qwen 32B\) zero\-shot transfer to other surface residuals \(§[4\.3](https://arxiv.org/html/2609.22230#S4.SS3)\)\.
### 3\.4Controls and interventions
Before intervening, we check two simple behavioral alternatives to a shared input bottleneck: whetherrandom\_stringsis harder thancommon\_wordson all seven models, and whether Qwen 32B answers the item count or the whitespace\-chunk count when those differ \(n=40n\{=\}40perNN\)\. A filler\-padding check on Qwen 32B is appendix\-only and is not used as a dilution refutation \(Appendix[G](https://arxiv.org/html/2609.22230#A7)\)\.
The causal analysis then asks whether models that make the same odd\-to\-even mistake share a late\-MLP magnitude handle\. We follow a fixed discovery chain rather than an open search: logit lens points to a late flip, block decomposition attributes most of the wrong\-direction update to the MLP, and a small per\-failure sweep nominates scaling that MLP output by−0\.5\-0\.5\(reverse\_half\)\. We report selection\-corrected held\-out fix rates over all probe\-cell odd\-length failures for Qwen 32B, Qwen 14B, and Gemma 27B \(2020split\-half rounds\)\. Qwen 32B also includes class\-mean steering and a global scale sweep that exposes the even/odd trade\-off\. Details are in Appendix[G](https://arxiv.org/html/2609.22230#A7)\.
## 4Results
We start with behavior on the shared prompt lattice, then ask what remains readable when the answer is wrong, then whether a late\-MLP magnitude fix transfers, and finally how modes change across released post\-training stages\. Unless noted, main\-eval numbers use identical prompts across models, greedy decoding, and the even−\-odd gap on lengths\{10,12,14\}\\\{10,12,14\\\}vs\.\{11,13,15\}\\\{11,13,15\\\}\(300300prompts/side; bootstrap95%95\\%CIs\)\. Full grids are in Appendices[C](https://arxiv.org/html/2609.22230#A3)–[D](https://arxiv.org/html/2609.22230#A4)\.
### 4\.1Error modes differ across models
The first question is whether failures look alike\. On identical prompts they do not: wrong answers form distinct modes rather than one shared length\-driven pattern\. Table[3](https://arxiv.org/html/2609.22230#S4.T3)summarizes the panel: Qwen 14B/32B, R1\-Distill, and Gemma 27B show a large even−\-odd gap \(parity drop;\+0\.24\+0\.24to\+0\.40\+0\.40\); Qwen 72B weakens it \(\+0\.10\+0\.10\) and shifts top wrongs off the even neighbors; OLMo has no reliable gap \(CI crosses zero\) and concentrates wrongs on a few mid\-range integers rather than flipping parity; Llama 70B inverts the pooled gap and under\-counts\. We treat “mode class” as an operational label from three observables \(even−\-odd gap \(sign/CI\),P\(y^<y∣P\(\\hat\{y\}\{<\}y\\midwrong\)\), and top\-wrong mass\), not as a claim that any particular wrong integer is lineage\-stable: OLMo/Llama both under\-count \(0\.950\.95–0\.990\.99\) while Qwen 32B’s parity wrongs are mixed \(0\.460\.46\), which separates under\-count from parity\-flip even when their top\-wrong lists overlap\. Pooled gaps \(n=300n\{=\}300/side\) carry these claims; Qwen 32B adjacent cells likewise separate \(e\.g\.L=10L\{=\}10:1\.001\.00vs\.L=11L\{=\}11:0\.050\.05; Appendix[C](https://arxiv.org/html/2609.22230#A3)\), with directional signed errors \(Table[9](https://arxiv.org/html/2609.22230#A3.T9)\)\. Per\-condition length curves are in Appendix[J](https://arxiv.org/html/2609.22230#A10)\. Architecture and scale alone do not determine the shape: same\-scale OLMo lacks Gemma 27B’s drop, and R1 distillation does not remove Qwen’s\. Within\-family scale checks \(Appendix[B](https://arxiv.org/html/2609.22230#A2)\) support a Llama under\-count class at 8B \(bf16\) and 70B \(nf4; descriptive\)\. Gemma 9B does*not*reproduce Gemma 27B’s odd\-to\-even drop: it collapses onto a fixed mid\-range integer for longer lists, which makes some odd lengths accidentally correct and drives the even−\-odd gap negative \(−0\.30\-0\.30vs\.\+0\.25\+0\.25\)\. Lineage often predicts mode class better than matched scale \(OLMo vs\. Qwen at 32B\), but checkpoints within a lineage can shift class\.
Table 3:Error modes and tokenization control on identical prompts\. Gap==even−\-odd accuracy \(\{10,12,14\}\\\{10,12,14\\\}vs\.\{11,13,15\}\\\{11,13,15\\\};300300/side; bootstrap95%95\\%CIs\)\. Accc/ Accr==mean accuracy oncommon\_words/random\_strings;Δ=\\Delta\{=\}Accr−\{\}\_\{\\mathrm\{r\}\}\{\-\}Accc\(random uses∼4×\{\\sim\}4\\timesmore tokens atL=15L\{=\}15\)\.P\(CLOSEP\(u∣\\midw\)\)==P\(y^<y∣P\(\\hat\{y\}\{<\}y\\midwrong\)\)oncommon\_words\(separates mid\-range concentration from pure under\-count\)\. Top wrong==leading wrong emissions oncommon\_words\(lattice\-specific; not a length\-invariant attractor\)\. Full grids in Appendices[C](https://arxiv.org/html/2609.22230#A3)–[D](https://arxiv.org/html/2609.22230#A4)\.
### 4\.2Heavier tokenization is not harder on this benchmark
A natural follow\-up is whether those failures are just tokenization load\. If subword fragmentation caused them,random\_stringsshould be hardest\. Table[3](https://arxiv.org/html/2609.22230#S4.T3)shows the opposite on our lattice: Accr\>\{\}\_\{\\mathrm\{r\}\}\>Acccfor every model \(Δ\\Deltafrom\+0\.07\+0\.07to\+0\.24\+0\.24\) despite∼4×\{\\sim\}4\\timesmore tokens per item atL=15L\{=\}15\. Items are also more distinctive, so this rules out “more tokens⇒\\Rightarrowharder” without closing every input\-side story \(Limitations\)\. A multi\-word\-item control on Qwen 32B likewise fails a chunk\-counting prediction: at evenNN, answers favor the item count \(∼\\sim65–78%\) over the whitespace\-chunk count \(≤\\leq12\.5%\), and at oddNNthey fall onto even integer modes \(e\.g\.85%85\\%say “12” atN=11N\{=\}11\) rather than either count\. Parity oscillation is also hard to reconcile with monotone attention\-dilution stories; we do not treat filler\-padding as a causal dilution refutation \(Appendix[G](https://arxiv.org/html/2609.22230#A7)\)\. The breaks are therefore better read as model\-conditioned answer structure than as a generic input bottleneck\.
### 4\.3The true count usually survives when the answer is wrong
Modes and tokenization controls still leave open whether the count is missing inside the network\. When greedy decoding is wrong, a linearly recoverable true\-count signal is usually still present at the answer position\. Table[4](https://arxiv.org/html/2609.22230#S4.T4)shows last\-layer probe round\-accuracy near0\.920\.92–0\.980\.98across all seven models, while model accuracy ranges from roughly0\.300\.30to0\.890\.89on the samecommon\_wordsprompts\. The raw “probe right, model wrong” counts track the independence productprobe acc\.×\(1−model acc\.\)\\text\{probe acc\.\}\\times\(1\-\\text\{model acc\.\}\)within0\.020\.02for every model, so they are not themselves evidence of a special residual; the conditional controls in Table[5](https://arxiv.org/html/2609.22230#S4.T5)carry the claim\. On model\-wrong subsets, true\-count probe accuracy remains0\.850\.85–0\.960\.96against a permuted\-label baseline near0\.070\.07and a majority\-class baseline of0\.110\.11–0\.260\.26\(always predict the modal true count among wrongs\); a matched probe trained on the emitted integer reaches only0\.400\.40–0\.760\.76on the same residuals\. The contrast is cleanest for Qwen and weaker for OLMo/Llama, where wrong emissions are more concentrated \(Limitations\)\. Read\-only diagnostics agree on timing: atL=11L\{=\}11, after forcing the leading “1”, Qwen 32B assignsP\(“2”\)=0\.885P\(\\text\{\`\`2''\}\)=0\.885vs\.P\(“1”\)=0\.111P\(\\text\{\`\`1''\}\)=0\.111, and logit\-lens traces show mid\-network correct\-digit mass overwritten late \(Appendix[J](https://arxiv.org/html/2609.22230#A10)\)\. We read this as late emission failure with a preserved residual count, not as proof that the frozen unembedding uses the probe’s coordinates; the probe interpolates within the trained range but does not extrapolate beyond it \(Appendix[E](https://arxiv.org/html/2609.22230#A5)\), consistent with a bounded\-resolution count axis rather than an unbounded cardinality direction\.
Table 4:Cross\-model probe\-vs\-model accuracy on the same520520\-sequencecommon\_wordscell \(L∈\{3,…,15\}L\\\!\\in\\\!\\\{3,\\ldots,15\\\},40/L40/L\)\. “Probe right, model wrong” counts sequences where the linear probe atℓlast\\ell\_\{\\text\{last\}\}predicts the correct integer while the model emits the wrong digit\. The last column reports the share of those cases at odd lengths\{11,13,15\}\\\{11,13,15\\\}\.Table 5:Probe controls on the520520\-sequencecommon\_wordscell, last\-layer answer\-position residual, all seven models\.P\(✓∣×\)P\(\\checkmark\\mid\\times\)/P\(✓∣✓\)P\(\\checkmark\\mid\\checkmark\): held\-out true\-count probe round\-accuracy on model\-wrong / model\-correct subsets \(Wilson95%95\\%CIs\)\.*perm\.*: mean round\-accuracy under five label permutations \(chance≈0\.077\{\\approx\}0\.077\)\.*emit\-probe on errors*: probe trained on the model’s emitted integer, evaluated on model\-wrong residuals\. On the same residuals where the true count is decodable at0\.850\.85–0\.960\.96, the wrong emission is decodable at only0\.400\.40–0\.760\.76\.
### 4\.4A late\-MLP scale lever helps Qwen, not Gemma
If two models make the same odd\-to\-even mistake, it is natural to reuse the same late\-layer fix\. We test that transfer for one concrete intervention: scaling a late MLP output \(reverse\_half\)\. The informative result is not the modest fix rate but its zero\-sum structure: the same surface drop does not imply a shared magnitude handle\. Figure[2](https://arxiv.org/html/2609.22230#S4.F2)summarizes the full\-failure protocol: selection\-corrected held\-out recovery is0\.20±0\.060\.20\\pm 0\.06on Qwen 32B and0\.25±0\.050\.25\\pm 0\.05on Qwen 14B, but only0\.055±0\.0220\.055\\pm 0\.022on Gemma 27B, despite Gemma’s matching odd\-to\-even drop and high last\-layer probe \(∼0\.92\{\\sim\}0\.92\)\. Within Qwen, a continuous late\-MLP scale sweep \(module output×s\{\\times\}s;s=1s\{=\}1is the unsteered baseline\) raises even\-target accuracy from0\.7550\.755to0\.8650\.865while odd\-target accuracy falls from0\.6530\.653to0\.5810\.581, andreverse\_halffixes8484cases while breaking182182\. Class\-mean residual steering can move all three models with that drop \(best cliff\-fix∼0\.33\{\\sim\}0\.33/0\.430\.43/0\.280\.28on Qwen 32B / Qwen 14B / Gemma 27B\), but only as a zero\-sum trade\-off \(Gemma correct\-control accuracy1\.0→0\.351\.0\\to 0\.35\)\. What fails to transfer is therefore the MLP\-*magnitude*lever, not every residual intervention; small\-sample vignettes overstated both families \(Qwen6/86/8; Gemma≈0\.69\{\\approx\}0\.69onn=8n\{=\}8\), and the full\-failure protocol is the claim\. This is consistent with a small late write being geometrically cheap: adjacent digit unembeddings are highly aligned \(most aligned pair\(1,2\)\(1,2\)atcos=0\.865\\cos\{=\}0\.865; Appendix[J](https://arxiv.org/html/2609.22230#A10)\)\. The reading is a miscalibrated prior write in Qwen that magnitude scaling can partially reverse, and a different usable direction in Gemma: matching surface mode, non\-matching magnitude lever\.
Figure 2:Causal MLP scale interventions\.Top:per\-layer fraction of odd\-length failures fixed by scaling that layer’s MLP output bys∈\{−0\.5,0\}s\\in\\\{\-0\.5,0\\\}over the late\-layer candidate set \(Appendix[G](https://arxiv.org/html/2609.22230#A7); small\-sample panels are upper bounds\)\. Selection\-corrected full\-failure held\-out rates:0\.20±0\.060\.20\\pm 0\.06\(Qwen 32B\),0\.25±0\.050\.25\\pm 0\.05\(Qwen 14B\),0\.055±0\.0220\.055\\pm 0\.022\(Gemma 27B\)\.Bottom:even\- vs\. odd\-target accuracy as the MLP output scalesssweeps on Qwen 32B \(s=1s\{=\}1is baseline\)\. Layer indices differ by analysis \(discovery vs\. split\-half vs\. continuous sweep\); see Appendix[G](https://arxiv.org/html/2609.22230#A7)\.
### 4\.5Developmental and protocol analyses
The last question is whether modes are already present before instruction tuning, or strengthen across released stages; we also check how fragile they are to how the count is asked\. On Qwen 14B/32B, matched base checkpoints already show acommon\_wordseven−\-odd gap \(bootstrap95%95\\%CIs exclude zero:0\.470\.47\[0\.33,0\.60\]\[0\.33,0\.60\]at 32B;0\.220\.22\[0\.05,0\.38\]\[0\.05,0\.38\]at 14B\), and instruction tuning widens those gaps to0\.770\.77and0\.530\.53\. Wrong\-emission mass on the Qwen even modes\{8,12,14\}\\\{8,12,14\\\}likewise rises under instruct \(0\.45→0\.720\.45\{\\to\}0\.72at 32B;0\.45→0\.650\.45\{\\to\}0\.65at 14B; endpoint CIs exclude each other\)\. For OLMo 2–32B we evaluate the released post\-training ladder on the samecommon\_wordslattice: base lacks an odd\-to\-even drop and, when wrong, prefers small integers \(55/66/44; modal\-wrong mass on a mid\-range integer0\.070\.07\); the intermediate DPO stage raises that mass to0\.170\.17\(top wrongs still mixed\); the final instruct checkpoint reaches0\.280\.28with that mid\-range peak clearest on this lattice\. The released SFT checkpoint is not usable here: almost all generations are non\-integer “write a Python solution…” refusals, so we exclude it from the mass series\. A5050k\-document OLMo\-mix integer\-frequency proxy correlates with accuracy overall \(r=0\.66r\{=\}0\.66\) but not forN≥8N\\geq 8\(r=−0\.12r\{=\}\{\-\}0\.12\); mining chosen\-vs\-rejected integers inallenai/olmo\-2\-0325\-32b\-preference\-mixlikewise yields no overweight of that mid\-range integer\. A wrong\-neighbor emission\-mass proxy is stronger for Qwen 32B \(r=−0\.86r\{=\}\{\-\}0\.86\) than for OLMo \(r=−0\.59r\{=\}\{\-\}0\.59\)\. We therefore treat the OLMo concentration as late\-post\-training sensitive rather than a pretrain parity prior, without claiming an installing document statistic or a single causal stage\.
Modes are also protocol\-conditioned\. Ask paraphrases \(count\_entries,how\_many,cardinality;2020lists×\\timeslengths\{10,…,15\}\\\{10,\\ldots,15\\\}\) leave the odd\-to\-even drop intact: Gemma 27B gaps stay in\[0\.27,0\.33\]\[0\.27,0\.33\], Qwen 14B in\[0\.38,0\.52\]\[0\.38,0\.52\], Qwen 32B in\[0\.80,0\.92\]\[0\.80,0\.92\]\. Forced enumeration before the integer \(enumerate\_then\_answer\) abolishes direct\-ask failures where we tested it: Qwen 32B and Gemma 27B both reach≈0\.98\{\\approx\}0\.98–1\.001\.00even and odd accuracy \(gap00;n=120n\{=\}120\); OLMo 32B likewise jumps from floor under ask paraphrases to even1\.001\.00/ odd0\.970\.97\(gap≈0\.03\{\\approx\}0\.03;n=120n\{=\}120\)\. We did not run the enumeration template on R1\-Distill; R1’s retained drop under the default ask is already evidence that reasoning distillation alone does not remove the mode\. As a minimal ecological check, the same item lists under a tool\-argument constraint \(“output the integer number of arguments you will pass”\) preserve the drop on both Qwen 32B \(gap0\.900\.90vs\.0\.880\.88on the bracket control;n=120n\{=\}120\) and Gemma 27B \(gap0\.200\.20vs\.0\.270\.27\); odd\-length wrongs remain nearby even integers\. A warehouse\-inventory paragraph frame is mixed: Gemma keeps a gap \(0\.230\.23\), while Qwen collapses on both even and odd lengths \(gap0\.130\.13with even accuracy only0\.130\.13\)\. Prose embedding can change absolute difficulty without removing Gemma’s drop or Qwen’s tool\-constraint drop\. As a held\-out concentration→\\toreverse\_halfcheck, Mistral\-Small\-24B\-Instruct\-2501 shows a parity cliff \(gap≈0\.28\{\\approx\}0\.28\) with diffuse neighbor mass \(≈0\.37\{\\approx\}0\.37\), so the pre\-registered rule predicts near\-null recovery, yet selection\-corrected held\-out fix rate is0\.28±0\.060\.28\\pm 0\.06\(a miss\)\. Concentration alone does not forecast the lever; boundaries for interpretation are collected in Limitations\.
## 5Discussion
The results relocate list\-cardinality failure from a shared input bottleneck to family\-conditioned late emission\. Tokenization\-load and monotone dilution accounts predict shared, input\-driven degradation\([Singh and Strouse, 2024](https://arxiv.org/html/2609.22230#bib.bib5);[Zhang et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib11);[Veličković et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib19);[Nakanishi, 2025](https://arxiv.org/html/2609.22230#bib.bib20)\); family\-conditioned modes, an inverted surface\-form ordering, and a surviving residual count jointly push against that package on this task\. None of this proves the frozen unembedding reads the probe’s coordinates\.
We attribute the split to family\-level factors—pretraining lineage, tokenizer, template, and post\-training—not architecture or scale alone\. Modes strengthen across released stages, survive a tool\-argument framing, and vanish under forced enumeration: they behave like emission biases\([McCoy et al\., 2024](https://arxiv.org/html/2609.22230#bib.bib27)\), not a missing counting algorithm\. We have not identified the installing training signal\.
Mechanistically, what is “wrong in the model” is not one shared late\-MLP overwrite\. Qwen has a partially reversible late MLP prior write \(zero\-sum under scale\); Gemma shares the surface cliff and a residual steering handle, but not that MLP magnitude lever \(§[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\)\. Read\-only diagnostics can show mid\-network correct\-digit mass in both families without implying that the bad update is an MLP output you can reverse\-scale\. Concentration of wrong mass is at best an in\-sample correlate ofreverse\_halfusefulness, not a law \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\. Together with the aligned digit geometry and bounded residual count axis reported in §[4\.3](https://arxiv.org/html/2609.22230#S4.SS3)–[4\.4](https://arxiv.org/html/2609.22230#S4.SS4), the Qwen odd\-length collapse reads as an ordinally brittle output head plus a late family\-local write\.
## 6Conclusion
On list cardinality, counting failure is not one phenomenon\. Lineage\-conditioned integer modes appear on identical prompts, survive ask paraphrases, and \(for Qwen/Gemma\) survive a tool\-argument constraint framing; stated tokenization\-load and chunk\-counting accounts fail their tests; a true\-count signal is usually still linearly recoverable when emission is wrong; and models that share a parity cliff need not share a late\-MLP*magnitude*lever \(Qwen recovers0\.200\.20–0\.250\.25underreverse\_halfwhile Gemma is near null; residual steering can still move both, only zero\-sum\)\. Forced enumeration abolishes the Qwen 32B, Gemma 27B, and OLMo 32B direct\-ask gaps; R1\-Distill retains a parity cliff under the default ask, evidence that the mode is emission/protocol\-conditioned on this benchmark\. Error shape remains a first\-class object for transfer checks: a simple concentration→\\tolever predictor misses on size\-matched Mistral\-Small\-24B\.
## Limitations
The story this paper tells is about a specific ask: an English, no\-CoT request for the length of a bracketed list\. Ask paraphrases leave the Qwen and Gemma cliffs intact, and a tool\-argument constraint preserves them on Qwen 32B and Gemma 27B, so the failure is not only a bracket\-quiz artifact; forced enumeration abolishes the same gaps \(§[4\.5](https://arxiv.org/html/2609.22230#S4.SS5)\)\. Mid\-generation and multi\-turn settings remain out of scope\.
On the representation side, a high probe score means the count is linearly present in the residual, not that the frozen unembedding reads that direction\. The true\-count probe still beats permutation and majority\-class baselines on model\-wrong subsets, but the emit\-versus\-true contrast is weaker for OLMo and Llama than for Qwen and Gemma, so the cleanest “knows but won’t say” reading is family\-conditioned rather than universal\.
The late\-MLP magnitude lever helps Qwen and is near\-null on Gemma; residual steering can still move both, but only as a zero\-sum trade\. Wrong\-mass concentration looked like a predictor of that lever in\-sample and then missed on size\-matched Mistral\-Small\-24B, so we treat concentration as a correlate, not a transfer rule\.
Finally, what we call lineage mixes tokenizer, chat template, and post\-training, so we cannot say which of those sets the mode class—and that class can even change within one line \(Gemma 9B; Appendix[B](https://arxiv.org/html/2609.22230#A2)\)\. We also never found a training\-document count that explains why those attractor integers win\. Llama 70B and Qwen 72B results are nf4 only \(descriptive\); multi\-word and filler checks are Qwen\-32B\-only; and withn=20n\{=\}20per cell, the main claims use pooled contrasts\.
## References
- Alain and Bengio \(2017\)G\. Alain and Y\. BengioUnderstanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644\.Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Behrenset al\.\(2025\)F\. Behrens, L\. Biggio, and L\. ZdeborováCounting in small transformers: the delicate interplay between attention and feed\-forward layers\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 3500–3532\.Note:Also arXiv:2407\.11542External Links:[Link](https://proceedings.mlr.press/v267/behrens25a.html)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Belinkov \(2022\)Y\. BelinkovProbing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Belroseet al\.\(2023\)N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. SteinhardtEliciting latent predictions from transformers with the tuned lens\.arXiv preprint arXiv:2303\.08112\.External Links:[Link](https://arxiv.org/abs/2303.08112)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- Burnset al\.\(2023\)C\. Burns, H\. Ye, D\. Klein, and J\. SteinhardtDiscovering latent knowledge in language models without supervision\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2212.03827)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Chang and Bisk \(2025\)Y\. Chang and Y\. BiskLanguage models need inductive biases to count inductively\.InInternational Conference on Learning Representations,Note:Also arXiv:2405\.20131External Links:[Link](https://openreview.net/forum?id=s3IBHTTDYl)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Dattaet al\.\(2026\)A\. Datta, M\. Marreddy, A\. Mehler, Z\. Zhao, and R\. MamidiFrom early encoding to late suppression: interpreting LLMs on character counting tasks\.Note:arXiv:2604\.00778, first posted 2026\-04\-01\. Concurrent: character counting; correct count encoded early/mid, suppressed late\.External Links:2604\.00778,[Link](https://arxiv.org/abs/2604.00778)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- DeepSeek\-AIet al\.\(2025\)DeepSeek\-AI, D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang,et al\.DeepSeek\-R1: incentivizing reasoning capability in LLMs via reinforcement learning\.External Links:2501\.12948,[Link](https://arxiv.org/abs/2501.12948)Cited by:[§3\.1](https://arxiv.org/html/2609.22230#S3.SS1.p1.1)\.
- Dziriet al\.\(2023\)N\. Dziri, X\. Lu, M\. Sclar, X\. L\. Li, L\. Jiang, B\. Y\. Lin, P\. West, C\. Bhagavatula, R\. Le Bras, J\. D\. Hwang, S\. Sanyal, S\. Welleck, X\. Ren, A\. Ettinger, Z\. Harchaoui, and Y\. ChoiFaith and fate: limits of transformers on compositionality\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:[Link](https://arxiv.org/abs/2305.18654)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p1.1)\.
- Fuet al\.\(2024\)T\. Fu, R\. Ferrando, J\. Conde, C\. Arriaga, and P\. ReviriegoWhy do large language models \(LLMs\) struggle to count letters?\.External Links:2412\.18626,[Link](https://arxiv.org/abs/2412.18626)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p1.1),[§1](https://arxiv.org/html/2609.22230#S1.p2.1),[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Garcia \(2026\)G\. GarciaThe right answer, the wrong direction: why transformers fail at counting and how to fix it\.Note:arXiv:2605\.03258, first posted 2026\-05\-05\. Concurrent: count linearly recoverable; geometric misalignment with digit unembedding\.External Links:2605\.03258,[Link](https://arxiv.org/abs/2605.03258)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- Gemma Teamet al\.\(2024\)Gemma Team, M\. Rivière, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju,et al\.Gemma 2: improving open language models at a practical size\.External Links:2408\.00118,[Link](https://arxiv.org/abs/2408.00118)Cited by:[§3\.1](https://arxiv.org/html/2609.22230#S3.SS1.p1.1)\.
- Giulianelliet al\.\(2018\)M\. Giulianelli, J\. Harding, F\. Mohnert, D\. Hupkes, and W\. ZuidemaUnder the hood: using diagnostic classifiers to investigate and improve how language models track agreement information\.InProceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP,pp\. 240–248\.External Links:[Link](https://aclanthology.org/W18-5426/)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Golkaret al\.\(2024\)S\. Golkar, A\. Bietti, M\. Pettee, M\. Eickenberg, M\. Cranmer, K\. Hirashima, G\. Krawezik, N\. Lourie, M\. McCabe, R\. Morel, R\. Ohana, L\. H\. Parker, B\. Régaldo\-Saint Blancard, K\. Cho, and S\. HoContextual counting: a mechanistic study of transformers on a quantitative task\.External Links:2406\.02585,[Link](https://arxiv.org/abs/2406.02585)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle,et al\.The Llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.1](https://arxiv.org/html/2609.22230#S3.SS1.p1.1)\.
- Gurnee and Tegmark \(2024\)W\. Gurnee and M\. TegmarkLanguage models represent space and time\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2310.02207)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Hahn \(2020\)M\. HahnTheoretical limitations of self\-attention in neural sequence models\.Transactions of the Association for Computational Linguistics8,pp\. 156–171\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00306),[Link](https://aclanthology.org/2020.tacl-1.11/)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Hasaniet al\.\(2026\)H\. Hasani, A\. Izadi, F\. Askari, M\. Bagherian, S\. Mohammadian, M\. Izadi, and M\. S\. BaghshahUnderstanding counting mechanisms in large language and vision\-language models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 5125–5133\.Note:Also arXiv:2511\.17699External Links:[Link](https://arxiv.org/abs/2511.17699)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- Heinzerling and Inui \(2024\)B\. Heinzerling and K\. InuiMonotonic representation of numeric attributes in language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 175–195\.External Links:[Link](https://aclanthology.org/2024.acl-short.18/)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Hewitt and Liang \(2019\)J\. Hewitt and P\. LiangDesigning and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 2733–2743\.Note:Control tasks and selectivity: the methodological standard for ruling out probe memorisation\.External Links:[Link](https://aclanthology.org/D19-1275/)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2023\)K\. Li, O\. Patel, F\. Viégas, H\. Pfister, and M\. WattenbergInference\-time intervention: eliciting truthful answers from a language model\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:[Link](https://arxiv.org/abs/2306.03341)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- McCoyet al\.\(2024\)R\. T\. McCoy, S\. Yao, D\. Friedman, M\. D\. Hardy, and T\. L\. GriffithsEmbers of autoregression show how large language models are shaped by the problem they are trained to solve\.Proceedings of the National Academy of Sciences121\(41\)\.Note:PNAS published version\. Preprint arXiv:2309\.13638\. Output\-prior framing of LLM failure\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2322420121),[Link](https://www.pnas.org/doi/10.1073/pnas.2322420121)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.22230#S5.p2.1)\.
- Mistral AI \(2025\)Mistral AIMistral small 3\.Note:Mistral AI blogCheckpointmistralai/Mistral\-Small\-24B\-Instruct\-2501External Links:[Link](https://mistral.ai/news/mistral-small-3/)Cited by:[§3\.1](https://arxiv.org/html/2609.22230#S3.SS1.p1.1)\.
- Nakanishi \(2025\)K\. M\. NakanishiScalable\-softmax is superior for attention\.Note:SSMax: rescales attention logits by log\(n\) to recover sharp selectivity at long context\.External Links:2501\.19399,[Link](https://arxiv.org/abs/2501.19399)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p2.1),[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.22230#S5.p1.1)\.
- Nogueiraet al\.\(2021\)R\. Nogueira, Z\. Jiang, and J\. LinInvestigating the limitations of transformers with simple arithmetic tasks\.arXiv preprint arXiv:2102\.13019\.External Links:[Link](https://arxiv.org/abs/2102.13019)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p1.1)\.
- nostalgebraist \(2020\)nostalgebraistInterpreting GPT: the logit lens\.Note:LessWrongExternal Links:[Link](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- Orgadet al\.\(2025\)H\. Orgad, M\. Toker, Z\. Gekhman, R\. Reichart, I\. Szpektor, H\. Kotek, and Y\. BelinkovLLMs know more than they show: on the intrinsic representation of LLM hallucinations\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://arxiv.org/abs/2410.02707)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Razeghiet al\.\(2022\)Y\. Razeghi, R\. L\. Logan IV, M\. Gardner, and S\. SinghImpact of pretraining term frequencies on few\-shot numerical reasoning\.InFindings of the Association for Computational Linguistics: EMNLP 2022,pp\. 840–854\.External Links:[Link](https://aclanthology.org/2022.findings-emnlp.59)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p1.1)\.
- Simset al\.\(2025\)A\. Sims, T\. Foster, K\. Kaleb, T\. H\. Nguyen, J\. Lee, J\. N\. Foerster, Y\. W\. Teh, and C\. LuStochasTok: improving fine\-grained subword understanding in LLMs\.External Links:2506\.01687,[Link](https://arxiv.org/abs/2506.01687)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p1.1),[§1](https://arxiv.org/html/2609.22230#S1.p2.1),[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Singh and Strouse \(2024\)A\. K\. Singh and D\. J\. StrouseTokenization counts: the impact of tokenization on arithmetic in frontier LLMs\.arXiv preprint arXiv:2402\.14903\.External Links:[Link](https://arxiv.org/abs/2402.14903)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p2.1),[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.22230#S5.p1.1)\.
- Stolfoet al\.\(2023\)A\. Stolfo, Y\. Belinkov, and M\. SachanA mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 7035–7052\.Note:Causal\-mediation localisation of arithmetic answer carriers in pretrained LMs; closest mech\-interp precedent for our late\-MLP analysis\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.435/)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- Team OLMoet al\.\(2025\)Team OLMo, P\. Walsh, L\. Soldaini, D\. Groeneveld, K\. Lo, S\. Arora,et al\.2 OLMo 2 furious\.Note:arXiv:2501\.00656; first posted 2024\-12\-31External Links:2501\.00656,[Link](https://arxiv.org/abs/2501.00656)Cited by:[§3\.1](https://arxiv.org/html/2609.22230#S3.SS1.p1.1)\.
- Veličkovićet al\.\(2025\)P\. Veličković, C\. Perivolaropoulos, F\. Barbero, and R\. PascanuSoftmax is not enough \(for sharp size generalisation\)\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 61190–61211\.Note:Also arXiv:2410\.01104\. Cited for the attention\-dilution / sharpness account of length\-generalisation failure\.External Links:[Link](https://proceedings.mlr.press/v267/velickovic25a.html)Cited by:[Appendix G](https://arxiv.org/html/2609.22230#A7.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.22230#S1.p2.1),[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.22230#S5.p1.1)\.
- Venkatesh \(2026\)S\. VenkateshRepeated\-token counting reveals a dissociation between representations and outputs\.Note:arXiv:2605\.09239, first posted 2026\-05\-10\. Concurrent: repeated\-token lists; late MLP overwrite of a correct residual count\.External Links:2605\.09239,[Link](https://arxiv.org/abs/2605.09239)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px3.p1.1)\.
- Wallaceet al\.\(2019\)E\. Wallace, Y\. Wang, S\. Li, S\. Singh, and M\. GardnerDo NLP models know numbers? probing numeracy in embeddings\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 5307–5315\.External Links:[Link](https://aclanthology.org/D19-1534)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2024\)A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei,et al\.Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[§3\.1](https://arxiv.org/html/2609.22230#S3.SS1.p1.1)\.
- Yehudaiet al\.\(2024\)G\. Yehudai, H\. Kaplan, G\. Dar, R\. Rassin, A\. Ghandeharioun, M\. Geva, and A\. GlobersonWhen can transformers count to n?\.External Links:2407\.15160,[Link](https://arxiv.org/abs/2407.15160)Cited by:[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2024\)X\. Zhang, J\. Cao, and C\. YouCounting ability of large language models and impact of tokenization\.External Links:2410\.19730,[Link](https://arxiv.org/abs/2410.19730)Cited by:[§1](https://arxiv.org/html/2609.22230#S1.p2.1),[§2](https://arxiv.org/html/2609.22230#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.22230#S5.p1.1)\.
## Acknowledgments
Code and figure drafting used AI coding assistants; all reported numbers come from the released scripts and CSVs\. The authors reviewed and take responsibility for the final text and results\.
## Contents of the appendix
A§[A](https://arxiv.org/html/2609.22230#A1): benchmark construction and examples\. B§[B](https://arxiv.org/html/2609.22230#A2): within\-family scale checks\. C§[C](https://arxiv.org/html/2609.22230#A3): Qwen 32B accuracy and signed\-error grids\. D§[D](https://arxiv.org/html/2609.22230#A4): remaining models’ accuracy grids\. E–F§[E](https://arxiv.org/html/2609.22230#A5)–[F](https://arxiv.org/html/2609.22230#A6): probe extrapolation and per\-length detail\. G§[G](https://arxiv.org/html/2609.22230#A7): causal intervention protocols\. H§[H](https://arxiv.org/html/2609.22230#A8): tokenization\-control note\. I§[I](https://arxiv.org/html/2609.22230#A9): error\-mode separability\. J§[J](https://arxiv.org/html/2609.22230#A10): additional figures\. K§[K](https://arxiv.org/html/2609.22230#A11): seeds, compute, and release\.
## Appendix ABenchmark construction and examples
This is a controlled synthetic benchmark, not a scraped corpus\. Every list is generated by sampling items from a fixed pool withnumpy\.random\.default\_rng\(seed\)\(Table[19](https://arxiv.org/html/2609.22230#A11.T19)\), then joining them with spaces inside the shared instruction template of §[3\.2](https://arxiv.org/html/2609.22230#S3.SS2)\. Sampling is without replacement whenLLis at most the pool size, and with replacement otherwise\. The same seeds \(and therefore the same1,3001\{,\}300strings\) are used for every model\.
#### Pools\.
common\_words:4848frequent English nouns \(e\.g\.apple,table,window\)\.rare\_words:4040low\-frequency English words \(e\.g\.perspicacious,sesquipedalian\)\.multilingual\_single:2020everyday nouns each in Spanish, Russian, Chinese, and Arabic; each list is drawn from one language\.multilingual\_mixed: the union of those four lexicons in one pool\.random\_strings:200200pseudowords of length44–66over consonants and digits \(seed303303\), built to fragment under BPE\.
#### Design scope\.
The benchmark is a controlled stress test with fixed seeds, closed pools, identical prompts across models, and a tokenization condition that multiplies tokens per item by∼4×\{\\sim\}4\\times\. It is not a naturally occurring corpus with human annotation or ecological coverage of how people ask counting questions\.n=20n\{=\}20per cell supports the even–odd aggregate \(n=300n\{=\}300/side\) but is thin for single\-cell point estimates; we therefore lean on gaps and CIs in the main text\.
Table 6:Example lists from the released eval \(Qwen tokenizer token counts at right\)\. Same strings are shown to every model; multilingual rows are abbreviated here for layout \(full Unicode lists ship in the CSV release\)\.
## Appendix BWithin\-family scale checks
Gemma 9B and Llama 8B are evaluated on the same1,3001\{,\}300\-prompt lattice as the core panel \(main\-eval only; no second probe/causal stack\)\. Table[7](https://arxiv.org/html/2609.22230#A2.T7)reportscommon\_wordseven−\-odd gaps\. Llama 8B matches the Llama 70B under\-count class\. Gemma 9B does not reproduce Gemma 27B’s odd\-to\-even drop: its negative gap comes from collapsing onto a fixed mid\-range integer on longer lists, so some odd true lengths are accidentally correct\.
Table 7:Within\-family behavioral scale checks on the same1,3001\{,\}300\-prompt lattice \(common\_wordseven−\-odd gap overL∈\{10,…,15\}L\\in\\\{10,\\ldots,15\\\}; bootstrap95%95\\%CIs\)\.P\(u∣w\)=P\(y^<y∣P\(\\mathrm\{u\}\\mid\\mathrm\{w\}\)=P\(\\hat\{y\}\{<\}y\\midwrong\)\)\. Llama 8B matches Llama 70B’s under\-count class\. Gemma 9B’s negative gap is a fixed mid\-range collapse on this lattice, not a flipped 27B\-style odd\-to\-even drop\.
## Appendix CFull accuracy grid \(Qwen 32B\)
The main text reports pooled even−\-odd gaps\. Table[8](https://arxiv.org/html/2609.22230#A3.T8)gives the underlying per\-condition×\\timesper\-length accuracy grid for the reference model Qwen 2\.5–32B \(2020lists per cell;1,3001\{,\}300prompts total\)\. Bold columns mark odd lengths\{11,13,15\}\\\{11,13,15\\\}, wherecommon\_wordsaccuracy collapses\. Table[9](https://arxiv.org/html/2609.22230#A3.T9)reports mean signed error \(predicted−\-true\) on the same lattice: oncommon\_words, errors atL=11L\{=\}11tend toward\+1\+1\(say “12”\) and atL=15L\{=\}15toward−1\-1\(say “14”\), whileml\_mixedunder\-counts severely\. Grids for the other six models are in Appendix[D](https://arxiv.org/html/2609.22230#A4)\.
Table 8:Full accuracy grid for the reference model Qwen 2\.5–32B,L∈\{3,…,15\}L\\in\\\{3,\\ldots,15\\\}and the five surface conditions\.2020sequences per cell;1,3001\{,\}300total\. Bold marks odd lengths\{11,13,15\}\\\{11,13,15\\\}\.Table 9:Mean signed error \(predicted−\-true\) for Qwen 2\.5–32B\. Oncommon\_words,L=11L\{=\}11errors are\+1\+1on average \(model says “12”\);L=15L\{=\}15errors are−1\-1on average \(model says “14”\)\.ml\_mixedundercounts severely\.
## Appendix DPer\-model accuracy grids
Tables[10](https://arxiv.org/html/2609.22230#A4.T10)–[15](https://arxiv.org/html/2609.22230#A4.T15)give the same per\-condition×\\timesper\-length accuracy grid for the six non\-reference models on the shared1,3001\{,\}300\-prompt eval\. The Qwen 32B reference grid is Table[8](https://arxiv.org/html/2609.22230#A3.T8)above\. Every cell is mean accuracy over2020sequences; seeds and lists are identical across models\.
Table 10:Qwen 2\.5–14B accuracy grid\. Odd\-length collapse at\{11,13,15\}\\\{11,13,15\\\}is sharp on every condition;random\_stringsremains highest\.Table 11:Qwen 2\.5–72B \(4\-bit nf4\) accuracy grid\. The odd\-length collapse at\{11,13,15\}\\\{11,13,15\\\}is much weaker than at 14B/32B; the odd–even gap is partly preserved onrare\_wordsandml\_mixedbut absent oncommon\_words\.Table 12:Gemma 2–27B accuracy grid\. Overall accuracy is lower across the board; the odd\-length collapse is most visible oncommon\_words,rare\_words, andml\_mixed\.Table 13:DeepSeek\-R1\-Distill\-Qwen\-32B accuracy grid \(reasoning enabled\)\. The odd\-length collapse at\{11,13,15\}\\\{11,13,15\\\}is preserved through R1 distillation; even–odd gap \(0\.39\) is essentially identical to its base Qwen 2\.5–32B \(0\.40\)\.Table 14:OLMo 2–32B accuracy grid\. Overall accuracy is lower than the Qwen/Gemma row even at short lengths, but there is no odd\-length collapse: odd lengths are no worse than adjacent even lengths on most conditions\. This is the no\-collapse baseline that the cross\-family argument in §[4\.1](https://arxiv.org/html/2609.22230#S4.SS1)rests on\.Table 15:Llama 3\.1–70B \(4\-bit nf4\) accuracy grid\. Oncommon\_words, the model is accurate at short lengths and then collapses aboveL=8L\\\!=\\\!8to a small\-integer under\-counting attractor rather than to the Qwen/Gemma odd\-length collapse\.
## Appendix ELinear count probe: extrapolation falsifier
Figure[3](https://arxiv.org/html/2609.22230#A5.F3)shows the length\-extrapolation probe of §[4\.3](https://arxiv.org/html/2609.22230#S4.SS3)in detail\. The probe is a ridge regressor \(α=10\\alpha=10\) fit on the 320common\_wordssequences withL∈\{3,…,10\}L\\in\\\{3,\\ldots,10\\\}and evaluated on the 200 held\-out sequences withL∈\{11,…,15\}L\\in\\\{11,\\ldots,15\\\}\. We plot the mean predicted count, by held\-out true count, at four representative layers \(ℓ4,ℓ30,ℓ51,ℓ63\\ell\_\{4\},\\ell\_\{30\},\\ell\_\{51\},\\ell\_\{63\}\)\. At every layer the predicted mean for every held\-out count flattens to roughly the largest trained count \(∼9\.9\\sim\\\!9\.9\); round\-accuracy is0\.000\.00at every test length and every layer\. By contrast, the leave\-one\-count\-out probe of §[4\.3](https://arxiv.org/html/2609.22230#S4.SS3), which trains on every count except a single held\-out one, achieves round\-accuracy≥0\.70\\geq 0\.70on the held\-out count forL⋆∈\{11,12,13,14\}L^\{\\star\}\\in\\\{11,12,13,14\\\}\. Together these falsify the claim that the residual contains a single linear “cardinality direction” that can be extended indefinitely, while preserving the claim that the count is approximately linearly decodable within the trained range\.
Figure 3:Length\-extrapolation falsifier\. Mean probe\-predicted count, by held\-out true count, for a probe trained onL≤10L\\leq 10\. At every layer \(only four shown\), the predicted mean collapses to the train\-set ceiling \(∼9\.9\\sim\\\!9\.9\) for every held\-out count; round\-accuracy is0\.000\.00\. The dashed line isy=xy=x, the prediction a single global cardinality axis would produce\.
## Appendix FLinear count probe: per\-length detail
Table[16](https://arxiv.org/html/2609.22230#A6.T16)reports the per\-length mean predicted count and round\-accuracy of the layer\-6363ridge probe \(α=1\\alpha=1,55\-fold CV\) on the520520\-sequencecommon\_wordscell, alongside the reference Qwen\-32B model’s own greedy accuracy on the same prompts\. The probe is at or above0\.850\.85round\-accuracy at every length exceptL=12L\{=\}12\(0\.750\.75\); the model collapses to0\.050\.05/0\.250\.25/0\.000\.00at odd lengths\{11,13,15\}\\\{11,13,15\\\}\. The gap atL=11L\{=\}11is0\.85−0\.05=0\.800\.85\-0\.05=0\.80\.
Table 16:Per\-length probe vs\. model accuracy at the answer\-position residual \(layer6363, Qwen\-32B\)\. The probe column is the held\-out \(5\-fold CV\) ridge prediction; the model column is the model’s own greedy emit on the same prompts\. Bold rows are the odd\-length collapse\. Overall:r=0\.997r=0\.997,MAE=0\.20\\text\{MAE\}=0\.20, round\-acc0\.9230\.923\.
## Appendix GCausal intervention protocols
#### Block decomposition\.
For a single failure example\(x,ycorrect,ywrong\)\(x,y\_\{\\text\{correct\}\},y\_\{\\text\{wrong\}\}\)withxxthe answer\-position residual entering blockℓ\\ell, letaℓa\_\{\\ell\}andmℓm\_\{\\ell\}denote the attention and MLP outputs of that block, so that the block update isΔℓ=aℓ\+mℓ\\Delta\_\{\\ell\}=a\_\{\\ell\}\+m\_\{\\ell\}\. We define the \(correct−\-wrong\) digit\-margin direction in residual space asΔu=WU\[:,ycorrect\]−WU\[:,ywrong\]\\Delta\_\{u\}=W\_\{U\}\[:,y\_\{\\text\{correct\}\}\]\-W\_\{U\}\[:,y\_\{\\text\{wrong\}\}\], whereWUW\_\{U\}is the unembedding matrix\. Projections are⟨aℓ,Δu⟩\\langle a\_\{\\ell\},\\Delta\_\{u\}\\rangleand⟨mℓ,Δu⟩\\langle m\_\{\\ell\},\\Delta\_\{u\}\\rangle; norms areℓ2\\ell\_\{2\}norms in the residual basis\. Layerℓ=62\\ell\{=\}62was selected as the late layer with the largest negative MLP projection onΔu\\Delta\_\{u\}over the eight near\-miss failures\.
#### Per\-failure layer search \(§[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\)\.
For each of the eight near\-miss failures we replacemlpℓ\(x\)\\text\{mlp\}\_\{\\ell\}\(x\)with one of\{0,−12mlpℓ\(x\)\}\\\{0,\\,\-\\tfrac\{1\}\{2\}\\text\{mlp\}\_\{\\ell\}\(x\)\\\}\("zero" / "reverse\-half"\) at one layerℓ\\ellat a time, sweepingℓ∈\{34,41,46,47,49,50,51,52,54,57–63\}\\ell\\in\\\{34,41,46,47,49,50,51,52,54,57\\text\{\-\-\}63\\\}, and re\-running the forward pass\. A failure is counted as "fixed atℓ\\ell" when both the argmax digit becomes the correct digit and the \(correct−\-wrong\) margin becomes positive\. Layer5252reverse\_halfis the single\(ℓ,intervention\)\(\\ell,\\text\{intervention\}\)pair that fixes the largest number \(6/86/8\) of these failures\.
#### Split\-half layer search \(selection\-corrected\)\.
For the Qwen\-32B population estimate, the samereverse\_halfintervention is applied at each of1616late MLP layers \(ℓ∈\{34,41,46,47,49,50,51,52,54,57,…,63\}\\ell\\in\\\{34,41,46,47,49,50,51,52,54,57,\\ldots,63\\\}\) to every one of the102102common\_wordsodd\-length failures \(L∈\{11,13,15\}L\\in\\\{11,13,15\\\}, model wrong\), one\(ℓ,failure\)\(\\ell,\\text\{failure\}\)cell at a time\. For each of2020random halvings of the failure set, the layer with the highest fix rate on the selection half is evaluated on the disjoint evaluation half\. Reported statistics are the distribution of selected layers and the mean±\\pmSD of the evaluation\-half fix rate\. The Qwen\-14B estimate uses the identical protocol on all103103probe\-cell odd\-length failures over the last1616MLP layers \(ℓ∈\{32,…,47\}\\ell\\in\\\{32,\\ldots,47\\\}for the4848\-layer stack\), yielding held\-out0\.25±0\.050\.25\\pm 0\.05withℓ33\\ell\_\{33\}selected in16/2016/20splits\. Gemma\-27B uses the same full\-failure protocol on all112112probe\-cell odd\-length failures over the last1616MLP layers of the4646\-layer stack \(ℓ∈\{30,…,45\}\\ell\\in\\\{30,\\ldots,45\\\}\), yielding held\-out0\.055±0\.0220\.055\\pm 0\.022; an earliern=8n\{=\}8vignette that suggested∼0\.69\{\\sim\}0\.69was selection bias and is not used as an estimate\.
#### Filler\-padding capacity check \(not a dilution refutation\)\.
Softmax\-dilution accounts concern resolution over theNNcounted items\([Veličković et al\., 2025](https://arxiv.org/html/2609.22230#bib.bib19)\)\. Prepending neutral filler at fixedNNchanges context length, not the number of competing count\-relevant targets, so this is only an answer\-position capacity check\. ForN∈\{5,13\}N\\in\\\{5,13\\\}andn=40n\{=\}40freshcommon\_wordslists per cell on Qwen 32B, the user message is either the standard prompt or the same prompt preceded by4040repetitions of a neutral filler sentence \(∼760\{\\sim\}760tokens\), inside a single user turn \(∼55→∼815\{\\sim\}55\\to\{\\sim\}815tokens\)\. Greedy decoding:N=5N\{=\}5accuracy moves from1\.001\.00to0\.900\.90andN=13N\{=\}13from0\.050\.05to0\.000\.00\. Long\-context padding is not the odd\-length bottleneck; we do not treat this result as evidence against softmax dilution\.
#### Multi\-word item control\.
Lists ofNNitems in which exactly one item is a familiar two\-word phrase \(e\.g\. “new york”\), so item countNNand whitespace\-chunk countN\+1N\{\+\}1differ\.n=40n\{=\}40lists perNN,N∈\{10,…,14\}N\\in\\\{10,\\ldots,14\\\}; we report the fraction of answers equal toNN\(item counting\) vs\.N\+1N\{\+\}1\(chunk counting\)\.
#### Class\-mean steering\.
From the cached probe\-cell residuals we compute per\-count class meansμc\\mu\_\{c\}at layerℓ∈\{49,52\}\\ell\\in\\\{49,52\\\}\. For each odd\-length failure with true countNNand wrong emissiony^\\hat\{y\}, a forward hook addsβ\(μN−μy^\)\\beta\\,\(\\mu\_\{N\}\-\\mu\_\{\\hat\{y\}\}\)\(withβ∈\{2,4,8\}\\beta\\in\\\{2,4,8\\\}\) to the last\-position residual at layerℓ\\ellon every forward pass during greedy decoding\. Controls: a norm\-matched fixed random direction per example, a same\-day unsteered baseline, and a4040\-sequence model\-correct sample steered away from its nearest attractor to measure collateral breakage\. The probe\-hyperplane projection variant \(replacing the residual’s probe read\-out value with the target count\) is reported as a null\.
#### Global scale sweep\.
A forward hook multiplies the chosen layer’s MLP*module output*by a scalarssbefore the residual add, i\.e\.mlpℓ\(x\)←s⋅mlpℓ\(x\)\\mathrm\{mlp\}\_\{\\ell\}\(x\)\\leftarrow s\\cdot\\mathrm\{mlp\}\_\{\\ell\}\(x\), fors∈\{0\.00,0\.25,0\.50,0\.75,1\.00,1\.25\}s\\in\\\{0\.00,\\,0\.25,\\,0\.50,\\,0\.75,\\,1\.00,\\,1\.25\\\}\. Thuss=1s\{=\}1is the identity \(unsteered baseline\) ands=0s\{=\}0zeros that MLP contribution; valuess≠1s\\neq 1are the intervention\. The continuous trade\-off quoted in §[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\(even0\.755→0\.8650\.755\{\\to\}0\.865, odd0\.653→0\.5810\.653\{\\to\}0\.581\) uses the full main eval at the discovery layerℓ=52\\ell\{=\}52\. Figure[2](https://arxiv.org/html/2609.22230#S4.F2)\(bottom\) shows the same qualitative even/odd trade\-off; axis labels may name a nearby late layer chosen by a later candidate\-layer search \(e\.g\.ℓ=58\\ell\{=\}58on a cliff subsample\)\. Both are late\-MLP magnitude sweeps, not a claim that a single index is uniquely causal\. “Even\-target” / “odd\-target” accuracy average correctness over even / odd true counts in\{3,…,15\}\\\{3,\\ldots,15\\\}\.
#### Layer\-index map \(Qwen\-32B\)\.
Different analyses pick different late layers, and we do not equate them: block\-decomposition margin projection onn=8n\{=\}8near\-misses highlightsℓ=62\\ell\{=\}62; the same vignette’s best single\-layerreverse\_halfisℓ=52\\ell\{=\}52\(6/86/8\); selection\-corrected full\-failure split\-half most often selectsℓ=49\\ell\{=\}49; the continuous magnitude sweep above usesℓ=52\\ell\{=\}52\(full eval\) or a nearby late candidate in the figure panel\. Top panels of Figure[2](https://arxiv.org/html/2609.22230#S4.F2)plot the late\-layer candidate subset used in each model’s search, not every transformer layer\.
#### SwiGLU feature decomposition \(§[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\)\.
A SwiGLU MLP block computesmlp\(h\)=Wdown\(SiLU\(Wgateh\)⊙Wuph\)=Wdownz\\text\{mlp\}\(h\)=W\_\{\\text\{down\}\}\\bigl\(\\text\{SiLU\}\(W\_\{\\text\{gate\}\}h\)\\odot W\_\{\\text\{up\}\}h\\bigr\)=W\_\{\\text\{down\}\}\\,z, wherez∈ℝdffz\\in\\mathbb\{R\}^\{d\_\{\\text\{ff\}\}\}is the SwiGLU hidden vector anddff=27,648d\_\{\\text\{ff\}\}\{=\}27\{,\}648for Qwen\-32B\. For the failure example in §[4\.4](https://arxiv.org/html/2609.22230#S4.SS4), we computed the per\-index contribution to the \(correct−\-wrong\) margin,ck=zk⋅⟨Wdown\[:,k\],Δu⟩c\_\{k\}=z\_\{k\}\\cdot\\langle W\_\{\\text\{down\}\}\[:,k\],\\,\\Delta\_\{u\}\\rangle, and selectedk⋆=argminkckk^\{\\star\}=\\arg\\min\_\{k\}c\_\{k\}\. Single\-feature patches replacezk⋆z\_\{k^\{\\star\}\}with one of\{0,12zk⋆,−12zk⋆,−zk⋆\}\\\{0,\\,\\tfrac\{1\}\{2\}z\_\{k^\{\\star\}\},\\,\-\\tfrac\{1\}\{2\}z\_\{k^\{\\star\}\},\\,\-z\_\{k^\{\\star\}\}\\\}while leaving every other index ofzzuntouched, and the forward pass is re\-run from layer5252onward\.
## Appendix HTokenization control across models
A pure tokenization\-load account predicts that lists whose items fragment into more subword tokens should be harder to count\. Table[3](https://arxiv.org/html/2609.22230#S4.T3)in the main text tests that prediction: every model is more accurate onrandom\_stringsthan oncommon\_wordsdespite roughly4×4\\timesmore tokens atL=15L\{=\}15\(§[4\.2](https://arxiv.org/html/2609.22230#S4.SS2)\)\.
## Appendix IError\-mode separability \(not model fingerprinting\)
We ask how well counting answers on this fixed benchmark separate models\. Features are answer integers, signed errors, length, and condition; splits are by prompt text so no prompt appears in both train and test\. Chance is1/71/7for model identity and1/41/4for family \(Qwen, Gemma, OLMo, Llama; R1\-Distill is grouped with Qwen\)\.
Table 17:Error\-mode separability on the shared benchmark\. Single answers are weak identifiers; bags of wrong answers separate models well in this closed seven\-model set\. Family identity separates almost perfectly at the bag level\. This is descriptive structure, not a claim that counting errors fingerprint models in the wild\.Table 18:Wrong\-emission modes on the1,3001\{,\}300\-prompt eval\.*Even share*is the fraction of wrong emissions that are even integers\. Qwen\-family models are even\-heavy; among OLMo’s wrong answers, mass peaks at “1111”; Llama is odd\-heavy / under\-counting\.Single answers are weak identifiers \(model accuracy0\.270\.27–0\.390\.39\)\. Bags of wrong answers separate the closed seven\-model set well \(0\.930\.93\), and family identity is essentially perfect at the bag level\. Residual confusions are mostly inside the Qwen family\. We report this as descriptive separability on this benchmark, not as a fingerprinting method for arbitrary models or prompts\.
## Appendix JAdditional Visual Diagnostics
These plots support claims already stated in the main text; they are optional reading once the grids and intervention protocols above are clear\.
Figure 4:Cross\-model overwrite maps for sampled odd\-length failures\. Each row is one prompt; color showsP\(model’s wrong units digit\)−P\(correct units digit\)P\(\\text\{model's wrong units digit\}\)\-P\(\\text\{correct units digit\}\)under a layerwise logit\-lens readout\. Green regions indicate layers where the correct digit is favored; red regions indicate layers where the eventual wrong digit dominates\. Qwen14, Qwen32, and Gemma27B all show a mid\-to\-late transition toward the wrong digit, a*read\-only*diagnostic of late failure\. Causalreverse\_halfconfirms a usable late\-MLP lever only in Qwen \(§[4\.4](https://arxiv.org/html/2609.22230#S4.SS4)\)\.Figure 5:Accuracy vs\. list length by surface condition \(odd lengths\{11,13,15\}\\\{11,13,15\\\}shaded\)\. Read for the cross\-panel*shape*split \(parity oscillation vs\. flat vs\. under\-count\), not per\-condition zig\-zags; main\-text summary is Table[3](https://arxiv.org/html/2609.22230#S4.T3)and Fig\.[1](https://arxiv.org/html/2609.22230#S1.F1)\.Figure 6:Reference\-model odd\-length collapse and signed\-error diagnostics for Qwen 2\.5–32B\. The odd\-length failures are directional rather than diffuse:L=11L=11tends upward toward “12”, whileL=15L=15tends downward toward “14”\.Figure 7:Qwen 2\.5–32B logit\-lens trajectories\. Correct\-digit mass appears in middle layers and is suppressed in late layers, consistent with a late overwrite rather than an absent count signal\.Figure 8:Cross\-model second\-digit logit lens on odd\-length failures\. For each of Qwen 14B, Qwen 32B, and Gemma 2 27B, we force the leading “11” on length\-\{11,13,15\}\\\{11,13,15\\\}failures and read off the per\-layer probability of the correct vs\. wrong units digit\. In every model the correct digit is favored at intermediate depth and the wrong digit dominates only late, with mean finalP\(wrong\)∈\[0\.77,0\.92\]P\(\\text\{wrong\}\)\\in\[0\.77,0\.92\]\. Read\-only evidence of late failure, not by itself a claim that the same MLP lever is causal in every family \(Fig\.[2](https://arxiv.org/html/2609.22230#S4.F2)\)\.Figure 9:Digit\-unembedding PCA for Qwen 2\.5–32B\. Digit geometry is ordinal; even/odd is not cleanly separated, supporting the claim that the parity pattern is not a simple head\-side parity direction\.Figure 10:Probe\-basis per\-failure flip\-layer analysis for Qwen 2\.5–32B\. The plot tracks when the probe agrees with the true count versus the model’s wrong emitted count on probe\-right/model\-wrong examples\. It is a probe\-basis diagnostic, not the model’s own logit\-lens basis, and complements the logit\-lens trajectories in Figure[7](https://arxiv.org/html/2609.22230#A10.F7)\.Figure 11:Model\-output prior proxy: marginal wrong\-emission distributions by model\. Each family concentrates wrong answers on a small set of integers, but the attractor set differs across families\.Figure 12:Accuracy versus model\-output\-prior proxy\. For Qwen 32B and OLMo, per\-count accuracy is lower when neighboring integers have high wrong\-emission mass\. This is a proxy for learned output prior, not a direct pretraining\-frequency measurement\.
## Appendix KReproducibility
This section records the seeds, hardware, and release plan needed to regenerate the main numbers\. Benchmark construction details are in Appendix[A](https://arxiv.org/html/2609.22230#A1); intervention protocols are in Appendix[G](https://arxiv.org/html/2609.22230#A7)\.
#### Per\-condition seeds\.
All sampling in the main and fine\-grained evals usesnumpy\.random\.default\_rng\(seed\)with the seeds in Table[19](https://arxiv.org/html/2609.22230#A11.T19)\. The same seeds are used across all seven models, so the per\-model accuracy grids cover identical1,3001\{,\}300\+200200sequences\. The random\-string pool of200200pseudowords is built once under seed303303and shared across both evals\.
Table 19:Per\-condition data\-generation seeds\. The same seeds are used across all seven models so each per\-model accuracy grid is computed on the same1,3001\{,\}300\+200200sequences\. The random\-string pool of200200pseudowords is built once under seed303303and shared across both evals\.
#### Compute\.
All bf16 evaluations were run on a single NVIDIA H100 80GB\. 4\-bit evaluations \(Qwen\-72B and Llama\-70B\) used the same hardware\. Approximate wall\-clock per model:∼\\sim1515minutes for the1,3001\{,\}300\-row main eval \(non\-reasoning\),∼\\sim22hours for R1\-Distill with thinking enabled \(long generations\),∼\\sim55minutes for the residual\-capture pass that feeds the probe,∼\\sim22minutes for the per\-layer probe fit on CPU\. The full panel \(seven models, main \+ fine eval \+ residuals\) fits comfortably in a single working day on one H100\. Mechanistic interventions on Qwen\-32B \(block decomposition, scale sweep over sixssvalues, feature\-level patches\) add∼\\sim22hours total\.
#### Code and data release\.
Evaluation scripts, model stubs, probe and intervention drivers, raw result CSVs, and figure\-generation code will be released upon acceptance\. Reproducing model evaluations requires accepting the original HuggingFace licenses for each checkpoint\.Similar Articles
Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts
This paper demonstrates that post-training quantization can silently alter how large language models reason, even when task accuracy is preserved, through a taxonomy-based analysis of 30,000 chain-of-thought outputs across multiple models and benchmarks.
Ten Failure Modes That Define Multimodal AI Systems
The article catalogs ten documented failure modes in multimodal AI systems where models generate fluent answers that break correspondence with actual inputs, based on benchmark papers and research studies.
Pigeonholing: Bad prompts hurt models to collapse and make mistakes
This paper introduces 'pigeonholing,' a phenomenon where bad prompts cause LLMs to collapse and repeat errors, leading to a 38-40% performance drop. Experiments across 10 tasks and 10 models show worsening with more conversation turns, and propose RLVR with synthetic errors as a mitigation.
Where MCP tool-selection actually breaks: retrieval-based fixes cap at ~23% of failures
Recent analysis reveals that retrieval-based tool selection for LLM agents caps out at recovering ~23% of failures, while readout-side interventions addressing attention biases recover 59-91% of failures, indicating that the real bottleneck is in the model's output processing rather than input filtering.
Your agent isn't failing because of the model, it's failing because nobody built a stop button
The article argues that the primary failure point for AI agents in production is not the model itself, but the lack of infrastructure such as stop buttons, billing oversight, and traceability for tool calls.