A new transformer passes its hidden state to the next token instead of recomputing it (25 minute read)

TLDR AI Papers

Summary

The paper introduces LIFT (Latent Information Feedback Transformer), which propagates deep-layer hidden states across generation steps instead of relying solely on decoded tokens, using teacher-forced pretraining with states derived from an off-the-shelf LM's next-token distribution. Experiments on 135M–1B models show LIFT outperforms standard Transformers on language modeling, reasoning, and procedural tasks under token-matched budgets, with code and models released.

Researchers built LIFT, a transformer that feeds its internal state back into the next generation step instead of squeezing everything through the single token it outputs. Training stays parallel because the states come precomputed from an off-the-shelf model's next-token predictions, while at inference the model feeds back its own. LIFT models from 135M to 1B parameters beat token-matched standard transformers on language modeling, reasoning, and procedural tasks, so the next question is whether the gains hold at frontier scale.
Original Article
View Cached Full Text

Cached at: 10/01/26, 02:55 PM

# Pretraining Latent Information Feedback Transformers with Teacher Supervision
Source: [https://arxiv.org/html/2609.38149](https://arxiv.org/html/2609.38149)
Dor TiroshAffiliation:Blavatnik School of Computer Science and AI, Tel Aviv UniversityEmail:[dortirosh@mail](mailto:dortirosh@mail)Mor GevaAffiliation:Blavatnik School of Computer Science and AI, Tel Aviv University

###### Abstract

Transformer language models \(LMs\) are feed\-forward: deep\-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token\. This narrow channel forces models to recompute intermediate results and to discard alternative continuations\. In this work, we remove this bottleneckduring pretraining, introducing theLIFT\(Latent Information Feedback Transformer\) architecture and training method which enable LMs to propagate state across generation\. We achieve this by turning recurrent\-state learning into a teacher\-forced prediction problem: each input token is paired with an information\-dense state, derived from the next\-token distribution of an off\-the\-shelf pretrained LM\. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state\. As the input states are precomputed, pretraining remains fully parallel across positions\. At inference, the model’s own predicted states are fed back, with a minor computational overhead that decreases with model size\. Experiments with pretrained models ranging from 135M to 1B parameters show thatLIFTconsistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token\-matched budget, while being on par with or ahead of compute\-matched Transformers\. Moreover, a controlled study on a state\-tracking task shows that a tinyLIFToutperforms same\-size Transformers trained on8×8\\timesmore data, even when trained with the states of a Transformer that fails the task\. Overall, we show that LMs can learn to exploit deep\-to\-shallow feedback during pretraining via scalable teacher supervision\.111We release our code and trained models at[https://github\.com/dortirosh1/LIFT/](https://github.com/dortirosh1/LIFT/)\.

## 1Introduction

Scaling language models \(LMs\) has been driven by major improvements to the Transformer architecture\([Fedus et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib22);[DeepSeek\-AI et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib18);[Qiu et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib51)\)\. Still, the design is feed\-forward: deeper layer representations are never fed back to shallower ones\. The only pathway through which information propagates “downward” is the decoded token\. This creates a narrow communication channel that bottlenecks the computation, leading the model to recompute previous computations\([Fan et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib21);[Biran et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib6)\)or drop possible continuations\([Zhu et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib90)\)\.

Several efforts have attempted to mitigate this bottleneck at post\-training, by teaching models to generate compressed representations of their reasoning traces\. During inference, these methods propagate a continuous hidden representation at each step instead of a single token\([Hao et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib30);[Shen et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib61);[Wei et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib75)\)\. Although promising, such methods require expensive sequential latent rollouts during training, which may not suffice for rewiring the model to this new input distribution of richer continuous representations\([Rizvi\-Martel et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib56);[Wu et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib80)\)\.

A natural question that arises is whether we can teach LMs to exploit such deep\-to\-shallow feedbackat pretraining, while avoiding infeasible training costs resulting from sequential processing\. Recent work tackles this by approximating full recurrence via Jacobi\-style iterations, updating positions in parallel through multiple forward passes\([Zeng et al\., 2026a](https://arxiv.org/html/2609.38149#bib.bib84);[Cai et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib12);[Wang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib72);[Huang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib35);[Zhang & Ren, 2026](https://arxiv.org/html/2609.38149#bib.bib86)\)\. However, this approach has an inherent quality–efficiency tradeoff: a better approximation requires more iterations and thus more training cost\.

Figure 1:We remove the information flow bottleneck in Transformers \(left\) by introducing the LIFT architecture and pretraining approach \(middle\), which allows state propagation at inference \(right\)\.In this work, we proposeLIFT\(Latent Information Feedback Transformer\), an alternative approach for pretraining LMs with state propagation, which converts recurrent\-state learning into a teacher\-forced prediction problem \(Figure[1](https://arxiv.org/html/2609.38149#S1.F1)\)\. During training, the model receives at every position a tokenxix\_\{i\}and state𝐬i\\mathbf\{s\}\_\{i\}, and is trained to predict both the next tokenxi\+1x\_\{i\+1\}and the next state𝐬i\+1\\mathbf\{s\}\_\{i\+1\}\. To keep training efficient, we capitalize on the prevalence of publicly available LMs and obtain states fromnext\-token distributionsof an external ‘‘teacher’’ model222Unlike in knowledge distillation\([Hinton et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib33)\), the teacher’s role here is to supply the student’s*inputs*, not only its targets; at inference the teacher is removed and the student consumes its own states\.which runs in parallel on the same texts\. At inference, the model feeds its own predicted states back through this channel, alongside its generated tokens\. Importantly, the states used in training are informative from the first training step and produced independently of the trained model, preserving parallelism across sequence positions\.

We evaluateLIFTthrough comprehensive experiments, demonstrating its added expressivity and consistent performance gains\. First, we considerS5S\_\{5\}permutation composition, a state\-tracking task that fixed\-depth Transformers cannot solve as sequence lengthNNgrows\([Merrill et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib46)\)\. Training two\-layer Transformer models on this task shows that, while a vanilla Transformer fails on sequences ofN≥12N\\geq 12steps,LIFTmodels solve 12\-step sequences \(100%100\\%accuracy\) even when the states come from a Transformer teacher that itself fails at that length \(3%3\\%\)\. Thus, the feedback channel lets a fixed\-depth model solve longer sequences than a Transformer of the same depth can, and parallel training on fixed\-depth Transformer states is sufficient for learning it\.

Next, we pretrainLIFTLMs from135135M to11B parameters, on up to100100B tokens at the11B scale\([Hoffmann et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib34), about5×5\\timesthe Chinchilla budget;\), and evaluate them on language modeling and a wide suite of downstream benchmarks\. Across evaluations,LIFToutperforms token\-matched Transformers at every scale, as well as distillation and Jacobi\-iteration feedback baselines, and is level with or ahead of compute\-matched Transformers\. The largest gains are on procedural tasks, such as multi\-operand arithmetic and pattern continuation, whereLIFTmatches or beats a Transformer trained on2×2\{\\times\}more tokens at every scale\.

Overall, our work shows that Transformer LMs can learn to exploit deep\-to\-shallow feedback through efficient, fully parallel pretraining with teacher\-forced states, improving both language modeling and reasoning\. Our contributions can be summarized as follows:

- •We introduce, to our knowledge, the first method for pretraining a Transformer LM with a teacher\-forced latent feedback channel: a pretrained LM supplies the states during training, keeping training fully parallel, and the model generates its own states at inference\.
- •We developLIFT, which uses next\-token distributions as states, aligning state supervision with the model’s native prediction task\.
- •On a synthetic state\-tracking task, we show that the feedback channel lets a fixed\-depthLIFTsolve the task at lengths where a Transformer of the same depth fails\.
- •We demonstrate the advantages ofLIFTLMs of 135M to 1B parameters, which achieve better language modeling capabilities and downstream performance in token\-matched and compute\-matched settings, compared to a standard Transformer and to feedback variants\.

## 2Information Flow Bottleneck in Transformer LMs

Consider anLL\-layer autoregressive Transformer LM,pθp\_\{\\theta\}, with hidden dimensionddand vocabulary𝒱\\mathcal\{V\}\([Vaswani et al\., 2017](https://arxiv.org/html/2609.38149#bib.bib70)\)\. For an input sequence of tokensx1,…,xnx\_\{1\},\\dots,x\_\{n\}, let𝐡iℓ∈ℝd\\mathbf\{h\}\_\{i\}^\{\\ell\}\\in\\mathbb\{R\}^\{d\}denote the representation at positioniiafter layerℓ\\ell, with𝐡i0\\mathbf\{h\}\_\{i\}^\{0\}being the input token embedding\. At each position, the output projectionU∈ℝ\|𝒱\|×dU\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times d\}produces the next\-token distribution:

𝐩i=softmax⁡\(𝐨i\);𝐨i=U​𝐡iL,\\mathbf\{p\}\_\{i\}=\\operatorname\{softmax\}\(\\mathbf\{o\}\_\{i\}\)\\;\\;;\\;\\;\\mathbf\{o\}\_\{i\}=U\\mathbf\{h\}\_\{i\}^\{L\},\(1\)where𝐨i\\mathbf\{o\}\_\{i\}are the logits\. At inference, information moves across positions via two channels:

Attention channel:\\displaystyle\\text\{Attention channel:\}𝐡i≤jℓ→𝐡jℓ\+1\\displaystyle\\mathbf\{h\}^\{\\ell\}\_\{i\\leq j\}\\bm\{\\rightarrow\}\\mathbf\{h\}\_\{j\}^\{\\ell\+1\}\(2\)Token feedback channel:\\displaystyle\\text\{Token feedback channel:\}𝐡iL→𝐩i→xi\+1→𝐡i\+10\.\\displaystyle\\mathbf\{h\}\_\{i\}^\{L\}\\bm\{\\rightarrow\}\\mathbf\{p\}\_\{i\}\\rightarrow x\_\{i\+1\}\\bm\{\\rightarrow\}\\mathbf\{h\}\_\{i\+1\}^\{0\}\.Communicationupwardis mediated by attention, which lets later positionsj\>ij\>iread the representation at a previous positionii\. Token feedback is the only way to move informationdownward; a representation can influence a shallower hidden representation in subsequent stepsonlyat decoding time, and only by encoding information in the current position’s output, which re\-enters as input in the next step\. Crucially, this channel is extremely narrow, limited to a single decoded tokenxi\+1x\_\{i\+1\}, which can carry at mostlog2⁡\|𝒱\|\\log\_\{2\}\|\\mathcal\{V\}\|bits of information \(≈17\\approx\{17\}bits for\|𝒱\|=100,000\|\\mathcal\{V\}\|=100\{,\}000\)\.

Allowing lower layers to access previously computed higher\-layer representations, i\.e\.,𝐡iL→𝐡i\+10\\mathbf\{h\}\_\{i\}^\{L\}\\bm\{\\rightarrow\}\\mathbf\{h\}\_\{i\+1\}^\{0\}, could potentially enhance the model’s computation\. First, later positions could reuse results computed in deep layers\. Without such access, an intermediate result that emerges deep in the model is out of reach of lower layers that need it, so the model must recompute it earlier or fail\([Fan et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib21)\)\. For instance, when a model resolves the first hop of a multi\-hop query “too late” for its remaining layers to complete the second, routing that late representation back into an earlier layer corrects up to66%66\\%of the failures\([Biran et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib6)\)\. Second, the model could carry several continuations forward, rather than committing to the sampled token\. This matters when the model must explore alternatives: in graph reachability, continuous thoughts that hold many search paths at once let a two\-layer Transformer find the answer in as many steps as the graph’s diameter, instead of following a single path, which requires quadratically more steps\([Zhu et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib90)\)\. Third, the computation’s depth could grow with the sequence, as in a recurrent network\. This matters for tasks whose number of sequential steps grows with the input, such as state tracking over a long sequence, which fixed\-depth Transformers cannot solve\([Merrill et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib46);[Merrill & Sabharwal, 2025](https://arxiv.org/html/2609.38149#bib.bib44)\)\.

## 3Latent Information Feedback Transformers \(LIFT\)

Motivated by the problem above, we proposeLIFT, a scalable architecture and training method for deep\-to\-shallow state propagation across positions, effectively removing the information flow bottleneck \(Figure[1](https://arxiv.org/html/2609.38149#S1.F1)\)\. We describe the modified architecture that fuses states with input tokens \(§[3\.1](https://arxiv.org/html/2609.38149#S3.SS1)\), and explain our approach for parallel training with teacher\-forced states \(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\)\.

### 3\.1LIFTArchitecture for Next Token\-State Prediction

Our goal is to let LMs propagate an additional state alongside the predicted token, carrying more information from deep layers\.LIFTimplements this by predicting a state along with every token:

xi\+1,𝐬i\+1=pθ​\(x≤i,𝐬≤i\)x\_\{i\+1\},\\mathbf\{s\}\_\{i\+1\}=p\_\{\\theta\}\(\\,x\_\{\\leq i\},\\mathbf\{s\}\_\{\\leq i\}\)\(3\)
##### Token and state prediction

InLIFTmodels, bothxi\+1x\_\{i\+1\}and𝐬i\+1\\mathbf\{s\}\_\{i\+1\}are derived from the model’s output logits𝐨i\\mathbf\{o\}\_\{i\}\. The next tokenxi\+1x\_\{i\+1\}is sampled as usual \(Eq\.[1](https://arxiv.org/html/2609.38149#S2.E1)\), while𝐬i\+1\\mathbf\{s\}\_\{i\+1\}is obtained by truncating the output distribution to its top\-kktokens: thekklargest logits are kept, the rest are set to−∞\-\\infty, and a softmax with temperatureτ\\taurenormalizes over the kept tokens:

𝐨i′=topk⁡\(𝐨i\)\\displaystyle\\mathbf\{o\}^\{\\prime\}\_\{i\}=\\operatorname\{top\}\_\{k\}\(\\mathbf\{o\}\_\{i\}\)\(4\)𝐬i\+1=softmax⁡\(𝐨i′/τ\)∈ℝ\|𝒱\|,\\displaystyle\\mathbf\{s\}\_\{i\+1\}=\\operatorname\{softmax\}\(\\mathbf\{o\}^\{\\prime\}\_\{i\}/\\tau\)\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\},so that𝐬i\+1\\mathbf\{s\}\_\{i\+1\}has at mostkknon\-zero entries\. Building the state from the model’s output distribution lets the model carry substantially more information than the sampled token: it ranks the alternative continuations and retains much of the context that produced it\([Morris et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib49)\)\.

Notably, a possible alternative for states could be the model’s hidden representations before the output projection\([Hao et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib30);[Shen et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib61);[Cai et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib12)\)\. We use distributions as they are the more natural choice in our case, where the model learns to predict the teacher’s states \(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\)\. For any two models that share a tokenizer, these distributions live in the same space, where coordinates correspond to token probabilities, and predicting them is a well\-established objective\([Hinton et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib33)\)\. The student’s own predictions, which replace the teacher’s states at inference, then live in that same space too\. Hidden states have no shared coordinates across models\([Sriram et al\., 2018](https://arxiv.org/html/2609.38149#bib.bib63)\), so the student would have to regress onto another model’s internal representation, which need not align with its own and which, in our early exploration, we found harder to predict\.

##### Token and state fusion

To process the joint input of tokens and states, we design a dedicatedfusionlayer that transforms them into a single representation, which feeds the standard Transformer blocks\. Specifically,𝐬i\\mathbf\{s\}\_\{i\}is first projected through the token embedding matrixE∈ℝ\|𝒱\|×dE\\in\\mathbb\{R\}^\{\|\\mathcal\{V\}\|\\times d\}, yielding the empirical mean induced by𝐬i\\mathbf\{s\}\_\{i\}, and then concatenated with𝐡i0\\mathbf\{h\}\_\{i\}^\{0\}\. The concatenation is fused through a SwiGLU block\([Shazeer, 2020](https://arxiv.org/html/2609.38149#bib.bib60)\)with a residual connection to the token embedding:

𝐮i=\[RMSNorm⁡\(𝐡i0\);RMSNorm⁡\(Ws​E⊤​𝐬i\)\]\\displaystyle\\mathbf\{u\}\_\{i\}=\\big\[\\mathrm\{RMSNorm\}\(\\mathbf\{h\}\_\{i\}^\{0\}\)\\;;\\ \\mathrm\{RMSNorm\}\(W\_\{s\}\\,E^\{\\top\}\\mathbf\{s\}\_\{i\}\)\\big\]\(5\)𝐡^i0=𝐡i0\+Wd​\(SiLU⁡\(Wg​𝐮i\)⊙Wu​𝐮i\)\\displaystyle\\hat\{\\mathbf\{h\}\}^\{0\}\_\{i\}=\\mathbf\{h\}\_\{i\}^\{0\}\+W\_\{d\}\\big\(\\mathrm\{SiLU\}\(W\_\{g\}\\mathbf\{u\}\_\{i\}\)\\odot W\_\{u\}\\mathbf\{u\}\_\{i\}\\big\)whereWs∈ℝd×dW\_\{s\}\\in\\mathbb\{R\}^\{d\\times d\},Wg,Wu∈ℝ2​d×2​dW\_\{g\},W\_\{u\}\\in\\mathbb\{R\}^\{2d\\times 2d\}andWd∈ℝd×2​dW\_\{d\}\\in\\mathbb\{R\}^\{d\\times 2d\}are the fusion layer’s parameters\. At the first position, where no previous output exists, the projected stateE⊤​𝐬1E^\{\\top\}\\mathbf\{s\}\_\{1\}is replaced by a learned vector\. The fusion layer holds all the parametersLIFTadds:11​d211d^\{2\}weights and the initial\-state vector,3%3\\%of our 1B model\. Since the Transformer stack grows withL​d2Ld^\{2\}, this share shrinks with depth\. The fused representation𝐡^i0\\hat\{\\mathbf\{h\}\}^\{0\}\_\{i\}is then fed to the standard Transformer stack, interacting with the context, which in turn contains information about previous tokens and states by the same mechanism\.

### 3\.2LIFTTraining with Teacher\-Forced States

Decoding in autoregressive Transformers is already sequential, so feeding back the state \(Eq\.[3](https://arxiv.org/html/2609.38149#S3.E3)\) adds only a minor cost\. However, at training time, such recurrent computations forfeit the parallelism that enables scalable training of Transformer LMs\. To tackle this, we observe that this dependency disappears if the input states are computedin advance: instead of obtaining them via autoregressive generation, the states can be supplied by an external teacher, throughteacher forcing\. This reduces the problem to finding a teacher capable of producing informative states\. Here, we capitalize on the growing availability of open pretrained LMs\([Wolf et al\., 2020](https://arxiv.org/html/2609.38149#bib.bib77)\), which offers many capable teachers, recycling their pre\-invested compute into a higher\-quality training signal\.

##### Training with teacher states

Given a training sequence𝐱=⟨x1,…,xn⟩\\mathbf\{x\}=\\langle x\_\{1\},\\dots,x\_\{n\}\\rangle, the teacher first processes𝐱\\mathbf\{x\}in parallel, emitting for every positioniiits own output logits𝐨i⋆\\mathbf\{o\}^\{\\star\}\_\{i\}\. These logits are then transformed according to Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)to yield teacher states𝐬i\+1⋆\\mathbf\{s\}^\{\\star\}\_\{i\+1\}\. Notably, as the teacher is frozen, it admits common training optimizations that require fixed weights\([Xiao et al\., 2023](https://arxiv.org/html/2609.38149#bib.bib81);[Aminabadi et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib2), e\.g\.,\)\. Once the states are all available, they are paired with their matching tokens, i\.e\.,\(xi\+1,𝐬i\+1⋆\)\(x\_\{i\+1\},\\mathbf\{s\}^\{\\star\}\_\{i\+1\}\)\. The pairs are then passed to the model’s fusion layerin parallel\(Eq\.[5](https://arxiv.org/html/2609.38149#S3.E5)\), and forwarded through the full backbone as in standard causal\-LM training,withoutpropagating gradients to the teacher\. The state channel is thus teacher\-forced, same as the token channel, and full parallelism is maintained\.

At inference, states are produced by the model autoregressively, making input processing sequential\. To allow prompt prefilling in one parallel pass without states, we employprefix state dropoutduring training: the states of each training sequence are replaced with probabilityppby a learned bias𝐛∈ℝd\\mathbf\{b\}\\in\\mathbb\{R\}^\{d\}\(§[A\.1](https://arxiv.org/html/2609.38149#A1.SS1)\)\. The model thus learns to operate both with and without states, allowing prefilling\.

##### Mitigating train–inference state distribution shifts

Because the model is trained on teacher states but consumes its own states at inference, a mismatch between the two distributions can accumulate over decoding steps\([Ranzato et al\., 2016](https://arxiv.org/html/2609.38149#bib.bib54)\)\. We mitigate this in two ways\. First, we supervise the model’s state predictions with the teacher’s, so that the states it generates at inference remain close to those it was trained on\. To this end, we add a state\-alignment term to the training objective:

ℒ=∑i\(−log⁡𝐩i​\(xi\+1\)⏟language modeling\+λ​𝒟k\(𝐨i⋆∥𝐨i\)⏟state alignment\),\\mathcal\{L\}=\\sum\_\{i\}\\Big\(\\underbrace\{\-\\log\\mathbf\{p\}\_\{i\}\(x\_\{i\+1\}\)\}\_\{\\text\{language modeling\}\}\\;\+\\;\\lambda\\,\\underbrace\{\\mathcal\{D\}\_\{k\}\\big\(\\mathbf\{o\}\_\{i\}^\{\\star\}\\,\\\|\\,\\mathbf\{o\}\_\{i\}\\big\)\}\_\{\\text\{state alignment\}\}\\Big\),\(6\)whereλ≥0\\lambda\\geq 0weights the two terms and𝒟k\\mathcal\{D\}\_\{k\}is a forward\-KL divergence\([Hinton et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib33)\)adapted to top\-kkdistributions \(§[A\.2](https://arxiv.org/html/2609.38149#A1.SS2)\)\.

Second, in the last10%10\\%of training, we adapt the model to its own states, dropping the teacher and the alignment term\. Each training step executes one or two initial forward passes \(with probability50%50\\%\), to produce the model’s own states: the first pass reads the input tokens with the bias𝐛\\mathbf\{b\}at every position\. The second \(optional\) pass is fed the states predicted by the first\. To train the model, we pass the states produced at the end of the process above instead of those typically supplied by the teacher, without propagating gradients back to the initial iterations \(further details are in §[A\.3](https://arxiv.org/html/2609.38149#A1.SS3)\)\.

## 4Evaluation of State Tracking Capabilities

We start by evaluatingLIFTon state tracking, showing the expressive advantage of its architecture over a standard Transformer, and how it can learn from both strong and weak teachers\.

Figure 2:S5S\_\{5\}accuracy at increasing task lengthsNN\(linear forN≤16N\{\\leq\}16, logarithmic forN\>16N\{\>\}16\), of Transformers \(solid lines\) versusLIFTmodels \(dashed lines\)\. Standard error is over seeds\.##### TheS5S\_\{5\}task

In this task, a model must predict the composition of a sequence of permutations, which requires state tracking along the sequence\([Liu et al\., 2023](https://arxiv.org/html/2609.38149#bib.bib43)\)\. LetS5S\_\{5\}be the group of permutations of five elements \(120 in total\)\. Given an initial states0∈S5s\_\{0\}\\in S\_\{5\}and permutationsa1,…,aN∈S5a\_\{1\},\\dots,a\_\{N\}\\in S\_\{5\}, the model must predict the final statesNs\_\{N\}obtained by applying them tos0s\_\{0\}:

sN=aN∘⋯∘a1∘s0,si=ai∘si−1\.s\_\{N\}=a\_\{N\}\\circ\\cdots\\circ a\_\{1\}\\circ s\_\{0\},\\qquad s\_\{i\}=a\_\{i\}\\circ s\_\{i\-1\}\.\(7\)We encode all the elements inS5S\_\{5\}as single tokens; thus, the model is given the sequence of tokens⟨s0,a1,…,aN⟩\\langle s\_\{0\},a\_\{1\},\\dots,a\_\{N\}\\rangleand should predict a single token\. Importantly, this task is𝖭𝖢1\\mathsf\{NC\}^\{1\}\-complete: unless𝖳𝖢0=𝖭𝖢1\\mathsf\{TC\}^\{0\}=\\mathsf\{NC\}^\{1\}, a Transformer of fixed depth cannot solve it for growingNN, and the number of layers must grow asΩ⁡\(log⁡N\)\\Omega\(\\log N\)\([Merrill & Sabharwal, 2025](https://arxiv.org/html/2609.38149#bib.bib44)\)\. Conversely, a recurrent model needs only a constant number of layers, as its computation depth grows with the sequence\([Merrill et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib46)\)\.

##### Experimental setup

We train two\-layer Transformers with22M parameters from scratch;LIFTadds its fusion layer \(7%7\\%of the parameters\)\. Input sequences of lengthsN∈\[1,16\]N\\in\[1,16\]are drawn from1212fixed generators ofS5S\_\{5\}at every training step, with a random initial state\. We report accuracy on2,0002\{,\}000held\-out sequences per length, up toN=64N\{=\}64, averaged accuracy over 3 seeds \(see §[B\.1](https://arxiv.org/html/2609.38149#A2.SS1)for more details\)\. We compare the following models:

1. 1\.Transformer: A token\-matched vanilla Transformer, trained for5,0005\{,\}000steps\. Following[Liu et al\. \(2023\)](https://arxiv.org/html/2609.38149#bib.bib43), the loss is computed on the correct statesis\_\{i\}at every positionii, not only on the finalsNs\_\{N\}\. At test time, it answers in one parallel pass\.
2. 2\.Transformer×\\times8: Same as the above baseline, except that the model is trained8×8\\timeslonger \(40,00040\{,\}000steps\)\. Notably, larger training budgets do not yield substantial improvements \(see §[C\.1](https://arxiv.org/html/2609.38149#A3.SS1)\)\.
3. 3\.LIFT\-RNN: TheLIFTarchitecture trained sequentially for40,00040\{,\}000steps: at every position the output distribution is embedded with the token and passed to the next position, and gradients flow across steps using only the answer loss\.
4. 4\.LIFT\(Transformer×\\times8 teacher\): Trained in parallel for5,0005\{,\}000steps with this teacher\.
5. 5\.LIFT\(RNN teacher\): Trained in parallel for5,0005\{,\}000steps withLIFT\-RNN as a teacher\.

##### Results

Figure[2](https://arxiv.org/html/2609.38149#S4.F2)shows that the fixed\-depth Transformer struggles with this task, as theory predicts: even with8×8\\timesthe training budget, it solves only lengths up toN=6N\{=\}6and is at chance fromN=12N\{=\}12on\. Training sequentially with latent feedback \(LIFT\-RNN\) removes this bottleneck, allowing a model of the same depth to solve every length in the training range\. When trained fully in parallel on the states of this recurrent model,LIFTinherits its unbounded computation depth: it matches the teacher in the training range and even extrapolates beyond it, achieving96%96\\%atN=32N\{=\}32and65%65\\%atN=64N\{=\}64, although trained only on lengths up toN=16N\{=\}16\. Interestingly,LIFTshows such extrapolation even when trained with a depth\-bounded non\-recurrent teacher\.LIFTsolves every length up toN=12N\{=\}12and half of the sequences atN=16N\{=\}16, whereas its teacher, a same\-size Transformer trained with8×8\\timesthe budget, is at chance beyondN=8N\{=\}8\. We stress that, trained this way,LIFTdoes not acquire unbounded depth and fails beyond the training range\. Still, learning from the teacher’s states yields a large gain over a standard Transformer, even one trained much longer\.

## 5Experiments with Pretrained LMs

We conduct comprehensive evaluations ofLIFTLMs in two setups, benchmarking against vanilla Transformer and distillation baselines on a wide set of benchmarks \(§[5\.2](https://arxiv.org/html/2609.38149#S5.SS2)\), and comparing with feedback Transformers that approximate recurrent states through repeated parallel passes \(§[5\.3](https://arxiv.org/html/2609.38149#S5.SS3)\)\.

### 5\.1Experimental Setting

##### Pretrained LMs

We pretrainLIFTLMs with 135M, 350M and 1B parameters, using the OLMo 2 architecture, codebase and pretraining data\([Walsh et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib71)\), each for5×5\\timesits Chinchilla token budget\([Hoffmann et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib34)\)\. Models at 135M scale are trained with three seeds, and larger models with one seed \(due to compute constraints\)\. The teacher at each scale is a same\-size model pretrained for10×10\\timesits Chinchilla budget; stronger teachers, either larger or trained for much longer, change the results only marginally \(§[D\.4](https://arxiv.org/html/2609.38149#A4.SS4)\)\.LIFTusesk=1024k\{=\}1024,τ=1\.5\\tau\{=\}1\.5andλ=1\.5\\lambda\{=\}1\.5at every scale; see §[D](https://arxiv.org/html/2609.38149#A4)for ablations and §[B\.2](https://arxiv.org/html/2609.38149#A2.SS2)for additional implementation details\. We compare against the following baselines, all evaluated after an identical learning\-rate annealing phase\([Wen et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib76)\):

1. 1\.Vanilla Transformer: A standard OLMo 2 Transformer, identical toLIFTexcept for the fusion layer\. We report results for a model trained on the same token budget and on the same compute budget as ours, accounting for the teacher’s forward pass \(§[F\.1](https://arxiv.org/html/2609.38149#A6.SS1.SSS0.Px1)\)\.
2. 2\.Distillation: The vanilla Transformer trained with the forward\-KL distillation term\([Hinton et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib33)\)in Eq\.[6](https://arxiv.org/html/2609.38149#S3.E6), isolating the effect of the distillation objective from that of the latent feedback\.
3. 3\.LIFTw/o states: Our trainedLIFTLM with the channel “switched off”: in place of a state, the fusion layer receives the learned bias𝐛\\mathbf\{b\}of prefix state dropout \(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\)\. No state is fed back, so the model runs as a standard Transformer in one parallel pass\.

##### Comparison with feedback Transformers

We adapt the setup of T2MLR\([Cai et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib12)\): a Transformer whose layers read a state computed by deeper layers at earlier positions, pretrained by refining all positions in parallel over repeated passes \(Jacobi iterations\) instead of sequentially\. All models share the SmolLM2\-135M backbone\([Allal et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib1)\)and are trained on a 10B\-token FineWeb\-Edu sample\([Penedo et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib50)\)\. We compareLIFTwithT2MLR, both our own training run and the published checkpoint\. In addition, we compare with the concurrent models of[Wang et al\. \(2026\)](https://arxiv.org/html/2609.38149#bib.bib72), which feed the hidden state at the previous position back to the input and train with a few passes per token\. At the time of writing, the official code and models are not publicly available, so we implemented the method ourselves, using the schedule of the authors’ larger runs; we call this adaptationMulti\-pass Transformer\. We also compare against a vanilla Transformer\. Models are compared under \(a\) matched tokens, each trains on about 10B tokens, and \(b\) matched compute, each receives the training FLOPs of T2MLR’s 10B\-token run, so cheaper methods train on more tokens, revisiting the data once exhausted\. See §[B\.3](https://arxiv.org/html/2609.38149#A2.SS3)for implementation details\.

##### Evaluation benchmarks

We use the following evaluations for both sets of experiments:

1. 1\.Language modeling: Perplexity on held\-out text of the training corpus: for the three\-scale experiments, the 1M\-token C4\([Raffel et al\., 2020](https://arxiv.org/html/2609.38149#bib.bib52)\)validation subset of the OLMo 2 evaluation set\([Walsh et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib71)\); for the feedback comparison, a held\-out FineWeb\-Edu sample of 127K tokens\. Perplexity on the other sources of the OLMo 2 evaluation set is reported in §[C](https://arxiv.org/html/2609.38149#A3)\.
2. 2\.Downstream tasks: Twenty benchmark suite run through OLMES\([Gu et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib29)\), covering the base\-model evaluation families of OLMo 2 and OLMo 3\([Team Olmo et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib69)\)\. Eight of these are at chance or floor \(under 3%\) for all models at these scales; the other twelve form three groups, each reported as the mean over its benchmarks \(the full list is in §[B\.4](https://arxiv.org/html/2609.38149#A2.SS4)\): - •*MC*: Multiple\-choice questions on science, commonsense and world knowledge\. Each answer option is scored by its likelihood given the question and the highest wins\. - •*Gen\.*: Generative tasks on reading comprehension, knowledge recall and multi\-step reasoning\. The model generates the answer, scored by F1 or exact match against the reference\. - •*RC*: Reference\-completion tasks in science, knowledge, math and code\. The reference answer is scored by its likelihood under the model, reported in bits\-per\-byte \(BPB\), lower is better\.
3. 3\.Arithmetic: Our own synthetic task, testing multi\-step computation carried in the model’s latent state \(no chain of thought\)\. Each sample adds and subtracts 2–10 single\-digit numbers, e\.g\.2 \+ 0 \- 6 \+ 7 \+ 5 =\(answer8\); we report the correct answer’s BPB \(more metrics in §[B\.4](https://arxiv.org/html/2609.38149#A2.SS4)\)\.

### 5\.2Evaluating Language Modeling, Knowledge, and Reasoning Capabilities

##### LIFToutperforms Transformer baselines at matched tokens and at matched compute

Table[1](https://arxiv.org/html/2609.38149#S5.T1)presents the main results, with a per\-benchmark breakdown in §[C](https://arxiv.org/html/2609.38149#A3)\. At matched tokens,LIFTis ahead on every metric at every scale: perplexity drops by55–5\.5%5\.5\\%, and it wins 33 of the 36 downstream benchmark comparisons\. At matched compute, where the Transformer trains on over40%40\\%more data,LIFTkeeps the lead on every metric at 350M and 1B and is level with or ahead of it at 135M\. These gains are consistent over the three 135M\-scale training runs \(see §[C\.4](https://arxiv.org/html/2609.38149#A3.SS4)\)\. Importantly, the gains come from the added channel rather than from the training signal alone:LIFTw/o states, the same weights run without the channel, and the distillation baseline, trained with the same state\-alignment objective but without the channel, trailLIFTon perplexity and on most of the benchmarks at every scale\. Further ablations reinforce this \(§[D](https://arxiv.org/html/2609.38149#A4)\): trainingLIFTwith a narrower channel, without the state\-alignment objective, or without the teacher at all significantly erodes the gains\.

##### LIFT’s token advantage grows with training

Figure[3\(b\)](https://arxiv.org/html/2609.38149#S5.F3.sf2)compares the cross\-entropy on the evaluation set at matched checkpoints along training, showing how many more training tokens the Transformer baseline needs to reach the performance ofLIFT\. At every scale, this multiplier grows along training: the longer both models train, the more tokens the Transformer needs to catch up, suggesting that the token efficiency ofLIFTincreases with training\.

##### The largest gains are on procedural tasks

Across all evaluations,LIFTachieves the largest gains on procedural tasks—our arithmetic task and the pattern continuation task from the OLMES basic\-skills suite—which require a step\-by\-step answer derivation\. On arithmetic,LIFTis ahead of the Transformer baseline at every scale, outperforming its own teacher, which was trained on2×2\\timesmore tokens \(11–5%5\\%fewer bits on average; §[C\.5](https://arxiv.org/html/2609.38149#A3.SS5)\)\. These gains persist across increasing numbers of operands \(Figure[3\(a\)](https://arxiv.org/html/2609.38149#S5.F3.sf1)\)\. Pattern continuation asks for the next element of a short sequence that follows a rule, over numbers, letters, symbols or dates \(e\.g\.,2 4 6 8→\\to10,E G I K→\\toM\)\.LIFTbeats the token\-matched Transformer by66–10%10\\%in bits per byte at every scale, and the2×2\\times\-token Transformer by44–5%5\\%at 350M and 1B \(a tie at 135M; §[C\.6](https://arxiv.org/html/2609.38149#A3.SS6)\)\.

##### Performance with parallel prefill

LIFTmodels can process the input prompt sequentially \(with state predictions\) or in parallel \(with the bias state;LIFTw/o states in Table[1](https://arxiv.org/html/2609.38149#S5.T1)\)\. The results show thatLIFTw/o states is comparable to the token\-matched Transformer baseline across all scales, indicatingLIFTmaintains strong capabilities without sequential processing of the input\. In §[E](https://arxiv.org/html/2609.38149#A5), we compare additional methods for performing parallel prefill withLIFTmodels, which vary in efficiency and quality\. We find that iteratively refining the prompt, as in the adaptation phase \(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\), recovers much of the gain of sequential processing within a few iterations\.

Table 1:Performance ofLIFTLMs against non\-feedback Transformer baselines at three scales\.\(a\)Arithmetic\(b\)Language Modeling
Figure 3:LIFTperformance on arithmetic and language modeling\.\(a\)LIFTshows a consistent advantage on arithmetic that persists as task complexity increases, outperforming a baseline trained on2×2\\timesmore tokens\.\(b\)Throughout training,LIFTis more token\-efficient than a vanilla Transformer, and the advantagegrowswith training: each horizontal segment marks the additional tokens the Transformer needs to reach the same loss\.

### 5\.3Comparison with Jacobi\-Iteration Feedback Transformers

Table[2](https://arxiv.org/html/2609.38149#S5.T2)shows the average results over three training seeds, with seed noise and additional results reported in §[C\.7](https://arxiv.org/html/2609.38149#A3.SS7)\. At matched tokens,LIFTshows the best average performance on every metric, tied with T2MLR on generation tasks, though we note that in this setting all methods are within noise\. At matched compute,LIFTis on par with or better than the vanilla Transformer on all tasks, outperforming the latent feedback baselines\. On the arithmetic task, performance is within the noise estimate compared to our reproduction of T2MLR, though this is mainly due to the relatively high variance of the latter\. Notably, the feedback baselines fall behind the vanilla Transformer on every metric \(except arithmetic\) once compute is matched\.

Crucially, pretraining with multiple iterations introduces a tradeoff between approximation quality and training efficiency\. More iterations improve state approximation, providing a better training signal per token, but they make each token more expensive to process\([Cai et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib12)\)\. Our results reflect this tradeoff: with 16 refinement passes, T2MLR achieves lower perplexity than the Multi\-pass Transformer at matched tokens, but at matched compute, their order reverses as the Multi\-pass Transformer’s fewer passes let it train on more tokens\.LIFTobtains its training states from a teacher, avoiding this iterative approximation, and is equal to or better than both methods on every metric under either budget\. Multi\-pass training also increases memory use because activations from the passes used for backpropagation must be stored\.LIFTuses less peak memory than both feedback baselines at the measured batch sizes \(see §[F\.2](https://arxiv.org/html/2609.38149#A6.SS2)for memory usage comparison\)\.

Table 2:Comparison with Jacobi\-iteration feedback Transformers at matched tokens and matched compute\. We report results for T2MLR using the official checkpoint \(“HF”\) and our reproduction\.

## 6Related Work

##### Latent feedback in post\-training

A growing body of work attempts to post\-train LMs to use forms of latent feedback, typically focusing on reasoning tasks\.[Hao et al\. \(2025\)](https://arxiv.org/html/2609.38149#bib.bib30)train LMs to reason by treating the last hidden state as a latent token, processed for a fixed number of steps\. This approach, and several follow\-up efforts that adapt the learning signal\([Shen et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib61);[Wei et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib75)\), rely on sequential rollouts that are difficult to scale\. To reduce this cost,[Wu et al\. \(2025\)](https://arxiv.org/html/2609.38149#bib.bib79)replace the sequential rollouts with Jacobi iterations\. Rather than replacing tokens with latents,[Tang & Lu \(2026\)](https://arxiv.org/html/2609.38149#bib.bib68)keep the decoded token and route the previous token’s higher\-layer states to lower layers, fine\-tuning the model with sequential processing over groups of tokens\. Alternative approaches feed back soft tokens, the probability\-weighted average of token embeddings, similar toLIFT\. These soft tokens are used either without any training\([Zhang et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib88);[Zhuang et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib91)\)or with reinforcement learning\([Butt et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib11)\)\. Finally, some works teacher\-force the latent inputs using supervision from gold reasoning traces\([Tan et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib67);[Amos et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib3)\)or from an oracle\([Gozeten et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib28)\); this enables parallel training but requires annotated reasoning\.

##### Latent feedback in pretraining

More related to our work are methods focusing on latent feedback during pretraining\. Early approaches compute the states with exact recurrence, which makes training sequential over positions\([Fan et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib21)\)or over segments\([Bulatov et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib8)\)\. More recent feedback LMs instead approximate Jacobi iterations, following fixed\-point iteration methods developed to parallelize RNN training\([Lim et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib42);[Danieli et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib17)\)\. These LMs feed back either keys and values\([Zhang & Ren, 2026](https://arxiv.org/html/2609.38149#bib.bib86)\)or hidden states\([Zeng et al\., 2026a](https://arxiv.org/html/2609.38149#bib.bib84);[Huang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib35);[Cai et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib12);[Wang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib72)\), and trade training cost for approximation quality, a tradeoff we examine in §[5\.3](https://arxiv.org/html/2609.38149#S5.SS3)\. In all of these methods, states are self\-generated at training time, requiring sequential rollouts or an approximation, unlikeLIFT, which teacher\-forces the states with a trained LM\. In[Kumar & Isola \(2026\)](https://arxiv.org/html/2609.38149#bib.bib38), an RNN is similarly pretrained with teacher\-forced memory, but using an encoder trained in tandem\.[Tack et al\. \(2026\)](https://arxiv.org/html/2609.38149#bib.bib65)use a pretrained teacher, but only as a source of targets: the model predicts concepts extracted from it and mixes them into its hidden states\. We provide an extended discussion and detailed comparison in §[G](https://arxiv.org/html/2609.38149#A7)\.

## 7Conclusion

We introduceLIFT, a Transformer LM enhanced with deep\-to\-shallow feedback during pretraining, realized by a state constructed from the next\-token output distribution\. Unlike prior methods that rely on expensive sequential rollouts or state approximations,LIFTleverages an external teacher model to provide informative states, supporting parallel training\. During pretraining,LIFTis aligned to predict its own states, removing the reliance on the teacher at inference time\. In a controlled setting on anS5S\_\{5\}state\-tracking task, a two\-layerLIFTsuccessfully solves the task at lengths where its own teacher, trained for8×8\\timeslonger, fails to exceed chance performance\. Through extensive LM pretraining experiments at the 135M–1B parameter scale, we showLIFToutperforms a standard Transformer on a variety of tasks, when matched in tokens and compute, and displays training token efficiency that grows as training progresses\. Compared against alternative feedback methods,LIFTis comparable to or better than all baselines considered in both token and compute matched setting\.

Notably, our work focuses on a certain choice for the teacher that shares the architecture of the student\. Studying the effects of teachers with other sizes and possibly different vocabularies is an attractive avenue for future work\. Another is trainingLIFTfar beyond its teacher’s training budget, a regime in which its performance and training dynamics remain open questions\.

#### Acknowledgments

We thank Oren Pereg and Jonathan Mamou for their support and useful discussions in early stages of this work\. We also thank Ziyang Cai and Xingyu Zhu for kindly sharing their trained T2MLR models\. This research was supported in part by Intel and a Moonshot Grant by the Planning and Budgeting Committee of the Council for Higher Education in Israel\. It was made possible through a GPU compute resource grant funded by the Association of University Heads, the Council for Higher Education and the AI Research Compute Center\.

## References

- Allal et al\. \(2025\)Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Agustín Piqueres Lajarín, Hynek Kydlíček, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan\-Son Nguyen, Ben Burtenshaw, Clémentine Fourrier, Haojun Zhao, Hugo Larcher, Mathieu Morlon, Cyril Zakka, Colin Raffel, Leandro von Werra, and Thomas Wolf\.SmolLM2: When smol goes big — data\-centric training of a fully open small language model\.In*Second Conference on Language Modeling*, 2025\.URL[https://openreview\.net/forum?id=3JiCl2A14H](https://openreview.net/forum?id=3JiCl2A14H)\.
- Aminabadi et al\. \(2022\)Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He\.DeepSpeed\-Inference: Enabling efficient inference of Transformer models at unprecedented scale\.In*SC22: International Conference for High Performance Computing, Networking, Storage and Analysis*, pp\. 1–15\. IEEE, November 2022\.doi:10\.1109/SC41404\.2022\.00051\.
- Amos et al\. \(2026\)Ido Amos, Avi Caciularu, Mor Geva, Amir Globerson, Jonathan Herzig, Lior Shani, and Idan Szpektor\.Latent reasoning with supervised thinking states, 2026\.URL[https://arxiv\.org/abs/2602\.08332](https://arxiv.org/abs/2602.08332)\.
- Austin et al\. \(2021\)Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton\.Program synthesis with large language models\.*arXiv preprint arXiv:2108\.07732*, 2021\.
- Bengio et al\. \(2015\)Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer\.Scheduled sampling for sequence prediction with recurrent neural networks\.In C\. Cortes, N\. Lawrence, D\. Lee, M\. Sugiyama, and R\. Garnett \(eds\.\),*Advances in Neural Information Processing Systems*, volume 28\. Curran Associates, Inc\., 2015\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2015/file/e995f98d56967d946471af29d7bf99f1\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2015/file/e995f98d56967d946471af29d7bf99f1-Paper.pdf)\.
- Biran et al\. \(2024\)Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson\.Hopping too late: Exploring the limitations of large language models on multi\-hop queries\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen \(eds\.\),*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pp\. 14113–14130, Miami, Florida, USA, November 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.emnlp\-main\.781\.URL[https://aclanthology\.org/2024\.emnlp\-main\.781/](https://aclanthology.org/2024.emnlp-main.781/)\.
- Bisk et al\. \(2020\)Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi\.PIQA: Reasoning about physical commonsense in natural language\.*Proceedings of the AAAI Conference on Artificial Intelligence*, 34\(05\):7432–7439, April 2020\.ISSN 2159\-5399\.doi:10\.1609/aaai\.v34i05\.6239\.URL[https://ojs\.aaai\.org/index\.php/AAAI/article/view/6239](https://ojs.aaai.org/index.php/AAAI/article/view/6239)\.
- Bulatov et al\. \(2022\)Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev\.Recurrent memory Transformer\.In S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(eds\.\),*Advances in Neural Information Processing Systems*, volume 35, pp\. 11079–11091\. Curran Associates, Inc\., 2022\.doi:10\.52202/068431\-0805\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/file/47e288629a6996a17ce50b90a056a0e1\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2022/file/47e288629a6996a17ce50b90a056a0e1-Paper-Conference.pdf)\.
- Burns et al\. \(2024\)Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu\.Weak\-to\-strong generalization: Eliciting strong capabilities with weak supervision\.In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp \(eds\.\),*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 4971–5012\. PMLR, 21–27 Jul 2024\.URL[https://proceedings\.mlr\.press/v235/burns24b\.html](https://proceedings.mlr.press/v235/burns24b.html)\.
- Busbridge et al\. \(2025\)Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, and Russell Webb\.Distillation scaling laws\.In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste\-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu \(eds\.\),*Proceedings of the 42nd International Conference on Machine Learning*, volume 267 of*Proceedings of Machine Learning Research*, pp\. 5977–6045\. PMLR, 13–19 Jul 2025\.URL[https://proceedings\.mlr\.press/v267/busbridge25a\.html](https://proceedings.mlr.press/v267/busbridge25a.html)\.
- Butt et al\. \(2026\)Natasha Butt, Ariel Kwiatkowski, Ismail Labiad, Julia Kempe, and Yann Ollivier\.Soft tokens, hard truths\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 114650–114675, 2026\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/ba9f3bcc11c65c9591a563254e863c10\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/ba9f3bcc11c65c9591a563254e863c10-Paper-Conference.pdf)\.
- Cai et al\. \(2026\)Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He, and Sanjeev Arora\.T2MLR: Transformer with temporal middle\-layer recurrence, 2026\.URL[https://arxiv\.org/abs/2607\.15178](https://arxiv.org/abs/2607.15178)\.
- Chen et al\. \(2021\)Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al\.Evaluating large language models trained on code\.*arXiv preprint arXiv:2107\.03374*, 2021\.
- Clark et al\. \(2019\)Christopher Clark, Kenton Lee, Ming\-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova\.BoolQ: Exploring the surprising difficulty of natural yes/no questions\.In Jill Burstein, Christy Doran, and Thamar Solorio \(eds\.\),*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pp\. 2924–2936, Minneapolis, Minnesota, June 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/N19\-1300\.URL[https://aclanthology\.org/N19\-1300/](https://aclanthology.org/N19-1300/)\.
- Clark et al\. \(2018\)Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord\.Think you have solved question answering? Try ARC, the AI2 reasoning challenge\.*arXiv preprint arXiv:1803\.05457*, 2018\.
- Cobbe et al\. \(2021\)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\.Training verifiers to solve math word problems\.*arXiv preprint arXiv:2110\.14168*, 2021\.
- Danieli et al\. \(2026\)Federico Danieli, Pau Rodriguez, Miguel Sarabia, Xavier Suau, and Luca Zappella\.ParaRNN: Unlocking parallel training of nonlinear RNNs for large language models\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 53548–53579, 2026\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/57d8ebf4c2f050a6485f370d47656a9e\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/57d8ebf4c2f050a6485f370d47656a9e-Paper-Conference.pdf)\.
- DeepSeek\-AI et al\. \(2024\)DeepSeek\-AI et al\.DeepSeek\-V3 technical report, 2024\.URL[https://arxiv\.org/abs/2412\.19437](https://arxiv.org/abs/2412.19437)\.
- Dehghani et al\. \(2019\)Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser\.Universal Transformers\.In*International Conference on Learning Representations*, 2019\.URL[https://openreview\.net/forum?id=HyzdRiR9Y7](https://openreview.net/forum?id=HyzdRiR9Y7)\.
- Dua et al\. \(2019\)Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner\.DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs\.In Jill Burstein, Christy Doran, and Thamar Solorio \(eds\.\),*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pp\. 2368–2378, Minneapolis, Minnesota, June 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/N19\-1246\.URL[https://aclanthology\.org/N19\-1246/](https://aclanthology.org/N19-1246/)\.
- Fan et al\. \(2021\)Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar\.Addressing some limitations of transformers with feedback memory, 2021\.URL[https://arxiv\.org/abs/2002\.09402](https://arxiv.org/abs/2002.09402)\.
- Fedus et al\. \(2022\)William Fedus, Barret Zoph, and Noam Shazeer\.Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity\.*Journal of Machine Learning Research*, 23\(120\):1–39, 2022\.URL[http://jmlr\.org/papers/v23/21\-0998\.html](http://jmlr.org/papers/v23/21-0998.html)\.
- Feng et al\. \(2023\)Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang\.Towards revealing the mystery behind chain of thought: A theoretical perspective\.In A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(eds\.\),*Advances in Neural Information Processing Systems*, volume 36, pp\. 70757–70798\. Curran Associates, Inc\., 2023\.doi:10\.52202/075280\-3100\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/dfc310e81992d2e4cedc09ac47eff13e\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/dfc310e81992d2e4cedc09ac47eff13e-Paper-Conference.pdf)\.
- Furlanello et al\. \(2018\)Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar\.Born again neural networks\.In Jennifer Dy and Andreas Krause \(eds\.\),*Proceedings of the 35th International Conference on Machine Learning*, volume 80 of*Proceedings of Machine Learning Research*, pp\. 1607–1616\. PMLR, 10–15 Jul 2018\.URL[https://proceedings\.mlr\.press/v80/furlanello18a\.html](https://proceedings.mlr.press/v80/furlanello18a.html)\.
- Geiping et al\. \(2025\)Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein\.Scaling up test\-time compute with latent reasoning: A recurrent depth approach\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 41340–41391\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-1380\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/3b01972cf31e6fa0fe29e4b8b5c2a0a1\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/3b01972cf31e6fa0fe29e4b8b5c2a0a1-Paper-Conference.pdf)\.
- Gemma Team \(2024\)Gemma Team\.Gemma 2: Improving open language models at a practical size, 2024\.URL[https://arxiv\.org/abs/2408\.00118](https://arxiv.org/abs/2408.00118)\.
- Gloeckle et al\. \(2024\)Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez\-Paz, and Gabriel Synnaeve\.Better & faster large language models via multi\-token prediction\.In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp \(eds\.\),*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 15706–15734\. PMLR, 21–27 Jul 2024\.URL[https://proceedings\.mlr\.press/v235/gloeckle24a\.html](https://proceedings.mlr.press/v235/gloeckle24a.html)\.
- Gozeten et al\. \(2026\)Alperen Gozeten, Muhammed Ildiz, Xuechen Zhang, Hrayr Harutyunyan, Ankit Singh Rawat, and Samet Oymak\.Continuous chain of thought enables parallel exploration and reasoning\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 82881–82914, 2026\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/85ec26ef94c3acb4c195e905df1ff4f7\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/85ec26ef94c3acb4c195e905df1ff4f7-Paper-Conference.pdf)\.
- Gu et al\. \(2025\)Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi\.OLMES: A standard for language model evaluations\.In Luis Chiruzzo, Alan Ritter, and Lu Wang \(eds\.\),*Findings of the Association for Computational Linguistics: NAACL 2025*, pp\. 5020–5048, Albuquerque, New Mexico, April 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-195\-7\.doi:10\.18653/v1/2025\.findings\-naacl\.282\.URL[https://aclanthology\.org/2025\.findings\-naacl\.282/](https://aclanthology.org/2025.findings-naacl.282/)\.
- Hao et al\. \(2025\)Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason E Weston, and Yuandong Tian\.Training large language models to reason in a continuous latent space\.In*Second Conference on Language Modeling*, 2025\.URL[https://openreview\.net/forum?id=Itxz7S4Ip3](https://openreview.net/forum?id=Itxz7S4Ip3)\.
- Hendrycks et al\. \(2021a\)Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt\.Measuring massive multitask language understanding\.In*International Conference on Learning Representations*, 2021a\.URL[https://openreview\.net/forum?id=d7KBjmI3GmQ](https://openreview.net/forum?id=d7KBjmI3GmQ)\.
- Hendrycks et al\. \(2021b\)Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the MATH dataset\.In J\. Vanschoren and S\. Yeung \(eds\.\),*Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks*, volume 1, 2021b\.URL[https://datasets\-benchmarks\-proceedings\.neurips\.cc/paper\_files/paper/2021/file/be83ab3ecd0db773eb2dc1b0a17836a1\-Paper\-round2\.pdf](https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/be83ab3ecd0db773eb2dc1b0a17836a1-Paper-round2.pdf)\.
- Hinton et al\. \(2015\)Geoffrey Hinton, Oriol Vinyals, and Jeff Dean\.Distilling the knowledge in a neural network, 2015\.URL[https://arxiv\.org/abs/1503\.02531](https://arxiv.org/abs/1503.02531)\.
- Hoffmann et al\. \(2022\)Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karén Simonyan, Erich Elsen, Oriol Vinyals, Jack Rae, and Laurent Sifre\.An empirical analysis of compute\-optimal large language model training\.In S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(eds\.\),*Advances in Neural Information Processing Systems*, volume 35, pp\. 30016–30030\. Curran Associates, Inc\., 2022\.doi:10\.52202/068431\-2176\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/file/c1e2faff6f588870935f114ebe04a3e5\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2022/file/c1e2faff6f588870935f114ebe04a3e5-Paper-Conference.pdf)\.
- Huang et al\. \(2026\)Zeyi Huang, Xuehai He, LiLiang Ren, Yiping Wang, Baolin Peng, Hao Cheng, Shuohang Wang, Pengcheng He, Jianfeng Gao, Yong Jae Lee, and Yelong Shen\.Latent recurrent Transformer: Architecture exploration, training strategies, and scaling behavior, 2026\.URL[https://arxiv\.org/abs/2605\.26797](https://arxiv.org/abs/2605.26797)\.
- Joshi et al\. \(2017\)Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer\.TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension\.In Regina Barzilay and Min\-Yen Kan \(eds\.\),*Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 1601–1611, Vancouver, Canada, July 2017\. Association for Computational Linguistics\.doi:10\.18653/v1/P17\-1147\.URL[https://aclanthology\.org/P17\-1147/](https://aclanthology.org/P17-1147/)\.
- Kim & Rush \(2016\)Yoon Kim and Alexander M\. Rush\.Sequence\-level knowledge distillation\.In Jian Su, Kevin Duh, and Xavier Carreras \(eds\.\),*Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing*, pp\. 1317–1327, Austin, Texas, November 2016\. Association for Computational Linguistics\.doi:10\.18653/v1/D16\-1139\.URL[https://aclanthology\.org/D16\-1139/](https://aclanthology.org/D16-1139/)\.
- Kumar & Isola \(2026\)Akarsh Kumar and Phillip Isola\.Pretraining recurrent networks without recurrence, 2026\.URL[https://arxiv\.org/abs/2606\.06479](https://arxiv.org/abs/2606.06479)\.
- Kwiatkowski et al\. \(2019\)Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming\-Wei Chang, Andrew M\. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov\.Natural Questions: A benchmark for question answering research\.*Transactions of the Association for Computational Linguistics*, 7:452–466, 2019\.doi:10\.1162/tacl\_a\_00276\.URL[https://aclanthology\.org/Q19\-1026/](https://aclanthology.org/Q19-1026/)\.
- Li et al\. \(2024\)Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma\.Chain of thought empowers Transformers to solve inherently serial problems\.In B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(eds\.\),*International Conference on Learning Representations*, volume 2024, pp\. 11911–11943, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/3309b4112c9f04a993f2bbdd0274bba1\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/3309b4112c9f04a993f2bbdd0274bba1-Paper-Conference.pdf)\.
- Lightman et al\. \(2024\)Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(eds\.\),*International Conference on Learning Representations*, volume 2024, pp\. 39578–39601, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/aca97732e30bcf1303bc22ac3924fd16\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/aca97732e30bcf1303bc22ac3924fd16-Paper-Conference.pdf)\.
- Lim et al\. \(2024\)Yi Heng Lim, Qi Zhu, Joshua Selfridge, and Muhammad Firmansyah\.Parallelizing non\-linear sequential models over the sequence length\.In B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(eds\.\),*International Conference on Learning Representations*, volume 2024, pp\. 55334–55360, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/f3bfbd65743e60c685a3845bd61ce15f\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/f3bfbd65743e60c685a3845bd61ce15f-Paper-Conference.pdf)\.
- Liu et al\. \(2023\)Bingbin Liu, Jordan T\. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang\.Transformers learn shortcuts to automata\.In*The Eleventh International Conference on Learning Representations*, 2023\.URL[https://openreview\.net/forum?id=De4FYqjFueZ](https://openreview.net/forum?id=De4FYqjFueZ)\.
- Merrill & Sabharwal \(2025\)Will Merrill and Ashish Sabharwal\.A little depth goes a long way: The expressive power of log\-depth Transformers\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 95315–95339\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-3186\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/88dd7aa6979e352fda7c4952ca8eac59\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/88dd7aa6979e352fda7c4952ca8eac59-Paper-Conference.pdf)\.
- Merrill & Sabharwal \(2024\)William Merrill and Ashish Sabharwal\.The expressive power of Transformers with chain of thought\.In B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(eds\.\),*International Conference on Learning Representations*, volume 2024, pp\. 7690–7706, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/1f59721c106ea80f613299039112f651\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/1f59721c106ea80f613299039112f651-Paper-Conference.pdf)\.
- Merrill et al\. \(2024\)William Merrill, Jackson Petty, and Ashish Sabharwal\.The illusion of state in state\-space models\.In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp \(eds\.\),*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 35492–35506\. PMLR, 21–27 Jul 2024\.URL[https://proceedings\.mlr\.press/v235/merrill24a\.html](https://proceedings.mlr.press/v235/merrill24a.html)\.
- Mihaylov et al\. \(2018\)Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal\.Can a suit of armor conduct electricity? A new dataset for open book question answering\.In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii \(eds\.\),*Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing*, pp\. 2381–2391, Brussels, Belgium, October\-November 2018\. Association for Computational Linguistics\.doi:10\.18653/v1/D18\-1260\.URL[https://aclanthology\.org/D18\-1260/](https://aclanthology.org/D18-1260/)\.
- Mihaylova & Martins \(2019\)Tsvetomila Mihaylova and André F\. T\. Martins\.Scheduled sampling for Transformers\.In Fernando Alva\-Manchego, Eunsol Choi, and Daniel Khashabi \(eds\.\),*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop*, pp\. 351–356, Florence, Italy, July 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/P19\-2049\.URL[https://aclanthology\.org/P19\-2049/](https://aclanthology.org/P19-2049/)\.
- Morris et al\. \(2024\)John X\. Morris, Wenting Zhao, Justin Chiu, Vitaly Shmatikov, and Alexander Rush\.Language model inversion\.In B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(eds\.\),*International Conference on Learning Representations*, volume 2024, pp\. 35863–35883, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/999fcab97007ebef0cda9949550b4a9e\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/999fcab97007ebef0cda9949550b4a9e-Paper-Conference.pdf)\.
- Penedo et al\. \(2024\)Guilherme Penedo, Hynek Kydlíček, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf\.The FineWeb datasets: Decanting the web for the finest text data at scale\.In A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(eds\.\),*Advances in Neural Information Processing Systems*, volume 37, pp\. 30811–30849\. Curran Associates, Inc\., 2024\.doi:10\.52202/079017\-0970\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/370df50ccfdf8bde18f8f9c2d9151bda\-Paper\-Datasets\_and\_Benchmarks\_Track\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/370df50ccfdf8bde18f8f9c2d9151bda-Paper-Datasets_and_Benchmarks_Track.pdf)\.
- Qiu et al\. \(2025\)Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, Jingren Zhou, and Junyang Lin\.Gated attention for large language models: Non\-linearity, sparsity, and attention\-sink\-free\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 100092–100118\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-3345\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/904e89bb4e632e75fb47f093b620b257\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/904e89bb4e632e75fb47f093b620b257-Paper-Conference.pdf)\.
- Raffel et al\. \(2020\)Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J\. Liu\.Exploring the limits of transfer learning with a unified text\-to\-text transformer\.*Journal of Machine Learning Research*, 21\(140\):1–67, 2020\.URL[http://jmlr\.org/papers/v21/20\-074\.html](http://jmlr.org/papers/v21/20-074.html)\.
- Rajpurkar et al\. \(2016\)Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang\.SQuAD: 100,000\+ questions for machine comprehension of text\.In Jian Su, Kevin Duh, and Xavier Carreras \(eds\.\),*Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing*, pp\. 2383–2392, Austin, Texas, November 2016\. Association for Computational Linguistics\.doi:10\.18653/v1/D16\-1264\.URL[https://aclanthology\.org/D16\-1264/](https://aclanthology.org/D16-1264/)\.
- Ranzato et al\. \(2016\)Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba\.Sequence level training with recurrent neural networks\.In*International Conference on Learning Representations*, 2016\.URL[https://arxiv\.org/abs/1511\.06732](https://arxiv.org/abs/1511.06732)\.
- Reddy et al\. \(2019\)Siva Reddy, Danqi Chen, and Christopher D\. Manning\.CoQA: A conversational question answering challenge\.*Transactions of the Association for Computational Linguistics*, 7:249–266, 2019\.doi:10\.1162/tacl\_a\_00266\.URL[https://aclanthology\.org/Q19\-1016/](https://aclanthology.org/Q19-1016/)\.
- Rizvi\-Martel et al\. \(2026\)Michael Rizvi\-Martel, Guillaume Rabusseau, and Marius Mosbach\.The illusion of superposition? A principled analysis of latent thinking in language models\.In*Third Conference on Language Modeling*, 2026\.URL[https://arxiv\.org/abs/2604\.06374](https://arxiv.org/abs/2604.06374)\.
- Sakaguchi et al\. \(2021\)Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi\.WinoGrande: An adversarial Winograd schema challenge at scale\.*Communications of the ACM*, 64\(9\):99–106, August 2021\.ISSN 1557\-7317\.doi:10\.1145/3474381\.
- Sap et al\. \(2019\)Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi\.Social IQa: Commonsense reasoning about social interactions\.In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan \(eds\.\),*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\)*, pp\. 4463–4473, Hong Kong, China, November 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/D19\-1454\.URL[https://aclanthology\.org/D19\-1454/](https://aclanthology.org/D19-1454/)\.
- Saunshi et al\. \(2025\)Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J\. Reddi\.Reasoning with latent thoughts: On the power of looped Transformers\.In Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(eds\.\),*International Conference on Learning Representations*, volume 2025, pp\. 14855–14881, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/file/2676109d49d1eb26d6bc584a8f556305\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2025/file/2676109d49d1eb26d6bc584a8f556305-Paper-Conference.pdf)\.
- Shazeer \(2020\)Noam Shazeer\.GLU variants improve transformer, 2020\.URL[https://arxiv\.org/abs/2002\.05202](https://arxiv.org/abs/2002.05202)\.
- Shen et al\. \(2025\)Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He\.CODI: Compressing chain\-of\-thought into continuous space via self\-distillation\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng \(eds\.\),*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 677–693, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.doi:10\.18653/v1/2025\.emnlp\-main\.36\.URL[https://aclanthology\.org/2025\.emnlp\-main\.36/](https://aclanthology.org/2025.emnlp-main.36/)\.
- Snell et al\. \(2025\)Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar\.Scaling LLM test\-time compute optimally can be more effective than scaling parameters for reasoning\.In Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(eds\.\),*International Conference on Learning Representations*, volume 2025, pp\. 10131–10165, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/file/1b623663fd9b874366f3ce019fdfdd44\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2025/file/1b623663fd9b874366f3ce019fdfdd44-Paper-Conference.pdf)\.
- Sriram et al\. \(2018\)Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates\.Cold fusion: Training Seq2Seq models together with language models\.In*Interspeech 2018*, pp\. 387–391, 2018\.doi:10\.21437/Interspeech\.2018\-1392\.
- Suzgun et al\. \(2023\)Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei\.Challenging BIG\-bench tasks and whether chain\-of\-thought can solve them\.In Anna Rogers, Jordan Boyd\-Graber, and Naoaki Okazaki \(eds\.\),*Findings of the Association for Computational Linguistics: ACL 2023*, pp\. 13003–13051, Toronto, Canada, July 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.findings\-acl\.824\.URL[https://aclanthology\.org/2023\.findings\-acl\.824/](https://aclanthology.org/2023.findings-acl.824/)\.
- Tack et al\. \(2026\)Jihoon Tack, Jack Lanchantin, Jane Dwivedi\-Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason E Weston, and Xian Li\.LLM pretraining with continuous concepts\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 54214–54231, 2026\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/59056767478c7df64e6250eadfeb0a04\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/59056767478c7df64e6250eadfeb0a04-Paper-Conference.pdf)\.
- Talmor et al\. \(2019\)Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant\.CommonsenseQA: A question answering challenge targeting commonsense knowledge\.In Jill Burstein, Christy Doran, and Thamar Solorio \(eds\.\),*Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\)*, pp\. 4149–4158, Minneapolis, Minnesota, June 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/N19\-1421\.URL[https://aclanthology\.org/N19\-1421/](https://aclanthology.org/N19-1421/)\.
- Tan et al\. \(2025\)Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Ruihua Song, and Jian Luan\.Think silently, think fast: Dynamic latent compression of LLM reasoning chains\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 4646–4668\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-0164\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/0706261aedab63814a2b73c32564b4c4\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/0706261aedab63814a2b73c32564b4c4-Paper-Conference.pdf)\.
- Tang & Lu \(2026\)Mohan Tang and Sidi Lu\.Turbo Connection: Reasoning as information flow from higher to lower layers\.In*Forty\-third International Conference on Machine Learning*, 2026\.URL[https://openreview\.net/forum?id=pLIJx2cFib](https://openreview.net/forum?id=pLIJx2cFib)\.
- Team Olmo et al\. \(2025\)Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng, Shane Arora, Shashank Gupta, Taira Anderson, Teng Xiao, Tyler Murray, Tyler Romero, Victoria Graf, Akari Asai, Akshita Bhagia, Alexander Wettig, Alisa Liu, Aman Rangapur, Chloe Anastasiades, Costa Huang, Dustin Schwenk, Harsh Trivedi, Ian Magnusson, Jaron Lochner, Jiacheng Liu, Lester James V\. Miranda, Maarten Sap, Malia Morgan, Michael Schmitz, Michal Guerquin, Michael Wilson, Regan Huff, Ronan Le Bras, Rui Xin, Rulin Shao, Sam Skjonsberg, Shannon Zejiang Shen, Shuyue Stella Li, Tucker Wilde, Valentina Pyatkin, Will Merrill, Yapei Chang, Yuling Gu, Zhiyuan Zeng, Ashish Sabharwal, Luke Zettlemoyer, Pang Wei Koh, Ali Farhadi, Noah A\. Smith, and Hannaneh Hajishirzi\.Olmo 3, 2025\.URL[https://arxiv\.org/abs/2512\.13961](https://arxiv.org/abs/2512.13961)\.
- Vaswani et al\. \(2017\)Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin\.Attention is all you need\.In I\. Guyon, U\. Von Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(eds\.\),*Advances in Neural Information Processing Systems*, volume 30\. Curran Associates, Inc\., 2017\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa\-Paper\.pdf](https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf)\.
- Walsh et al\. \(2025\)Evan Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Allyson Ettinger, Michal Guerquin, David Heineman, Hamish Ivison, Pang Wei Koh, Jiacheng Liu, Saumya Malik, William Merrill, Lester James Validad Miranda, Jacob Morrison, Tyler Murray, Crystal Nam, Jake Poznanski, Valentina Pyatkin, Aman Rangapur, Michael Schmitz, Sam Skjonsberg, David Wadden, Christopher Wilhelm, Michael Wilson, Luke Zettlemoyer, Ali Farhadi, Noah A\. Smith, and Hannaneh Hajishirzi\.2 OLMo 2 furious \(COLM’s version\)\.In*Second Conference on Language Modeling*, 2025\.URL[https://openreview\.net/forum?id=2ezugTT9kU](https://openreview.net/forum?id=2ezugTT9kU)\.
- Wang et al\. \(2026\)Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, and John Langford\.Full\-bandwidth transformer, 2026\.URL[https://arxiv\.org/abs/2608\.08888](https://arxiv.org/abs/2608.08888)\.
- Wang et al\. \(2023\)Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H\. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou\.Self\-consistency improves chain of thought reasoning in language models\.In*The Eleventh International Conference on Learning Representations*, 2023\.URL[https://openreview\.net/forum?id=1PL1NIMMrw](https://openreview.net/forum?id=1PL1NIMMrw)\.
- Wei et al\. \(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou\.Chain\-of\-thought prompting elicits reasoning in large language models\.In S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(eds\.\),*Advances in Neural Information Processing Systems*, volume 35, pp\. 24824–24837\. Curran Associates, Inc\., 2022\.doi:10\.52202/068431\-1800\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf)\.
- Wei et al\. \(2026\)Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Jiaqi Wang, Xipeng Qiu, and Dahua Lin\.SIM\-CoT: Supervised implicit chain\-of\-thought\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 56721–56742, 2026\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/5d087955ee13fe9a7402eedec879b9c3\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/5d087955ee13fe9a7402eedec879b9c3-Paper-Conference.pdf)\.
- Wen et al\. \(2025\)Kaiyue Wen, Zhiyuan Li, Jason Wang, David Hall, Percy Liang, and Tengyu Ma\.Understanding warmup\-stable\-decay learning rates: A river valley loss landscape view\.In Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(eds\.\),*International Conference on Learning Representations*, volume 2025, pp\. 42840–42885, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/file/6a1fe80a9e2dcda0b3e5fd0fd87eb097\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2025/file/6a1fe80a9e2dcda0b3e5fd0fd87eb097-Paper-Conference.pdf)\.
- Wolf et al\. \(2020\)Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush\.Transformers: State\-of\-the\-art natural language processing\.In Qun Liu and David Schlangen \(eds\.\),*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pp\. 38–45, Online, October 2020\. Association for Computational Linguistics\.doi:10\.18653/v1/2020\.emnlp\-demos\.6\.URL[https://aclanthology\.org/2020\.emnlp\-demos\.6/](https://aclanthology.org/2020.emnlp-demos.6/)\.
- Wu & Tu \(2024\)Haoyi Wu and Kewei Tu\.Layer\-condensed KV cache for efficient inference of large language models\.In Lun\-Wei Ku, Andre Martins, and Vivek Srikumar \(eds\.\),*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 11175–11188, Bangkok, Thailand, August 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.acl\-long\.602\.URL[https://aclanthology\.org/2024\.acl\-long\.602/](https://aclanthology.org/2024.acl-long.602/)\.
- Wu et al\. \(2025\)Haoyi Wu, Zhihao Teng, and Kewei Tu\.Parallel continuous chain\-of\-thought with Jacobi iteration\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng \(eds\.\),*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 914–926, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.doi:10\.18653/v1/2025\.emnlp\-main\.47\.URL[https://aclanthology\.org/2025\.emnlp\-main\.47/](https://aclanthology.org/2025.emnlp-main.47/)\.
- Wu et al\. \(2026\)Junhong Wu, Jinliang Lu, Zixuan Ren, Gangqiang Hu, Zhi Wu, Dai Dai, and Hua Wu\.LLMs are single\-threaded reasoners: Demystifying the working mechanism of Soft Thinking\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 4533–4548, 2026\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/0895de150660b6c2aef9b0f098cb1dcd\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/0895de150660b6c2aef9b0f098cb1dcd-Paper-Conference.pdf)\.
- Xiao et al\. \(2023\)Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han\.SmoothQuant: Accurate and efficient post\-training quantization for large language models\.In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett \(eds\.\),*Proceedings of the 40th International Conference on Machine Learning*, volume 202 of*Proceedings of Machine Learning Research*, pp\. 38087–38099\. PMLR, 23–29 Jul 2023\.URL[https://proceedings\.mlr\.press/v202/xiao23c\.html](https://proceedings.mlr.press/v202/xiao23c.html)\.
- Yue et al\. \(2025\)Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang, Zhen Qin, Jinsung Yoon, Lanyu Shang, Jiawei Han, and Dong Wang\.Hybrid latent reasoning via reinforcement learning\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 5501–5530\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-0195\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/087f7678af2abbe0c37ecc30bd7d1adf\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/087f7678af2abbe0c37ecc30bd7d1adf-Paper-Conference.pdf)\.
- Zellers et al\. \(2019\)Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi\.HellaSwag: Can a machine really finish your sentence?In Anna Korhonen, David Traum, and Lluís Màrquez \(eds\.\),*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics*, pp\. 4791–4800, Florence, Italy, July 2019\. Association for Computational Linguistics\.doi:10\.18653/v1/P19\-1472\.URL[https://aclanthology\.org/P19\-1472/](https://aclanthology.org/P19-1472/)\.
- Zeng et al\. \(2026a\)Boyi Zeng, He Li, Shixiang Song, Yixuan Wang, Zitong Wang, Ziwei He, Xinbing Wang, and Zhouhan Lin\.PonderLM\-2: Pretraining LLM with latent thoughts in continuous space\.In*Forty\-third International Conference on Machine Learning*, 2026a\.URL[https://openreview\.net/forum?id=yVFxjNzCQm](https://openreview.net/forum?id=yVFxjNzCQm)\.
- Zeng et al\. \(2026b\)Boyi Zeng, Shixiang Song, Siyuan Huang, Yixuan Wang, He Li, Ziwei He, Xinbing Wang, Zhiyu Li, and Zhouhan Lin\.PonderLM: Pretraining language models to ponder in continuous space\.In C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(eds\.\),*International Conference on Learning Representations*, volume 2026, pp\. 91874–91893, 2026b\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2026/file/94572a66c6b0710c5263f15d8ef6b463\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2026/file/94572a66c6b0710c5263f15d8ef6b463-Paper-Conference.pdf)\.
- Zhang & Ren \(2026\)Wenbo Zhang and Xiang Ren\.WhiteMatter: All\-to\-all cross\-layer connections via KV mixing, 2026\.URL[https://arxiv\.org/abs/2608\.18486](https://arxiv.org/abs/2608.18486)\.
- Zhang et al\. \(2026\)Xiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin, Yu Cheng, and Junchi Yan\.NITP: Next implicit token prediction for LLM pre\-training\.In*Forty\-third International Conference on Machine Learning*, 2026\.URL[https://openreview\.net/forum?id=oPQmCKS1tV](https://openreview.net/forum?id=oPQmCKS1tV)\.
- Zhang et al\. \(2025\)Zhen Zhang, Xuehai He, Weixiang Yan, Ao Shen, Chenyang Zhao, and Xin Wang\.Soft Thinking: Unlocking the reasoning potential of LLMs in continuous concept space\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 168990–169012\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-5629\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/f7396d1c54d51416958d63e285377103\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/f7396d1c54d51416958d63e285377103-Paper-Conference.pdf)\.
- Zhong et al\. \(2024\)Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan\.AGIEval: A human\-centric benchmark for evaluating foundation models\.In Kevin Duh, Helena Gomez, and Steven Bethard \(eds\.\),*Findings of the Association for Computational Linguistics: NAACL 2024*, pp\. 2299–2314, Mexico City, Mexico, June 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.findings\-naacl\.149\.URL[https://aclanthology\.org/2024\.findings\-naacl\.149/](https://aclanthology.org/2024.findings-naacl.149/)\.
- Zhu et al\. \(2025\)Hanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao, Stuart J Russell, and Yuandong Tian\.Reasoning by superposition: A theoretical perspective on chain of continuous thought\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 79931–79963\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-2670\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/72c363c2a573ca2128bd176d3317696b\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/72c363c2a573ca2128bd176d3317696b-Paper-Conference.pdf)\.
- Zhuang et al\. \(2025\)Yufan Zhuang, Liyuan Liu, Chandan Singh, Jingbo Shang, and Jianfeng Gao\.Mixture of Inputs: Text generation beyond discrete token sampling\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(eds\.\),*Advances in Neural Information Processing Systems*, volume 38, Main Conference, pp\. 26363–26387\. Curran Associates, Inc\., 2025\.doi:10\.52202/085713\-0889\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/25c61012c845908bff2ea6223cdc6192\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/25c61012c845908bff2ea6223cdc6192-Paper-Conference.pdf)\.

## Appendix AAdditional Method Details

### A\.1Prefix State Dropout

For each training sequence, with probabilitypp, we draw a prefix lengthm∼𝒰​\{0,…,n\}m\\sim\\mathcal\{U\}\\\{0,\\dots,n\\\}and feed the fusion layer \(Eq\.[5](https://arxiv.org/html/2609.38149#S3.E5)\)

𝐞~i=\{𝐛1<i≤mE⊤​𝐬i⋆i\>m\\tilde\{\\mathbf\{e\}\}\_\{i\}=\\begin\{cases\}\\mathbf\{b\}&1<i\\leq m\\\\ E^\{\\top\}\\mathbf\{s\}^\{\\star\}\_\{i\}&i\>m\\end\{cases\}\(8\)in place of the projected stateE⊤​𝐬i⋆E^\{\\top\}\\mathbf\{s\}^\{\\star\}\_\{i\}\. The first position always keeps the learned initial\-state vector, since no previous output exists there at inference either\. Becausemmranges up to the full length, the model also learns to process an entire sequence without states, the mode evaluated asLIFTw/o states \(§[5\.2](https://arxiv.org/html/2609.38149#S5.SS2)\)\. The same augmentation is applied to the loss\-carrying pass of the adaptation phase \(§[A\.3](https://arxiv.org/html/2609.38149#A1.SS3)\)\.

### A\.2The State\-Alignment Loss

The alignment term𝒟k\\mathcal\{D\}\_\{k\}of Eq\.[6](https://arxiv.org/html/2609.38149#S3.E6)compares the teacher’s and the student’s distributions at the state temperatureτ\\tauon thekktokens of𝐬i\+1⋆\\mathbf\{s\}^\{\\star\}\_\{i\+1\}, aggregating the remaining probability mass into a single tail outcome\. Formally, letSSbe thekktokens of𝐬i\+1⋆\\mathbf\{s\}^\{\\star\}\_\{i\+1\}, let𝐩τ⋆=softmax⁡\(𝐨i⋆/τ\)\\mathbf\{p\}^\{\\star\}\_\{\\tau\}=\\operatorname\{softmax\}\(\\mathbf\{o\}^\{\\star\}\_\{i\}/\\tau\)and𝐩τ=softmax⁡\(𝐨i/τ\)\\mathbf\{p\}\_\{\\tau\}=\\operatorname\{softmax\}\(\\mathbf\{o\}\_\{i\}/\\tau\)be the teacher’s and the student’s full\-vocabulary distributions at temperatureτ\\tau, and letPPandQQbe the mass each puts onSS\. Since the state renormalizes the teacher’s probabilities onSS\(Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)\),pτ⋆​\(v\)=P​si\+1⋆​\(v\)p^\{\\star\}\_\{\\tau\}\(v\)=P\\,s^\{\\star\}\_\{i\+1\}\(v\)forv∈Sv\\in S\. Then

𝒟k\(𝐨i⋆∥𝐨i\)=τ2\[∑v∈SPsi\+1⋆\(v\)logP​si\+1⋆​\(v\)pτ​\(v\)\+\(1−P\)log1−P1−Q\],\\mathcal\{D\}\_\{k\}\\big\(\\mathbf\{o\}^\{\\star\}\_\{i\}\\,\\\|\\,\\mathbf\{o\}\_\{i\}\\big\)=\\tau^\{2\}\\Big\[\\sum\_\{v\\in S\}P\\,s^\{\\star\}\_\{i\+1\}\(v\)\\log\\frac\{P\\,s^\{\\star\}\_\{i\+1\}\(v\)\}\{p\_\{\\tau\}\(v\)\}\\;\+\\;\(1\-P\)\\log\\frac\{1\-P\}\{1\-Q\}\\Big\],\(9\)whereτ2\\tau^\{2\}is the usual distillation scaling\([Hinton et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib33)\)\. Keeping only the terms onSSwould push the student to lower the probability of every token outsideSS, a constant pressure to sharpen its distribution\. With the tail, the loss is minimized exactly when the student matches the teacher onSSand puts mass1−P1\-Poutside it; and as a KL between coarsened distributions, it is a lower bound on the full\-vocabulary KL that does not decrease askkgrows\.

### A\.3Adaptation Phase

In the last10%10\\%of training, each microbatch draws its number of passesR∈\{2,3\}R\\in\\\{2,3\\\}with equal probability and runs:

1. 1\.Pass 1\(no gradient\)\. Every position after the first is fed the bias𝐛\\mathbf\{b\}, which carries no information about the context; the first position keeps its learned initial state \(§[3\.1](https://arxiv.org/html/2609.38149#S3.SS1)\)\. The logits at every position give a state \(Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)\)\.
2. 2\.Pass 2\(no gradient, only whenR=3R=3\)\. Every position after the first is fed the state that pass 1 predicted at the preceding position, over the whole sequence, yielding refined states\.
3. 3\.Final pass\(with gradient\)\. Each sequence independently draws, with probabilitypp, a prefix lengthm∼𝒰​\{0,…,n\}m\\sim\\mathcal\{U\}\\\{0,\\dots,n\\\}\(Eq\.[8](https://arxiv.org/html/2609.38149#A1.E8)\)\. Positions1<i≤m1<i\\leq mare fed𝐛\\mathbf\{b\}, and all later positions the states of the previous pass\. This is the only pass that carries a loss: the language\-modeling term of Eq\.[6](https://arxiv.org/html/2609.38149#S3.E6)over every position, with the alignment term off \(λ=0\\lambda=0\)\.

The fed\-back states are detached, so no gradient flows between passes, and the no\-gradient passes add1\.51\.5forward passes per step on average \(Table[18](https://arxiv.org/html/2609.38149#A6.T18)\)\. The teacher is not needed in this phase\.

## Appendix BExperimental Setup

### B\.1S5S\_\{5\}State Tracking: Setup

##### Models and training\.

All models of §[4](https://arxiv.org/html/2609.38149#S4)use the OLMo 2 block \(§[B\.2](https://arxiv.org/html/2609.38149#A2.SS2)\), slightly adapted to their tiny size and small vocabulary: two layers, width256256, four attention heads of dimension6464, a SwiGLU MLP of hidden size512512, RoPE withθ=10,000\\theta=10\{,\}000, and untied input and output embeddings over a synthetic vocabulary of1,0471\{,\}047tokens \(padded to1,1521\{,\}152\) that includes the120120elements ofS5S\_\{5\}\. They are trained in fp32 with AdamW \(β=\(0\.9,0\.95\)\\beta=\(0\.9,0\.95\),ϵ=10−8\\epsilon=10^\{\-8\}, weight decay0\.10\.1on weight matrices\), a peak learning rate of10−310^\{\-3\}with200200warm\-up steps and a cosine decay to0\.01×0\.01\\timesthe peak, gradient clipping at1\.01\.0, and batches of512512sequences\. Each step draws a fresh batch of a single length, chosen uniformly from the training lengthsN∈\{1,2,3,4,6,8,12,16\}N\\in\\\{1,2,3,4,6,8,12,16\\\}\. The initial state is uniform overS5S\_\{5\}, and eachaia\_\{i\}is drawn uniformly from1212fixed generators, which include a transposition and a55\-cycle and thus generateS5S\_\{5\}\. The input is⟨BOS,s0,a1,…,aN,=⟩\\langle\\texttt\{BOS\},s\_\{0\},a\_\{1\},\\dots,a\_\{N\},\\texttt\{=\}\\rangle, and the answer is predicted at the last position\.

##### LIFTonS5S\_\{5\}\.

These models differ from the recipe of §[3](https://arxiv.org/html/2609.38149#S3)in several respects\. The fusion layer is a single linear map,𝐡^i0=Wf​\[𝐡i0;We⊤​𝐬i\]\\hat\{\\mathbf\{h\}\}^\{0\}\_\{i\}=W\_\{f\}\\,\[\\mathbf\{h\}^\{0\}\_\{i\}\\,;\\,W\_\{e\}^\{\\top\}\\mathbf\{s\}\_\{i\}\]withWf∈ℝd×2​dW\_\{f\}\\in\\mathbb\{R\}^\{d\\times 2d\}, in place of Eq\.[5](https://arxiv.org/html/2609.38149#S3.E5), and the state keeps the topk=256k=256tokens atτ=1\\tau=1\.LIFT\-RNN is trained on the answer loss alone, backpropagating through all positions\. The teacher\-fed models add the forward KL to the teacher’s full output distribution at every position before the answer, with weight11, and use no prefix state dropout\. The teacher’s states are built from its logits, computed in one parallel pass for the Transformer×\\times8 teacher and sequentially, as at inference, for theLIFT\-RNN teacher\. In the last10%10\\%of steps, the teacher\-fed models are instead fed their own states, computed sequentially with no gradient through the states, and the losses are unchanged\.

##### Evaluation\.

We report the accuracy of the argmax prediction at the answer position\.LIFTmodels are evaluated sequentially, as at inference, each position fed the state predicted at the preceding one; the Transformer answers in one parallel pass\. Each length has one evaluation set, shared by all models and seeds and excluded from training:2,0002\{,\}000sequences, except atN=1N\{=\}1\(200200\) andN=2N\{=\}2\(1,7281\{,\}728\), which have only1,4401\{,\}440and17,28017\{,\}280distinct sequences\. All models are trained with three seeds \(00,11and22\), and each teacher\-fedLIFTmodel uses the teacher of its own seed, giving three distinct student–teacher pairs\.

### B\.2Implementation Details

##### Training budgets and teachers\.

The three\-scale experiments use training budgets of5×5\\timesthe Chinchilla token budget\([Hoffmann et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib34)\): 13\.4B, 34\.9B, and 107\.4B tokens for the 135M, 350M, and 1B models, respectively \(rounded in Table[3](https://arxiv.org/html/2609.38149#A2.T3)\)\. At each scale, we use a pretrained vanilla baseline of the same size as one teacher and a larger pretrained model as the other\. The larger teachers for the 135M and 350M models are pretrained vanilla baselines; for the 1B model, we use pretrained OLMo 2\-7B\([Walsh et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib71)\)\. Table[16](https://arxiv.org/html/2609.38149#A4.T16)lists the teacher sizes and checkpoint training budgets\. The compute\-matched budgets are given in §[F\.1](https://arxiv.org/html/2609.38149#A6.SS1.SSS0.Px1)\.

##### Model configurations\.

All three models use the OLMo 2 block\([Walsh et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib71)\): RoPE \(θ=500,000\\theta=500\{,\}000\), RMSNorm with QK\-norm, a SwiGLU MLP\([Shazeer, 2020](https://arxiv.org/html/2609.38149#bib.bib60)\)with ratio88, a head dimension of128128, and the dolma2 tokenizer \(vocabulary100,278100\{,\}278, padded to100,352100\{,\}352\)\. The 1B model is OLMo 2’s 1B configuration unchanged; the 135M and 350M models scale down its width and depth, and tie the input and output embeddings\. Table[3](https://arxiv.org/html/2609.38149#A2.T3)lists the configurations; model names refer to non\-embedding parameter counts\.

Table 3:Model and training configurations\. Parameter counts exclude the RMSNorm vectors; the 1B counts both embedding matrices, the smaller models one tied matrix\.
##### Optimization\.

We use AdamW withβ=\(0\.9,0\.95\)\\beta=\(0\.9,0\.95\),ϵ=10−8\\epsilon=10^\{\-8\}, weight decay0\.10\.1\(not applied to the embeddings\), gradient clipping at a norm of1\.01\.0, and a z\-loss with weight10−510^\{\-5\}, all as in OLMo 2\. Training runs in bf16 mixed precision with the cross\-entropy loss computed in fp32\. The state input, the teacher’s top\-kkdistribution, is quantized to 8 bits\. Within each model size,LIFTand its vanilla control share the configuration, the random seed and therefore the exact data order\.

##### LIFThyperparameters\.

The three\-scale experiments use the sameLIFTsettings at every scale\. The state keeps the topk=1024k=1024tokens at temperatureτ=1\.5\\tau=1\.5\(Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)\); the alignment term uses the samekkandτ\\tau, with weightλ=1\.5\\lambda=1\.5\(Eq\.[6](https://arxiv.org/html/2609.38149#S3.E6)\)\. Prefix state dropout is applied to each training sequence with probabilityp=0\.5p=0\.5\(§[A\.1](https://arxiv.org/html/2609.38149#A1.SS1)\)\. The adaptation phase covers the last10%10\\%of training, coinciding with the learning\-rate anneal, and draws one or two no\-gradient passes per step with equal probability \(§[A\.3](https://arxiv.org/html/2609.38149#A1.SS3)\)\. The comparison of §[5\.3](https://arxiv.org/html/2609.38149#S5.SS3)uses its own settings \(§[B\.3](https://arxiv.org/html/2609.38149#A2.SS3)\)\.

##### Pretraining data\.

The training corpus is sampled from the OLMo 2 stage\-1 pretraining mix: DCLM\-baseline, StarCoder, peS2o, Proof\-Pile\-2 and OLMo\-Mix, tokenized with the dolma2 tokenizer\. Our stratified versions sample whole files per source domain so that the source mixture is preserved: approximately93\.8%93\.8\\%DCLM,2\.6%2\.6\\%StarCoder,2\.0%2\.0\\%peS2o,1\.3%1\.3\\%Proof\-Pile\-2 and0\.4%0\.4\\%OLMo\-Mix\. We use a 400B\-token selection for the 350M and 1B runs and a 40B\-token selection for the 135M runs; every run consumes its budget from a single pass over a random order of its selection\.

##### Learning\-rate schedule\.

All models follow the WSD protocol of[Wen et al\. \(2025\)](https://arxiv.org/html/2609.38149#bib.bib76)\. The learning rate warms up linearly from zero over the steps listed in Table[3](https://arxiv.org/html/2609.38149#A2.T3), is then held constant at the peak value for the first90%90\\%of the training budget \(4\.5×4\.5\\timesChinchilla\), and is annealed over the last10%10\\%\(0\.5×0\.5\\timesChinchilla\), branching from the trunk checkpoint at90%90\\%\. The anneal uses the inverse\-proportional decay shape of[Wen et al\. \(2025\)](https://arxiv.org/html/2609.38149#bib.bib76), in which the reciprocal of the learning rate rises linearly, and ends at0\.1×0\.1\\timesthe peak value\. All results reported in §[5](https://arxiv.org/html/2609.38149#S5)are after an identical annealing over the last 10% of training to an LR of0\.1×0\.1\\timesthe starting value\.

### B\.3Feedback\-Transformer Comparison: Setup

##### Setup\.

All models of §[5\.3](https://arxiv.org/html/2609.38149#S5.SS3)are trained with the T2MLR authors’ training code on the SmolLM2\-135M backbone\([Allal et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib1)\): 30 layers, width 576, 9 attention heads with 3 key\-value heads, tied embeddings, 135M parameters including embeddings, and the SmolLM2 tokenizer of 49,152 tokens\. The data is the 10B\-token FineWeb\-Edu sample \(sample\-10BT;[Penedo et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib50)\) packed into sequences of 2,048 tokens; one epoch is 19,073 steps of 256 sequences\. The training configuration is identical for all models: AdamW withβ=\(0\.9,0\.98\)\\beta=\(0\.9,0\.98\)and weight decay0\.010\.01, a peak learning rate of5×10−45\\times 10^\{\-4\}with a warm\-up of 954 steps \(5%5\\%\), a global batch of 256 sequences \(0\.50\.5M tokens per step\), gradient clipping at1\.01\.0, bf16 precision, and a WSD schedule with a linear decay to0\.1×0\.1\\timesthe peak over the last10%10\\%of steps; forLIFTthat last10%10\\%is the adaptation phase of §[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\. T2MLR runs in the authors’ configuration, T2MLR\(13,18\) with 16 forward and 4 backward refinement passes, and is evaluated with the exact sequential recurrence, as in their paper\.LIFTusesk=1024k=1024and is trained without the KL objective, so that the measured gain comes from state propagation and not from the distillation signal; its teacher is the last checkpoint of the vanilla baseline model \(the same model, trained on 23B tokens\)\. The compute\-matched budget is T2MLR’s 19,073 steps at2\.28×2\.28\\timesthe FLOPs of a vanilla step per token; at1\.36×1\.36\\timesper token in the trunk and1\.87×1\.87\\timesin the adaptation phase,LIFTreaches it after 16\.2B tokens and the vanilla model after 23\.0B\. The checkpoint the authors published was trained for 19,294 steps against our 19,073 \(1\.2%1\.2\\%more tokens\), with a cosine decay to0\.001×0\.001\\timesthe peak\. Parameter counts are 138\.2M forLIFT, 136\.2M for T2MLR, 135\.2M for the Multi\-pass Transformer and 134\.5M for the vanilla model\. Perplexity is measured in fp32 on 127K held\-out FineWeb\-Edu tokens; the arithmetic and LM\-eval protocols are those of §[B\.4](https://arxiv.org/html/2609.38149#A2.SS4)\.

##### Multi\-pass Transformer\.

We take from the paper\([Wang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib72)\)its feedback path, a gated linear unit that fuses the previous top\-layer state \(the value\) with the current token embedding \(the gate\); its pass schedule,75%75\\%of the batches with a single pass,22%22\\%with two and3%3\\%with three, which is the schedule the authors report for their larger runs; its prefix mixin, which leaves a random prefix of each sequence on plain embeddings in the extra passes, so that training matches the prompt\-then\-generation structure of inference; and its stabilizers, noise on the carried state and normalization of the fused input \(the third, tying the input and output embeddings, the shared backbone already does\)\. We could not take their optimizer, since all models here share one and changing it for a single model would confound the comparison, nor their depth scaling of the residual branches, which alters the backbone shared by all models here and which we replace by carrying the normalized top\-layer state\. We use the shared 135M, 10B\-token setup instead of their 1B model and 400B\-token dataset\. As with T2MLR, we report the model in its recurrent mode\.

### B\.4Evaluation Benchmarks

##### Downstream suite\.

The twenty suites are run with OLMES\([Gu et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib29)\)and its task configurations \(Table[4](https://arxiv.org/html/2609.38149#A2.T4)\);LIFTscores every prompt and answer with its own states\. The window is each model’s trained context \(1,024 / 2,048 / 4,096 tokens at 135M / 350M / 1B; 2,048 for the SmolLM2 models\), and at 135M, generation is capped at 256 tokens, the same for every 135M model, so its generative numbers compare within the size, not across sizes\. Cloze scoring appends each answer option to the question and scores it by its length\-normalized log\-likelihood; the prediction is the highest\-scoring option\. Bits per byte is−log2⁡P⁡\(gold answer\)\-\\log\_\{2\}P\(\\text\{gold answer\}\)divided by the answer’s UTF\-8 bytes\. The main table keeps a suite if, at 1B, at least one model is five points above chance \(multiple choice\) or above 5% \(generative\); bits\-per\-byte suites always enter\. This leaves out the letter\-format multiple choice \(MMLU, AGI\-Eval\), where every model is at chance, and generative math and code \(GSM8K, MATH, HumanEval, MBPP\), where every model is under 3%\. BoolQ is left out as well: a yes/no cloze whose majority class is 62%, which several models fall below and which moves by ten points across training seeds\. Every excluded suite is reported in Tables[7](https://arxiv.org/html/2609.38149#A3.T7)–[15](https://arxiv.org/html/2609.38149#A3.T15)\.

Table 4:The downstream suite\. Items: scored instances \(per task for the core tasks\); Shots: in\-context examples; Group: the Table[1](https://arxiv.org/html/2609.38149#S5.T1)column, or the reason a suite is reported only in the appendix\. Task sources are cited below the table\.SuiteTestsItemsShotsScoringGroupCore 8: ARC\-Easy, ARC\-Challenge, CommonsenseQA, HellaSwag, OpenBookQA, PIQA, SocialIQA, WinoGrandeScience and commonsense questions500–1,2675cloze acc\.MCMMLU57 academic subjects14,0425cloze acc\.MCBasic skills \(OLMES\)Arithmetic, string operations, pattern continuation, coding, logical reasoning, common knowledge5,9675cloze acc\.MCGenerative QA: CoQA, SQuAD, Jeopardy, Natural Questions, DROPConversational, extractive and closed\-book QA7,983 / 1,000 / 2,116 / 1,000 / 1,0005 \(CoQA 0\)F1Gen\.TriviaQAClosed\-book trivia7,9935F1Gen\.BBH27 reasoning tasks with chain of thought6,5113exact matchGen\.ARC, MMLU, basic skills; MATH; HumanEval; MBPPThe gold answer’s likelihood3,548 / 14,042 / 5,967 / 5,000 / 164 / 5005 / 5 / 5 / 4 / 3 / 3bits per byteRCBoolQYes/no questions on a passage1,0005cloze acc\.coin flipMMLU, letter formatAs above, scored by the option letter14,0425letter acc\.at chanceAGI\-Eval EnglishExam questions, letter format2,6461letter acc\.at chanceGSM8KGrade\-school math, chain of thought1,3195 and 8exact matchunder 3%MATH\-500, Minerva MATHCompetition math500, 5,0004exact matchunder 3%HumanEval, MBPPProgram synthesis, 20 samples atT=0\.8T\{=\}0\.8164, 5003pass@1under 3%
Sources: ARC\([Clark et al\., 2018](https://arxiv.org/html/2609.38149#bib.bib15)\), CommonsenseQA\([Talmor et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib66)\), HellaSwag\([Zellers et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib83)\), OpenBookQA\([Mihaylov et al\., 2018](https://arxiv.org/html/2609.38149#bib.bib47)\), PIQA\([Bisk et al\., 2020](https://arxiv.org/html/2609.38149#bib.bib7)\), SocialIQA\([Sap et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib58)\), WinoGrande\([Sakaguchi et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib57)\), MMLU\([Hendrycks et al\., 2021a](https://arxiv.org/html/2609.38149#bib.bib31)\), CoQA\([Reddy et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib55)\), SQuAD\([Rajpurkar et al\., 2016](https://arxiv.org/html/2609.38149#bib.bib53)\), Natural Questions\([Kwiatkowski et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib39)\), DROP\([Dua et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib20)\), TriviaQA\([Joshi et al\., 2017](https://arxiv.org/html/2609.38149#bib.bib36)\), BBH\([Suzgun et al\., 2023](https://arxiv.org/html/2609.38149#bib.bib64)\), MATH\([Hendrycks et al\., 2021b](https://arxiv.org/html/2609.38149#bib.bib32)\), HumanEval\([Chen et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib13)\), MBPP\([Austin et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib4)\), BoolQ\([Clark et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib14)\), AGI\-Eval\([Zhong et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib89)\), GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib16)\)\.

##### Arithmetic\.

Expressions are generated by drawingnnintegers uniformly from00to99andn−1n\-1operators uniformly from\{\+,−\}\\\{\+,\-\\\}, and are evaluated left to right; duplicates are discarded, which is why then=2n=2bin holds all 200 possible expressions\. The answer is an integer, possibly negative\. Every prompt consists of the same three worked examples \(3\-shot\), followed by the query, one expression per line, and the scored continuation is the answer with its leading space:

> 4 \- 6 = \-2 0 \+ 5 \+ 7 \- 1 \- 3 = 8 3 \- 3 \- 1 = \-1 2 \+ 0 \- 6 \+ 7 \+ 5 =

with answer8\. Bits per byte is−log2⁡P⁡\(answer\)\-\\log\_\{2\}P\(\\text\{answer\}\)summed over the set and divided by the total number of UTF\-8 bytes of the scored answers\. The exact\-match accuracy of the teacher\-forced argmax is recorded alongside it \(Table[12](https://arxiv.org/html/2609.38149#A3.T12)\): it is substantial only at two operands and a few percent from three operands on, for every model, which is why Table[1](https://arxiv.org/html/2609.38149#S5.T1)reports bits per byte\.

## Appendix CAdditional Results

### C\.1Transformer Training Budget onS5S\_\{5\}

§[4](https://arxiv.org/html/2609.38149#S4)attributes the two\-layer Transformer’s ceiling onS5S\_\{5\}to depth\. EveryS5S\_\{5\}model hasdmodel=256d\_\{\\text\{model\}\}=256and trains with batch size512512for5,0005\{,\}000steps unless stated otherwise\. A competing explanation is that the model is simply undertrained, so we trained the same baseline at four budgets, all else held fixed: the token\-matched5,0005\{,\}000steps,8×8\\times\(40,00040\{,\}000\),20×20\\times\(100,000100\{,\}000\) and40×40\\times\(200,000200\{,\}000\)\. Figure[4](https://arxiv.org/html/2609.38149#A3.F4)reports all four, with three seeds each, under the protocol of §[4](https://arxiv.org/html/2609.38149#S4)\.

The model gains nothing after20×20\\times, where the20×20\\timesand40×40\\timescurves coincide, and it remains at≈10%\\approx 10\\%atN=12N\{=\}12and at chance atN=16N\{=\}16and every longer length, in every seed of every budget\.

Figure 4:The same two\-layer Transformer onS5S\_\{5\}at four training budgets, darker lines for longer training\. Bands show standard error over three seeds\. Training beyond 100K steps \(20×20\\times\) brings no further gain\.
### C\.2Perplexity by Source and Evaluation Uncertainty

Table[5](https://arxiv.org/html/2609.38149#A3.T5)reports perplexity on every source of the OLMo 2 perplexity evaluation set: the full C4 validation subset and10%10\\%of each other source\.LIFTis ahead of the token\-matched, compute\-matched and distillation baselines on all eleven sources at every size\. Table[6](https://arxiv.org/html/2609.38149#A3.T6)gives, for the paper’s training seed, a paired bootstrap interval over the evaluation sequences for every perplexity gap of Table[1](https://arxiv.org/html/2609.38149#S5.T1); it measures how much a gap depends on the held\-out text at a fixed training run, not seed variance\.

Table 5:Perplexity by source on the OLMo 2 perplexity evaluation set: the full C4 validation slice \(971 sequences of 1,024 tokens\) and10%10\\%of each other source \(at least 50 sequences\)\. Columns as in Table[1](https://arxiv.org/html/2609.38149#S5.T1); lower is better, bold is the best in a row\.SourceLIFTw/o statesTr\. tokenTr\. computeDistillationSeqs135MC430\.6532\.4432\.4331\.0431\.77971Common Crawl31\.5533\.4433\.5532\.0132\.6350Books33\.5835\.9335\.9033\.8534\.8650Wikipedia19\.6420\.7020\.7819\.8720\.3950peS2o17\.8518\.9319\.0018\.1218\.5651S2ORC38\.8840\.9141\.5039\.6940\.3497Reddit33\.8935\.8335\.7434\.3734\.8850Stack \(code\)7\.107\.457\.597\.207\.3850Pile18\.2019\.2919\.3018\.5218\.8766ICE30\.5232\.9833\.9231\.2332\.2889WikiText\-10324\.6926\.7226\.8825\.2326\.3050350MC423\.8224\.9625\.1524\.3224\.52971Common Crawl24\.7025\.9026\.0225\.0625\.3250Books24\.2925\.6925\.9424\.7625\.1250Wikipedia15\.8516\.5216\.7516\.1116\.2950peS2o12\.4812\.9913\.0412\.6712\.7751S2ORC27\.4428\.6829\.3228\.4628\.1597Reddit26\.7928\.1728\.3027\.4927\.6550Stack \(code\)5\.946\.136\.256\.046\.0450Pile14\.0714\.7714\.9714\.4914\.6466ICE24\.5126\.0026\.7325\.6126\.4489WikiText\-10318\.4319\.5619\.5918\.8619\.07501BC418\.8119\.6219\.7919\.1719\.40971Common Crawl19\.6220\.3720\.5919\.9620\.1650Books17\.8218\.7119\.0518\.2718\.4450Wikipedia12\.3112\.7312\.8812\.4612\.6850peS2o10\.0710\.3910\.5310\.2610\.3351S2ORC22\.3823\.0523\.6923\.1223\.6397Reddit21\.7422\.6322\.7822\.2222\.3450Stack \(code\)4\.854\.975\.034\.894\.9350Pile11\.2811\.7111\.8711\.5011\.6166ICE19\.3420\.2320\.4319\.9119\.9689WikiText\-10313\.7914\.3614\.4513\.9614\.1750Table 6:Evaluation\-set uncertainty of the perplexity gaps of Table[1](https://arxiv.org/html/2609.38149#S5.T1), for the paper’s training seed: the gap ofLIFTto each Transformer baseline with a paired bootstrap95%95\\%interval over the evaluation sequences \(both models scored on the same text; 5,000 resamples\)\. This measures how much a gap depends on the held\-out text, at a fixed training run; it is not seed variance, which Table[10](https://arxiv.org/html/2609.38149#A3.T10)reports for 135M\. Negative is in favor ofLIFT\.
### C\.3Downstream Results per Suite

Tables[7](https://arxiv.org/html/2609.38149#A3.T7)–[9](https://arxiv.org/html/2609.38149#A3.T9)give every OLMES suite for the rows of Tables[1](https://arxiv.org/html/2609.38149#S5.T1)and[16](https://arxiv.org/html/2609.38149#A4.T16), including the suites left out of the main table\.

Table 7:OLMES results at 135M for the rows of Tables[1](https://arxiv.org/html/2609.38149#S5.T1)and[16](https://arxiv.org/html/2609.38149#A4.T16): the suites of the three Table[1](https://arxiv.org/html/2609.38149#S5.T1)groups, then the suites left out of it \(at chance or floor for every model, and BoolQ\)\. Accuracy % for the multiple\-choice, generative and excluded rows; bits per byte \(lower is better\) for the BPB rows; suites as in Table[4](https://arxiv.org/html/2609.38149#A2.T4)\. Gray: the 2×\\times\-token teacher, outside the bold rule\. Bold: best in the row\.Table 8:OLMES results at 350M for the rows of Tables[1](https://arxiv.org/html/2609.38149#S5.T1)and[16](https://arxiv.org/html/2609.38149#A4.T16): the suites of the three Table[1](https://arxiv.org/html/2609.38149#S5.T1)groups, then the suites left out of it \(at chance or floor for every model, and BoolQ\)\. Accuracy % for the multiple\-choice, generative and excluded rows; bits per byte \(lower is better\) for the BPB rows; suites as in Table[4](https://arxiv.org/html/2609.38149#A2.T4)\. Gray: the 2×\\times\-token teacher, outside the bold rule\. Bold: best in the row\.Table 9:OLMES results at 1B for the rows of Tables[1](https://arxiv.org/html/2609.38149#S5.T1)and[16](https://arxiv.org/html/2609.38149#A4.T16): the suites of the three Table[1](https://arxiv.org/html/2609.38149#S5.T1)groups, then the suites left out of it \(at chance or floor for every model, and BoolQ\)\. Accuracy % for the multiple\-choice, generative and excluded rows; bits per byte \(lower is better\) for the BPB rows; suites as in Table[4](https://arxiv.org/html/2609.38149#A2.T4)\. Gray: the 2×\\times\-token teacher, outside the bold rule\. Bold: best in the row\.
### C\.4Seed Variance at 135M

Table[10](https://arxiv.org/html/2609.38149#A3.T10)gives the 135M rows of Table[1](https://arxiv.org/html/2609.38149#S5.T1)over three training seeds\.

Table 10:The 135M rows of Table[1](https://arxiv.org/html/2609.38149#S5.T1)over three training seeds \(6198 / 6199 / 6200\): mean±\\pmsd of each column\. We regard a difference between two models as beyond seed noise when their mean±\\pmsd ranges do not overlap\. Pattern: the basic\-skills pattern\-continuation task \(534 questions\)\.
### C\.5Arithmetic by Number of Operands

Tables[11](https://arxiv.org/html/2609.38149#A3.T11)and[12](https://arxiv.org/html/2609.38149#A3.T12)break the arithmetic results down by the number of operands: bits per byte of the answer, and exact\-match accuracy of the teacher\-forced argmax\.

Table 11:Arithmetic by number of operands: bits per byte of the answer \(lower is better\), for the rows of Tables[1](https://arxiv.org/html/2609.38149#S5.T1)and[16](https://arxiv.org/html/2609.38149#A4.T16)plus the vanilla Transformer trained on 2×\\timesthe tokens \(Figure[3\(a\)](https://arxiv.org/html/2609.38149#S5.F3.sf1)\); 1,000 expressions per operand count \(200 at two operands\), three\-shot\. Avg\. is over all expressions \(total bits over total answer bytes\), the number reported in Table[1](https://arxiv.org/html/2609.38149#S5.T1)\. Bold: best in the column within a size\.Number of operandsnnModelTokens2345678910Avg\.135MLIFT13\.4B135M@29\.6B2\.032\.132\.122\.272\.272\.352\.362\.422\.422\.29LIFT13\.4B350M@38\.4B2\.232\.212\.172\.262\.252\.322\.352\.382\.412\.29LIFTw/o states13\.4B135M@29\.6B2\.242\.322\.262\.372\.372\.472\.492\.552\.562\.42LIFTw/o states13\.4B350M@38\.4B2\.332\.282\.242\.442\.452\.582\.592\.632\.642\.48Transformer, token\-match13\.4B–2\.022\.282\.292\.422\.452\.552\.572\.582\.592\.46Transformer, compute\-match19\.3B–2\.032\.212\.322\.482\.522\.612\.622\.622\.652\.50Transformer, 2×\\timestokens29\.6B–2\.202\.282\.292\.372\.372\.442\.452\.472\.502\.39Distillation13\.4B135M@29\.6B2\.032\.182\.222\.322\.332\.432\.452\.482\.492\.36350MLIFT34\.9B350M@77\.8B1\.561\.972\.002\.112\.132\.182\.192\.242\.242\.12LIFT34\.9B1B@111\.8B1\.501\.941\.992\.112\.162\.182\.222\.272\.282\.14LIFTw/o states34\.9B350M@77\.8B1\.642\.002\.072\.202\.212\.282\.302\.352\.362\.21LIFTw/o states34\.9B1B@111\.8B1\.552\.102\.162\.302\.332\.362\.392\.432\.442\.30Transformer, token\-match34\.9B–2\.042\.272\.242\.382\.452\.542\.592\.602\.622\.46Transformer, compute\-match49\.9B–1\.842\.062\.112\.252\.332\.422\.492\.492\.512\.33Transformer, 2×\\timestokens77\.8B–1\.642\.002\.062\.152\.242\.302\.382\.402\.442\.24Distillation34\.9B350M@77\.8B1\.832\.072\.082\.202\.202\.272\.312\.362\.412\.231BLIFT107\.4B1B@222B0\.741\.861\.931\.961\.992\.042\.072\.112\.131\.99LIFT107\.4B1B@4001B0\.711\.841\.911\.982\.022\.072\.122\.182\.212\.02LIFT107\.4B7B@701B0\.791\.791\.831\.881\.921\.982\.022\.052\.061\.92LIFTw/o states107\.4B1B@222B0\.931\.981\.982\.012\.032\.082\.122\.182\.192\.05LIFTw/o states107\.4B1B@4001B0\.771\.911\.941\.981\.992\.032\.072\.112\.121\.99LIFTw/o states107\.4B7B@701B1\.041\.871\.881\.931\.962\.012\.052\.082\.101\.97Transformer, token\-match107\.4B–1\.041\.861\.972\.012\.032\.072\.122\.152\.162\.03Transformer, compute\-match152\.4B–0\.951\.942\.012\.022\.052\.112\.172\.202\.212\.07Transformer, 2×\\timestokens222B–0\.781\.851\.941\.962\.032\.062\.102\.152\.152\.01Distillation107\.4B1B@222B0\.771\.731\.921\.931\.952\.002\.052\.072\.091\.95Table 12:Arithmetic by number of operands: exact\-match accuracy % of the teacher\-forced argmax, for the expressions and rows of Table[11](https://arxiv.org/html/2609.38149#A3.T11)\. Avg\. is over all expressions\. Bold: best in the column within a size\.Number of operandsnnModelTokens2345678910Avg\.135MLIFT13\.4B135M@29\.6B7\.05\.03\.42\.43\.22\.52\.23\.01\.93\.0LIFT13\.4B350M@38\.4B5\.54\.44\.43\.34\.02\.52\.22\.53\.63\.4LIFTw/o states13\.4B135M@29\.6B6\.04\.53\.83\.13\.12\.71\.73\.12\.83\.2LIFTw/o states13\.4B350M@38\.4B5\.04\.73\.53\.33\.92\.02\.32\.82\.33\.1Transformer, token\-match13\.4B–7\.04\.54\.13\.02\.63\.02\.42\.02\.93\.2Transformer, compute\-match19\.3B–5\.54\.63\.32\.22\.92\.73\.03\.02\.03\.0Transformer, 2×\\timestokens29\.6B–4\.54\.63\.42\.32\.83\.02\.82\.62\.83\.1Distillation13\.4B135M@29\.6B7\.05\.23\.92\.92\.62\.22\.52\.12\.43\.1350MLIFT34\.9B350M@77\.8B21\.56\.44\.24\.22\.42\.73\.32\.52\.54\.0LIFT34\.9B1B@111\.8B22\.55\.65\.34\.12\.52\.23\.03\.11\.93\.9LIFTw/o states34\.9B350M@77\.8B15\.06\.44\.33\.63\.11\.92\.23\.12\.63\.7LIFTw/o states34\.9B1B@111\.8B23\.05\.44\.93\.82\.02\.52\.52\.92\.33\.8Transformer, token\-match34\.9B–9\.55\.14\.33\.42\.83\.02\.22\.23\.13\.4Transformer, compute\-match49\.9B–8\.54\.94\.62\.91\.93\.32\.72\.52\.23\.3Transformer, 2×\\timestokens77\.8B–13\.06\.13\.33\.22\.92\.72\.93\.42\.73\.6Distillation34\.9B350M@77\.8B11\.05\.14\.13\.83\.42\.12\.72\.92\.83\.51BLIFT107\.4B1B@222B58\.58\.75\.35\.24\.93\.73\.03\.02\.65\.9LIFT107\.4B1B@4001B63\.012\.84\.44\.54\.34\.53\.42\.62\.06\.2LIFT107\.4B7B@701B63\.511\.45\.74\.53\.72\.93\.73\.62\.56\.2LIFTw/o states107\.4B1B@222B46\.58\.94\.64\.54\.12\.82\.61\.92\.35\.0LIFTw/o states107\.4B1B@4001B59\.510\.84\.43\.94\.13\.63\.03\.13\.55\.9LIFTw/o states107\.4B7B@701B47\.010\.85\.34\.32\.92\.73\.44\.53\.05\.6Transformer, token\-match107\.4B–46\.58\.14\.64\.14\.12\.13\.34\.02\.25\.1Transformer, compute\-match152\.4B–48\.59\.03\.83\.53\.72\.03\.03\.42\.24\.9Transformer, 2×\\timestokens222B–56\.013\.34\.15\.43\.83\.02\.43\.32\.36\.0Distillation107\.4B1B@222B56\.514\.16\.25\.14\.82\.52\.43\.42\.96\.4
### C\.6Pattern Continuation

Table[13](https://arxiv.org/html/2609.38149#A3.T13)gives the pattern\-continuation results discussed in §[5\.2](https://arxiv.org/html/2609.38149#S5.SS2)\.

Table 13:Pattern continuation, the basic\-skills task of OLMES \(534 five\-shot questions, e\.g\.,2 4 6 8→\\to10\): cloze accuracy % and bits per byte of the answer \(BPB, lower is better\), for the rows of Table[1](https://arxiv.org/html/2609.38149#S5.T1)at the paper’s training seed, and the Transformer trained on 2×\\timesthe tokens \(gray, outside the bold rule\)\. Three\-seed statistics at 135M are in Table[10](https://arxiv.org/html/2609.38149#A3.T10)\. Bold: best in column\.
### C\.7Feedback\-Transformer Comparison: Results

##### Seeds\.

Every model of Table[2](https://arxiv.org/html/2609.38149#S5.T2)except T2MLR \(HF\) is trained with three seeds \(42, 43, 44\)\. Table[14](https://arxiv.org/html/2609.38149#A3.T14)reports the mean and standard deviation of each column; Table[2](https://arxiv.org/html/2609.38149#S5.T2)shows the means\.

Table 14:The models of Table[2](https://arxiv.org/html/2609.38149#S5.T2)over three training seeds \(42 / 43 / 44\): mean±\\pmsd of each column\. T2MLR \(HF\) is not listed: it is the authors’ single released checkpoint, so it has no seed statistics\. We regard a difference between two models as beyond seed noise when their mean±\\pmsd ranges do not overlap\. T2MLR \(our run\) sets the compute budget, so its row is the same under both budgets\. Bold: the best mean within each budget\. Per\-suite results of seed 42 are in Table[15](https://arxiv.org/html/2609.38149#A3.T15)\.Table[15](https://arxiv.org/html/2609.38149#A3.T15)gives every OLMES suite for §[5\.3](https://arxiv.org/html/2609.38149#S5.SS3)\(seed 42; three\-seed statistics in Table[14](https://arxiv.org/html/2609.38149#A3.T14)\)\.

Table 15:OLMES results for the feedback\-Transformer comparison \(Table[2](https://arxiv.org/html/2609.38149#S5.T2)\), every suite of the scan; the T2MLR columns coincide across the two budgets\. Rows and units as in Table[7](https://arxiv.org/html/2609.38149#A3.T7)\. Bold: best within a budget\.

## Appendix DAblations

Beyond a Transformer,LIFTspends about one forward pass per training token, the teacher’s\. The compute\-matched Transformer spends the same compute on over40%40\\%more tokens instead, andLIFTis ahead of it at every scale \(Table[1](https://arxiv.org/html/2609.38149#S5.T1)\)\. Here we ask where that lead comes from, changing one thing at a time: how the teacher’s forward pass is used \(§[D\.1](https://arxiv.org/html/2609.38149#A4.SS1)\), how wide the fed\-back state is \(§[D\.2](https://arxiv.org/html/2609.38149#A4.SS2)\), whether the gain is built in the teacher\-forced trunk or in the adaptation phase \(§[D\.3](https://arxiv.org/html/2609.38149#A4.SS3)\), and how strong the teacher must be \(§[D\.4](https://arxiv.org/html/2609.38149#A4.SS4)\)\. Every variant keeps the architecture, recipe, learning\-rate schedule and data order of §[B\.2](https://arxiv.org/html/2609.38149#A2.SS2)\. Figures[5](https://arxiv.org/html/2609.38149#A4.F5)and[6](https://arxiv.org/html/2609.38149#A4.F6)plot each variant’s improvement over the compute\-matched Transformer, so the zero line stands for spending the extra pass on more data\. We lead with perplexity on the full C4 validation subset, which separates the variants well beyond seed noise, and report the downstream columns of Table[1](https://arxiv.org/html/2609.38149#S5.T1)after it\.

### D\.1Teacher Forcing: Inputs, Targets, or No Teacher

We compare four ways to spend the extra forward pass:

1. 1\.LIFT: the teacher’s states are both the inputs and the alignment target \(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\)\.
2. 2\.LIFT, teacher as input only:LIFTwithλ=0\\lambda=0throughout, so the teacher’s states are inputs and never targets\.
3. 3\.LIFT, no teacher \(1 Jacobi iteration\): no teacher at any point; the extra pass is the model’s own\. Each training step first runs a pass without gradient in which every position after the first receives the bias𝐛\\mathbf\{b\}, and the states it predicts \(Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)\) are fed to the pass that carries the loss, with prefix state dropout as usual andλ=0\\lambda=0\. This is the adaptation phase \(§[A\.3](https://arxiv.org/html/2609.38149#A1.SS3)\) run from the first step, one Jacobi iteration without gradient between the passes\. The last10%10\\%of training isLIFT’s adaptation phase, unchanged\.
4. 4\.Distillation: the teacher’s states are the target only, and nothing is fed back \(the baseline of Table[1](https://arxiv.org/html/2609.38149#S5.T1)\)\.

Every variant is trained at 135M with the three seeds of Table[10](https://arxiv.org/html/2609.38149#A3.T10)\.

Figure 5:Teacher ablation at 135M: spendingLIFT’s extra forward pass\. Each bar is a model’s perplexity improvement on the full C4 validation subset over the compute\-matched Transformer, which spends that compute on more tokens \(the zero line; right is better\): the mean over three training seeds of the per\-seed difference, with whiskers of±1\\pm 1sd\.##### Only the teacher’s states as inputs and targets beat the compute\-matched Transformer\.

At 135M,LIFTimproves perplexity over the compute\-matched Transformer by0\.47±0\.080\.47\\pm 0\.08\(Figure[5](https://arxiv.org/html/2609.38149#A4.F5)\)\. Trained on its own states, the same architecture ties it \(\+0\.01±0\.05\+0\.01\\pm 0\.05, within0\.070\.07at every seed\); the teacher thus accounts for98%98\\%ofLIFT’s gain over the compute\-matched Transformer\. Self\-generated states do improve on a Transformer trained on the same tokens, by1\.401\.40perplexity, but spending the extra pass on data instead improves on it by about as much,1\.391\.39\. With the teacher’s states as inputs only \(λ=0\\lambda=0\),LIFTfalls0\.19±0\.020\.19\\pm 0\.02behind the compute\-matched Transformer \(31\.23±0\.0331\.23\\pm 0\.03against31\.04±0\.0031\.04\\pm 0\.00\), and behind the no\-teacher variant too\. Without the alignment term the model makes less use of the teacher’s states, reaching30\.72±0\.0430\.72\\pm 0\.04when fed them against30\.25±0\.0530\.25\\pm 0\.05forLIFT, and its own states drift further from them: feeding it its own states instead costs0\.510\.51perplexity, against0\.320\.32forLIFT, and run without states it is a full point behind the token\-matched Transformer \(33\.4733\.47against32\.4432\.44\)\. The alignment term is thus what carries the teacher\-forced channel over to inference \(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\)\.

##### As a target only, the teacher is worth less than the data\.

Distillation falls behind the compute\-matched Transformer by0\.77±0\.040\.77\\pm 0\.04\. Its forward pass through the teacher is better spent on more tokens; used as both inputs and targets, the same teacher’s states giveLIFTits lead\.

##### Tuned distillation\.

The Distillation baseline of Table[1](https://arxiv.org/html/2609.38149#S5.T1)reusesLIFT’s support and temperature: the top1,0241\{,\}024tokens with an aggregated tail,τ=1\.5\\tau=1\.5,λ=1\.5\\lambda=1\.5, next\-token loss kept\. Three further variants at 135M, with the same teacher, seed, data order and schedule, do not close the gap\. The same baseline atτ=1\\tau=1ties it \(31\.80 against 31\.77\)\. The recipe of[Busbridge et al\. \(2025\)](https://arxiv.org/html/2609.38149#bib.bib10), a pure forward KL over the full vocabulary atτ=1\\tau=1without the next\-token loss, is behind even the token\-matched Transformer \(32\.66 against 32\.43,\[0\.19,0\.29\]\[0\.19,0\.29\]\), and annealing that student with the next\-token loss recovers only part of the gap \(32\.29\)\.LIFTleads the best of the four by1\.111\.11\(\[1\.06,1\.17\]\[1\.06,1\.17\]\)\.

##### Downstream columns\.

At 135M and matched compute, most downstream differences are within seed noise, in the sense of Table[10](https://arxiv.org/html/2609.38149#A3.T10)\(overlapping mean±\\pmsd ranges\)\. Two are beyond it:LIFT’s generative score \(15\.0±0\.615\.0\\pm 0\.6against13\.6±0\.613\.6\\pm 0\.6\) and Distillation’s reference\-completion loss \(1\.096±0\.0061\.096\\pm 0\.006against1\.085±0\.0051\.085\\pm 0\.005bits per byte\)\. Without a teacher no column moves beyond noise; on arithmetic the no\-teacher variant hasLIFT’s mean \(2\.282\.28bits per byte for both\), and both are within the compute\-matched Transformer’s spread \(2\.39±0\.092\.39\\pm 0\.09\)\.

##### Backpropagating through own states\.

Training on the model’s own states exactly means backpropagating through the recurrence, asLIFT\-RNN does onS5S\_\{5\}\(§[4](https://arxiv.org/html/2609.38149#S4)\)\. This makes training sequential over positions, removing the parallelism that pretraining relies on: each position waits for the state of the one before it, the backward pass traverses every step, and the attention inputs of every step must be stored, which grows quadratically with the sequence length in a direct implementation; activation checkpointing and fused attention kernels, built for parallel training, do not apply as they are\. Feedback models therefore truncate the recurrence to a few parallel passes and backpropagate through them, as T2MLR and the Multi\-pass Transformer do\. At matched compute both fall behind the vanilla Transformer, and both need more memory thanLIFT\(§[5\.3](https://arxiv.org/html/2609.38149#S5.SS3), Table[19](https://arxiv.org/html/2609.38149#A6.T19)\)\. The no\-teacher variant above is the cheapest point of this family, one extra pass without gradient, and ties the compute\-matched Transformer\. None of the variants that generate their own training states beats a Transformer given the same compute;LIFT, whose training states come from the teacher, does\.

### D\.2Channel Width

We varykk, the number of teacher tokens a state keeps \(Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)\), at 135M, in both the fed\-back state and the support of the alignment term, with everything else as in Table[1](https://arxiv.org/html/2609.38149#S5.T1), including its seed\. Atk=1k=1the state is the teacher’s top token scaled by its probability, a second token channel;k=1024k=1024is our setting\. Figure[6](https://arxiv.org/html/2609.38149#A4.F6)shows each model with its states and, for reference, the same model run without them \(LIFTw/o states in Table[1](https://arxiv.org/html/2609.38149#S5.T1)\)\.

Figure 6:Channel width at 135M: improvement in perplexity over the compute\-matched Transformer \(the zero line; up is better\) forkkfrom11to20482048, with the model’s own states fed back \(LIFT\) and for the same model run without states \(LIFTw/o states\)\. Dashed: the token\-matched Transformer\. Every step inkkis beyond the paired bootstrap95%95\\%interval over evaluation sequences \(Table[6](https://arxiv.org/html/2609.38149#A3.T6)’s protocol\)\.##### The gain grows with the width of the channel\.

Perplexity improves monotonically withkk\. The one\-token channel is1\.161\.16behind the compute\-matched Transformer andk=16k=16is0\.390\.39behind; onlyk≥128k\\geq 128is ahead of it, by0\.140\.14, by0\.380\.38atk=1024k=1024and by0\.520\.52atk=2048k=2048\. Measured against the token\-matched Transformer,k=1k=1keeps13%13\\%ofLIFT’s gain,k=16k=16keeps56%56\\%andk=128k=128keeps86%86\\%\. An extra token of input is thus not whatLIFT’s states contribute; their width is\. Arithmetic follows the same order:2\.502\.50,2\.352\.35,2\.312\.31and2\.292\.29bits per byte atk=1k=1,1616,128128and10241024, the one\-token channel at the level of the Transformers \(2\.462\.46token\-matched,2\.502\.50compute\-matched\)\.

##### The width acts through the states\.

Run without states, the models trained at everykkare within0\.480\.48perplexity of each other and within0\.360\.36of the token\-matched Transformer: widening the channel, which also widens the alignment term’s support, does not train a better Transformer\. What grows withkkis what the states carry at inference,0\.400\.40perplexity atk=1k=1,1\.361\.36atk=16k=16,1\.701\.70atk=128k=128,1\.781\.78atk=1024k=1024and1\.791\.79atk=2048k=2048\.

##### Choice ofkk\.

Returns diminish with width, from0\.770\.77perplexity betweenk=1k=1and1616and0\.530\.53between1616and128128to0\.240\.24between128128and10241024, still well beyond evaluation noise \(paired bootstrap95%95\\%interval\[0\.20,0\.29\]\[0\.20,0\.29\]\)\. Doubling tok=2048k=2048improves perplexity by a further0\.140\.14\(\[0\.10,0\.18\]\[0\.10,0\.18\]; one training seed, against a seed spread of0\.080\.08forLIFTat 135M\), but not through the states: what they carry has stopped growing \(1\.781\.78atk=1024k=1024,1\.791\.79atk=2048k=2048\), and the model run without states improves by about as much \(32\.3132\.31against32\.4432\.44\), within the spread of the without\-states models acrosskk\. The width costs almost nothing:kkenters the compute only through the soft\-token sumE⊤​𝐬iE^\{\\top\}\\mathbf\{s\}\_\{i\},0\.13%0\.13\\%of a training step at 1B fork=1024k=1024\(Table[18](https://arxiv.org/html/2609.38149#A6.T18)\)\. We usek=1024k=1024at every scale\.

### D\.3Trunk vs\. Adaptation Phase

##### The gain is built by the teacher\-forced trunk, not by the adaptation phase\.

Two further runs isolate the adaptation phase of §[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\.*Adaptation only*: we take the token\-matched Transformer at90%90\\%of training, add a freshly initialized fusion layer, and runLIFT’s adaptation phase over the last10%10\\%\. At 135M this reaches32\.5332\.53perplexity, slightly worse than the Transformer itself \(32\.4332\.43\); at 1B,19\.6819\.68against19\.7919\.79,12%12\\%ofLIFT’s gain over the Transformer \(18\.8118\.81\)\.*No adaptation*:LIFTat 1B trained with the teacher until the end keeps73%73\\%of that gain \(19\.0719\.07\) and is still ahead of the compute\-matched Transformer \(19\.1719\.17\); adapting over the last5%5\\%instead of10%10\\%recovers the rest \(18\.8118\.81\)\. The adaptation phase thus fits to the model’s own states a channel that the teacher\-forced trunk has already made useful\. Run over the whole of training, it is the no\-teacher variant above, which only ties the compute\-matched Transformer\.

### D\.4Teacher Strength

Table[16](https://arxiv.org/html/2609.38149#A4.T16)reportsLIFTtrained at each scale with the same\-size teacher of §[5\.1](https://arxiv.org/html/2609.38149#S5.SS1)and with stronger teachers, either larger or trained for much longer\.

Table 16:Teacher choice:LIFTandLIFTw/o states per teacher \(size@training tokens\) against the token\-matched vanilla Transformer\.*Tokens*: the training budget of every model in the row block,5×5\\timesthe Chinchilla budget of its size\.*Teacher PPL*: the teacher’s own perplexity on the same evaluation set\. The results hold on every metric, and barely change, with a larger teacher or one trained much longer\. Other columns as in Table[1](https://arxiv.org/html/2609.38149#S5.T1)\. Bold: best in column within a size\.ModelSizeTokensPPL↓\\downarrowArith\.↓\\downarrowMC %LIFT135M13\.4B135M@29\.6B29\.8030\.652\.2939\.4LIFT135M13\.4B350M@38\.4B24\.9130\.402\.2939\.5LIFTw/o states135M13\.4B135M@29\.6B29\.8032\.442\.4238\.7LIFTw/o states135M13\.4B350M@38\.4B24\.9132\.402\.4839\.0Transformer, token\-match135M13\.4B––32\.432\.4638\.1LIFT350M34\.9B350M@77\.8B23\.4923\.822\.1245\.0LIFT350M34\.9B1B@111\.8B19\.7223\.782\.1445\.3LIFTw/o states350M34\.9B350M@77\.8B23\.4924\.962\.2144\.2LIFTw/o states350M34\.9B1B@111\.8B19\.7225\.012\.3044\.2Transformer, token\-match350M34\.9B––25\.152\.4643\.2LIFT1B107\.4B1B@222B18\.6218\.811\.9952\.5LIFT1B107\.4B1B@4001B17\.3118\.772\.0251\.9LIFT1B107\.4B7B@701B14\.9718\.851\.9251\.7LIFTw/o states1B107\.4B1B@222B18\.6219\.622\.0551\.2LIFTw/o states1B107\.4B1B@4001B17\.3119\.641\.9950\.9LIFTw/o states1B107\.4B7B@701B14\.9719\.781\.9750\.5Transformer, token\-match1B107\.4B––19\.792\.0350\.5

## Appendix EPrefill Modes: Refining the Prompt’s Key\-Value Cache

During prefill nothing is generated: the pass over the prompt only builds the key\-value cache that decoding attends to, and the state entering the first answer token\. The evaluations of §[5](https://arxiv.org/html/2609.38149#S5)build this cache sequentially, feeding each prompt position the state predicted at the preceding one\. Here we ask whether the cache can be built in parallel instead, and at what price\. Decoding is sequential in every mode, as it is in any autoregressive Transformer: each answer token is fed the state predicted at the preceding position \(Eq\.[4](https://arxiv.org/html/2609.38149#S3.E4)\)\. Only the construction of the prompt’s cache differs\.

##### Prefill modes\.

W/o states: one parallel pass over the prompt with the learned bias𝐛\\mathbf\{b\}at every position after the first, as in a standard Transformer\.𝑹Rrefinements:RRfurther parallel passes over the prompt, each fed, at every position after the first, the state that the previous pass predicted at the preceding position; the first decoding step attends to the cache of the last pass\.Sequential: the prompt is processed one position at a time, as the answer is; this isLIFTas evaluated in §[5](https://arxiv.org/html/2609.38149#S5)\. The reference point isLIFTw/o states, which uses no states in prefill or decoding\.

##### Relative gain\.

For each groupggof MC, Gen\. and RC, the*states gain*isGg=xgLIFT−xgw/oG\_\{g\}=x\_\{g\}^\{\\textsc\{LIFT\}\{\}\}\-x\_\{g\}^\{\\text\{w/o\}\}, the difference betweenLIFT, with states through prefill and decoding, andLIFTw/o states\. The relative gain of a mode or modelXXis\(xgX−xgw/o\)/Gg\(x\_\{g\}^\{X\}\-x\_\{g\}^\{\\text\{w/o\}\}\)/G\_\{g\}, and we report its mean over the three groups:0%0\\%isLIFTw/o states and100%100\\%isLIFT\.

Figure 7:Relative gain ofLIFTwith parallel prefill \(blue points\) by the number of refinement passes over the prompt, againstLIFT\(100%100\\%\) andLIFTw/o states \(0%0\\%\)\. The token\- and compute\-matched Transformers are placed on the same scale; below0%0\\%, a model is behindLIFTw/o states\.
##### One refinement pass suffices for the key\-value cache\.

With one refinement pass,LIFTmatches sequential prefill to within0\.20\.2points on MC, Gen\. and RC at every scale, a relative gain of9393–104%104\\%; three and five passes change the scores by at most0\.20\.2points \(Figure[7](https://arxiv.org/html/2609.38149#A5.F7), Table[17](https://arxiv.org/html/2609.38149#A5.T17)\)\. Arithmetic, reported separately in Table[17](https://arxiv.org/html/2609.38149#A5.T17), follows the same pattern: one refinement recovers8989–102%102\\%of its states gain, and three passes essentially all of it\. The Full\-bandwidth Transformer shows the same pattern for its fused prefill passes, where most of the improvement comes from the first one\([Wang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib72)\)\. Refining the prompt’s cache thus costs one extra parallel forward pass rather thannnsequential steps\.

##### Without refinement,LIFTstays ahead of the Transformer baselines\.

Prefilling without states and decoding with them keeps a relative gain of3737–51%51\\%, at the prefill cost of a standard Transformer\. In this modeLIFTis ahead of the token\-matched Transformer at every scale, asLIFTw/o states already is, and it is ahead of the compute\-matched Transformer at 350M and level with it at 1B \(Figure[7](https://arxiv.org/html/2609.38149#A5.F7)\)\. On arithmetic this mode keeps less of the gain,1818–26%26\\%; it is still ahead of the compute\-matched Transformer at every scale and of the token\-matched one at 135M and 350M, and level with the latter at 1B\.

##### Cost\.

Prefill without states has no sequential dependency along the prompt: it is one parallel pass\. WithRRrefinements, positions depend on each other only across consecutive passes, so the prompt is read inR\+1R\+1parallel passes atR\+1R\+1times the FLOPs of one\. Sequential prefill costs the FLOPs of one pass, since the key\-value cache avoids recomputation, but reads a prompt ofnntokens innndependent steps\. A refined cache can thus be obtained either with one extra parallel pass, doubling the prefill FLOPs, or withnnsequential steps at the FLOPs of one pass\. Counting every forward pass, over the prompt or over one answer token, as one step, an answer ofDDtokens takesD\+1D\+1steps without refinement andD\+2D\+2with one: the extra pass adds1/\(D\+1\)1/\(D\+1\)of the steps, a share that shrinks as the answer grows\. At the mean answer lengths of our generative tasks, it adds17%17\\%on TriviaQA \(5 tokens\),13%13\\%on the generative QA suite \(6\.56\.5tokens\) and0\.5%0\.5\\%on chain\-of\-thought BBH \(209 tokens\)\.

Table 17:Prefill modes ofLIFTon arithmetic and the OLMES groups of Table[1](https://arxiv.org/html/2609.38149#S5.T1), with the token\- and compute\-matched Transformers of that table\. Prefill cost: forward passes to read a prompt ofnntokens, and FLOPs relative to one forward pass over the prompt\. Full propagation: every answer token is fed the state predicted at the preceding position; –: a model without states\. The prefill, decoding and cost columns describe how each model runs in deployment, where the answer is generated one token at a time\. They are not the cost of scoring these benchmarks, where a given answer, as on MC and RC, is scored in one parallel pass by the models without states\.

## Appendix FCost Analysis

### F\.1Training and Inference FLOPs

Table[18](https://arxiv.org/html/2609.38149#A6.T18)itemizes every additionLIFTmakes to a vanilla Transformer’s compute, excluding the teacher’s forward pass, for the11B model \(d=2048d=2048,L=16L=16,N=1\.07N=1\.07B non\-embedding parameters,k=1024k=1024,\|𝒱\|=100,352\|\\mathcal\{V\}\|=100\{,\}352\)\. A multiply\-add counts as two FLOPs\. Percentages are relative to the standard6​N6NFLOPs per training token \(2​N2Nforward,4​N4Nbackward\) over non\-embedding parameters; counting the output head and the attention scores at44k context lowers every percentage by about a quarter\. The scaling column usesN≈16​L​d2N\\approx 16Ld^\{2\}, the OLMo 2 block shape at11B \(N≈20​L​d2N\\approx 20Ld^\{2\}at3232B, where the fusion block is0\.85%0\.85\\%\)\.

Table 18:Compute added per token byLIFTat11B, excluding the teacher forward\. The trunk is the first90%90\\%of training; the adaptation phase is the last10%10\\%\(§[3\.2](https://arxiv.org/html/2609.38149#S3.SS2)\)\.ComponentFLOPs / token% of a6​N6NstepScalingFusion block, Eq\. \([5](https://arxiv.org/html/2609.38149#S3.E5)\), forward \+ backward66​d266d^\{2\}2772774\.3%4\.3\\%11​d2/N=11/\(16​L\)11d^\{2\}/N=11/\(16L\), falls with depthSoft\-token sumE⊤​𝐬iE^\{\\top\}\\mathbf\{s\}\_\{i\}, forward \+ backward4​k​d4kd880\.13%0\.13\\%k/\(24​L​d\)k/\(24Ld\)Forward KL on the top\-kksupport \+ tail, forward \+ backward≈5​\|𝒱\|\+20​k\\approx 5\|\\mathcal\{V\}\|\+20k0\.50\.5<0\.01%<0\.01\\%≈5​\|𝒱\|/\(96​L​d2\)\\approx 5\|\\mathcal\{V\}\|/\(96Ld^\{2\}\)Prefix state dropout, learned bias000000–Trunk total, per training token2862864\.4%4\.4\\%Adaptation phase:1\.51\.5extra no\-gradient forwards on average1\.5​\(2​N\+22​d2\+2​k​d\)1\.5\\,\(2N\+22d^\{2\}\+2kd\)3,3663\{,\}366\+52%\+52\\%of those stepsfixed by the scheduleWhole run \(0\.9×0\.9\\timestrunk\+0\.1×\+\\ 0\.1\\timesadaptation\)≈9\.6%\\approx 9\.6\\%Inference, per generated token22​d2\+2​k​d22d^\{2\}\+2kd\+ top\-kk97974\.5%4\.5\\%of a2​N2Nforward11/\(16​L\)11/\(16L\)Three remarks\. First, the KL term needs the student’s full\-vocabulary log\-partition at temperatureτ\\tau, one extra pass over logits that the language\-modeling loss already materializes; no\(B,T,\|𝒱\|\)\(B,T,\|\\mathcal\{V\}\|\)tensor is created beyond them\. Second, the adaptation passes run without gradient and their fed\-back distribution is detached, so they hold no activation memory; a step in that phase costs about1\.5×1\.5\\timesa trunk step\. Third, over a whole run the additions total under10%10\\%of the vanilla training FLOPs, and every term except the schedule\-fixed adaptation shrinks with model size\.

##### Compute\-matched budgets\.

The compute\-matched Transformers receive the training FLOPs of the wholeLIFTrun, including the teacher’s forward pass, the fusion layer and the extra passes of the adaptation phase: 19\.3B, 49\.9B and 152\.4B tokens at 135M, 350M and 1B, i\.e\.,1\.421\.42–1\.44×1\.44\\timesLIFT’s budget\. This accounting also charges a teacher forward pass during the adaptation phase, which our implementation mistakenly ran although its output is not used there \(§[A\.3](https://arxiv.org/html/2609.38149#A1.SS3)\); without it, the matched budgets would be 18\.8B, 48\.7B and 148\.8B tokens\. The compute\-matched baselines thus train on about2\.5%2\.5\\%more tokens than the method requires, which favors them\.

### F\.2Memory Footprint

Table[19](https://arxiv.org/html/2609.38149#A6.T19)reports the peak GPU memory ofLIFT, T2MLR and the Multi\-pass Transformer in the setting of §[5\.3](https://arxiv.org/html/2609.38149#S5.SS3), measured with the same code, in the same container and on the same hardware\. Setup: 2 NVIDIA B200 GPUs \(183 GB each\) with data parallelism, PyTorch 2\.9 \(NVIDIA container 25\.10\), bf16, sequences of 2,048 tokens, a global batch of 256 sequences reached by gradient accumulation, and the production configuration of each method: T2MLR\(13,18\) with 16 forward and 4 backward refinement passes, andLIFTin its trunk configuration, including the teacher forward pass and the top\-kk\(k=1024k=1024\) state construction, with the forward\-KL term enabled\. The latter makes this a conservative memory comparison for the KL\-free model of §[5\.3](https://arxiv.org/html/2609.38149#S5.SS3)\. Peak memory is the maximum memory allocated by PyTorch on one GPU\.

Table 19:Peak memory per GPU ofLIFT, T2MLR and the Multi\-pass Transformer at 135M \(SmolLM2 backbone, 2,048\-token sequences, B200\)\. Memory is the peak allocated by PyTorch\. The Multi\-pass Transformer’s peak is measured on a three\-pass batch, which determines the memory needed to run its full schedule\.LIFT’s memory beyond a vanilla step includes the teacher’s logits for the top\-kkselection and the dense soft\-token matrix, both transient\. T2MLR stores the recurrent cache and the activations of the refinement passes it differentiates through\. The Multi\-pass Transformer has the largest peak memory because its three\-pass batches retain activations from all three passes\.

## Appendix GExtended Related Work

##### Depth and sequential computation\.

Looped and recurrent\-depth Transformers reapply a block of layers to increase the computation spent on each token\([Dehghani et al\., 2019](https://arxiv.org/html/2609.38149#bib.bib19);[Saunshi et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib59);[Geiping et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib25)\); their inference cost grows with the number of iterations, and the depth available to a token is still bounded by the iterations run at that step\. Chain of thought instead adds serial computation through generated tokens\([Wei et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib74)\), which extends the expressivity of fixed\-depth Transformers\([Feng et al\., 2023](https://arxiv.org/html/2609.38149#bib.bib23);[Li et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib40);[Merrill & Sabharwal, 2024](https://arxiv.org/html/2609.38149#bib.bib45)\), and test\-time methods scale it further through sampling, verification and search\([Wang et al\., 2023](https://arxiv.org/html/2609.38149#bib.bib73);[Lightman et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib41);[Snell et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib62)\)\. Every such step, however, passes through the single decoded token\. InLIFT, the computation’s depth grows with the sequence at a fixed per\-token cost \(§[F\.1](https://arxiv.org/html/2609.38149#A6.SS1)\), and the approach is complementary to both\.

##### Teachers as targets and as inputs\.

Pretrained teachers usually serve as targets: knowledge distillation matches their distributions\([Hinton et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib33);[Kim & Rush, 2016](https://arxiv.org/html/2609.38149#bib.bib37)\), also in LM pretraining\([Gemma Team, 2024](https://arxiv.org/html/2609.38149#bib.bib26)\), and richer objectives supervise future tokens or latents\([Gloeckle et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib27);[Zhang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib87)\), leaving inference unchanged\. An LM’s distribution has also been fed to another model as an input, but one that remains at inference\([Sriram et al\., 2018](https://arxiv.org/html/2609.38149#bib.bib63)\)\. InLIFT, the teacher’s distribution is a training input that the model learns to replace with its own, and our distillation baseline isolates the target role \(§[5\.1](https://arxiv.org/html/2609.38149#S5.SS1)\)\. As with born\-again and weak\-to\-strong students\([Furlanello et al\., 2018](https://arxiv.org/html/2609.38149#bib.bib24);[Burns et al\., 2024](https://arxiv.org/html/2609.38149#bib.bib9)\), the student can surpass its teacher \(§[4](https://arxiv.org/html/2609.38149#S4)\); our adaptation phase is scheduled sampling\([Bengio et al\., 2015](https://arxiv.org/html/2609.38149#bib.bib5);[Mihaylova & Martins, 2019](https://arxiv.org/html/2609.38149#bib.bib48)\)on the state channel\.

##### Comparison with the closest methods\.

Table[20](https://arxiv.org/html/2609.38149#A7.T20)lists, for the methods closest toLIFT, what is fed back, where it enters the model, whether the sampled token is kept, how the fed\-back states are obtained during training, and at which stage the channel is learned\.LIFTis the only method whose training\-time states are supplied by an external pretrained LM, which is what allows a single parallel pass during pretraining\.

Table 20:Methods that feed information back across generation steps\.*Training\-time states*: how the fed\-back states are obtained during training;*own*means generated by the model being trained\.†arXiv preprint\.MethodFed\-back signalEnters atTokenTraining\-time statesStageFeedback Transformer\([Fan et al\., 2021](https://arxiv.org/html/2609.38149#bib.bib21)\)All layers’ states of past positionsAttention of every layerkeptown, sequentialpretrainingRMT\([Bulatov et al\., 2022](https://arxiv.org/html/2609.38149#bib.bib8)\)Output memory tokens of the previous segmentInput of the next segmentkeptown, sequential over segmentspretrainingLCKV\([Wu & Tu, 2024](https://arxiv.org/html/2609.38149#bib.bib78)\)Top\-layer keys and valuesAttention of every layerkeptown, iterative parallel passespretrainingTurbo Connection\([Tang & Lu, 2026](https://arxiv.org/html/2609.38149#bib.bib68)\)Higher\-layer states of the previous tokenLower layerskeptown, sequential in groupsfine\-tuningPonderLM\([Zeng et al\., 2026b](https://arxiv.org/html/2609.38149#bib.bib85)\)Top\-kksoft token of its own predictionInput of the same positionkeptown, extra passes per positionpretrainingPonderLM\-2\([Zeng et al\., 2026a](https://arxiv.org/html/2609.38149#bib.bib84)\)Last hidden stateAn added input positionkeptown, Jacobi passespretrainingT2MLR†\([Cai et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib12)\)Middle\-layer hidden stateAn earlier middle layerkeptown, Jacobi passespretrainingFull\-bandwidth Transformer†\([Wang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib72)\)Top\-layer hidden stateInput, gated with the tokenkeptown, one to three parallel passespretrainingLatent Recurrent Transformer†\([Huang et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib35)\)Hidden state of a source layerAttention and residual streamkeptown, interleaved parallel trainingpretrainingWhiteMatter†\([Zhang & Ren, 2026](https://arxiv.org/html/2609.38149#bib.bib86)\)Keys and values of all layersAttention of every layerkeptown, fixed\-point passespretrainingCoconut\([Hao et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib30)\)Last hidden stateInputreplacedown, sequentialpost\-trainingCODI\([Shen et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib61)\)Projected last hidden stateInputreplacedown, sequentialpost\-trainingPCCoT\([Wu et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib79)\)Last hidden stateInputreplacedown, Jacobi passespost\-trainingSoft Thinking, MoI\([Zhang et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib88);[Zhuang et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib91)\)Soft token of its own distributionInputreplaced, mixednone \(training\-free\)inference onlyHRPO\([Yue et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib82)\)Soft token of its own distributionInput, gated with the tokenkeptown, RL rolloutspost\-trainingCoT2\([Gozeten et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib28)\)Soft tokenInputreplacedteacher\-forced, from an oracletask trainingCoLaR\([Tan et al\., 2025](https://arxiv.org/html/2609.38149#bib.bib67)\)Compressed reasoning\-token embeddingsInputreplacedteacher\-forced, from gold tracespost\-trainingThinking States†\([Amos et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib3)\)Compressed reasoning of the previous chunkA shallow layerkeptteacher\-forced, from gold annotationspost\-trainingCoCoMix\([Tack et al\., 2026](https://arxiv.org/html/2609.38149#bib.bib65)\)Its own predicted concept vectorInterleaved with hidden stateskeptown, same passpretrainingSMT†\([Kumar & Isola, 2026](https://arxiv.org/html/2609.38149#bib.bib38)\)RNN memoryRNN state–teacher\-forced, from a jointly trained encoderpretraining \(RNN\)LIFT\(ours\)Top\-kknext\-token distributionInput, fused with the tokenkeptteacher\-forced, from a pretrained LMpretraining

Similar Articles

@FinanceYF5: Next token prediction is short-sighted. What if the Transformer learns to predict its own next hidden state? Jayden Teoh proposes Next-Latent Prediction (NextLat): a self-supervised learning method that teaches the Transformer to form...

X AI KOLs Following

Jayden Teoh proposes Next-Latent Prediction (NextLat), a self-supervised learning method that teaches the Transformer to learn to predict the next hidden state, thereby forming a compact world model for reasoning and planning, and achieves up to 3.3x inference speedup through self-speculative decoding.

Next-Latent Prediction Transformers [R]

Reddit r/MachineLearning

Microsoft Research introduces Next-Latent Prediction (NextLat), a self-supervised method that trains transformers to predict their own next latent state, enabling compact world models for reasoning and planning and achieving up to 3.3x faster inference via self-speculative decoding.

Next-Latent Prediction Transformers Learn Compact World Models

Papers with Code Trending

Introduces Next-Latent Prediction (NextLat), a self-supervised objective that trains transformers to predict their next latent state, encouraging compact internal world models and improving generalization across sequence modeling tasks.