Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

arXiv cs.AI Papers

Summary

This paper critiques conventional depth truncation methods for evaluating recursive language models and introduces the Depth Control Protocol (DCP) to disentangle and isolate factors affecting depth utilization, improving evaluation accuracy.

arXiv:2609.19934v1 Announce Type: new Abstract: Depth-recurrent language models iteratively apply a small layer stack, decoupling per-token compute from distinct parameter count. To determine whether such a model genuinely utilizes its depth, both recurrence and layer-pruning literatures rely on a shared evaluation: truncating depth at inference time, plotting quality against retained depth fraction, and reading off the slope. While cheap and training-free, this metric suffers from an unexamined flaw: it extracts a single scalar from an intervention that alters multiple model properties simultaneously. Depth truncation concurrently reduces the number of block applications, decreases the volume of distinct computation performed, and pushes the readout head onto an out-of-distribution residual stream. The observed slope conflates all three factors, yet is conventionally interpreted as reflecting solely the second. We propose the Depth Control Protocol (DCP), a diagnostic suite that disentangles these three quantities. DCP comprises three positive controls that isolate each factor while varying the others, a negative control applying the identical interventions to dense transformers to ensure the effect is not an artifact of the measurement protocol, and a controlled training intervention to verify causality. The linchpin control, running the full budget of block applications while executing only a single distinct iteration, is strictly realizable only in depth-wise weight-sharing architectures, since in a dense network repeating a layer yields an entirely different model rather than the same model in an alternative configuration.
Original Article
View Cached Full Text

Cached at: 09/18/26, 09:28 AM

# Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models
Source: [https://arxiv.org/html/2609.19934](https://arxiv.org/html/2609.19934)
Thanh Tung KhuatAffiliation:NuverxAI \- AI & Creative Innovation Company Limited,thanhtung\.khuat@nuverxai\.comNguyen Thanh DungAffiliation:Ho Chi Minh City University of Technology \(HCMUT\),thanhdungng04@gmail\.com

September 17, 2026

###### Abstract

Depth\-recurrent language models iteratively apply a small layer stack, decoupling per\-token compute from distinct parameter count\. To determine whether such a model genuinely utilizes its depth, both recurrence and layer\-pruning literatures rely on a shared evaluation: truncating depth at inference time, plotting quality against retained depth fraction, and reading off the slope\. While cheap and training\-free, this metric suffers from an unexamined flaw: it extracts a single scalar from an intervention that alters multiple model properties simultaneously\. Depth truncation concurrently reduces the number of block applications, decreases the volume of distinct computation performed, and pushes the readout head onto an out\-of\-distribution residual stream\. The observed slope conflates all three factors, yet is conventionally interpreted as reflecting solely the second\.

We propose the Depth Control Protocol \(DCP\), a diagnostic suite that disentangles these three quantities\. DCP comprises three positive controls that isolate each factor while varying the others, a negative control applying the identical interventions to dense transformers to ensure the effect is not an artifact of the measurement protocol, and a controlled training intervention to verify causality\. The linchpin control, running the full budget of block applications while executing only a single distinct iteration, is strictly realizable only in depth\-wise weight\-sharing architectures, since in a dense network repeating a layer yields an entirely different model rather than the same model in an alternative configuration\.

## 1Introduction

Depth\-recurrent language models iteratively apply the same small stack of layers to their own representations, rather than stacking numerous layers with unique parameters\([Dehghani et al\., 2018](https://arxiv.org/html/2609.19934#bib.bib1);[Geiping et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib2)\)\. Through depth\-wise weight sharing, per\-token computation is decoupled from the number of distinct parameters: the same set of weights can be unrolled for two iterations or thirty\-two iterations\. This property has been exploited along two primary applications\. The first is inference\-time compute scaling, spending additional iterations on challenging inputs without enlarging the model footprint\. The second is latent\-space reasoning, executing multiple transformation steps prior to token generation rather than externalizing reasoning traces into explicit text\. Both paradigms are especially appealing at modest scale: if depth can partially substitute for data, a model trained on far fewer than trillions of tokens can still attain nontrivial reasoning capabilities\.

Both directions necessitate answering the same empirical question: does a given model genuinely utilize its depth? At present, this question is answered via an evaluation shared across both recurrence and layer\-pruning literature\([Gromov et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib14);[Men et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib15)\): truncating depth at inference, plotting quality as a function of retained depth fraction, and interpreting the slope as the contribution of depth\. This metric is cheap and requires no retraining, which has cemented it as the default across both lines of work\. However, the shared limitation of current evaluations is that they collapse a multi\-factorial intervention into a single scalar, subsequently attributing that scalar entirely to a single attribute\. This paper asks what this metric actually measures\.

Figure 1:Summary of main results\.\(a\)Prefix truncation, the standard metric, yields a monotonic decrease in NLL as the number of retained iterations increases\. The repeat control executes the full88block applications at everykkwhile varying only the number of*distinct*iterations; it remains nearly flat acrossk=1,2,4k=1,2,4\. The gap between the two curves atk=1k=1therefore does not arise from distinct depth\.\(b\)Decomposition of the naive gap into four constituent components across two checkpoints; error bars denote95%95\\%bootstrap confidence intervals over5,0005\{,\}000document resamplings\. The distinct depth component is approximately zero in both checkpoints\. The suffix control is omitted from panel \(a\) for visual clarity; the complete three\-control suite is presented in Figure[2](https://arxiv.org/html/2609.19934#S5.F2)\.The root cause of this limitation is that depth truncation conflates at least three distinct quantities\. First is how many*block applications*write into the residual stream\. Second is how much*distinct computation*these applications execute\. Third is whether the readout head remains calibrated to the residual distribution it receives, as depth truncation shifts the readout into a residual distribution unseen during training\. The observed slope reflects the sum of all three, yet in practice it is interpreted as though it isolates only the second\.

#### Proposed Method\.

We propose the Depth Control Protocol \(DCP\), a diagnostic procedure that disentangles these three quantities and validates the underlying causal mechanisms\. DCP consists of three core components\. The first component is a suite of three*positive controls*, each holding one quantity fixed while manipulating the others, thereby attributing each portion of the apparent gap to its rightful source\. The linchpin control in this set runs the full budget of block applications while executing only a single distinct iteration; it is*strictly realizable*only under depth\-wise weight sharing, because in a dense network repeating thekk\-th layer eight times produces a different model rather than the same model under an alternative runtime configuration\. The second component is a*negative control*applying the exact same interventions to dense transformers, serving to distinguish genuine confounders from artifacts of the measurement protocol itself\. The third component is a*controlled training intervention*that perturbs the single suspected causal variable, thereby turning observed correlations into verified causal relationships\. Formal specifications of all three components are detailed in Section[3](https://arxiv.org/html/2609.19934#S3)\.

Applying DCP across five configurations reveals that naive truncation substantially overestimates the contribution of recurrent depth \(Figure[1](https://arxiv.org/html/2609.19934#S1.F1)\)\. The overestimated quantity is the NLL reduction on a held\-out set when unrolling additional iterations—the precise metric underpinning claims that “depth buys performance”\. On a 542\.8M parameter recurrent model trained for mathematical reasoning, naive truncation attributes a reduction of1\.42511\.4251nats to latent depth\. Our three positive controls reveal that47\.9%47\.9\\%of this gap stems from the number of block applications rather than the distinct computation they perform, with an additional25\.6%25\.6\\%attributable to readout miscalibration; additional distinct iterations within a reasoning block contribute−2\.8%\-2\.8\\%\[−3\.2,−2\.4\]\[\-3\.2,\-2\.4\], indicating that under this control they yield no NLL improvement\.

Negative controls and training interventions demonstrate that this finding is not idiosyncratic to a single model\. Applying the identical control suite to layer truncation on two standard transformers reverses the sign of the “block application” component, proving that these confounders are not measurement artifacts\. Furthermore, when evaluated on Huginn\-0125—a public recurrent model whose iterations were*sampled during training*and which features published claims of inference\-time compute scaling—the calibration confounder disappears entirely: its fitted temperature remains flat within0\.0450\.045across a32×32\\timesdepth range, compared to0\.570\.57across an8×8\\timesrange in the fixed\-depth model\. Across four models, the confounder tracks exactly one variable: whether the readout head witnessed more than a single recurrent depth during training, independent of model family or parameter scale\. The vulnerability identified by these results thus resides not in the recurrent architecture itself, but in the*training schedule*: a model trained strictly at a single depth develops a readout specialized to that depth, causing any truncation evaluation to conflate specialization drift with the genuine contribution of depth\.

These two components possess different scopes, which we explicitly delineate\. The*application count*confounder is unique to weight\-sharing architectures: its sign inverts when applied to dense transformers\. The*calibration*confounder, by contrast, is not: evaluated on Qwen2\.5\-Math\-1\.5B with general text where the model is well calibrated \(T=1\.03T=1\.03at full depth\), layer truncation still inflates the fitted temperature to1\.911\.91, accounting for16\.5%16\.5\\%of the apparent gap\. Consequently, any study drawing conclusions from performance curves across retained layers without recalibration reports a metric contaminated by this component, even for non\-recurrent architectures\.

These findings yield two practical implications\. Randomly sampling recurrent depth during training eliminates this confounder at zero inference cost\. Until this practice becomes standard, depth truncation studies on models trained at fixed depth should report DCP alongside naive curves, as naive curves alone exaggerate the depth effect by roughly a factor of two\.

#### Contributions\. Our main contributions in this research include:

1. 1\.We propose the Depth Control Protocol \(DCP\), a diagnostic suite comprising three positive controls that decompose depth truncation into its constituent quantities, alongside a negative control and a training intervention \(Section[3](https://arxiv.org/html/2609.19934#S3)\)\. We formally state its applicability conditions, prove the decomposition is exact and exhaustive, and outline five identification limits, including path dependency within the configuration space \(Section[3\.3](https://arxiv.org/html/2609.19934#S3.SS3)\)\.
2. 2\.Using the three positive controls, we identify and quantify two primary confounders in depth truncation: block application count and calibration drift \(Section[5](https://arxiv.org/html/2609.19934#S5)\)\.
3. 3\.We decompose the1\.42511\.4251nat naive depth effect into47\.9%47\.9\\%block applications,−2\.8%\-2\.8\\%distinct depth, and54\.8%54\.8\\%inter\-block composition, with bootstrap confidence intervals excluding zero across all components\. The finding that distinct depth contributes approximately zero is replicated across three checkpoints spanning distinct training regimes, including a reinforcement\-learning checkpoint \(−0\.14%\-0\.14\\%, Section[5\.4](https://arxiv.org/html/2609.19934#S5.SS4)\)\.
4. 4\.We introduce a negative control on dense transformers demonstrating that these confounders are intrinsic properties of the evaluated models rather than measurement artifacts: applying identical interventions to layer truncation reverses the sign of the “block application” component \(Section[6](https://arxiv.org/html/2609.19934#S6)\)\.
5. 5\.We employ cross\-model comparisons to validate the underlying generative mechanism, specifically readout specialization to a single training depth\. On Huginn\-0125, where recurrence depth is sampled during training, the calibration confounder vanishes \(−0\.6%\-0\.6\\%vs\.25\.6%25\.6\\%\); across four models, this confounder tracks this operational variable rather than architecture family or parameter scale \(Section[7](https://arxiv.org/html/2609.19934#S7)\)\.
6. 6\.We introduce a controlled training intervention establishing causal attribution:1,0001\{,\}000steps of continued training with sampled depth reduces the calibration share from29\.9%29\.9\\%to5\.1%5\.1\\%, whereas a control branch trained for the identical budget at fixed depth remains virtually unchanged \(Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3)\)\.
7. 7\.We characterize the dynamics of the recurrent stack across five configurations \(Section[8](https://arxiv.org/html/2609.19934#S8)\)\. Only Huginn\-0125 functions as a contraction mapping, explaining its empirical resilience; the remaining four configurations, including both dense transformers, amplify perturbations at comparable rates\. Sampled\-depth training does not alter this dynamical property, demonstrating that calibration and operator dynamics are distinct mechanisms\.
8. 8\.We formulate a concrete remedial intervention—sampling recurrence depth during training across the target deployment configuration space—and recommend reporting DCP alongside any truncation\-based depth claims\. Our evaluation toolkit is made publicly available with this paper\.

## 2Related Work

#### Adaptive Computation and Recurrent Depth\.

Adaptive Computation Time \(ACT\)\([Graves, 2016](https://arxiv.org/html/2609.19934#bib.bib3)\)introduced variable per\-token compute via a learned halting distribution; PonderNet\([Banino et al\., 2021](https://arxiv.org/html/2609.19934#bib.bib4)\)reformulated halting through a probabilistic objective\. Universal Transformers\([Dehghani et al\., 2018](https://arxiv.org/html/2609.19934#bib.bib1)\)combined depth\-wise weight sharing with ACT\. Huginn\-0125\([Geiping et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib2)\)scaled a weight\-tied recurrent core to 3\.5B parameters, sampling recurrence steps during training—a pivotal operational feature for this study\. Our evaluated architecture differs by chaining*two*distinct reasoning blocks rather than repeating a single block, which renders the distinction between “distinct computation” and “repeated unrolling” empirically measurable\.

This research line has expanded rapidly in 2026\. Recent proposals investigate*what*to repeat\([Lin et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib5)\), residual normalization under tied weights\([Li et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib6)\), gated modulation to prevent representational collapse\([Hegazy et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib7)\), and retrofitting recurrent depth into pretrained dense models\([Shapiro, 2026](https://arxiv.org/html/2609.19934#bib.bib8)\)\. Crucially, all of these works report model quality as a function of iteration count\. This family of performance curves is precisely what this paper investigates\.

#### Dynamics of Recurrent Operators\.

A parallel inquiry explores*when*unrolling additional iterations is advantageous, answering through the dynamical properties of the trained operator\.[Viakhirev et al\. \(2026\)](https://arxiv.org/html/2609.19934#bib.bib9)categorized operators into settled, boundary, and drifting regimes, establishing sufficient conditions under which added depth preserves solution fidelity\.[Zhang et al\. \(2026\)](https://arxiv.org/html/2609.19934#bib.bib27)investigated when tied recurrence faithfully implements an algorithm, identifying a compute\-budget law linking execution speed to the training contract\. SCORE\([Godin, 2026](https://arxiv.org/html/2609.19934#bib.bib10)\)went further by enforcing contractivity by construction via an ODE\-style update\.

We explicitly clarify our conceptual overlap with these works\. Our contraction coefficient measurements in Section[8](https://arxiv.org/html/2609.19934#S8)and the regime classifications of[Viakhirev et al\. \(2026\)](https://arxiv.org/html/2609.19934#bib.bib9)probe the same underlying operator properties, arriving at concordant qualitative conclusions: contractive operators exhibit resilience to depth variations, whereas amplifying operators do not\. The divergence lies in the direction of the core inquiry: they ask what happens when depth is*added*beyond the training budget; we ask where the performance drop observed when depth is*removed*actually originates\. The primary contribution of this work lies not in dynamical characterization, but in decomposing the apparent performance gap into constituent quantities—a separation that dynamical analysis alone cannot achieve\.

#### Diagnostics for Recurrent Computation\.

Closest in philosophical posture is[Lam\-Muir \(2026\)](https://arxiv.org/html/2609.19934#bib.bib11), which directly asked whether reported latent reasoning represents genuine computation or an*instrument artifact*, utilizing pre\-registered probe gates and simultaneous measurements across readout and latent channels\. They concluded the phenomenon was genuine and localized at the readout\.[Lin et al\. \(2026\)](https://arxiv.org/html/2609.19934#bib.bib5)introduced the Iteration Transfer Ratio \(ITR\) to quantify the non\-redundant contribution of each iteration\.

These metrics and DCP address fundamentally different questions, a distinction we emphasize as critical\. ITR identifies*which iterations warrant unrolling*for architectural design\.[Lam\-Muir \(2026\)](https://arxiv.org/html/2609.19934#bib.bib11)interrogated whether a specific empirical phenomenon is real\. DCP asks whether a*widely adopted evaluation protocol*correctly attributes credit, answering by decomposing the apparent gap into interventional components\. Notably, all three converge on the same locus: the readout\. In their work, the readout is where genuine computation crystallizes; in ours, the readout is where calibration drift accumulates and is mistakenly counted as the contribution of depth\.

#### Latent Reasoning\.

Quiet\-STaR\([Zelikman et al\., 2024](https://arxiv.org/html/2609.19934#bib.bib12)\)trains latent rationales via REINFORCE objectives; Coconut\([Hao et al\., 2024](https://arxiv.org/html/2609.19934#bib.bib13)\)conducts multi\-step reasoning continuous latent space rather than language token space\. These paradigms share the goal of non\-verbalized internal computation with recurrent models, but vary the number of latent*tokens*rather than the number of*block applications*of a layer stack; hence, our identified confounders do not manifest in the same form\.

#### Layer Pruning and Early Exit\.

Depth truncation is standard across layer pruning, where the metric of interest is residual performance after layer ablation\([Gromov et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib14);[Men et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib15)\), and across early\-exit models, where inference halts at intermediate layers conditioned on inputs\([Elbayad et al\., 2019](https://arxiv.org/html/2609.19934#bib.bib16);[Schuster et al\., 2022](https://arxiv.org/html/2609.19934#bib.bib17)\)\. Both paradigms manipulate*which*and*how many*layers execute\. This literature is likewise undergoing self\-scrutiny:[Wang et al\. \(2025\)](https://arxiv.org/html/2609.19934#bib.bib18)showed that pruning just one or two layers shatters inference\-time compute scaling, while[Shi et al\. \(2026\)](https://arxiv.org/html/2609.19934#bib.bib19)traced catastrophic collapse to sharp transitions in decision representations\. Both draw conclusions from performance curves over remaining layers\.

The fundamental divergence from our work lies here\. In a dense network, each layer is a unique function; hence, no configuration of “LLapplications, one distinct computation” exists: repeating layerkkcreates a different model rather than the same model under an alternate configuration\. Our repeat control \(Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)\) requires depth\-wise weight sharing and cannot be constructed within the setting of those works\. This is why our protocol isolates application count from distinct computation, a separation that dense layer pruning cannot achieve in principle\.

#### Training with Stochastic Depth\.

Stochastic depth\([Huang et al\., 2016](https://arxiv.org/html/2609.19934#bib.bib20)\)and LayerDrop\([Fan et al\., 2019](https://arxiv.org/html/2609.19934#bib.bib21)\)randomly drop layers during training, with LayerDrop explicitly targeting training\-free inference pruning\. The intervention we explore in Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3)—sampling recurrence depth during training—belongs to this conceptual family\.

We make no claim of novelty regarding this training technique\. The contribution of this paper operates on a different plane: prior works propose a*training technique*to induce prunability, whereas we interrogate whether an*evaluation metric*is interpretable, demonstrating that fixed\-depth training invalidates that metric in a quantifiable manner\. Consequently, the two lines converge: the technique originally proposed to enhance prunability is precisely what restores the validity of the evaluation metric\. An open cell in our design—evaluating our control suite on a dense model trained with LayerDrop—is discussed in Section[11](https://arxiv.org/html/2609.19934#S11)\.

#### Calibration\.

Temperature scaling\([Guo et al\., 2017](https://arxiv.org/html/2609.19934#bib.bib22)\)is the standard single\-parameter post\-hoc calibration technique\. We employ it to*quantify*the magnitude of miscalibration, not to patch the model: the fitted temperature measures how far the output distribution departs from the distribution it was trained upon\. Early\-exit literature also examines calibration at intermediate exit points\([Schuster et al\., 2022](https://arxiv.org/html/2609.19934#bib.bib17)\), but treats calibration as a*control signal*to govern halting decisions\. Here, calibration serves as a*measurement probe*, and the specific profile we observe \(fitted temperature remaining flat across truncated depths before dropping sharply to≈1\\approx 1strictly at the training depth—a step function rather than a gradual drift\) acts as an empirical diagnostic signature rather than generic miscalibration\.

#### Position Embeddings and Evaluation Data\.

The evaluated model utilizes rotary position embeddings \(RoPE\)\([Su et al\., 2024](https://arxiv.org/html/2609.19934#bib.bib24)\)with a dual\-stream variant detailed in Section[4](https://arxiv.org/html/2609.19934#S4)\. We evaluate on held\-out mathematical text; benchmark contamination in this domain is an established concern\([Brown et al\., 2020](https://arxiv.org/html/2609.19934#bib.bib25);[Touvron et al\., 2023](https://arxiv.org/html/2609.19934#bib.bib26)\), and our evaluation corpus was curated following a rigorous decontamination audit\. All reported confidence intervals represent percentile bootstrap intervals\([Tibshirani and Efron, 1993](https://arxiv.org/html/2609.19934#bib.bib23)\)\.

## 3Proposed Method: Depth Control Protocol \(DCP\)

This section details the Depth Control Protocol \(DCP\), our proposed diagnostic procedure\. The problem DCP addresses is: given a quality\-versus\-depth curve obtained via depth truncation, disentangle the apparent gap into its constituent causal sources and determine which source genuinely reflects the contribution of depth\. All three positive controls intervene exclusively at inference time over frozen weights; their computational cost is of the same order as the naive truncation they augment\.

#### Notation\.

Consider a recurrent model comprisingnbn\_\{b\}reasoning blocks, each appliednin\_\{i\}times, yielding a total budget ofN=nb⋅niN=n\_\{b\}\\cdot n\_\{i\}block applications\. Standard prefix truncation retains the firstk≤Nk\\leq Napplications and records performanceQ⁡\(k\)Q\(k\); the apparent gap isΔnaive=Q⁡\(1\)−Q⁡\(N\)\\Delta\_\{\\text\{naive\}\}=Q\(1\)\-Q\(N\)\. To ensure learned per\-iteration embeddings remain aligned with training, the original global iteration indexgidx=b⋅ni\+tg\_\{\\text\{idx\}\}=b\\cdot n\_\{i\}\+tis preserved for blockbbat iterationtt, rather than re\-indexing from zero\.

### 3\.1Component 1: Three Positive Controls

The three interventions below each isolate one quantity while holding the remaining quantities constant, thereby attributing each portion ofΔnaive\\Delta\_\{\\text\{naive\}\}to its rightful source\.

#### Repeat Control \(repeat\)\.

Execute the firstkkdistinct applications, then repeat thekk\-th application until completingNNtotal block applications\. This control holds*block applications*constant, preserving both the norm and the empirical distribution of the residual stream feeding into the readout head exactly as seen during training, while varying solely the volume of*distinct*computation\. This represents the linchpin control of DCP and is*strictly realizable*only in depth\-wise weight\-sharing architectures: in a dense network, repeating layerkkmultiple times constructs a different function rather than evaluating the same model under an alternate execution schedule\.

#### Suffix Control \(suffix\)\.

Execute the*final*kkapplications rather than the initial ones\. This control decouples “how many iterations” from “which specific iterations”, thereby verifying whether iteration count is indeed the governing causal variable\.

#### Temperature Calibration Control\.

Refit a single logit temperatureTTfor each depthkkon a held\-out calibration set and report post\-calibration performance\. Because the final layer normalization and readout head are trained on the residual statistics of full forward passes, truncation miscalibrates them for reasons unrelated to reasoning capacity; recalibration isolates this distributional drift from underlying computational capability\. We fitTTvia L\-BFGS with strong\-Wolfe line search onlog⁡T\\log T\([Guo et al\., 2017](https://arxiv.org/html/2609.19934#bib.bib22)\), resolving optimal values to<10−3<10^\{\-3\}; a coarse grid search is inadequate here because temperature drifts compared across architectures differ by an order of magnitude\. We emphasize that temperature scaling is used here to*quantify*the extent of miscalibration, not to modify model predictions\.

### 3\.2Component 2: Decomposition of the Apparent Gap

Using the three controls above,Δnaive\\Delta\_\{\\text\{naive\}\}is decomposed into four constituent components, each defined as the performance difference between two conditions differing in exactly one operational attribute:

Δapply\\displaystyle\\Delta\_\{\\text\{apply\}\}=Qprefix​\(1\)−Qrepeat​\(1\),\\displaystyle=Q\_\{\\text\{prefix\}\}\(1\)\-Q\_\{\\text\{repeat\}\}\(1\),\(1\)Δdistinct\\displaystyle\\Delta\_\{\\text\{distinct\}\}=Qrepeat​\(1\)−Qrepeat​\(ni\),\\displaystyle=Q\_\{\\text\{repeat\}\}\(1\)\-Q\_\{\\text\{repeat\}\}\(n\_\{i\}\),\(2\)Δcompose\\displaystyle\\Delta\_\{\\text\{compose\}\}=Qrepeat​\(ni\)−Qrepeat​\(N\),\\displaystyle=Q\_\{\\text\{repeat\}\}\(n\_\{i\}\)\-Q\_\{\\text\{repeat\}\}\(N\),\(3\)Δcalib\\displaystyle\\Delta\_\{\\text\{calib\}\}=Δnaive−\(QprefixT​\(1\)−QprefixT​\(N\)\),\\displaystyle=\\Delta\_\{\\text\{naive\}\}\-\\bigl\(Q^\{T\}\_\{\\text\{prefix\}\}\(1\)\-Q^\{T\}\_\{\\text\{prefix\}\}\(N\)\\bigr\),\(4\)whereQTQ^\{T\}denotes performance after temperature recalibration at eachkk\. The relative share of each component is defined as its ratio toΔnaive\\Delta\_\{\\text\{naive\}\}\. The calibration component by definition overlaps with the first three components, as it is evaluated across the same boundary conditions \(k=1k=1andk=Nk=N\); we report it alongside the structural decomposition rather than summing into the total\.

Because all components are differences across conditions and all shares are ratios of these differences, confidence intervals computed independently per condition would be misleading: conditions are evaluated on identical documents and are strongly correlated\. DCP therefore mandates resampling document indices*once per bootstrap iteration*and propagating that identical resample across all conditions and derived components\([Tibshirani and Efron, 1993](https://arxiv.org/html/2609.19934#bib.bib23)\)\.

### 3\.3Applicability Conditions and Identification Limits

This section explicitly formalizes the assumptions of DCP, identifying which quantities are identifiable and which are not\. An empirical evaluation protocol is only sound when its failure modes are rigorously characterized\.

#### Property 1 \(Exact and Exhaustive Decomposition\)\.

The three components in \([1](https://arxiv.org/html/2609.19934#S3.E1)\)–\([3](https://arxiv.org/html/2609.19934#S3.E3)\) form a telescoping sequence along the pathprefix​\(1\)→repeat​\(1\)→repeat​\(ni\)→repeat​\(N\)\\text\{prefix\}\(1\)\\rightarrow\\text\{repeat\}\(1\)\\rightarrow\\text\{repeat\}\(n\_\{i\}\)\\rightarrow\\text\{repeat\}\(N\)within configuration space\. BecauseQrepeat​\(N\)=Qprefix​\(N\)Q\_\{\\text\{repeat\}\}\(N\)=Q\_\{\\text\{prefix\}\}\(N\)by definition \(both executing the full forward pass\), the sum of the three components collapses algebraically toQprefix​\(1\)−Qprefix​\(N\)=ΔnaiveQ\_\{\\text\{prefix\}\}\(1\)\-Q\_\{\\text\{prefix\}\}\(N\)=\\Delta\_\{\\text\{naive\}\}\. The decomposition leaves no unassigned residual and requires no additivity assumptions regarding underlying effects\.

#### Property 2 \(Forced Path, Not Arbitrary Choice\)\.

A natural critique is that attribution shares depend upon the path traversed through configuration space, rendering individual labels arbitrary\. This critique fails here: the geometry of realizable configurations forces exactly one path whose segments are strictly univariate\.

Assigning coordinates\(a,d\)\(a,d\)to each configuration, whereaadenotes block applications anddddenotes distinct iterations, the two truncation regimes map to two distinct loci:

prefix​\(k\)\\displaystyle\\text\{prefix\}\(k\)⟼\(a,d\)=\(k,k\),\\displaystyle\\;\\longmapsto\\;\(a,d\)=\(k,k\),\(5\)repeat​\(k\)\\displaystyle\\text\{repeat\}\(k\)⟼\(a,d\)=\(N,k\)\.\\displaystyle\\;\\longmapsto\\;\(a,d\)=\(N,k\)\.\(6\)Prefix truncation*couples*aaanddd; repeat controls fixa=Na=Nand vary onlydd\. Consequently:

- •A step varying onlyaarequires two configurations with identicalddand differentaa\. The unique valid pair isprefix​\(k\)\\text\{prefix\}\(k\)andrepeat​\(k\)\\text\{repeat\}\(k\)at matchingkk\.
- •A step varying onlyddrequires two configurations with identicalaaand differentdd\. The unique valid pair isrepeat​\(k1\)\\text\{repeat\}\(k\_\{1\}\)andrepeat​\(k2\)\\text\{repeat\}\(k\_\{2\}\)\.

Any decomposition into strictly univariate steps can only utilize these two step types, and the path in \([1](https://arxiv.org/html/2609.19934#S3.E1)\)–\([3](https://arxiv.org/html/2609.19934#S3.E3)\) is the unique sequence connectingprefix​\(1\)\\text\{prefix\}\(1\)torepeat​\(N\)\\text\{repeat\}\(N\)\. The path is not merely one among many valid choices; it is the sole valid choice\.

#### Limit 1 \(Confounded Paths Yield Meaningless Labels\)\.

This does not prevent one from constructing alternative paths; it simply means such paths must contain multivariate steps\. The consequences are severe\. Consider the pathprefix​\(1\)→prefix​\(ni\)→repeat​\(ni\)→repeat​\(N\)\\text\{prefix\}\(1\)\\rightarrow\\text\{prefix\}\(n\_\{i\}\)\\rightarrow\\text\{repeat\}\(n\_\{i\}\)\\rightarrow\\text\{repeat\}\(N\): while its sum still telescopes toΔnaive\\Delta\_\{\\text\{naive\}\}, its first step transitions from\(1,1\)\(1,1\)to\(ni,ni\)\(n\_\{i\},n\_\{i\}\), varying*both*coordinates simultaneously\. Attributing that step’s difference entirely to distinct depth conflatesddwith the simultaneous shift inaa\.

Table[1](https://arxiv.org/html/2609.19934#S3.T1)demonstrates the numerical consequences on our calibration grid\.

Table 1:Why univariate steps are mandatory\. The right columns trace the pathprefix​\(1\)→prefix​\(4\)→repeat​\(4\)→repeat​\(8\)\\text\{prefix\}\(1\)\\rightarrow\\text\{prefix\}\(4\)\\rightarrow\\text\{repeat\}\(4\)\\rightarrow\\text\{repeat\}\(8\), where the first step varies both block applications and distinct iterations simultaneously\. While the sum remains exact, the “distinct” component absorbs nearly the entire gap, while the “application” component receives negative values in two branches, rendering the labels meaningless\. Values evaluated at step6,0006\{,\}000\.Under the confounded path, the component labeled “distinct depth” surges from−1\.60%\-1\.60\\%to88\.72%88\.72\\%in branch C, while “block applications” turns negative in two of three branches\. A decomposition where an application share drops below−49%\-49\\%does not represent an alternative interpretation; it signals that labels have lost their physical meaning\. We highlight this because we initially constructed such a path during exploratory robustness checks; its88\.72%88\.72\\%figure appeared deceptively plausible before checking coordinate orthogonality\.

#### Limit 2 \(Composition Component is Inherently Mixed\)\.

The third step,repeat​\(ni\)→repeat​\(N\)\\text\{repeat\}\(n\_\{i\}\)\\rightarrow\\text\{repeat\}\(N\), maintainsa=Na=Nwhile increasingddfromnin\_\{i\}toNN\. In architectures withnb≥2n\_\{b\}\\geq 2, this transition simultaneously engages the second reasoning block\. This step is therefore*not*strictly univariate in the narrow sense, which is why we designate it “inter\-block composition” rather than distinct depth\. Disentangling these two factors would require a configuration runningNNapplications withdddistinct iterations*confined entirely within the first block*—an execution schedule not realizable under standard loop unrolling\.

#### Condition C1 \(Depth\-Wise Weight Sharing\)\.

The repeat control assumes the same functionfθf\_\{\\theta\}is appliedNNtimes, such thatht\+1=fθ​\(ht\)h\_\{t\+1\}=f\_\{\\theta\}\(h\_\{t\}\)with sharedθ\\theta\. In dense networks, each layer is an independent functionfθtf\_\{\\theta\_\{t\}\}; repeating layerkkconstructs an*entirely different model*rather than the same model under an alternate schedule\. When C1 is violated, DCP is restricted to prefix, suffix, and calibration controls; this is why the negative control in Section[6](https://arxiv.org/html/2609.19934#S6)lacks the repeat component\.

#### Condition C2 \(Loop Index Preservation\)\.

If a model employs learned per\-iteration embeddings, the repeat control must preserve the original global iteration indexgidx=b⋅ni\+tg\_\{\\text\{idx\}\}=b\\cdot n\_\{i\}\+t\. Re\-indexing would introduce a second confounded variable, preventing attribution to distinct computation\.

#### Condition C3 \(Absence of Early Exit\)\.

DCP assumes that depth truncation evaluates*the same model under an alternate runtime configuration*\. If an architecture incorporates early\-exit classifiers or halting gates that alter behavior upon truncation, this assumption breaks and components become uninterpretable\. The architecture examined here uses an*update*gate rather than a halting probability \(Equation \([7](https://arxiv.org/html/2609.19934#S4.E7)\)\), satisfying C3\.

#### Condition C4 \(Requirement of at Least Two Distinct Blocks\)\.

The composition componentΔcompose\\Delta\_\{\\text\{compose\}\}is defined strictly whennb≥2n\_\{b\}\\geq 2\. For single\-block recurrent cores,ni=Nn\_\{i\}=Nand this component is identically zero, reducing the decomposition to two components\. This explains why the full decomposition does not transfer to Huginn\-0125, even though all other controls transfer seamlessly\.

#### Limit 3 \(Calibration Component is a Lower Bound\)\.

Temperature scaling is a single\-parameter family\. If miscalibration departs from a pure temperature shift—such as mean logit drift—fittingTTcaptures only the variance explainable by temperature\. Therefore,Δcalib\\Delta\_\{\\text\{calib\}\}must be interpreted as a*lower bound*on total miscalibration\. Furthermore, it overlaps with the structural components by definition \(measured across the same endpoints\) and must not be added to their sum\.

#### Limit 4 \(Scope of the Negative Control\)\.

The negative control supports the conclusion that confounders are not artifacts of the measurement protocol\. It matches the truncation schedule, temperature fitting routine, and evaluation data, but does*not*match parameter scale, pretraining corpus, or architectural family\. It cannot rule out hypotheses involving unobserved variables that co\-vary with the depth schedule\.

#### Limit 5 \(Intervention Establishes Sufficiency, Not Necessity\)\.

The controlled intervention in Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3)perturbs exactly one variable against a control branch receiving an identical training budget\. This proves the depth schedule is*sufficient*to shift calibration share\. It does not prove necessity, nor does it exclude the possibility of alternative interventions producing similar effects\.

#### Condition C5 \(Quality Metric Selection\)\.

The metricQQmust exhibit sufficiently low variance to resolve differences between adjacent configurations\. On undertrained checkpoints, single\-digit generation accuracy exhibits variance that overwhelms the underlying effect size\. Consequently, all quantities in this paper are likelihood\-based\. DCP does not strictly requireQQto be a likelihood metric, but mandates that its variance be small relative to effect differences—a condition that must be verified prior to adopting task\-level metrics\.

### 3\.4Component 3: Negative Control on Dense Transformers

The three positive controls characterize how the apparent gap decomposes, but do not establish whether that decomposition is an artifact of the evaluation protocol\. Any truncation protocol could potentially induce confounders, for instance if removing layers consistently shifts residual distributions in a uniform direction\. DCP therefore requires a negative control: applying identical interventions to layer truncation on standard dense transformers, where the hypothesized confounder mechanisms should not operate\. If the measured shares on the negative control differ in sign or magnitude from the target model, the confounders represent genuine properties of the evaluated model rather than measurement artifacts\. In dense transformers, repeat controls are unrealizable; hence, the negative control executes prefix, suffix, and calibration sweeps\.

### 3\.5Component 4: Controlled Training Intervention

The preceding components provide correlational evidence: confounders appear in some models and are absent in others\. To establish causal attribution, DCP mandates a training intervention modifying*exactly one suspected causal variable*against a control branch receiving an identical step budget without the intervention\. In this paper, that variable is the depth schedule: the intervention branch trains with uniformly sampled recurrence depth, while the control branch trains at fixed depth\. Both branches are subsequently evaluated using the identical positive controls\. This paired design eliminates the competing hypothesis that observed shifts merely reflect additional training steps\.

## 4Experimental Setup

#### Model Under Study\.

The primary empirical subject of this paper is Sona, a depth\-recurrent language model we trained for mathematical reasoning\. Sona is not the method proposed in this paper; the proposed method is DCP \(Section[3](https://arxiv.org/html/2609.19934#S3)\), and Sona serves as one of five configurations evaluated under DCP\. We utilize our own model as a primary benchmark because DCP requires precise knowledge of the depth schedule during pretraining—an operational detail seldom released with model weights\.

Sona is a 542\.8M parameter decoder\-only model featuring a two\-phase layer stack:1616perception layers, followed bynb=2n\_\{b\}=2reasoning blocks, each appliedni=4n\_\{i\}=4times, yielding a total of88block applications per forward pass\. The hidden dimension is15361536with2424attention heads \(44KV heads\) and an FFN intermediate dimension of43524352\. A latent “thought” bank of3232tokens is carried forward across recurrence steps\. Attention between the thought bank and the textual sequence is restricted: thought→\\rightarrowthought is unmasked, thought→\\rightarrowtext and text→\\rightarrowthought are blocked, and text→\\rightarrowtext is strictly causal; the two streams communicate exclusively through a causal bridge\. Positional encodings utilize rotary position embeddings \(RoPE\)\([Su et al\., 2024](https://arxiv.org/html/2609.19934#bib.bib24)\)via a dual\-stream variant\. Each reasoning block computes an energy\-based gateg=σ\(−\(E\+α\|ΔE\|\)/τ\)g=\\sigma\\\!\\left\(\-\(E\+\\alpha\|\\Delta E\|\)/\\tau\\right\), applied per\-position and per\-iteration:

xout=g⋅h\+\(1−g\)⋅xresidual,x\_\{\\text\{out\}\}=g\\cdot h\+\(1\-g\)\\cdot x\_\{\\text\{residual\}\},\(7\)establishing thatggis an*update*gate governing the proportion of block output written into the residual stream, rather than a halting probability; the architecture features no early\-exit mechanism, departing from adaptive computation schemes governed by learned halting distributions\([Graves, 2016](https://arxiv.org/html/2609.19934#bib.bib3);[Banino et al\., 2021](https://arxiv.org/html/2609.19934#bib.bib4)\)\. Equation \([7](https://arxiv.org/html/2609.19934#S4.E7)\) is transcribed directly from the released implementation codebase, not from design notes\.

#### Checkpoints\.

We evaluate three frozen checkpoints of the same architecture: a pretraining checkpoint at step34,00034\{,\}000, a downstream checkpoint at step4,0004\{,\}000of subsequent fine\-tuning, and, for replication in Section[5\.4](https://arxiv.org/html/2609.19934#S5.SS4), a second pretraining checkpoint at step40,00040\{,\}000trained on the decontaminated corpus with repaired document boundaries\. All three checkpoints are undertrained relative to modern compute\-optimal small models; absolute task performance is not the object of study\. The downstream checkpoint underwent preference optimization following supervised fine\-tuning\. Its task accuracy is modest, which is precisely why all evaluations in this work rely on likelihood metrics rather than discrete generation metrics: at single\-digit accuracy, generative variance completely dominates the structural effects under investigation\.

#### Evaluation Data\.

All measurements are conducted on a held\-out set of200200problem–solution documents drawn from the English mathematical corpus used to train Sona, absent from all training splits\. This corpus was reconstructed following a decontamination audit detailed separately; benchmark contamination in mathematical reasoning represents an acknowledged risk\([Brown et al\., 2020](https://arxiv.org/html/2609.19934#bib.bib25);[Touvron et al\., 2023](https://arxiv.org/html/2609.19934#bib.bib26)\), so the held\-out set was audited against public benchmarks prior to use\. Section[9\.2](https://arxiv.org/html/2609.19934#S9.SS2)replicates the primary evaluation on out\-of\-domain text to verify that conclusions do not depend upon this domain choice\.

#### Metrics\.

The quality metricQQutilized throughout is negative log\-likelihood \(NLL\) under teacher forcing, evaluated over the solution segment of each document, with next\-token accuracy reported as a secondary metric\. Each configuration is evaluated over exactly125,100125\{,\}100tokens, identical across all experimental sweeps, so that all cross\-condition differences represent paired comparisons over the identical token set\. Likelihood is preferred over generation metrics because, at the performance regime of these checkpoints, generative sampling variance would obscure the structural effects under analysis\.

## 5Depth Truncation is Confounded

### 5\.1Naive Truncation

Standard depth truncation restricts recurrence to the firstkkblock applications out of the total88, reporting performance as a function ofkk\. To ensure learned per\-iteration embeddings remain aligned with pretraining, we preserve the original global iteration indexgidx=b⋅ni\+tg\_\{\\text\{idx\}\}=b\\cdot n\_\{i\}\+tfor blockbbat iterationtt, rather than re\-indexing from zero\.

Figure[2](https://arxiv.org/html/2609.19934#S5.F2)and Table[11](https://arxiv.org/html/2609.19934#A2.T11)display this measurement\. NLL decreases monotonically withkkacross both evaluations, achieving a1\.431\.43nat drop on the downstream checkpoint—a figure easily read as compelling evidence that latent recurrent depth performs substantial reasoning\.

Figure 2:Held\-out NLL as a function of retained iterations under three truncation modes on the pretraining checkpoint at step40,00040\{,\}000\. Prefix truncation drops monotonically\. The repeat control unrolls the full88block applications at everykk, varying only the number of*distinct*iterations; it remains nearly flat acrossk=1,2,4k=1,2,4\. Complete data in Tables[11](https://arxiv.org/html/2609.19934#A2.T11)and[12](https://arxiv.org/html/2609.19934#A2.T12)\.
### 5\.2Applying the Controls

We apply the three positive controls of DCP \(Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)\) to the identical checkpoint and held\-out documents used for naive truncation above\. All three intervene strictly at inference time over frozen weights; hence, all observed differences represent differences in how the same model is probed, not differences in model weights\.

### 5\.3Decomposition

Figure[2](https://arxiv.org/html/2609.19934#S5.F2)and Table[12](https://arxiv.org/html/2609.19934#A2.T12)decompose the1\.42511\.4251nat naive gap \(k=1→k=8k\{=\}1\\rightarrow k\{=\}8, prefix, raw\)\. Because all components are differences across conditions and all shares are ratios of these differences, confidence intervals computed independently per condition would be misleading: conditions are evaluated on identical documents and are strongly correlated\. We therefore resample document indices*once per bootstrap replication*and propagate that identical sample across all conditions and derived components \(n=200n=200documents,5,0005\{,\}000replications\)\. All four component confidence intervals exclude zero:

- •Block applications / residual distribution:0\.68310\.6831nats,47\.9%47\.9\\%\[46\.6,49\.346\.6,49\.3\]\.Prefixk=1k\{=\}1\(2\.64262\.6426\) versus repeatk=1k\{=\}1\(1\.95951\.9595\)\. The volume of distinct computation is identical; only the number of block applications differs\. Nearly half of the apparent depth effect is not depth\.
- •Distinct depth within a block:−0\.0392\-0\.0392nats,−2\.8%\-2\.8\\%\[−3\.2,−2\.4\-3\.2,\-2\.4\]\.Repeatk=1k\{=\}1\(1\.95951\.9595\) versus repeatk=4k\{=\}4\(1\.99871\.9987\)\. Four distinct iterations of the first reasoning block are not merely “no better” than a single iteration repeated: the confidence interval strictly excludes zero on the negative side, establishing that they are reliably*worse*, albeit by a modest margin\.
- •Inter\-block composition:0\.78120\.7812nats,54\.8%54\.8\\%\[53\.3,56\.353\.3,56\.3\]\.Repeatk=4k\{=\}4\(1\.99871\.9987\) versusk=8k\{=\}8\(1\.21731\.2173\), the unique configuration where both distinct reasoning blocks execute in their trained sequential order\.
- •Calibration:0\.36450\.3645nats,25\.6%25\.6\\%\[24\.5,26\.624\.5,26\.6\]\(overlapping with the first component\)\. The post\-calibration gap is1\.05991\.0599nats compared to1\.42511\.4251nats raw, and fitted temperature drifts from1\.451\.45atk=1k\{=\}1down to0\.910\.91atk=8k\{=\}8\(amplitude0\.570\.57\)\. HereTTis fixed to its full\-sample optimum, so this interval reflects document sampling rather than uncertainty inTTitself\.

The suffix control reinforces this conclusion\. It exhibits non\-monotonic behavior: unrolling the four iterations of the second block alone \(2\.35792\.3579\) is*worse*than unrolling two of them \(1\.93761\.9376\), and worse than unrolling four iterations of the first block \(2\.11782\.1178\)\. Iteration count is not the governing causal variable\.

#### Interpretation\.

On this model and at this training stage, latent recurrent depth itself contributes negligible measurable quality; the effect attributed to depth by naive truncation is dominated by residual statistics, readout calibration, and the engagement of a distinct second reasoning block\. We emphasize this is an empirical finding for a specific undertrained checkpoint, not an indictment of recurrent architectures in general; Section[7](https://arxiv.org/html/2609.19934#S7)shows that a recurrent model trained with sampled depth behaves entirely differently\. Section[5\.4](https://arxiv.org/html/2609.19934#S5.SS4)replicates this full decomposition on a second checkpoint: the ranking of the three primary components is preserved, while distinct depth remains near zero but flips sign—an observation that leads us to frame our conclusion in a more conservative, robust form\.

### 5\.4Replication Across Two Independent Checkpoints

The decomposition above evaluates a single checkpoint, which is vulnerable to the critique that attribution shares may be idiosyncratic to that model\. We replicated the*entire*protocol \(identical held\-out set, identicaln=200n=200documents, identical5,0005\{,\}000bootstrap replications, and identical CLI flags transcribed directly from original run logs\) on a pretraining checkpoint at step40,00040\{,\}000, trained further on the decontaminateden\_math\_v4corpus with repaired document boundaries\.

Figure[3](https://arxiv.org/html/2609.19934#S5.F3)and Table[3](https://arxiv.org/html/2609.19934#S5.T3)present both decompositions side by side\. The new checkpoint shares lineage with the pretraining checkpoint in Section[4](https://arxiv.org/html/2609.19934#S4)but diverged in pretraining data later in training, and is separated from the downstream checkpoint by thousands of training steps\. This serves as a test of*robustness*rather than full independent replication\.

#### A Third Checkpoint, Post\-Reinforcement Learning\.

The replication above spans two pretraining checkpoints of shared lineage\. To evaluate under a wider divergence, we applied the repeat control to a checkpoint that traversed the complete downstream post\-training pipeline: supervised fine\-tuning, preference optimization, and5050steps of RLOO reinforcement learning on GSM8K problems\. This checkpoint differs in data distribution, loss formulation, and optimization algorithm; if distinct depth shares were an artifact of pretraining, they should diverge substantially here\.

Table 2:Repeat control on the post\-reinforcement learning checkpoint\. Total block applications are fixed at88across all columns; only the number of*distinct*iterations varies\. Increasing from one to four distinct iterations shifts NLL by\+0\.0022\+0\.0022nats \(a slight degradation\)\.Table[2](https://arxiv.org/html/2609.19934#S5.T2)presents the results\. Four distinct iterations yield no performance gain over a single repeated iteration; the difference is\+0\.0022\+0\.0022nats towards degradation\. Full decomposition on this checkpoint attributes39\.2%39\.2\\%to block applications,60\.9%60\.9\\%to inter\-block composition, and−0\.14%\-0\.14\\%to distinct depth\. Three checkpoints across three disparate training regimes yield the identical conclusion: under the repeat control, additional distinct iterations buy no measurable performance in this architecture\.

Figure 3:Decomposition of the naive gap into four constituent components across two checkpoints using the identical evaluation protocol\. Error bars denote95%95\\%bootstrap confidence intervals over5,0005\{,\}000document resamplings\. The distinct depth component is near zero in both checkpoints and reverses sign between them\.Table 3:Decomposition across two checkpoints under identical protocol\. The rank ordering and sign of the three dominant components are strictly preserved; distinct depth remains near zero while reversing sign\.
#### Repeat Control\.

The repeat control produces matching qualitative results\. Holding total applications constant at88and varying only the number of*distinct*iterations, NLL is2\.29302\.2930,2\.29312\.2931, and2\.28512\.2851for11,22, and44distinct iterations, respectively\. Expanding distinct depth from11to44alters NLL by merely0\.00790\.0079nats\. On the identical model, increasing*block applications*from11to88while fixing distinct depth to11\(prefixk=1k\{=\}1,2\.92022\.9202, versus repeatk=1k\{=\}1,2\.29302\.2930\) improves NLL by0\.62720\.6272nats—nearly an80×80\\timeslarger shift\.

#### Calibration Step Function Replicates\.

Fitted temperatures are1\.5111\.511,1\.4771\.477, and1\.5051\.505atk=1,2,4k=1,2,4, dropping to1\.0191\.019atk=8k=8; on the downstream checkpoint, they are1\.4471\.447,1\.4831\.483,1\.4351\.435, dropping to0\.9060\.906\. Both profiles remain flat across all truncated depths before dropping sharply to≈1\\approx 1strictly at full depth, matching the diagnostic signature described in Section[9\.1](https://arxiv.org/html/2609.19934#S9.SS1)\.

#### Non\-Replicating Sign and Implications for Confidence Intervals\.

The sign of the distinct depth component reverses:−2\.8%\-2\.8\\%on the downstream checkpoint versus\+0\.5%\+0\.5\\%here, with both confidence intervals strictly excluding zero and not overlapping\.

We argue the proper interpretation is not that one measurement is invalid, but that bootstrap intervals*underestimate*true operational uncertainty\. Bootstrap resampling over documents with a frozen model captures only document sampling variance; it fails to capture checkpoint\-to\-checkpoint variance\. Here, inter\-checkpoint variation is3\.33\.3percentage points—over twenty times the width of either bootstrap interval\. In absolute magnitude, both values are negligible:−0\.0392\-0\.0392and\+0\.0079\+0\.0079nats out of an overall gap of≈1\.45\\approx 1\.45nats\.

We therefore frame our conclusion in a more robust form: across checkpoints, the distinct depth component falls within\[−3%,\+1%\]\[\-3\\%,\+1\\%\], indistinguishable from zero, with an unstable sign\. This claim is more modest than “distinct depth harms performance”, but does not depend upon fragile details and is the sole claim supported by the data\.

#### A Note on Naive Interpretation\.

On the newer checkpoint, automated evaluation routines would conclude that “recurrence works” across all three regimes, because NLL drops monotonically withkkunder prefix truncation \(2\.9202→2\.6338→2\.3937→1\.43902\.9202\\rightarrow 2\.6338\\rightarrow 2\.3937\\rightarrow 1\.4390\)\. DCP decomposition reveals that the share attributable to distinct depth is a meager0\.5%0\.5\\%\. The gulf between these two interpretations constitutes the core thesis of this paper\.

### 5\.5The Gate Learns a Non\-Degenerate Allocation Schedule

Although distinct depth fails to improve quality, the update gate in Equation \([7](https://arxiv.org/html/2609.19934#S4.E7)\) is non\-degenerate\. Table[13](https://arxiv.org/html/2609.19934#A2.T13)\(Appendix\) reports gate distributions per iteration across6060held\-out documents\. The gate regularizer targets a global mean of0\.50\.5; the model satisfies this globally while learning a monotonically increasing schedule, updating more aggressively in later iterations, with substantial within\-iteration variance\.

### 5\.6What a Depth Value Head Can “See”

A natural application for depth\-recurrent architectures is depth\-wise credit assignment: attaching a value head at each iteration and utilizing its incremental gain as an advantage estimate\. We test whether the informational signal required by such a head exists\. For each held\-out document, we construct a corrupted counterpart \(preserving the reasoning trace but perturbing the final answer\), extract latent representations at each iteration, and train a cross\-validated linear probe to discriminate valid from corrupted solutions\.

Probe AUC remains near chance and*degrades*with depth:0\.549→0\.5250\.549\\rightarrow 0\.525\(pretraining\) and0\.548→0\.5350\.548\\rightarrow 0\.535\(fine\-tuning\) from iteration00to77\. This is an architectural consequence rather than a training failure: the input to the verifier isthought\_state\+W​x¯prefix\\text\{thought\\\_state\}\+W\\,\\overline\{x\}\_\{\\text\{prefix\}\}, wherex¯prefix\\overline\{x\}\_\{\\text\{prefix\}\}denotes the*mean*across causal prefix representations, and the thought bank cannot attend into the textual sequence\. A perturbation in a single answer token shifts a multi\-hundred\-token average negligibly\. Consequently, in this architecture, any depth value head must hook into textual stream representations rather than the latent thought bank\.

## 6Negative Control: Layer Truncation on Standard Transformers

This section executes the third component of DCP \(Section[3\.4](https://arxiv.org/html/2609.19934#S3.SS4)\)\. The potential failure mode probed here is not a vulnerability of the model, but a potential artifact of*the measurement protocol itself*: any truncation procedure could conceivably induce confounders, for instance if removing compute systematically shifts residual stream distributions in a uniform direction\. Were this the case, the two confounders quantified in Section[5](https://arxiv.org/html/2609.19934#S5)would reveal nothing about Sona, reflecting merely an artifact of the evaluation suite\.

The negative control discriminates between these two hypotheses, which constitutes its primary advantage over reporting positive controls alone: it provides an empirical baseline where the hypothesized confounder mechanism should not operate; hence, any non\-zero share observed in this group must be attributed to measurement artifacts\. Specifically, if the confounders in Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)were measurement artifacts rather than intrinsic properties of recurrent depth, they should equally manifest when applying identical interventions to layer truncation in standard transformers—the exact evaluation widely employed in layer\-pruning and early\-exit literature\. We therefore replicate the protocol across two public models, truncating to the firstkklayers out ofLL\(*prefix*\) and, as a control, executing the firstkklayers and repeating thekk\-th layer until completingLLapplications \(*repeat*\)\.

#### The Two Components Do Not Transfer Identically\.

The preliminary negative control ran without temperature recalibration, informing solely the application count component\. Re\-evaluating with temperature scaling reveals a fundamental divergence between the two confounders\.

The*block application*component reverses sign in dense networks, as theoretically predicted:−31\.7%\-31\.7\\%on Qwen2\.5\-Math\-1\.5B and−21\.2%\-21\.2\\%on Qwen3\-1\.7B\. This effect is unique to weight\-sharing architectures and does not transfer to dense models\.

The*calibration*component, by contrast, transfers robustly\. On Qwen2\.5\-Math\-1\.5B evaluated on general text where the model is well\-calibrated \(T=1\.03T=1\.03at full depth\), fitted temperature drifts from1\.911\.91atk=4/28k=4/28to1\.031\.03at full depth, accounting for16\.5%16\.5\\%of the apparent gap\. The baselineT=1\.03T=1\.03at full depth is the crucial sanity check: it confirms the model is well\-calibrated on that domain*prior*to truncation, ensuring all subsequent drift is induced by layer removal rather than domain shift\.

We verified this condition by evaluating on specialized mathematics text, where the identical model exhibitsT=1\.12T=1\.12at full depth and a calibration share of only8\.0%8\.0\\%\. The divergence between domains demonstrates why evaluation data must be chosen where the baseline model is well\-calibrated: evaluating on a misaligned domain conflates background drift, paradoxically*attenuating*the measured calibration share\. On Qwen3\-1\.7B, full\-depth temperature is1\.291\.29even on general text; the baseline calibration condition is imperfectly satisfied, so we treat its28\.9%28\.9\\%calibration share as secondary evidence\.

The implications for the broader literature are substantial\. The calibration confounder is not idiosyncratic to recurrent models trained at fixed depth; it manifests fully when pruning layers in standard dense transformers\. Studies interpreting quality curves over retained layer counts without recalibration\([Gromov et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib14);[Men et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib15);[Wang et al\., 2025](https://arxiv.org/html/2609.19934#bib.bib18);[Shi et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib19)\)report a quantity contaminated by this component\. We do not claim their empirical findings are invalid; we show that their reported metric conflates two distinct mechanisms that our control suite successfully disentangles\.

#### Sanity Checks\.

Because custom unrolling bypasses standard library forward passes, we verified that atk=Lk=L, manual layer unrolling bitwise reproduces the standard reference forward pass\. Across both models, the maximum absolute logit divergence is00with100%100\\%argmax consensus, confirming the exactness of our attention masking and rotary embedding implementations\.

Figure 4:ε=NLL​\(truncated\)−NLL​\(full\)\\varepsilon=\\text\{NLL\}\(\\text\{truncated\}\)\-\\text\{NLL\}\(\\text\{full\}\), computed within each respective model to enable cross\-tokenizer comparison\. Log–log scale\. The full\-depth point is omitted sinceε=0\\varepsilon=0by definition\. The two recurrent models remain below0\.20\.2nats across most of the depth range; the two dense transformers exceed11nat even at the mildest measurable truncation\. Data in Tables[14](https://arxiv.org/html/2609.19934#A2.T14)and[15](https://arxiv.org/html/2609.19934#A2.T15)\.Table 4:Cross\-architecture decomposition\. The “application count” confounder reverses sign on standard transformers, and the calibration confounder is substantially attenuated\. Both represent structural characteristics of weight\-tied recurrent depth\.
#### Results\.

Neither confounder transfers unchanged\. The application count share turns negative across both dense models \(−35\.9%\-35\.9\\%and−39\.9%\-39\.9\\%\): repeating a layer degrades standard transformers across all depths, except when within two layers of the full stack\. The calibration share is3\.3%3\.3\\%and14\.5%14\.5\\%, compared to25\.6%25\.6\\%on the depth\-recurrent architecture\.

#### Mechanism\.

In a standard transformer, each layer is an independent function optimized strictly for its specific index within the network stack; repeating an intermediate layer shifts the residual stream severely out of distribution: atk=14/28k=14/28, repetition drives perplexity from1\.3×1041\.3\\times 10^\{4\}to3\.7×1063\.7\\times 10^\{6\}on Qwen2\.5\-Math\. In a depth\-recurrent model, by contrast, the block is explicitly designed for iterative execution, rendering application count a free variable that can be held fixed while manipulating distinct compute\. Consequently, the controls in Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)are not generic prescriptions for all depth truncation; they are tailored diagnostic controls for recurrent models, and this experiment delineates their valid domain of application\.

#### Caveat\.

Absolute NLL values are not cross\-comparable across architectures: truncating a standard transformer to half\-depth induces catastrophic failure \(perplexity\>104\>10^\{4\}\), whereas the recurrent model degrades smoothly\. The comparable quantities are the sign and proportional attribution shares of each decomposed component, not the absolute magnitude of the performance drop\.

## 7Vulnerability Analysis: Training at Fixed Depth

Huginn\-0125\([Geiping et al\., 2026](https://arxiv.org/html/2609.19934#bib.bib2)\)is a 3\.5B parameter depth\-recurrent model comprising22prefix layers, a44\-layer recurrent core, and22suffix layers\. Crucially, its recurrence depth was*sampled during pretraining*\(mean\_recurrence=32=32, following a Poisson–log\-normal distribution\), whereas Sona was trained exclusively at fixed depth88\. Consequently, Huginn’s readout head observed residual states originating from diverse recurrence depths rather than a single fixed depth, providing a direct empirical test of the generative mechanism behind calibration drift\.

#### What Does Not Transfer, and Why\.

Huginn’s recurrent core consists of a*single weight\-tied block*; hence, unrolling it forkkiterations*is*repeating it: application count and distinct compute are identical variables\. The structural decomposition in Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)is therefore*formally undefined*here, rather than merely unmeasured\. Only the calibration control transfers, which we report below\.

#### Results\.

Fitted temperature remains tightly bounded in\[0\.94,0\.99\]\[0\.94,0\.99\]across a32×32\\timesdepth range \(drift amplitude0\.0450\.045\), and recalibration accounts for merely−0\.6%\-0\.6\\%of the1\.841\.84nat depth gap\. By contrast, the fixed\-depth model exhibits a drift amplitude of0\.570\.57across an8×8\\timesrange and a25\.6%25\.6\\%calibration share—an order of magnitude larger drift\.

This finding isolates the underlying mechanism, which Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3)verifies via a controlled training intervention\. The confounder is linked specifically to*training at fixed depth*, not recurrent depth per se: a readout head trained on residuals from a single recurrence depth suffers calibration drift when evaluated at other depths\. This directly suggests a practical remedy: randomly sampling recurrence depth during training\. It likewise explains why standard transformers \(also trained at fixed layer counts\) in Section[6](https://arxiv.org/html/2609.19934#S6)retain modest calibration effects \(3\.3%3\.3\\%and14\.5%14\.5\\%\), whereas the depth\-sampled model shows none\.

#### Two Ancillary Observations\.

Huginn’s performance saturates aroundr≈16r\\approx 16\(1\.2899→1\.28881\.2899\\rightarrow 1\.2888fromr=16r=16tor=32r=32,Δ=0\.0009\\Delta=0\.0009nats\), despite being trained with a mean recurrence of3232\. Moreover, it demonstrates a substantial, monotonic, unconfounded depth effect of1\.831\.83nats\. The appropriate takeaway from Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)is therefore not that “recurrent depth cannot help” \(a well\-trained recurrent model unambiguously leverages its depth\), but that*this specific model*, trained at fixed depth and undertrained, fails to utilize its own depth\.

#### Caveat\.

Huginn and Sona differ in parameter scale \(3\.5B vs\. 542\.8M\), tokenizers, and pretraining corpora; hence, this provides mechanistic evidence rather than a controlled experiment\. The fully controlled counterpart—training the identical architecture twice under fixed versus sampled depth—is executed in Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3)\.

#### Pattern Across Four Models\.

Table[5](https://arxiv.org/html/2609.19934#S7.T5)synthesizes all calibration evaluations across this paper\. Across four models, calibration share tracks whether training employed fixed or sampled depth: two standard transformers and one recurrent model, all trained at a single fixed depth, yield calibration shares between3\.3%3\.3\\%and25\.6%25\.6\\%; the single model trained with sampled depth yields−0\.6%\-0\.6\\%\.

Table 5:Calibration confounder tracks depth variation during training, invariant to architectural family and parameter scale\.

## 8Dynamics of Recurrence

Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)demonstrated that additional distinct iterations fail to improve quality, but did not reveal why\. This section directly tracks the*latent state trajectories*through successive block applications\. All three dynamical measurements require only forward passes without retraining, enabling evaluation across all models in this study\.

#### Gate Magnitudes Do Not Measure What They Seem\.

Table[13](https://arxiv.org/html/2609.19934#A2.T13)reports update gate valuesγ\(g\)\\gamma^\{\(g\)\}, increasing monotonically from0\.340\.34to0\.700\.70\. However,γ\\gammagoverns only the*mixture ratio*between block output and residual stream; if block output is aligned with incoming residual states, the hidden state remains nearly invariant even with a wide\-open gate\. The proper metric is relative update magnitude‖x\(g\)−x\(g−1\)‖/‖x\(g−1\)‖\\\|x^\{\(g\)\}\-x^\{\(g\-1\)\}\\\|/\\\|x^\{\(g\-1\)\}\\\|, which on the fixed\-depth recurrent model is*non\-monotonic*:0\.260\.26,0\.350\.35,0\.400\.40,0\.330\.33,0\.290\.29,0\.390\.39,0\.520\.52,0\.410\.41\. The gate schedule and true update trajectory tell divergent stories\.

#### Iterations Compute Distinct Transformations\.

Stacking normalized updatesΔ​x\(g\)\\Delta x^\{\(g\)\}into a matrix and evaluating the participation ratio\(∑iσi\)2/∑iσi2\(\\sum\_\{i\}\\sigma\_\{i\}\)^\{2\}/\\sum\_\{i\}\\sigma\_\{i\}^\{2\}of singular values provides a soft estimate of the subspace dimensionality spanned\. The recurrent model achieves6\.246\.24out of a maximum88\(78%78\\%\), andcos⁡\(Δ​x\(g\),Δ​x\(1\)\)\\cos\(\\Delta x^\{\(g\)\},\\Delta x^\{\(1\)\}\)drops from0\.550\.55to0\.060\.06by the fourth iteration\. The iterations are non\-redundant: they execute geometrically distinct transformations\.

Juxtaposed with the repeat control, this sharpens our primary finding rather than explaining it away\. The model computes eight largely non\-redundant transformations, yet replacing four distinct iterations with a single repeated iteration alters NLL by merely0\.00790\.0079nats\. Computational diversity*exists internally*but is*not decoded by the readout*\.

#### Perturbation Response\.

We inject Gaussian noise with norm matching10%10\\%of state norm into an intermediate application, tracking‖δ\(g\)‖/‖δ\(0\)‖\\\|\\delta^\{\(g\)\}\\\|/\\\|\\delta^\{\(0\)\}\\\|across subsequent applications\. This represents the sole metric among the three directly comparable between recurrent and dense models, as it is scale\-invariant: in dense networks, early layers operate on unnormalized embeddings where‖Δ​x‖/‖x‖\\\|\\Delta x\\\|/\\\|x\\\|reaches2525–3333, on a completely different scale from recurrent steps\.

Table 6:Perturbation response\. A ratio below11indicates perturbation attenuation and contraction towards a fixed point; above11indicates perturbation amplification\. The first four configurations expand perturbations at comparable rates, irrespective of recurrent or dense architecture\. Huginn is the sole exception\. Evaluated on100100documents for Sona,4040–6060for Huginn and dense transformers\.Figure 5:\(a\)Per\-iteration update norms, weight\-tied models only; dense models differ in scale\. In Huginn, update norms decay exponentially from0\.920\.92to0\.0190\.019; in Sona, they do not attenuate\.\(b\)Perturbation response on log scale across all five configurations\. Only Huginn lies below the unity threshold\.
#### Contractivity Does Not Track Depth Schedule\.

Prior to evaluating dense transformers, a natural hypothesis was to attribute contractivity to sampled\-depth training, since Huginn possesses both traits\. Two empirical measurements refute this reading\.

First, the depth\-sampling intervention in Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3), which reduces calibration share from29\.9%29\.9\\%to5\.1%5\.1\\%on this identical model, shifts the noise ratio only from1\.0691\.069to1\.0531\.053\(remaining\>1\>1\)\. Depth sampling remedies calibration drift but*fails*to alter operator dynamics; the two mechanisms are orthogonal\.

Second, both dense transformers also amplify perturbations \(1\.0111\.011and1\.0601\.060\), with Qwen3\-1\.7B compounding perturbations over80×80\\timesacross2828layers\. Perturbation amplification is a pervasive characteristic of deep layer stacks, not an idiosyncratic defect of our evaluated model: it falls squarely within the range of fully converged production transformers\.

The accurate characterization is therefore that*Huginn learned a contraction mapping while the remaining four configurations did not*, and we*cannot*attribute this dynamical property to the depth schedule\. Potential explanatory factors include model scale \(3\.6B vs\. 543M and 1\.5–1\.7B\), pretraining token volume, or recurrent core architecture\. We report the empirical observation while leaving the causal origin open\.

#### Explaining Huginn’s Empirical Resilience\.

Section[6](https://arxiv.org/html/2609.19934#S6)observed that truncating Huginn to half\-depth incurs merely0\.00180\.0018nats, without an obvious explanation\. The dynamical data provides a direct answer: beyond iteration1616, update magnitudes fall below0\.030\.03\. Truncating that tail is virtually free because negligible computation is discarded\. In the fixed\-depth model, by contrast, truncation*does*discard substantial computation, yet performance still fails to improve\.

#### Limitations\.

These five configurations differ simultaneously across parameter scale, tokenizers, pretraining corpora, and optimization recipes; hence, cross\-model comparisons remain observational\. The sole strictly interventional comparison is the pair of\+1,000\+1\{,\}000steps \(fixed vs\. sampled depth\), initiated from the identical checkpoint with matching data and random seeds, which demonstrates that depth schedule is*not*the governing variable for contractivity\. We also do not know how this coefficient evolves with extended training: our evaluated checkpoint reached40,00040\{,\}000out of110,000110\{,\}000planned steps\. Because this measurement takes roughly two minutes, it represents an inexpensive diagnostic metric to track across future training milestones\.

## 9Direct Verification and Generality

The preceding sections established a correlational pattern: the calibration confounder manifests in fixed\-depth models and vanishes in sampled\-depth models\. This section executes the fourth component of DCP \(Section[3\.5](https://arxiv.org/html/2609.19934#S3.SS5)\)—a controlled training intervention—to establish causal attribution, accompanied by two tests of empirical generality\.

### 9\.1Calibration Drift is a Step Function, Not a Gradual Drift

Figure[6](https://arxiv.org/html/2609.19934#S9.F6)demonstrates that fitted temperature does not drift gradually with depth\. It remains virtually constant in the range1\.381\.38–1\.631\.63across all truncated depths, before dropping sharply to≈1\.00\\approx 1\.00strictly at full depth\. BecauseT=1\.0T=1\.0indicates that*no calibration adjustment is needed*, the proper reading is: the readout head is calibrated accurately at the precise depth it was trained upon, and miscalibrated to an approximately uniform degree at all other depths\. This pattern holds across all four evaluated configurations, including out\-of\-domain text\.

Figure 6:Fitted logit temperature as a function of retained depth fraction\. In fixed\-depth models,TTremains nearly constant across all truncated depths before dropping sharply to≈1\\approx 1strictly at full depth\. In sampled\-depth models,TThovers around11across the entire range\. Data in Table[16](https://arxiv.org/html/2609.19934#A2.T16)\.
### 9\.2Confounders Are Not Confined to Mathematical Text

Replicating the entire evaluation across200200non\-mathematical documents \(a FineWeb slice of the identical corpus\) yields a calibration share of41\.4%41\.4\\%, compared to29\.9%29\.9\\%on mathematical text on the same checkpoint\. The confounder not only persists out\-of\-domain, but is*amplified*; hence, it is not an idiosyncratic artifact of mathematical syntax\.

### 9\.3Depth Sampling Intervention

To causally verify the generative mechanism identified in Section[7](https://arxiv.org/html/2609.19934#S7)on the identical model exhibiting confounding, we trained three continued branches from the same checkpoint, using matching data, seeds, step budget \(1,0001\{,\}000steps\), and constant learning rate\. The sole divergence between branches is the pretraining depth schedule\.

#### Sampling Distribution Must Match Evaluated Deployment Space\.

Our first attempt sampled recurrence iterations*within each block independently*; hence, the model always executed both blocks, and total block applications strictly took values in\{2,4,6,8\}\\\{2,4,6,8\\\}\. During evaluation, prefix truncation withk≤4k\\leq 4executes exclusively the first block: configurations evaluated atk=1,2,4k=1,2,4*never occurred*during pretraining\. As theoretically predicted, calibration share did not attenuate \(Figure[7](https://arxiv.org/html/2609.19934#S9.F7)and Table[7](https://arxiv.org/html/2609.19934#S9.T7), row v1\)\. Our second attempt sampledk∼𝒰​\{1\.\.8\}k\\sim\\mathcal\{U\}\\\{1\.\.8\\\}*total block applications*, executing the exact prefix unrolling of lengthkk—matching the configuration family probed during evaluation\.

Figure 7:Depth sampling intervention\. Left: NLL versus retained iterations\. Right: Fitted temperature\. The sampled branch flattens the temperature profile; the control branch receiving the identical training budget at fixed depth does not\.Table 7:Depth sampling intervention,n=200n=200documents\. Branch v2 samples across the deployment schedule probed during evaluation and collapses the calibration confounder; the control branch, receiving an identical1,0001\{,\}000steps, remains virtually unchanged from baseline\.
#### Results\.

Branch v2 precisely reproduces the diagnostic signature of sampled\-depth models: fitted temperature remains bounded in\[1\.029,1\.069\]\[1\.029,1\.069\]across the entire range \(drift amplitude0\.0400\.040, matching the0\.0450\.045of Huginn\-0125 and compared to0\.4530\.453in this model prior to intervention\)\. Calibration share collapses from29\.9%29\.9\\%to5\.1%5\.1\\%\. Because the control branch received the identical1,0001\{,\}000training steps yet shifted only from29\.9%29\.9\\%to27\.3%27\.3\\%, the causal driver is unambiguously the depth schedule rather than additional optimization\.

This represents causal evidence on the identical model, identical data, and identical compute budget, perturbing a single operational variable\. It elevates the finding in Section[7](https://arxiv.org/html/2609.19934#S7)from a cross\-model observation to a verified causal relationship\.

#### Two Trade\-Offs, and One Implication\.

First, full\-depth performance degrades: NLL atk=8k=8increases from1\.44211\.4421to1\.57971\.5797\. Multi\-depth training trades peak performance for robust generalizability across depths, as theoretically expected\. Second, the apparent depth gap largely vanishes \(1\.3955→0\.12241\.3955\\rightarrow 0\.1224nats\): post\-intervention, executing one application trails eight applications by merely0\.120\.12nats\. Coupled with the finding in Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)that distinct depth contributes−2\.8%\-2\.8\\%, this reinforces the conclusion that recurrence in this architecture buys negligible performance: calibration confounders can be remedied, but the residual effect size is minimal\.

#### Limitations\.

The intervention spanned1,0001\{,\}000steps on a single model, single seed, initiated from an undertrained checkpoint\. We have not established whether training with sampled depth from initialization avoids the peak performance penalty\.

## 10Benchmarking the Protocol on Known Standards

All preceding results applied DCP to models where the ground\-truth attribution was*unknown a priori*\. This leaves open a critical methodological challenge: what guarantees the protocol measures ground truth rather than generating internally consistent numbers? This section addresses this by constructing three models where the ground truth is known*by construction*, verifying whether DCP faithfully recovers the known true states\.

#### Experimental Design\.

Three branches share identical architecture, corpora, random seeds, and training step budgets\. The sole manipulated variable is the pretraining depth schedule:

Branches A and B test*sensitivity*: does the protocol detect the confounder when present by construction, and remain silent when absent? Branch C tests*specificity*, representing the most discriminative test\. It disentangles two competing hypotheses that Branch B alone cannot separate: “whether depth was sampled” versus “whether the readout has observed the specific runtime configurations being evaluated”\. If our hypothesized mechanism holds, Branch C must exhibit confounding at depths outside its training support, despite being trained under variable depth\.

#### Pre\-Registered Predictions\.

The mechanistic theory makes a sharp structural prediction: in Branch C, fitted temperature must hover near1\.01\.0atk=5,…,8k=5,\\dots,8and drift atk=1,2,4k=1,2,4—namely,*the step discontinuity must align with the sampling support boundary*\. This is a prediction regarding the*spatial location*of a discontinuity, rather than a loose difference in signs\. That boundary is not fitted to empirical data; it is determined by the pretraining schedule prior to evaluation\. The decision threshold for calibration share was fixed at5%5\\%, midway between−0\.6%\-0\.6\\%on Huginn\-0125 and25\.6%25\.6\\%on fixed\-depth models\. All three predictions and thresholds were hard\-coded into synthesis scripts prior to executing the branches\.

#### Results\.

Table[8](https://arxiv.org/html/2609.19934#S10.T8)and Figure[8](https://arxiv.org/html/2609.19934#S10.F8)present the complete calibration grid\. All three branches strictly conform to pre\-registered predictions\.

Table 8:Calibration grid\. Three branches share identical architecture, corpora, seeds, and training step budgets, differing strictly in depth schedule\. The pre\-registered5%5\\%threshold for calibration share was established prior to run execution\. At maximum training budget: Branch A yields23\.5%23\.5\\%\(predicted YES\), Branch B yields−3\.7%\-3\.7\\%\(predicted NO\), Branch C yields11\.8%11\.8\\%\(predicted YES\)\.Figure 8:Calibration grid across three models differing in a single operational variable: pretraining depth schedule\.\(a\)Calibration share over training budget\. Branch A steadily ascends, Branch B hovers around zero, Branch C lies in between\. Confounders accumulate with training rather than being warm\-up artifacts\.\(b\)Temperature drift amplitudes mirror this ranking; Branch B matches the0\.0450\.045observed in Huginn\-0125\.\(c\)Fitted temperature per depth at step6,0006\{,\}000\. Branch A exhibits a step function, dropping to≈1\\approx 1strictly at training depth\. Branch B remains flat at≈1\\approx 1\. Branch C drops as it enters its trained support \(shaded region\)\. Because the grid samplesk∈\{1,2,4,8\}k\\in\\\{1,2,4,8\\\}, the exact boundary transition profile is unresolved\.Table[9](https://arxiv.org/html/2609.19934#S10.T9)decomposes the four components with paired95%95\\%bootstrap confidence intervals at step6,0006\{,\}000\. We pair the resampling: re\-drawing documents once per bootstrap replicate \(n=200n=200documents,5,0005\{,\}000iterations\) and propagating that exact sample across all conditions\. All twelve confidence intervals exclude zero\.

Table 9:Calibration grid decomposition with paired95%95\\%bootstrap confidence intervals over200200documents and5,0005\{,\}000iterations\. Raw distances:1\.29761\.2976nats for Branch A,0\.05740\.0574for Branch B, and1\.23021\.2302for Branch C\. Twelve out of twelve intervals exclude zero\.The distinct iteration component*reverses sign*depending on the pretraining schedule, and both signs are statistically reliable\. In the fixed\-depth branch, it is negative:−4\.81%\-4\.81\\%\[−5\.21,−4\.40\]\[\-5\.21,\-4\.40\]; in the fully sampled branch, it is positive:\+4\.81%\+4\.81\\%\[\+4\.24,\+5\.42\]\[\+4\.24,\+5\.42\]\. The two confidence intervals do not overlap\. Branch C, which was sampled but omitted the measured evaluation depths, sits intermediate and remains negative\.

We state this finding strictly within the bounds of what the data support\. What has been established is that the*sign*of the depth contribution is governed by the pretraining schedule, rather than being an intrinsic property of the recursive architecture\. What has*not*been established is that depth purchases substantial quality when trained properly: the raw distance of Branch B is only0\.05740\.0574nats, so\+4\.81%\+4\.81\\%corresponds to roughly0\.00280\.0028nats in absolute terms, compared to0\.06240\.0624nats for the negative component in Branch A\. The sign reversal is real; its magnitude is modest\.

Because all three branches match pre\-registered predictions, DCP transitions from a merely self\-consistent diagnostic protocol into a calibrated instrument evaluated against known ground\-truth standards, with measurable sensitivity and specificity\. Had any branch deviated, we would have reported it accordingly: thresholds and predictions were pre\-registered, leaving no degrees of freedom for post\-hoc adjustments\.

## 11Future Work

#### Pretraining with Sampled Depth from Scratch\.

Section[9\.3](https://arxiv.org/html/2609.19934#S9.SS3)performed an intervention on a checkpoint pretrained at fixed depth, incurring a performance penalty at full depth\. Whether training with sampled depth from the initial step avoids this penalty remains to be systematically verified\. Furthermore, the four\-model comparison in Section[7](https://arxiv.org/html/2609.19934#S7)remains observational: the depth\-sampled model differed from the other three models simultaneously in parameter scale, tokenizer, pretraining corpus, and training recipe\. A definitive controlled experiment would pretrain a single architecture twice on identical data, varying solely whether recurrence depth is fixed or sampled, and then compare the calibration share between branches\. At small scale, this requires several GPU\-days and directly separates observational correlations from causal effects of depth schedules\. We report observational results here because they reflect available computational resources, making this limitation explicit rather than leaving it to reader inference\.

#### Does Depth Utilization Emerge Over Training?

Both checkpoints evaluated here are substantially undertrained relative to modern standards; hence, the conclusion that distinct iterations contribute near zero may reflect the training phase rather than architectural limits\. A natural test is to evaluate the complete control suite across checkpoints along an extended pretraining trajectory and trace the repeat\-control distance over training steps\. A widening distance signifies emerging depth utilization; a flat trajectory suggests that recurrence remains unexploited at this scale\.

#### Is the Architecture Computationally Competitive?

Counting parameters along the computation path, eight block applications consume1\.4701\.470GFLOPs/token in the forward pass compared to1\.1841\.184GFLOPs/token for the same parameter set applied once, a ratio of1\.24×1\.24\\times\. A non\-recurrent FLOP\-matched baseline would therefore require only1\.24×1\.24\\timesthe token count\. This paper makes no claim regarding architectural superiority, and none of our findings depend on such a claim; a FLOP\-matched baseline is a prerequisite for asserting superiority, which we explicitly refrain from doing\.

## 12Limitations

The limitations below pertain to two distinct subjects, which we explicitly separate: limitations regarding the*scope of DCP*\(the proposed diagnostic protocol\), and limitations regarding the*empirical conclusions*drawn from applying DCP to the five evaluated configurations\. The formal conditions of applicability and the five named limitations of the protocol itself are detailed in Section[3\.3](https://arxiv.org/html/2609.19934#S3.SS3)and not repeated here\.

- •The confirmatory positive depth measurements were performed on a single recursive architecture at a single parameter scale \(542\.8542\.8M\), evaluated across three checkpoints from the same pretraining lineage\. Section[6](https://arxiv.org/html/2609.19934#S6)bounds the claim from below using two standard Transformers, but evaluating a second*depth\-recurrent*architecture remains essential \(Section[11](https://arxiv.org/html/2609.19934#S11)\)\.
- •All three checkpoints are substantially undertrained relative to modern compute standards\. The finding that distinct iterations contribute≈0\\approx 0may reflect the training stage rather than architectural capability\.
- •The bootstrap confidence intervals in this paper capture document sampling noise, but do not account for variance across checkpoints or random training seeds\. Section[5\.4](https://arxiv.org/html/2609.19934#S5.SS4)demonstrates that checkpoint\-to\-checkpoint variance can exceed the confidence interval width by twenty\-fold; thus, all reported intervals must be interpreted as lower bounds on true uncertainty\.
- •Depth is partially conflated with block identity: with22reasoning blocks iterated44times each,k≤4k\\leq 4only exercises the first block\. An intra\-block ablation schedule would disentangle these factors\.
- •Quality is measured via teacher\-forced NLL and next\-token accuracy rather than downstream generation task accuracy\. Generative evaluations were omitted because the available checkpoints achieve only single\-digit accuracy on targeted benchmarks, where high variance overwhelms diagnostic signal\.
- •The four\-model comparison in Section[7](https://arxiv.org/html/2609.19934#S7)is observational: the depth\-sampled model also differs in scale, tokenizer, and training corpus\. Section[11](https://arxiv.org/html/2609.19934#S11)outlines a controlled counterfactual experiment that remains to be executed\.
- •Evaluations are conducted exclusively on held\-out mathematical text\. Whether calibration share varies across diverse domain distributions remains unverified\.
- •The dynamical measurements in Section[8](https://arxiv.org/html/2609.19934#S8)compare five configurations differing concurrently in scale, tokenization, and training distribution\. We established that depth schedules*do not*explain operator contractivity, but what mechanisms do explain it remains unidentified\. We also do not know how this coefficient evolves under longer training\.

## 13Conclusion

On the recursive language model investigated here, naive depth truncation overestimates the genuine contribution of depth by approximately a factor of two\. Roughly47\.9%47\.9\\%of the apparent performance gap stems from repeated block applications, and an additional25\.6%25\.6\\%arises from readout calibration drift at depths unseen during training\. After controlling for both factors, additional distinct iterations within a reasoning block contribute−2\.8%\-2\.8\\%\[−3\.2,−2\.4\]\[\-3\.2,\-2\.4\]\.

This effect does not manifest across all recurrent models\. Layer pruning on standard Transformers produces no such confounding artifact, nor does a recursive model whose depth was sampled during pretraining\. Across four distinct models, calibration share systematically tracks whether the readout head observed variable depths during training\.

We verified this mechanism via a direct training intervention:1,0001\{,\}000steps of depth\-sampled continual training reduced calibration share from29\.9%29\.9\\%to5\.1%5\.1\\%on the very model where confounding was observed, whereas an identical compute budget of fixed\-depth training left calibration drift intact\. Sampled\-depth training incurs zero inference overhead\. Crucially, it only eliminates confounding when the training distribution spans the exact configuration family evaluated at inference; our initial attempt violated this condition and proved ineffective\. All positive and negative controls introduced in this paper are computationally lightweight and require no retraining\.

These findings do not imply that recursive depth is unhelpful\. The depth\-sampled model in our evaluation demonstrates substantial depth benefits free of calibration confounding\. What our findings demonstrate is that depth truncation, when applied to a model trained at fixed depth, substantially overestimates the genuine contribution of depth\.

## Reproducibility

All experimental measurements are generated by the scripts accompanying this paper:exp\_depth\_ablation\.py\(Section[5](https://arxiv.org/html/2609.19934#S5)\),exp\_depth\_ablation\_hf\.py\(Section[6](https://arxiv.org/html/2609.19934#S6)\),exp\_depth\_huginn\.py\(Section[7](https://arxiv.org/html/2609.19934#S7)\),exp\_decomposition\_ci\.py\(confidence intervals\), andcalibration\.py\(shared temperature fitting\)\. Each script outputs a self\-containedresults\.jsonrecording configurations, checkpoint steps, and raw per\-iteration evaluation metrics\.

## References

- Baninoet al\.\(2021\)A\. Banino, J\. Balaguer, and C\. BlundellPondernet: learning to ponder\.arXiv preprint arXiv:2107\.05407\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.19934#S4.SS0.SSS0.Px1.p2.2)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px8.p1.1),[§4](https://arxiv.org/html/2609.19934#S4.SS0.SSS0.Px3.p1.1)\.
- Dehghaniet al\.\(2018\)M\. Dehghani, S\. Gouws, O\. Vinyals, J\. Uszkoreit, and Ł\. KaiserUniversal transformers\.arXiv preprint arXiv:1807\.03819\.Cited by:[§1](https://arxiv.org/html/2609.19934#S1.p1.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p1.1)\.
- Elbayadet al\.\(2019\)M\. Elbayad, J\. Gu, E\. Grave, and M\. AuliDepth\-adaptive transformer\.arXiv preprint arXiv:1910\.10073\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px5.p1.1)\.
- Fanet al\.\(2019\)A\. Fan, E\. Grave, and A\. JoulinReducing transformer depth on demand with structured dropout\.arXiv preprint arXiv:1909\.11556\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px6.p1.1)\.
- Geipinget al\.\(2026\)J\. Geiping, S\. McLeish, N\. Jain, J\. Kirchenbauer, S\. Singh, B\. Bartoldson, B\. Kailkhura, A\. Bhatele, and T\. GoldsteinScaling up test\-time compute with latent reasoning: a recurrent depth approach\.Advances in Neural Information Processing Systems38,pp\. 41340–41391\.Cited by:[§1](https://arxiv.org/html/2609.19934#S1.p1.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p1.1),[§7](https://arxiv.org/html/2609.19934#S7.p1.1)\.
- Godin \(2026\)G\. GodinSCORE: replacing layer stacking with contractive recurrent depth\.arXiv preprint arXiv:2603\.10544\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px2.p1.1)\.
- Graves \(2016\)A\. GravesAdaptive computation time for recurrent neural networks\.arXiv preprint arXiv:1603\.08983\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.19934#S4.SS0.SSS0.Px1.p2.2)\.
- Gromovet al\.\(2025\)A\. Gromov, K\. Tirumala, H\. Shapourian, P\. Glorioso, and D\. A\. RobertsThe unreasonable ineffectiveness of the deeper layers\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 81906–81920\.Cited by:[§1](https://arxiv.org/html/2609.19934#S1.p2.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.19934#S6.SS0.SSS0.Px1.p5.1)\.
- Guoet al\.\(2017\)C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. WeinbergerOn calibration of modern neural networks\.InInternational Conference on Machine Learning \(ICML\),pp\. 1321–1330\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px7.p1.1),[§3\.1](https://arxiv.org/html/2609.19934#S3.SS1.SSS0.Px3.p1.1)\.
- Haoet al\.\(2024\)S\. Hao, S\. Sukhbaatar, D\. Su, X\. Li, Z\. Hu, J\. Weston, and Y\. TianTraining large language models to reason in a continuous latent space\.arXiv preprint arXiv:2412\.06769\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px4.p1.1)\.
- Hegazyet al\.\(2026\)A\. Hegazy, A\. Alanwar, and M\. ElhoushiRecurrentGPT: expressive depth through recurrent modulation in transformers\.arXiv preprint arXiv:2608\.15062\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p2.1)\.
- Huanget al\.\(2016\)G\. Huang, Y\. Sun, Z\. Liu, D\. Sedra, and K\. Q\. WeinbergerDeep networks with stochastic depth\.InEuropean conference on computer vision,pp\. 646–661\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px6.p1.1)\.
- Lam\-Muir \(2026\)S\. Lam\-MuirThe ignition is real, and it lives at the readout: latent composition, difficulty\-clocked ignition, and the interface\-constituted commit in a recurrent\-depth reasoner\.arXiv preprint arXiv:2608\.03263\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px3.p2.1)\.
- Liet al\.\(2026\)S\. Li, Y\. Zhang, J\. Guo, Q\. Gu, and M\. WangDeepLoop: depth scaling for looped transformers\.arXiv preprint arXiv:2607\.13491\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p2.1)\.
- Linet al\.\(2026\)R\. Lin, Y\. Guo, R\. Zhu, H\. Ye, and J\. K\. EshraghianAllocating recurrent compute in looped language models\.arXiv preprint arXiv:2608\.18230\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p2.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px3.p1.1)\.
- Menet al\.\(2025\)X\. Men, M\. Xu, Q\. Zhang, Q\. Yuan, B\. Wang, H\. Lin, Y\. Lu, X\. Han, and W\. ChenShortgpt: layers in large language models are more redundant than you expect\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 20192–20204\.Cited by:[§1](https://arxiv.org/html/2609.19934#S1.p2.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.19934#S6.SS0.SSS0.Px1.p5.1)\.
- Schusteret al\.\(2022\)T\. Schuster, A\. Fisch, J\. Gupta, M\. Dehghani, D\. Bahri, V\. Tran, Y\. Tay, and D\. MetzlerConfident adaptive language modeling\.Advances in Neural Information Processing Systems35,pp\. 17456–17472\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px7.p1.1)\.
- Shapiro \(2026\)M\. ShapiroRetrofitting recurrent depth into a pretrained language model: installation, extrapolation, transfer, and retention at two parameter budgets\.arXiv preprint arXiv:2608\.11233\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px1.p2.1)\.
- Shiet al\.\(2026\)B\. Shi, C\. Liu, C\. Gao, X\. Yang, and X\. GengUnderstanding performance collapse in layer\-pruned large language models via decision representation transitions\.arXiv preprint arXiv:2605\.07271\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.19934#S6.SS0.SSS0.Px1.p5.1)\.
- Suet al\.\(2024\)J\. Su, Y\. Lu, S\. Pan, A\. Murtadha, B\. Wen, and Y\. LiuRoFormer: enhanced transformer with rotary position embedding\.Neurocomputing568,pp\. 127063\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px8.p1.1),[§4](https://arxiv.org/html/2609.19934#S4.SS0.SSS0.Px1.p2.1)\.
- Tibshirani and Efron \(1993\)R\. J\. Tibshirani and B\. EfronAn introduction to the bootstrap\.Monographs on statistics and applied probability57\(1\),pp\. 1–436\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px8.p1.1),[§3\.2](https://arxiv.org/html/2609.19934#S3.SS2.p2.1)\.
- Touvronet al\.\(2023\)H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar, A\. Rodriguez, A\. Joulin, E\. Grave, and G\. LampleLLaMA: open and efficient foundation language models\.arXiv\.Note:arXiv:2302\.13971 \[cs\]Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px8.p1.1),[§4](https://arxiv.org/html/2609.19934#S4.SS0.SSS0.Px3.p1.1)\.
- Viakhirevet al\.\(2026\)I\. Viakhirev, K\. Borodin, A\. Almutairi, S\. Barannikov, M\. Abramov, and G\. MkrtchianThink shallow, solve deep: controlling recurrent dynamics for reliable test\-time depth\.arXiv preprint arXiv:2608\.18222\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px2.p2.1)\.
- Wanget al\.\(2025\)K\. Wang, T\. Lyu, G\. Su, L\. Yin, M\. Canini, J\. Geiping, and S\. LiuWhen fewer layers break more chains: layer pruning harms test\-time scaling in llms\.arXiv preprint arXiv:2510\.22228\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.19934#S6.SS0.SSS0.Px1.p5.1)\.
- Zelikmanet al\.\(2024\)E\. Zelikman, G\. Harik, Y\. Shao, V\. Jayasiri, N\. Haber, and N\. D\. GoodmanQuiet\-star: language models can teach themselves to think before speaking\.arXiv preprint arXiv:2403\.09629\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px4.p1.1)\.
- Zhanget al\.\(2026\)T\. Zhang, J\. Hu, Y\. Peng, and T\. XieWhen does recurrence become an algorithm? convergence selection in weight\-tied looped transformers\.arXiv preprint arXiv:2607\.20594\.Cited by:[§2](https://arxiv.org/html/2609.19934#S2.SS0.SSS0.Px2.p1.1)\.

## Appendix AArchitectural Details

Table 10:Configuration of the investigated model\. Parameter breakdown:409\.0409\.0M in perception layers,69\.269\.2M in reasoning blocks,49\.249\.2M in tied embeddings, and64\.664\.6M auxiliary \(latent thought bank, tree memory, bridges, gates\)\.
## Appendix BDetailed Metric Tables

The tables below provide the numerical data underlying Figures[2](https://arxiv.org/html/2609.19934#S5.F2),[4](https://arxiv.org/html/2609.19934#S6.F4),[6](https://arxiv.org/html/2609.19934#S9.F6), and[7](https://arxiv.org/html/2609.19934#S9.F7)\.

Table 11:Naive prefix truncation\. Both checkpoints exhibit smooth, monotonic improvement with depth\. Section[3\.1](https://arxiv.org/html/2609.19934#S3.SS1)demonstrates that the majority of this improvement is unattributable to genuine depth\.Table 12:Held\-out NLL under three controls on the fine\-tuned checkpoint \(n=200n=200documents,125,100125\{,\}100scored tokens per cell\)\. In the*repeat*column, block applications are held constant at88while distinct iteration count varies; this column is essentially flat fork∈\{1,2,4\}k\\in\\\{1,2,4\\\}\.Table 13:Per\-iteration update gate statistics\. Gate values increase monotonically and cross\-layer variance remains substantial, indicating that the model actively modulates write intensity across depth, even though this modulation fails to translate into measurable quality improvements\.Table 14:Held\-out NLL under layer pruning across both baseline models \(L=28L=28layers,n=200n=200documents\)\. The*repeat*column holds total layer applications atLLwhile varying distinct compute\. In contrast to depth\-recurrent architectures, repeating layers is consistently*worse*than simple truncation, unlesskkis within two layers of full depth\.Table 15:Huginn\-0125 evaluated across a32×32\\timesrecursive depth sweep \(n=200n=200held\-out documents, sharing identical evaluation documents andmax\_lenwith all other experiments\)\. Fitted temperature remains flat within a0\.0450\.045band; recalibration shifts depth distance by merely−0\.6%\-0\.6\\%\.Table 16:Fitted temperature across depth \(n=200n=200documents;95%95\\%bootstrap CI width<0\.015<0\.015across all cells\)\. Values near≈1\.0\\approx 1\.0appear exclusively atk=8k=8, exactly matching the pretraining depth\.
## Appendix CReproduction Commands

```
# Naive truncation + gate + probing
python exp_depth_ablation.py --ckpt <CKPT> \
    --n_teacher 200 --n_probe 150 --n_gate 60 --iters 1,2,4,8

# Controls
python exp_depth_ablation.py --ckpt <CKPT> --mode suffix   ...
python exp_depth_ablation.py --ckpt <CKPT> --mode repeat   ...
python exp_depth_ablation.py --ckpt <CKPT> --calibrate     ...

# Negative controls on standard Transformers
python exp_depth_ablation_hf.py --model Qwen/Qwen2.5-Math-1.5B \
    --n_docs 200 --fracs 0.25,0.5,0.75,0.86,0.93,1.0 --calibrate

# Pretraining corpus audit + document lengths
python verify_pretrain_bin.py --bin <TRAIN_BIN>
```

Similar Articles

Prediction Dynamics in Depth-Recurrent Language Models

arXiv cs.CL

This paper analyzes the dynamics of predictions in depth-recurrent language models, deriving a decomposition of score changes to explain how intermediate latent updates can preserve final answers despite fluctuations in scores.

Recursive Language Models Generalize Out of Domain

arXiv cs.CL

The paper demonstrates that recursive language models improve out-of-domain generalization by isolating subtasks, avoiding shortcuts used by standard Chain of Thought, and highlighting limitations in classical learning theory for true reasoning.