Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training
摘要
The paper introduces CD-RFT, a method to decouple the shared control bottleneck in RL post-training by regularizing a novel control coefficient, improving multi-task capability on models like Qwen2.5-7B and Llama-3.2-3B.
查看缓存全文
缓存时间: 2026/08/11 08:10
# Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training
Source: [https://arxiv.org/html/2608.08224](https://arxiv.org/html/2608.08224)
Binwen Tan1\\equalcontrib, Jingchao Wang3\\equalcontrib, Dengzhe Hou1,2, Lingyu Jiang1, Zeyuan Wu4, Yunhan Shen1, Fangzhou Lin5,6, Kazunori Yamada1,2, Atsushi Koike1
###### Abstract
Reinforcement learning post\-training unlocks complex reasoning in large language models\. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits a finite capability across competing tasks\. A representative line of work in mechanistic interpretability attributes the success of reinforcement\-learning fine\-tuning to stronger and more diverse circuit activation\. Complementing this view, we separate activation from control: an activated circuit need not control the post\-training reward gain\. Adapting Metabolic Control Analysis, we define the Post\-training Control Coefficient to measure component control over the reward gain and arrange these coefficients by task family into a control matrix, paired with an activation\-magnitude matrix\. We call cross\-task control concentration the Shared Control Bottleneck and the difference between activation and control concentration the Activation–Control Gap\. We show that highly shared activations can coexist with task\-specific control, while a small gap indicates that control concentrates along a shared direction and loses task specificity\. To reduce this concentration, we regularize the post\-training loss with the Shared Control Bottleneck and propose Control\-Diverse Reinforcement Fine\-Tuning \(CD\-RFT\)\. The exact regularizer gradient requires second\-order automatic differentiation incompatible with flash attention, so we derive a first\-order proxy with worst\-case overhead below eight percent\. On Qwen2\.5\-7B, CD\-RFT achieves the largest control decoupling and improves multi\-task capability over its matched GRPO recipe across all three domains\. The no\-KL variant leads on pass@11, and the KL\-penalized variant leads on the large\-kkpass@kkcoverage that KL otherwise degrades\. Together, these results show that the Shared Control Bottleneck serves as a mechanistic diagnostic and a training regularizer, and that both control decoupling and capability gains transfer to Llama\-3\.2\-3B\.
## 1Introduction
Figure 1:The control view at a glance\.Top:task\-wise reward fluxJmJ\_\{m\}\(Eq\. \([1](https://arxiv.org/html/2608.08224#S3.E1)\)\)\.Middle:attribution patching\(Syedet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib12)\)through a differentiable gate on each component yields the control coefficientCm,kC\_\{m,k\}\(Eq\. \([4](https://arxiv.org/html/2608.08224#S3.E4)\)\); A and B show that activation does not imply control\.Bottom:CD\-RFT keeps control task\-specific across families, loweringBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)and raising pass@11/pass@kk\.The complex reasoning of large language models is unlocked primarily during post\-training, where reinforcement learning maximizes a reward while holding the policy near its base model\(Ouyanget al\.[2022](https://arxiv.org/html/2608.08224#bib.bib1)\)\. As post\-training moves from single\-task to multi\-task settings, a central question emerges: what does it actually change inside the model, and how does it distribute a finite capability across competing tasks?
Benchmark scores answer only whether a model has improved, not where the change occurs\. Mechanistic interpretability quantifies, through attribution patching, the causal role of each component\(Syedet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib12)\); along this line, a representative work attributes the success of reinforcement\-learning fine\-tuning to stronger and more diverse circuit activation\(Zhanget al\.[2026](https://arxiv.org/html/2608.08224#bib.bib22)\)\. We complement this view: activation tells us which circuits fire, but not which ones control the post\-training reward gain, and that account is established on activation alone, within a single domain \(mathematical reasoning\)\.
To separate which components fire from which ones control the gain, we draw on Metabolic Control Analysis\(Kacser and Burns[1973](https://arxiv.org/html/2608.08224#bib.bib24); Heinrich and Rapoport[1974](https://arxiv.org/html/2608.08224#bib.bib25); Fell[1992](https://arxiv.org/html/2608.08224#bib.bib23)\), the systems\-biology framework for asking which local steps control a network’s overall flux, and define the reward flux as the mean log\-likelihood ratio margin by which the policy raises a reference target above the base model; a differentiable gate on each component then yields the Post\-training Control Coefficient, read out by a single backward pass, measuring which components control the gain\. Task\-wise coefficients and component activation magnitudes form the control and activation matrices, respectively\. We call cross\-task control concentration the Shared Control Bottleneck and the difference between activation and control concentration the Activation–Control Gap\. Highly shared activations can retain task\-specific control; a small gap indicates that control concentrates along a shared direction and loses task specificity\.
We turn this diagnosis into CD\-RFT, which regularizes the post\-training loss with the Shared Control Bottleneck\.111Code is available athttps://github\.com/tttbw/cd\-rft\.The exact regularizer gradient requires second\-order differentiation incompatible with flash attention, so we derive a first\-order proxy requiring one backward pass and less than eight percent overhead\. On Qwen2\.5\-7B, CD\-RFT attains the largest control decoupling\. Over its matched GRPO recipe, the no\-KL variant improves pass@11and the KL\-penalized variant improves large\-kkpass@kkcoverage\. Beyond Qwen2\.5\-7B, the same pattern of control decoupling and capability gain carries over to Llama\-3\.2\-3B\.
Our contributions are as follows\.
- •We define the Post\-training Control Coefficient, which separates activation contribution from control responsibility\.
- •We characterize excessive cross\-task control sharing using the Shared Control Bottleneck and the Activation–Control Gap, and derive an optimizable first\-order proxy for the bottleneck\.
- •We propose CD\-RFT, which connects mechanistic diagnosis to training by decoupling the control structure; reducing the diagnosed bottleneck improves multi\-task pass@11and pass@kkon Qwen2\.5\-7B, with the same pattern on Llama\-3\.2\-3B\.
## 2Related Work
##### Reinforcement learning post\-training paradigm for LLM\.
Reference\-constrained post\-training \(reinforcement learning from human feedback\(Ouyanget al\.[2022](https://arxiv.org/html/2608.08224#bib.bib1)\)and direct preference optimization\(Rafailovet al\.[2023](https://arxiv.org/html/2608.08224#bib.bib2)\)\) shares the log\-ratio structure on which our reward flux is based\. Reinforcement learning with verifiable rewards has become the dominant paradigm for reasoning: group relative policy optimization removes the value network\(Shaoet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib3)\), pure reinforcement learning elicits reasoning behaviours\(Guoet al\.[2025](https://arxiv.org/html/2608.08224#bib.bib4)\), whose gains over the base model, especially at large\-kkpass@k\\textup\{pass@\}k, are recently questioned\(Yueet al\.[2025](https://arxiv.org/html/2608.08224#bib.bib6)\), and later work extends it along training stability\(Yuet al\.[2025](https://arxiv.org/html/2608.08224#bib.bib5)\)and data coverage\(Luoet al\.[2025](https://arxiv.org/html/2608.08224#bib.bib9); Chenget al\.[2025](https://arxiv.org/html/2608.08224#bib.bib10)\)\. Closest to our setting, multi\-task GRPO jointly trains a single policy across several task families, re\-weighting tasks to balance worst\-task performance\(Ramesh and others[2026](https://arxiv.org/html/2608.08224#bib.bib11)\)\. Whether acting on the reward, the data, or the multi\-task objective, however, such methods stay at the external\-signal level and do not characterize how a finite capability is distributed inside the model, or how concentrated that distribution is; we instead lower the Shared Control Bottleneck directly in the control space\.
##### Mechanistic interpretability and causal attribution of fine\-tuning\.
Mechanistic interpretability quantifies the causal role of components by attribution patching\(Syedet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib12); Nanda[2023](https://arxiv.org/html/2608.08224#bib.bib14)\), an idea inherited from the circuit framework\(Elhageet al\.[2021](https://arxiv.org/html/2608.08224#bib.bib13)\)and circuit discovery\(Conmyet al\.[2023](https://arxiv.org/html/2608.08224#bib.bib15); Wanget al\.[2023](https://arxiv.org/html/2608.08224#bib.bib16); Kramáret al\.[2024](https://arxiv.org/html/2608.08224#bib.bib17)\)and made more faithful by integrated gradients\(Sundararajanet al\.[2017](https://arxiv.org/html/2608.08224#bib.bib18); Hannaet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib19)\); on fine\-tuning, one line shows that it amplifies or reuses existing mechanisms rather than reconstructing circuits\(Prakashet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib20); Jainet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib21)\)\. Most relevant to us,Zhanget al\.\([2026](https://arxiv.org/html/2608.08224#bib.bib22)\)attribute the success of reinforcement\-learning fine\-tuning to increased activation intensity and diversity, measured within a single domain \(mathematical reasoning\) by three statistics of the same EAP edge\-magnitude tensor: its mean \(activation intensity\), its entropy \(information complexity\), and its kurtosis \(distribution kurtosis\)\. All three describe how large and how spread the attributions are, not which components the reward gain depends on, a distinction Lemma[1](https://arxiv.org/html/2608.08224#Thmlemma1)makes precise: a circuit being activated does not entail that it controls the gain\. That line of work is moreover purely diagnostic\. We therefore redirect attribution to the Post\-training Control Coefficient over the reward flux, characterize the over\-sharing of control for the first time through the cross\-task Shared Control Bottleneck, and close the loop by writing that quantity into the training objective and measuring the resulting multi\-task gain\.
## 3The Post\-Training Control Coefficient
### 3\.1Post\-Training Reward Flux
To separate activation from control into checkable quantities, we take as our basic object the log\-likelihood\-ratio margin of the policy relative to the reference model on a reference target\. Our main experiments use reinforcement learning with verifiable rewards\(RLVR; Lambertet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib7)\), in which the reference targety⋆y^\{\\star\}is a verified\-correct completion; for a task familymmwith samples\(x,y⋆\)∼𝒟m\(x,y^\{\\star\}\)\\sim\\mathcal\{D\}\_\{m\}, we define the reward flux
Jm\(θ\)=𝔼\(x,y⋆\)∼𝒟m\[logπθ\(y⋆∣x\)−logπ0\(y⋆∣x\)\]\.J\_\{m\}\(\\theta\)=\\mathbb\{E\}\_\{\(x,y^\{\\star\}\)\\sim\\mathcal\{D\}\_\{m\}\}\\\!\\left\[\\log\\pi\_\{\\theta\}\(y^\{\\star\}\\mid x\)\-\\log\\pi\_\{0\}\(y^\{\\star\}\\mid x\)\\right\]\.\(1\)The integrandlog\(πθ/π0\)\\log\(\\pi\_\{\\theta\}/\\pi\_\{0\}\)in Eq\. \([1](https://arxiv.org/html/2608.08224#S3.E1)\) is the common structure of the family of reference\-constrained post\-training objectives: RLVR takes verifiable correctness as the reward and raises this log\-likelihood ratio on correct targets, DPO writes the implicit reward asβlog\(πθ/π0\)\\beta\\log\(\\pi\_\{\\theta\}/\\pi\_\{0\}\)\(Rafailovet al\.[2023](https://arxiv.org/html/2608.08224#bib.bib2)\), and RLHF targets a trade\-off between the reward andKL\(πθ∥π0\)\\mathrm\{KL\}\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{0\}\)\(Ouyanget al\.[2022](https://arxiv.org/html/2608.08224#bib.bib1)\)\. HenceJmJ\_\{m\}is not an artifact constructed for the analysis but this shared log\-ratio structure evaluated at the reference target\.
### 3\.2Metabolic Control Analysis
Metabolic Control Analysis \(MCA\) originates in systems biology\(Kacser and Burns[1973](https://arxiv.org/html/2608.08224#bib.bib24); Heinrich and Rapoport[1974](https://arxiv.org/html/2608.08224#bib.bib25)\)and studies the share of control that a single local step exerts over the overall flux of a network\. Its central object, the flux control coefficient, is the scaled sensitivity \(log–log derivative\) of the overall fluxJJto a local rateviv\_\{i\},
CiJ=∂lnJ∂lnvi=viJ∂J∂vi,C^\{J\}\_\{i\}=\\frac\{\\partial\\ln J\}\{\\partial\\ln v\_\{i\}\}=\\frac\{v\_\{i\}\}\{J\}\\,\\frac\{\\partial J\}\{\\partial v\_\{i\}\},\(2\)and its summation theorem states that in a closed conservative network the control coefficients sum to one, so control is systematically distributed rather than monopolized\(Fell[1992](https://arxiv.org/html/2608.08224#bib.bib23)\)\. We transplant this “how local rates control the overall flux” viewpoint to post\-training, taking the reward flux as the overall flux and the component gate as the local rate\. The residual structure breaks the required homogeneity, so∑kCm,k≠1\\sum\_\{k\}C\_\{m,k\}\\neq 1here, as Appendix Proposition[B\.1](https://arxiv.org/html/2608.08224#A2.Thmproposition1)shows; the distribution of control does not rely on this conservation law, however, and Theorem[1](https://arxiv.org/html/2608.08224#Thmtheorem1)proves independently that residual\-stream superposition alone keeps the control support from collapsing to a single gate\. We therefore borrow only the “control share” viewpoint of MCA\.
### 3\.3Differentiable Gating and Attribution Patching
On the residual update of each layer, we multiply a set of internal components each by a differentiable scalar gategkg\_\{k\},
hℓ\+1=hℓ\+∑k:ℓ\(k\)=ℓgkfk\(hℓ;θ\),h^\{\\ell\+1\}=h^\{\\ell\}\+\\sum\_\{k:\\ell\(k\)=\\ell\}g\_\{k\}\\,f\_\{k\}\(h^\{\\ell\};\\theta\),\(3\)wherefkf\_\{k\}is the transformation of componentkkandoutk\\mathrm\{out\}\_\{k\}is its output written into the residual stream; when the gates take their nominal valueg≡𝟏g\\equiv\\mathbf\{1\}the network recovers the original model under Appendix Assumption[A\.2](https://arxiv.org/html/2608.08224#A1.Thmassumption2), so thatgkg\_\{k\}is a continuous ablation intervention at the component level \(0for full ablation,11for intact\), and we denote the gated reward flux byJm\(θ;g\)J\_\{m\}\(\\theta;g\)\. The gate granularity is a free choice: an attention head, an MLP neuron, or a circuit edge\. We use the sublayer granularity, the coarsest choice under which the spectral\-proxy optimization stays feasible for full 7B fine\-tuning; finer gate sets \(head, neuron, or edge\) are left to future work\.
To measure the causal effect of such an intervention on a target metric, the standard tool in mechanistic interpretability is*attribution patching*\(Syedet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib12); Nanda[2023](https://arxiv.org/html/2608.08224#bib.bib14)\), which linearizes an activation replacement by a first\-order Taylor expansion \(Appendix Eq\. \([16](https://arxiv.org/html/2608.08224#A1.E16)\)\), so a single backward pass gives the attributions of all components at once\. We take the gate as the intervention variable and the target log\-likelihood as the metric, and read out the control coefficient next\.
### 3\.4The Coefficient and Its Expansion
Combining the flux control coefficient with the gating above, we define the post\-training control coefficient as the scaled gate\-sensitivity of the target log\-likelihoodℓm\(θ\):=𝔼𝒟m\[logπθ\(y⋆∣x\)\]\\ell\_\{m\}\(\\theta\):=\\mathbb\{E\}\_\{\\mathcal\{D\}\_\{m\}\}\[\\log\\pi\_\{\\theta\}\(y^\{\\star\}\\mid x\)\], the policy term of the flux,
Cm,k=1ℓm\(θ\)∂ℓm\(θ;g\)∂gk\|g=𝟏\.C\_\{m,k\}=\\frac\{1\}\{\\ell\_\{m\}\(\\theta\)\}\\,\\frac\{\\partial\\ell\_\{m\}\(\\theta;g\)\}\{\\partial g\_\{k\}\}\\bigg\|\_\{g=\\mathbf\{1\}\}\.\(4\)The reference termlogπ0\\log\\pi\_\{0\}is gate\-independent, so∂ℓm/∂gk=∂Jm/∂gk\\partial\\ell\_\{m\}/\\partial g\_\{k\}=\\partial J\_\{m\}/\\partial g\_\{k\}: the coefficient reports the gate\-sensitivity of the reward flux, now normalized by the stableℓm\\ell\_\{m\}\(\|ℓm\|≥ϵ0\|\\ell\_\{m\}\|\\geq\\epsilon\_\{0\}, Appendix Assumption[A\.3](https://arxiv.org/html/2608.08224#A1.Thmassumption3)\) rather than by the sign\-indefinite, possibly\-vanishingJmJ\_\{m\}\(Appendix Remark[A\.2](https://arxiv.org/html/2608.08224#A1.Thmremark2)\)\. Arranging the control coefficients into the control matrixC=\[Cm,k\]∈ℝM×KC=\[C\_\{m,k\}\]\\in\\mathbb\{R\}^\{M\\times K\}withM=MfamM=M\_\{\\mathrm\{fam\}\}rows, one per task family: each row is the control vector of familymm, averaged over the family’snnprobe samples, as detailed in Appendix[A](https://arxiv.org/html/2608.08224#A1)\.
The explanatory power of the control coefficient comes from its first\-order dominance over the relative change of the reward flux under small gate perturbations\.
###### Proposition 1\(first\-order expansion of control\)\.
Under Appendix Assumptions[A\.1](https://arxiv.org/html/2608.08224#A1.Thmassumption1)–[A\.3](https://arxiv.org/html/2608.08224#A1.Thmassumption3), and writing the gate log\-perturbationu=loggu=\\log g, the relative change of the target log\-likelihood of familymmsatisfies
ℓm\(θ;eu\)−ℓm\(θ\)ℓm\(θ\)=∑k=1KCm,kuk\+O\(∥u∥22\),u→𝟎\.\\frac\{\\ell\_\{m\}\(\\theta;e^\{u\}\)\-\\ell\_\{m\}\(\\theta\)\}\{\\ell\_\{m\}\(\\theta\)\}=\\sum\_\{k=1\}^\{K\}C\_\{m,k\}\\,u\_\{k\}\+O\\\!\\left\(\\lVert u\\rVert\_\{2\}^\{2\}\\right\),\\qquad u\\to\\mathbf\{0\}\.\(5\)
This is a first\-order Taylor expansion aboutu=𝟎u=\\mathbf\{0\}divided byℓm\(θ\)≠0\\ell\_\{m\}\(\\theta\)\\neq 0, proven in Appendix Proposition[A\.1](https://arxiv.org/html/2608.08224#A1.Thmproposition1)\.
Equation \([5](https://arxiv.org/html/2608.08224#S3.E5)\) makes the control coefficient the first\-order dominant coefficient of the relative change in gate log\-space, measuring the control of componentkkover the gain of familymm\. This differs from the activation magnitude tracked by prior analyses\(Zhanget al\.[2026](https://arxiv.org/html/2608.08224#bib.bib22)\), the average usage of componentkkon familymm,
Fm,k=𝔼𝒟m∥outk\(x;θ\)∥2,F=\[Fm,k\]∈ℝM×K\.F\_\{m,k\}=\\mathbb\{E\}\_\{\\mathcal\{D\}\_\{m\}\}\\lVert\\mathrm\{out\}\_\{k\}\(x;\\theta\)\\rVert\_\{2\},\\qquad F=\[F\_\{m,k\}\]\\in\\mathbb\{R\}^\{M\\times K\}\.\(6\)Because the reward flux is read out along a fixed directionwmw\_\{m\}defined in Appendix[A](https://arxiv.org/html/2608.08224#A1), the first\-order effect of a component depends on the part of its write that survives propagation towmw\_\{m\}, not on its norm; large usage and strong control thus decouple\.
###### Lemma 1\(the activation axis and the control axis are not interchangeable\)\.
For the gating of Eq\. \([3](https://arxiv.org/html/2608.08224#S3.E3)\) there are networks and task families in whichCm,⋅C\_\{m,\\cdot\}is not a scalar multiple ofFm,⋅F\_\{m,\\cdot\}\.
Appendix Lemma[A\.1](https://arxiv.org/html/2608.08224#A1.Thmlemma1)constructs one\.
Lemma[1](https://arxiv.org/html/2608.08224#Thmlemma1)has two consequences\. First, “post\-training enhances activation diversity” does not entail “post\-training improves the control structure”: the two axes are not interchangeable\. Second, optimizing on the activation axis changes the objective and cannot reduce control sharing \(Section[5\.3](https://arxiv.org/html/2608.08224#S5.SS3)\)\. We therefore treat the control matrix as a first\-class object\.
## 4The Shared Control Bottleneck
Distinguishing control from activation is what makes control\-aware training possible \(Section[5](https://arxiv.org/html/2608.08224#S5)\)\. This section builds that measurement: the Activation–Control Gap between the sharing of control and of activation\.
### 4\.1Cross\-Task Concentration
For a matrixX∈\{F,C\}X\\in\\\{F,C\\\}, its family Gram matrixGX=XX⊤∈ℝM×MG\_\{X\}=XX^\{\\top\}\\in\\mathbb\{R\}^\{M\\times M\}is symmetric positive semidefinite, and we characterize the cross\-task concentration by its normalized largest eigenvalue,
Bshared\(X\)=λmax\(GX\)tr\(GX\)∈\[1M,1\],B\_\{\\mathrm\{shared\}\}\(X\)=\\frac\{\\lambda\_\{\\max\}\(G\_\{X\}\)\}\{\\operatorname\{tr\}\(G\_\{X\}\)\}\\in\\Big\[\\tfrac\{1\}\{M\},1\\Big\],\(7\)withX≠0X\\neq 0so that the denominator is nonzero\. This quantity measures how much of the energy of the rows \(one per task family\) falls along a single shared direction\.
###### Proposition 2\(bounds, saturation, and invariance ofBsharedB\_\{\\mathrm\{shared\}\}\)\.
Let the eigenvalues ofGXG\_\{X\}beλ1≥⋯≥λM≥0\\lambda\_\{1\}\\geq\\dots\\geq\\lambda\_\{M\}\\geq 0\. Then: \(i\)1/M≤Bshared\(X\)≤11/M\\leq B\_\{\\mathrm\{shared\}\}\(X\)\\leq 1; \(ii\) the upper boundBshared=1B\_\{\\mathrm\{shared\}\}=1holds if and only ifrank\(X\)=1\\operatorname\{rank\}\(X\)=1, i\.e\. the row vectors are collinear; \(iii\) the lower boundBshared=1/MB\_\{\\mathrm\{shared\}\}=1/Mholds if and only if allλi\\lambda\_\{i\}are equal, i\.e\. the row vectors are pairwise orthogonal and of equal norm; \(iv\) it is invariant under magnitude scalingX↦cXX\\mapsto cX\(c≠0c\\neq 0\) and under orthogonal reparameterization of the gate coordinatesX↦XQX\\mapsto XQ\(QQ⊤=IKQQ^\{\\top\}=I\_\{K\}\)\.
Appendix Proposition[B\.2](https://arxiv.org/html/2608.08224#A2.Thmproposition2)proves these\.
Property \(iv\) therefore makesBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)comparable across methods and training steps at a fixed granularity, the prerequisite for a cross\-method mechanistic comparison\.Bshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)measures direction sharing and the distribution of control energy jointly\.
### 4\.2The Activation–Control Gap
Applying the concentration to the activation axis and the control axis separately, their difference characterizes the degree to which control is decoupled from activation,
ACG=Bshared\(F\)−Bshared\(C\)\.\\mathrm\{ACG\}=B\_\{\\mathrm\{shared\}\}\(F\)\-B\_\{\\mathrm\{shared\}\}\(C\)\.\(8\)That multi\-task activations are shared across tasks is an expected property, withBshared\(F\)B\_\{\\mathrm\{shared\}\}\(F\)near its upper bound \(measured≈99\.6\\approx 99\.6\), and as a ceiling it is not a target\.
###### Proposition 3\(characterization of the gap\)\.
Under the activational ceilingBshared\(F\)→1B\_\{\\mathrm\{shared\}\}\(F\)\\to 1, the gap satisfiesACG∈\[0,1−1/M\]\\mathrm\{ACG\}\\in\[0,\\,1\-1/M\], andACG→0\\mathrm\{ACG\}\\to 0if and only ifBshared\(C\)→Bshared\(F\)B\_\{\\mathrm\{shared\}\}\(C\)\\to B\_\{\\mathrm\{shared\}\}\(F\), in which case control is as concentrated as activation;ACG\\mathrm\{ACG\}is large if and only ifBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)lies far below the ceiling, corresponding to families sharing activations while control stays task\-specific\.
Appendix Proposition[B\.3](https://arxiv.org/html/2608.08224#A2.Thmproposition3)proves this\.
Hence the target is not to suppressBshared\(F\)B\_\{\\mathrm\{shared\}\}\(F\)but, at a fixed ceiling, to lowerBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)and thereby enlarge the gap in Eq\. \([8](https://arxiv.org/html/2608.08224#S4.E8)\), which is precisely the objective of the method in Section[5\.1](https://arxiv.org/html/2608.08224#S5.SS1)\. Empirically, reinforcement learning already lowersBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)below the base, so our claim is controlled\-variable: on a matched recipe, our method lowersBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)the most, as Appendix Remark[B\.2](https://arxiv.org/html/2608.08224#A2.Thmremark2)discusses\.
### 4\.3Distributedness of Control
Taking the overall shape of the control matrix, rather than a single\-gate attribution, as the object of analysis is justified by the fact that control is systematically distributed over the residual network\. This distributedness does not rely on the summation theorem of Section[3\.2](https://arxiv.org/html/2608.08224#S3.SS2)\(which fails here, Appendix Proposition[B\.1](https://arxiv.org/html/2608.08224#A2.Thmproposition1)\) but stems from the more fundamental structure of residual\-stream superposition\.
###### Theorem 1\(no single\-gate monopoly under residual superposition\)\.
Suppose the reward flux is read out by projecting the final residual onto a fixed readout directionwmw\_\{m\}\. Then, for generic network weights,∂gkJm\|g=𝟏≠0\\partial\_\{g\_\{k\}\}J\_\{m\}\|\_\{g=\\mathbf\{1\}\}\\neq 0for everykk: the control support is not a proper subset of\{1,…,K\}\\\{1,\\dots,K\\\}\.
Intuitively, residual superposition mixes the write of every component into the readout, so a gate with no control would need its write to cancel exactly alongwmw\_\{m\}, a measure\-zero coincidence \(the propagation identity and the generic\-weight argument are in Appendix Theorem[B\.1](https://arxiv.org/html/2608.08224#A2.Thmtheorem1)\)\. The control support is therefore generically not a single gate, so a single\-gate scalar cannot characterize it and the right object is the cross\-task concentrationBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)of the full matrix \(Figure[2](https://arxiv.org/html/2608.08224#S5.F2)a\)\. Concentration also implies a deployment\-level fragility geometry, largerBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)giving larger perturbation variance, as Appendix Proposition[B\.4](https://arxiv.org/html/2608.08224#A2.Thmproposition4)shows\.
## 5Control\-Diverse Regularizer
### 5\.1Objective
On top of the backbone post\-training lossℒpost\\mathcal\{L\}\_\{\\mathrm\{post\}\}, we add a single control\-diverse regularizer,
ℒCD−RFT\(θ\)=ℒpost\(θ\)\+λBshared\(C\(θ\)\),λ\>0,\\mathcal\{L\}\_\{CD\-RFT\}\(\\theta\)=\\mathcal\{L\}\_\{\\mathrm\{post\}\}\(\\theta\)\+\\lambda\\,B\_\{\\mathrm\{shared\}\}\\\!\\big\(C\(\\theta\)\\big\),\\qquad\\lambda\>0,\(9\)so that loweringBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)at a fixed activational ceiling enlarges the gap in Eq\. \([8](https://arxiv.org/html/2608.08224#S4.E8)\)\. It attaches to any reference\-constrained objective; we use GRPO\(Shaoet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib3)\)\. Computing∇θBshared\(C\)\\nabla\_\{\\theta\}B\_\{\\mathrm\{shared\}\}\(C\)exactly requires differentiating the already first\-order control coefficient once more inθ\\theta, i\.e\. second\-order automatic differentiation: eager\-only, memory\-exploding, and incompatible with flash\-attention\(Daoet al\.[2022](https://arxiv.org/html/2608.08224#bib.bib37)\)and parameter sharding\. The next three subsections reduce itsθ\\theta\-gradient to one backward pass over a stop\-gradient control estimate \(operation count in Appendix Remark[C\.5](https://arxiv.org/html/2608.08224#A3.Thmremark5)\)\. Table[1](https://arxiv.org/html/2608.08224#S5.T1)checks which of the two components ofBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)the regularizer moves: both fall\.
Table 1:Decomposition of theBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)drop on a matched GRPO/CD\-RFT pair \(β=0\\beta\{=\}0; seven probe seeds, same\-seed paired differences\)\.Bdir\(C\)=λmax\(R\)/MB\_\{\\mathrm\{dir\}\}\(C\)=\\lambda\_\{\\max\}\(R\)/M, withRRthe matrix of row cosines, isBsharedB\_\{\\mathrm\{shared\}\}on a row\-normalizedCCand so measures direction sharing alone;Bnorm\(C\)B\_\{\\mathrm\{norm\}\}\(C\)is the valueBsharedB\_\{\\mathrm\{shared\}\}takes at exactly orthogonal rows\. Only the direction term falls on every seed \(Appendix[F\.3](https://arxiv.org/html/2608.08224#A6.SS3)\)\.
### 5\.2A Stable Spectral\-Moment Ratio Replacing the Spectral Extremum
Optimizing the spectral extremumλmax\\lambda\_\{\\max\}directly invokes an eigendecomposition ofGCG\_\{C\}that is non\-smooth at top\-eigenvalue degeneracyλ1=λ2\\lambda\_\{1\}=\\lambda\_\{2\}, where the eigenvector derivative blows up as1/\(λ1−λi\)1/\(\\lambda\_\{1\}\-\\lambda\_\{i\}\), as Appendix Proposition[C\.1](https://arxiv.org/html/2608.08224#A3.Thmproposition1)shows\. Borrowing from spectral graph convolution the idea of avoiding eigendecomposition by polynomial approximation\(Defferrardet al\.[2016](https://arxiv.org/html/2608.08224#bib.bib26); Kipf and Welling[2017](https://arxiv.org/html/2608.08224#bib.bib27)\), and writing the normalized spectrumλ^i=λi/∑jλj\\hat\{\\lambda\}\_\{i\}=\\lambda\_\{i\}/\\sum\_\{j\}\\lambda\_\{j\}, we substitute the stable spectral\-moment ratio
ℛ\(C\)=tr\(GC2\)tr\(GC\)2=∑i=1Mλ^i2\.\\mathcal\{R\}\(C\)=\\frac\{\\operatorname\{tr\}\(G\_\{C\}^\{2\}\)\}\{\\operatorname\{tr\}\(G\_\{C\}\)^\{2\}\}=\\sum\_\{i=1\}^\{M\}\\hat\{\\lambda\}\_\{i\}^\{2\}\.\(10\)This quantity is a pure polynomial \(matrix multiplication, no eigendecomposition\), isC∞C^\{\\infty\}atC≠0C\\neq 0with no1/\(λ1−λi\)1/\(\\lambda\_\{1\}\-\\lambda\_\{i\}\)singularity, and is co\-monotone withBshared=maxiλ^iB\_\{\\mathrm\{shared\}\}=\\max\_\{i\}\\hat\{\\lambda\}\_\{i\}under the majorization order, a partial order, sharing the bounds\{1,1/M\}\\\{1,1/M\\\}, as Appendix Proposition[C\.2](https://arxiv.org/html/2608.08224#A3.Thmproposition2)establishes\. ReplacingBsharedB\_\{\\mathrm\{shared\}\}byℛ\\mathcal\{R\}buys a smooth objective with bounded gradient at degeneracy; both share extrema, and we verify their agreement empirically during training \(Figure[2](https://arxiv.org/html/2608.08224#S5.F2)b\)\.
### 5\.3Closed\-Form Sensitivity and the Irreplaceability of the Control Axis
LetC¯\\bar\{C\}be the control matrix obtained from one ordinary backward pass with the gradient with respect toθ\\thetastopped, and writeG¯=C¯C¯⊤\\bar\{G\}=\\bar\{C\}\\bar\{C\}^\{\\top\}\. The sensitivity of the spectral\-moment ratio to the control matrix admits a closed form, with no eigendecomposition,
W:=∂ℛ∂C\|C¯=4tr\(G¯\)2\(G¯C¯−ℛ¯tr\(G¯\)C¯\)∈ℝM×KW:=\\frac\{\\partial\\mathcal\{R\}\}\{\\partial C\}\\bigg\|\_\{\\bar\{C\}\}=\\frac\{4\}\{\\operatorname\{tr\}\(\\bar\{G\}\)^\{2\}\}\\big\(\\bar\{G\}\\,\\bar\{C\}\-\\bar\{\\mathcal\{R\}\}\\,\\operatorname\{tr\}\(\\bar\{G\}\)\\,\\bar\{C\}\\big\)\\in\\mathbb\{R\}^\{M\\times K\}\(11\)Appendix Lemma[C\.1](https://arxiv.org/html/2608.08224#A3.Thmlemma1)gives the derivation\. By the chain rule, the true gradient is the projection of this stop\-gradient weightWWalong the direction of the control matrix,∇θℛ=∇θ⟨W,C\(θ\)⟩\\nabla\_\{\\theta\}\\mathcal\{R\}=\\nabla\_\{\\theta\}\\langle W,\\,C\(\\theta\)\\rangle; henceWWmust act on the control matrixC\(θ\)C\(\\theta\)rather than on the activation usageAk=𝔼∥outk∥2A\_\{k\}=\\mathbb\{E\}\\lVert\\mathrm\{out\}\_\{k\}\\rVert\_\{2\}\.
###### Theorem 2\(irreplaceability of the control axis\)\.
Substituting the activation usageAkA\_\{k\}forC\(θ\)C\(\\theta\)fails on two structural grounds: it violates the premiseF≠CF\\neq Cof Lemma[1](https://arxiv.org/html/2608.08224#Thmlemma1); and the per\-gate scalar sums out the family dimension, so it cannot express per\-task decoupling\.
Appendix Theorem[C\.1](https://arxiv.org/html/2608.08224#A3.Thmtheorem1)proves this\. Empirically, minimizing∑kwkAk\\sum\_\{k\}w\_\{k\}A\_\{k\}under fixed reward further cuts the high\-usage shared directions and forces control to reconcentrate, the concentration spike observed on large models by early implementations that act on the activation usage\.
### 5\.4First\-Order Proxy Gradient via Central Differences in Gate Space
Expanding the Frobenius inner product row by row into gate\-space directional derivatives gives
⟨W,C\(θ\)⟩=∑m1ℓmDWm,⋅ℓm,\\langle W,C\(\\theta\)\\rangle=\\sum\_\{m\}\\frac\{1\}\{\\ell\_\{m\}\}D\_\{W\_\{m,\\cdot\}\}\\ell\_\{m\},\(12\)whereDWm,⋅ℓm=⟨Wm,⋅,∂gℓm\|g=𝟏⟩D\_\{W\_\{m,\\cdot\}\}\\ell\_\{m\}=\\langle W\_\{m,\\cdot\},\\,\\partial\_\{g\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}\\rangleis the directional derivative ofℓm\\ell\_\{m\}alongWm,⋅W\_\{m,\\cdot\}\. Since the gate\-space dimensionKKis far smaller thandimθ\\dim\\theta, we estimate this directional derivative by a central difference in gate space\.
###### Theorem 3\(central\-difference form of the proxy gradient\)\.
WritingW^m,⋅=Wm,⋅/∥Wm,⋅∥2\\hat\{W\}\_\{m,\\cdot\}=W\_\{m,\\cdot\}/\\lVert W\_\{m,\\cdot\}\\rVert\_\{2\}, we have
⟨W,C\(θ\)⟩=∑m=1M∥Wm,⋅∥2ℓm⋅ℓm\(θ;𝟏\+ϵW^m,⋅\)−ℓm\(θ;𝟏−ϵW^m,⋅\)2ϵ\+O\(ϵ2\),\\displaystyle\\langle W,C\(\\theta\)\\rangle=\\sum\_\{m=1\}^\{M\}\\frac\{\\lVert W\_\{m,\\cdot\}\\rVert\_\{2\}\}\{\\ell\_\{m\}\}\\cdot\\frac\{\\ell\_\{m\}\(\\theta;\\mathbf\{1\}\+\\epsilon\\hat\{W\}\_\{m,\\cdot\}\)\-\\ell\_\{m\}\(\\theta;\\mathbf\{1\}\-\\epsilon\\hat\{W\}\_\{m,\\cdot\}\)\}\{2\\epsilon\}\+O\(\\epsilon^\{2\}\),
\(13\)where the two gated forward passes carry only constant gate scalings and are ordinarily differentiable inθ\\theta; hence∇θℒproxy\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}is obtained by one backward pass, with no second\-order graph, and agrees with∇θℛ\\nabla\_\{\\theta\}\\mathcal\{R\}to withinO\(ϵ2\)O\(\\epsilon^\{2\}\)\.
Appendix Theorem[C\.2](https://arxiv.org/html/2608.08224#A3.Thmtheorem2)proves this\.
Intuitively, the gradient needs only the projection ofC\(θ\)C\(\\theta\)along the fixedWW, and a directional derivative along a fixed gate direction is the slope of a scalar read out by two forward passes; differentiation with respect to the gates is absorbed into the forward passes, leaving a first\-order graph inθ\\theta\. Numerically we perturb along the unit direction and rescale by∥Wm,⋅∥2\\lVert W\_\{m,\\cdot\}\\rVert\_\{2\}to stay in the linear regime, as Appendix Remark[C\.2](https://arxiv.org/html/2608.08224#A3.Thmremark2)notes\.
### 5\.5Estimation Protocol
The control matrix is estimated on a fixed probe ofnnsamples per family, each row a family mean, soBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)is annn\-dependent estimator: under the rank\-one signal\-plus\-noise model of Appendix Proposition[C\.3](https://arxiv.org/html/2608.08224#A3.Thmproposition3),
𝔼\[tr\(GC\)\]=s2Mfam⏟signal\+1ntr\(ΣN\)⏟per\-row noise,\\mathbb\{E\}\\big\[\\operatorname\{tr\}\(G\_\{C\}\)\\big\]=\\underbrace\{s^\{2\}M\_\{\\mathrm\{fam\}\}\}\_\{\\text\{signal\}\}\+\\underbrace\{\\tfrac\{1\}\{n\}\\operatorname\{tr\}\(\\Sigma\_\{N\}\)\}\_\{\\text\{per\-row noise\}\},\(14\)so a small probe inflates the denominator and under\-reads concentration, while the leading direction is located at a sample size independent ofKK\(Davis–Kahan\(Davis and Kahan[1970](https://arxiv.org/html/2608.08224#bib.bib28)\)\)\. We therefore fixnnacross compared models and cut variance by offline seed averaging rather than a larger probe\.
Figure 2:Distributedness of control and proxy fidelity\.\(a\)Every family spreads control over many gates \(PR≫1\\mathrm\{PR\}\\gg 1\), equally at everyBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\), so no per\-gate scalar characterizes control\. Circles Qwen2\.5\-7B, triangles Llama\-3\.2\-3B \(5 seeds\)\.\(b\)ℛ\\mathcal\{R\}tracksBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)in training \(CD\-RFT,β=0\\beta\{=\}0\)\.
### 5\.6Implementation and Scalability
The soft coefficientλ\\lambdaalone cannot curb the transient concentration peak during the reinforcement\-learning warm\-up, so we enforce the near\-hard constraint
minθℒpost\(θ\)s\.t\.Bshared\(C\(θ\)\)≤τ,\\min\_\{\\theta\}\\ \\mathcal\{L\}\_\{\\mathrm\{post\}\}\(\\theta\)\\quad\\text\{s\.t\.\}\\quad B\_\{\\mathrm\{shared\}\}\\\!\\big\(C\(\\theta\)\\big\)\\leq\\tau,\(15\)by a per\-step inner loop that descends−∇θℒproxy\-\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}untilBshared≤τB\_\{\\mathrm\{shared\}\}\\leq\\tauor a capKmaxK\_\{\\max\}is reached; Appendix Algorithm[1](https://arxiv.org/html/2608.08224#alg1)and Definition[C\.1](https://arxiv.org/html/2608.08224#A3.Thmdefinition1)give the full step and the\(τ,Kmax\)\(\\tau,K\_\{\\max\}\)setting\. Reduced to a single backward pass by Theorem[3](https://arxiv.org/html/2608.08224#Thmtheorem3), the regularizer is compatible with flash\-attention and parameter sharding, scales to full fine\-tuning of 7B on a single GPU without low\-rank adaptation, and adds only a minor per\-step overhead, as Appendix Remark[C\.4](https://arxiv.org/html/2608.08224#A3.Thmremark4)shows\.
## 6Experiments
We organize the experiments around two questions: \(i\) whether the Shared Control BottleneckBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)is the right object for characterizing the internal changes wrought by reinforcement\-learning post\-training, namely whether it yields a reproducible mechanistic signature that activation\-level metrics cannot; and \(ii\) whether explicitly decoupling this bottleneck on top of the same recipe \(CD\-RFT\) predictably translates into gains in multi\-task capability\.
### 6\.1Experimental Setup
We train Qwen2\.5\-7B\(Yanget al\.[2024](https://arxiv.org/html/2608.08224#bib.bib40)\)from the base model with GRPO\(Shaoet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib3)\)as the backbone, under an RLVR setting\(Lambertet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib7)\), on a balanced mixture of three families of verifiable tasks \(mathematics, code, and logic\) drawn respectively from DeepScaleR\(Luoet al\.[2025](https://arxiv.org/html/2608.08224#bib.bib9)\), code\-r1\-12k\(Liu and Zhang[2025](https://arxiv.org/html/2608.08224#bib.bib8)\), and the GURU logic subset\(Chenget al\.[2025](https://arxiv.org/html/2608.08224#bib.bib10)\), each scored by a rule or test\-case verifier, with data and mixing ratios in Appendix[A](https://arxiv.org/html/2608.08224#A1)\. To isolate the single variable of adding the control\-diverse regularizer, we compare along two paired axes, without KL \(β=0\\beta\{=\}0\) and with KL \(β=0\.01\\beta\{=\}0\.01\), on each of which CD\-RFT and its matched GRPO share every setting except the regularizer; the untrained base serves as an anchor\.
Evaluation uses the nine held\-out benchmarks of Table[2](https://arxiv.org/html/2608.08224#S6.T2), reported with unbiasedpass@k\\textup\{pass@\}k\(Chenet al\.[2021](https://arxiv.org/html/2608.08224#bib.bib31)\)at a fixed temperature applied identically to all methods\. The full setup is deferred to the appendix\.
### 6\.2Mechanistic Diagnosis of the Shared Control Bottleneck
#### Activation\-Level Metrics Are Insufficient\.
Two of the three activation metrics ofZhanget al\.\([2026](https://arxiv.org/html/2608.08224#bib.bib22)\)reproduce robustly \(reinforcement\-learning fine\-tuning raises activation intensity and lowers distribution kurtosis across all methods and datasets\), but the third, information complexity \(the “activation diversity” metric\), isdirection\-unstable: as Figure[3](https://arxiv.org/html/2608.08224#S6.F3)shows, among models of comparable capability it stays high under CD\-RFT yet nearly collapses under the matched GRPO, in the same direction across all three domains and by several standard errors\. Without overturning the intensity and kurtosis effects that do replicate, the component singled out as evidence of*activation diversity*disagrees with itself on equally healthy models, so activation\-level metrics cannot on their own answer what reinforcement learning changes\. The object of analysis should move from which components are activated to which control the reward flux, the control matrix and its cross\-task concentrationBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)\.
Figure 3:Information complexity \(edge\-distribution entropy\) across the three domains, with bootstrap confidence intervals\.
#### The Shared Control Bottleneck Yields a Robust Signature\.
We measureBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)under the probe protocol detailed in the appendix, reported in Table[2](https://arxiv.org/html/2608.08224#S6.T2)\. The activation axisBshared\(F\)B\_\{\\mathrm\{shared\}\}\(F\)stays at the reference ceiling of Section[4\.2](https://arxiv.org/html/2608.08224#S4.SS2)across all methods, so we analyze the control axis only\.
Table[2](https://arxiv.org/html/2608.08224#S6.T2)shows a clean ordering, base\>\>the RL baselines\>\>the CD\-RFT variants: reinforcement learning already lowers control sharing, and the regularizer lowers it further and most, so both variants open a larger Activation–Control Gap than their matched baseline\. Because absolute values depend on probe sampling, the method claim is read as the same\-seed paired difference of Table[1](https://arxiv.org/html/2608.08224#S5.T1)and Appendix Figure[F\.2](https://arxiv.org/html/2608.08224#A6.F2): robust on the no\-KL axis and directionally consistent, if smaller, with KL, matching the coverage\-side gains of the latter\.
Across five base models spanning two families and several scales, Appendix Figure[F\.3](https://arxiv.org/html/2608.08224#A6.F3)showsBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)above the9999th percentile of a simulated random reference at this shape \(42\.8±3\.442\.8\\pm 3\.4, well above the bound1/M1/M; Appendix[F\.2](https://arxiv.org/html/2608.08224#A6.SS2)\), so the bottleneck is not a single\-model artifact\. The two models we intervene on \(Qwen2\.5\-7B and Llama\-3\.2\-3B; Section[6\.5](https://arxiv.org/html/2608.08224#S6.SS5)\) have the deepest bottlenecks, and hence the most decoupling headroom; absolute values are tokenizer\- and probe\-dependent, so the comparison is read across models by ordering\.
### 6\.3Multi\-Task Capability
We evaluate capability on nine benchmarks and summarize by domain in Table[2](https://arxiv.org/html/2608.08224#S6.T2)\. The paired comparison is a controlled test of the diagnosis: if the Shared Control Bottleneck is what matters, decoupling it on an otherwise identical recipe should move capability\. We read the net change of CD\-RFT relative to its matched GRPO on eachβ\\betaaxis, an ordering the appendix temperature and checkpoint sweeps confirm\.
Table 2:Mechanism and capability on Qwen2\.5\-7B\.BsharedB\_\{\\mathrm\{shared\}\}abbreviatesBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\); it andACG\\mathrm\{ACG\}are measured on the fixed probe over seven probe seeds \(mean±\\pmSE\), while the capability columns are benchmark accuracies\. Math = MATH500/AMC23/AIME24/AIME25/Minerva; code = HumanEval\+/MBPP; logic = ordering\-puzzle/Logic\-graph; hard = AIME24/AIME25/Minerva\. Each domain averages its sets; overall averages the three domains; all entries are percentages; full ladder in Appendix Table[G\.1](https://arxiv.org/html/2608.08224#A7.T1)\.Two trends stand out\. First, adding the control\-diverse regularizer improves both accuracy axes over its matched GRPO baseline on almost every benchmark\. As Appendix Figure[G\.1](https://arxiv.org/html/2608.08224#A7.F1)shows, the no\-KL variant improves greedypass@1\\textup\{pass@\}1on 8 of 9 benchmarks, and the with\-KL variant improvespass@k\\textup\{pass@\}kcoverage on 8 of 9, so the two variants are complementary \(one leads onpass@1\\textup\{pass@\}1, the other on coverage\), and both dominate the base across all three domains\. The gains are those of a well\-tuned strong\-baseline comparison, not of over\-tuning: thepass@1\\textup\{pass@\}1margin over the matched RL baseline stays within a few points, while the coverage advantage of the with\-KL variant grows steadily withkkand is largest on the hardest problems, as Appendix Figure[G\.2](https://arxiv.org/html/2608.08224#A7.F2)shows\.
Second, the coverage shortfall at largekkis a known failure of RLVR: reward maximization trades large\-kkcoverage forpass@1\\textup\{pass@\}1, and the boundary can fall to or below the base\(Yueet al\.[2025](https://arxiv.org/html/2608.08224#bib.bib6)\)\. It tracks the reward objective, both matched GRPO baselines converging to the base on the hardest sets \(Appendix Table[G\.1](https://arxiv.org/html/2608.08224#A7.T1)\) atβ=0\\beta\{=\}0andβ=0\.01\\beta\{=\}0\.01alike, so it lies in the reward flux of Eq\. \([1](https://arxiv.org/html/2608.08224#S3.E1)\) that the analysis targets\. CD\-RFT addresses it in control space, where the diagnosis locates the bottleneck: on both axes it restores coverage to the base level or above by holding control decoupled while the reward flux is optimized, the behavioural counterpart of the mechanism of Section[6\.2](https://arxiv.org/html/2608.08224#S6.SS2)and, by construction, visible only at largekk\.
### 6\.4Ablation: Net Effect of the Regularizer
To isolate the effect of the control regularizer from backbone differences, we toggle only the on/off of the regularizer under a fully balanced setting\. This on/off contrast is precisely the matched GRPO→\\toCD\-RFT pair of Table[2](https://arxiv.org/html/2608.08224#S6.T2), whose control and capability axes we now read jointly\.
Both balanced pairs trace the same pattern: adding the regularizer lowersBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)\(mechanistic decoupling\), and overallpass@1\\textup\{pass@\}1andpass@k\\textup\{pass@\}krise in tandem\. The chain is tightest where the compression is strongest \(the no\-KL axis, whose larger drop inBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)coincides with the largestpass@1\\textup\{pass@\}1gain\), while the with\-KL axis compresses less but gains more on coverage\. The regularizer thus acts on the control axis, and the capability gain moves with the decoupling rather than with the backbone or hyperparameters\.
### 6\.5Robustness and Generalization
We test that the main result holds along three axes over which one might worry it was tuned: the evaluation temperature, the method’s own target hyperparameter, and the model family\.
##### Sampling temperature\.
We repeat the evaluation over a per\-domain temperature sweep bracketing the main\-table working temperature of each domain\. On the math hard set and on code, the ordering “CD\-RFT≥\\geqmatched RL” holds at every temperature onpass@1\\textup\{pass@\}1and at nearly all onpass@k\\textup\{pass@\}k; lowering the temperature raisespass@1\\textup\{pass@\}1and raising it raisespass@k\\textup\{pass@\}k, but the method ordering never flips, so the main\-table conclusions are not an artifact of a chosen temperature\. The full win\-rate table, per\-temperature curves, and the per\-domain breakdown are in Appendix[H](https://arxiv.org/html/2608.08224#A8)\.
##### Proxy target\.
Sweeping the target concentrationτ∈\{55,65,70\}\\tau\\in\\\{55,65,70\\\}\(onlyτ\\tauchanged\) traces a smoothpass@1\\textup\{pass@\}1–pass@k\\textup\{pass@\}ktrade\-off rather than a sharp optimum: the two looser settings are adjacent optima that exchange a littlepass@1\\textup\{pass@\}1for coverage, and only over\-compression \(τ=55\\tau\{=\}55\) is clearly worse on both axes\. The method is thus insensitive toτ\\tauby design rather than tuning, with the full ladder and per\-τ\\taureading in Appendix Table[I\.1](https://arxiv.org/html/2608.08224#A9.T1)\.
##### Second model family\.
Both the mechanism and the capability gain also hold on Llama\-3\.2\-3B\(Grattafioriet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib41)\): the control\-sharing signature reproduces, and CD\-RFT leadspass@k\\textup\{pass@\}kcoverage on every benchmark andpass@1\\textup\{pass@\}1on four of five, so the effect is not specific to Qwen2\.5\-7B\. Full protocol, ladders, and the mechanism table are in Appendix[K](https://arxiv.org/html/2608.08224#A11)\.
Table 3:Transfer to Llama\-3\.2\-3B \(base,β=0\\beta\{=\}0\)\. Each shaded \+ CD\-RFT row is the add\-on over the GRPO above it; L\-ord = Logic\-ordering, L\-graph = Logic\-graph; all columns are single benchmarks except code, the mean of HumanEval\+ and MBPP\. All entries are percentages\.
## 7Conclusion
We have reframed the interpretability of reinforcement\-learning post\-training: the question is not which components are activated, but which control the reward gain, and whether that control stays task\-specific or collapses into one shared channel\. This shift from an activation view to a control view turns a diagnosis into a lever, writing the Shared Control Bottleneck into the loss, CD\-RFT decouples control at low overhead and improves multi\-task pass@11and pass@kkacross two model families\. We see the control view as the broader contribution: once post\-training is read as reshaping which components control the reward flux, control\-level objectives become a natural handle on how a finite capability is split among competing tasks, beyond the multi\-task rule\-verifier setting studied here\.
## References
- J\. Austin, A\. Odena, M\. Nye, M\. Bosma, H\. Michalewski, D\. Dohan, E\. Jiang, C\. Cai, M\. Terry, Q\. Le, and C\. Sutton \(2021\)Program synthesis with large language models\.arXiv preprint arXiv:2108\.07732\.Cited by:[Table E\.2](https://arxiv.org/html/2608.08224#A5.T2)\.
- M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. d\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.\(2021\)Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[Table E\.2](https://arxiv.org/html/2608.08224#A5.T2),[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p2.1)\.
- Revisiting reinforcement learning for llm reasoning from a cross\-domain perspective\.arXiv preprint arXiv:2506\.14965\.Cited by:[Table E\.1](https://arxiv.org/html/2608.08224#A5.T1.1.4.3.2),[Table E\.2](https://arxiv.org/html/2608.08224#A5.T2),[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2),[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p1.2)\.
- A\. Conmy, A\. N\. Mavor\-Parker, A\. Lynch, S\. Heimersheim, and A\. Garriga\-Alonso \(2023\)Towards automated circuit discovery for mechanistic interpretability\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- T\. Dao, D\. Y\. Fu, S\. Ermon, A\. Rudra, and C\. Ré \(2022\)FlashAttention: fast and memory\-efficient exact attention with IO\-awareness\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§5\.1](https://arxiv.org/html/2608.08224#S5.SS1.p1.6)\.
- C\. Davis and W\. M\. Kahan \(1970\)The rotation of eigenvectors by a perturbation\. III\.SIAM Journal on Numerical Analysis7\(1\),pp\. 1–46\.Cited by:[Proposition C\.3](https://arxiv.org/html/2608.08224#A3.Thmproposition3.p1.11.11),[§5\.5](https://arxiv.org/html/2608.08224#S5.SS5.p1.5)\.
- M\. Defferrard, X\. Bresson, and P\. Vandergheynst \(2016\)Convolutional neural networks on graphs with fast localized spectral filtering\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§5\.2](https://arxiv.org/html/2608.08224#S5.SS2.p1.5)\.
- N\. Elhage, N\. Nanda, C\. Olsson, T\. Henighan, N\. Joseph, B\. Mann, A\. Askell, Y\. Bai, A\. Chen, T\. Conerly,et al\.\(2021\)A mathematical framework for transformer circuits\.Note:Transformer Circuits Thread;https://transformer\-circuits\.pub/2021/framework/index\.htmlAccessed: 2026\-07\-27Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- D\. A\. Fell \(1992\)Metabolic control analysis: a survey of its theoretical and experimental development\.Biochemical Journal286\(2\),pp\. 313–330\.Cited by:[§1](https://arxiv.org/html/2608.08224#S1.p3.1),[§3\.2](https://arxiv.org/html/2608.08224#S3.SS2.p1.3)\.
- A\. Grattafiori, A\. Dubey,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§6\.5](https://arxiv.org/html/2608.08224#S6.SS5.SSS0.Px3.p1.2)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi,et al\.\(2025\)DeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2)\.
- M\. Hanna, S\. Pezzelle, and Y\. Belinkov \(2024\)Have faith in faithfulness: going beyond circuit overlap when finding model mechanisms\.InConference on Language Modeling \(COLM\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Heinrich and T\. A\. Rapoport \(1974\)A linear steady\-state treatment of enzymatic chains: general properties, control and effector strength\.European Journal of Biochemistry42\(1\),pp\. 89–95\.Cited by:[§1](https://arxiv.org/html/2608.08224#S1.p3.1),[§3\.2](https://arxiv.org/html/2608.08224#S3.SS2.p1.2)\.
- D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. Steinhardt \(2021\)Measuring mathematical problem solving with the MATH dataset\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,Cited by:[Table E\.2](https://arxiv.org/html/2608.08224#A5.T2)\.
- S\. Jain, R\. Kirk, E\. S\. Lubana, R\. P\. Dick, H\. Tanaka, E\. Grefenstette, T\. Rocktäschel, and D\. S\. Krueger \(2024\)Mechanistically analyzing the effects of fine\-tuning on procedurally defined tasks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Kacser and J\. A\. Burns \(1973\)The control of flux\.Symposia of the Society for Experimental Biology27,pp\. 65–104\.Cited by:[§1](https://arxiv.org/html/2608.08224#S1.p3.1),[§3\.2](https://arxiv.org/html/2608.08224#S3.SS2.p1.2)\.
- T\. N\. Kipf and M\. Welling \(2017\)Semi\-supervised classification with graph convolutional networks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§5\.2](https://arxiv.org/html/2608.08224#S5.SS2.p1.5)\.
- J\. Kramár, T\. Lieberum, R\. Shah, and N\. Nanda \(2024\)AtP\*: an efficient and scalable method for localizing LLM behaviour to components\.arXiv preprint arXiv:2403\.00745\.Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Lambert, J\. Morrison, V\. Pyatkin, S\. Huang, H\. Ivison, F\. Brahman, L\. J\. V\. Miranda, A\. Liu, N\. Dziri, S\. Lyu,et al\.\(2024\)Tülu 3: pushing frontiers in open language model post\-training\.arXiv preprint arXiv:2411\.15124\.Cited by:[§3\.1](https://arxiv.org/html/2608.08224#S3.SS1.p1.3),[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p1.2)\.
- A\. Lewkowycz, A\. Andreassen, D\. Dohan, E\. Dyer, H\. Michalewski, V\. Ramasesh, A\. Slone, C\. Anil, I\. Schlag, T\. Gutman\-Solo,et al\.\(2022\)Solving quantitative reasoning problems with language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Table E\.2](https://arxiv.org/html/2608.08224#A5.T2)\.
- J\. Liu, C\. S\. Xia, Y\. Wang, and L\. Zhang \(2023\)Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Table E\.2](https://arxiv.org/html/2608.08224#A5.T2)\.
- J\. Liu and L\. Zhang \(2025\)Code\-r1: reproducing r1 for code with reliable rewards\.Note:https://github\.com/ganler/code\-r1Datasetganler/code\-r1\-12k; accessed: 2026\-07\-27Cited by:[Table E\.1](https://arxiv.org/html/2608.08224#A5.T1.1.3.2.2),[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p1.2)\.
- M\. Luo, S\. Tan, J\. Wong, X\. Shi, W\. Y\. Tang, M\. Roongta, C\. Cai, J\. Luo, L\. E\. Li, R\. A\. Popa, and I\. Stoica \(2025\)DeepScaleR: surpassing o1\-preview with a 1\.5b model by scaling rl\.Note:Notion Blog;https://github\.com/agentica\-project/rllmAccessed: 2026\-07\-27Cited by:[Table E\.1](https://arxiv.org/html/2608.08224#A5.T1.1.2.1.2),[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2),[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p1.2)\.
- J\. R\. Magnus and H\. Neudecker \(2019\)Matrix differential calculus with applications in statistics and econometrics\.3rd edition,John Wiley & Sons\.Cited by:[Proposition C\.1](https://arxiv.org/html/2608.08224#A3.Ex5.1.m1.1.1),[Proposition C\.1](https://arxiv.org/html/2608.08224#A3.Ex5.m2.1.1)\.
- A\. W\. Marshall, I\. Olkin, and B\. C\. Arnold \(2011\)Inequalities: theory of majorization and its applications\.2nd edition,Springer\.Cited by:[Appendix C](https://arxiv.org/html/2608.08224#A3.1.p1.2)\.
- N\. Nanda \(2023\)Attribution patching: activation patching at industrial scale\.Note:https://www\.neelnanda\.io/mechanistic\-interpretability/attribution\-patchingAccessed: 2026\-07\-27Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.08224#S3.SS3.p2.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.\(2022\)Training language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2608.08224#S1.p1.1),[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2),[§3\.1](https://arxiv.org/html/2608.08224#S3.SS1.p1.7)\.
- N\. Prakash, T\. R\. Shaham, T\. Haklay, Y\. Belinkov, and D\. Bau \(2024\)Fine\-tuning enhances existing mechanisms: a case study on entity tracking\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, S\. Ermon, C\. D\. Manning, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2),[§3\.1](https://arxiv.org/html/2608.08224#S3.SS1.p1.7)\.
- S\. Rajbhandari, J\. Rasley, O\. Ruwase, and Y\. He \(2020\)ZeRO: memory optimizations toward training trillion parameter models\.InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis \(SC\),Cited by:[Remark C\.4](https://arxiv.org/html/2608.08224#A3.Thmremark4.p1.1.1)\.
- S\. S\. Rameshet al\.\(2026\)Multi\-task grpo: reliable llm reasoning across tasks\.arXiv preprint arXiv:2602\.05547\.Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\.K\. Li, Y\. Wu, and D\. Guo \(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2),[§5\.1](https://arxiv.org/html/2608.08224#S5.SS1.p1.6),[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p1.2)\.
- M\. Sundararajan, A\. Taly, and Q\. Yan \(2017\)Axiomatic attribution for deep networks\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Syed, C\. Rager, and A\. Conmy \(2024\)Attribution patching outperforms automated circuit discovery\.InProceedings of the 7th BlackboxNLP Workshop \(EMNLP\),Cited by:[§F\.1](https://arxiv.org/html/2608.08224#A6.SS1.p1.1),[Figure 1](https://arxiv.org/html/2608.08224#S1.F1),[§1](https://arxiv.org/html/2608.08224#S1.p2.1),[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.08224#S3.SS3.p2.1)\.
- K\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt \(2023\)Interpretability in the wild: a circuit for indirect object identification in GPT\-2 small\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li,et al\.\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§6\.1](https://arxiv.org/html/2608.08224#S6.SS1.p1.2)\.
- Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.\(2025\)DAPO: an open\-source llm reinforcement learning system at scale\.arXiv preprint arXiv:2503\.14476\.Cited by:[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2)\.
- Y\. Yue, Z\. Chen, R\. Lu, A\. Zhao, Z\. Wang, Y\. Yue, S\. Song, and G\. Huang \(2025\)Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[Appendix K](https://arxiv.org/html/2608.08224#A11.SS0.SSS0.Px3.p1.6),[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px1.p1.2),[§6\.3](https://arxiv.org/html/2608.08224#S6.SS3.p3.6)\.
- H\. Zhang, Q\. Hao, F\. Xu, and Y\. Li \(2026\)Reinforcement learning fine\-tuning enhances activation intensity and diversity in the internal circuitry of llms\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§F\.1](https://arxiv.org/html/2608.08224#A6.SS1.p1.1),[Table F\.1](https://arxiv.org/html/2608.08224#A6.T1),[§1](https://arxiv.org/html/2608.08224#S1.p2.1),[§2](https://arxiv.org/html/2608.08224#S2.SS0.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2608.08224#S3.SS4.p4.4),[§6\.2](https://arxiv.org/html/2608.08224#S6.SS2.SSSx1.p1.1)\.
- Y\. Zhao, A\. Gu, R\. Varma, L\. Luo, C\. Huang, M\. Xu, L\. Wright, H\. Shojanazeri, M\. Ott, S\. Shleifer,et al\.\(2023\)PyTorch FSDP: experiences on scaling fully sharded data parallel\.Proceedings of the VLDB Endowment16\(12\),pp\. 3848–3860\.Cited by:[Remark C\.4](https://arxiv.org/html/2608.08224#A3.Thmremark4.p1.1.1)\.
APPENDIX
###### Contents
1. [1Introduction](https://arxiv.org/html/2608.08224#S1)
2. [2Related Work](https://arxiv.org/html/2608.08224#S2)
3. [3The Post\-Training Control Coefficient](https://arxiv.org/html/2608.08224#S3)1. [3\.1Post\-Training Reward Flux](https://arxiv.org/html/2608.08224#S3.SS1) 2. [3\.2Metabolic Control Analysis](https://arxiv.org/html/2608.08224#S3.SS2) 3. [3\.3Differentiable Gating and Attribution Patching](https://arxiv.org/html/2608.08224#S3.SS3) 4. [3\.4The Coefficient and Its Expansion](https://arxiv.org/html/2608.08224#S3.SS4)
4. [4The Shared Control Bottleneck](https://arxiv.org/html/2608.08224#S4)1. [4\.1Cross\-Task Concentration](https://arxiv.org/html/2608.08224#S4.SS1) 2. [4\.2The Activation–Control Gap](https://arxiv.org/html/2608.08224#S4.SS2) 3. [4\.3Distributedness of Control](https://arxiv.org/html/2608.08224#S4.SS3)
5. [5Control\-Diverse Regularizer](https://arxiv.org/html/2608.08224#S5)1. [5\.1Objective](https://arxiv.org/html/2608.08224#S5.SS1) 2. [5\.2A Stable Spectral\-Moment Ratio Replacing the Spectral Extremum](https://arxiv.org/html/2608.08224#S5.SS2) 3. [5\.3Closed\-Form Sensitivity and the Irreplaceability of the Control Axis](https://arxiv.org/html/2608.08224#S5.SS3) 4. [5\.4First\-Order Proxy Gradient via Central Differences in Gate Space](https://arxiv.org/html/2608.08224#S5.SS4) 5. [5\.5Estimation Protocol](https://arxiv.org/html/2608.08224#S5.SS5) 6. [5\.6Implementation and Scalability](https://arxiv.org/html/2608.08224#S5.SS6)
6. [6Experiments](https://arxiv.org/html/2608.08224#S6)1. [6\.1Experimental Setup](https://arxiv.org/html/2608.08224#S6.SS1) 2. [6\.2Mechanistic Diagnosis of the Shared Control Bottleneck](https://arxiv.org/html/2608.08224#S6.SS2) 3. [6\.3Multi\-Task Capability](https://arxiv.org/html/2608.08224#S6.SS3) 4. [6\.4Ablation: Net Effect of the Regularizer](https://arxiv.org/html/2608.08224#S6.SS4) 5. [6\.5Robustness and Generalization](https://arxiv.org/html/2608.08224#S6.SS5)
7. [7Conclusion](https://arxiv.org/html/2608.08224#S7)
8. [References](https://arxiv.org/html/2608.08224#bib)
9. [AFull Specification of the Control Coefficient](https://arxiv.org/html/2608.08224#A1)
10. [BFull Specification of the Shared Control Bottleneck](https://arxiv.org/html/2608.08224#A2)
11. [CFull Specification of the Control\-Diverse Regularizer](https://arxiv.org/html/2608.08224#A3)
12. [DCD\-RFT Training Step](https://arxiv.org/html/2608.08224#A4)
13. [EFull Experimental Settings](https://arxiv.org/html/2608.08224#A5)
14. [FAdditional Mechanistic Results](https://arxiv.org/html/2608.08224#A6)1. [F\.1Activation\-Level Metrics](https://arxiv.org/html/2608.08224#A6.SS1) 2. [F\.2Random Reference forBsharedB\_\{\\mathrm\{shared\}\}](https://arxiv.org/html/2608.08224#A6.SS2) 3. [F\.3Direction Versus Row Norm](https://arxiv.org/html/2608.08224#A6.SS3) 4. [F\.4Paired Verification and Cross\-Model Generality](https://arxiv.org/html/2608.08224#A6.SS4)
15. [GAdditional Capability Results](https://arxiv.org/html/2608.08224#A7)1. [G\.1Full Per\-Benchmark pass@kkLadder](https://arxiv.org/html/2608.08224#A7.SS1) 2. [G\.2Per\-Benchmark Paired Gains and Coverage](https://arxiv.org/html/2608.08224#A7.SS2)
16. [HFull Temperature\-Sweep Results](https://arxiv.org/html/2608.08224#A8)
17. [IProxy\-Target Sensitivity](https://arxiv.org/html/2608.08224#A9)
18. [JCross\-Checkpoint Robustness](https://arxiv.org/html/2608.08224#A10)
19. [KSecond Model: Llama\-3\.2\-3B](https://arxiv.org/html/2608.08224#A11)
20. [LComputational Overhead of the Control Regularizer](https://arxiv.org/html/2608.08224#A12)
This appendix has two parts\. Appendices[A](https://arxiv.org/html/2608.08224#A1)–[D](https://arxiv.org/html/2608.08224#A4)complete the theory: the function\-space conventions, the full set of assumptions, the secondary propositions, and the proofs abbreviated in the main text, with Appendix[A](https://arxiv.org/html/2608.08224#A1)corresponding to the setting \(Section[3](https://arxiv.org/html/2608.08224#S3)\), Appendix[B](https://arxiv.org/html/2608.08224#A2)to the problem \(Section[4](https://arxiv.org/html/2608.08224#S4)\), Appendix[C](https://arxiv.org/html/2608.08224#A3)to the method \(Section[5](https://arxiv.org/html/2608.08224#S5)\), and Appendix[D](https://arxiv.org/html/2608.08224#A4)giving one training step in full\. Appendices[E](https://arxiv.org/html/2608.08224#A5)–[L](https://arxiv.org/html/2608.08224#A12)complete the experiments: Appendix[E](https://arxiv.org/html/2608.08224#A5)fixes the settings shared by every reported number, Appendix[F](https://arxiv.org/html/2608.08224#A6)the mechanistic results deferred from Section[6\.2](https://arxiv.org/html/2608.08224#S6.SS2), Appendix[G](https://arxiv.org/html/2608.08224#A7)the per\-benchmark capability results deferred from Section[6\.3](https://arxiv.org/html/2608.08224#S6.SS3), Appendix[H](https://arxiv.org/html/2608.08224#A8)the temperature sweep and Appendix[I](https://arxiv.org/html/2608.08224#A9)the proxy\-target sweep of Section[6\.5](https://arxiv.org/html/2608.08224#S6.SS5), Appendix[J](https://arxiv.org/html/2608.08224#A10)the cross\-checkpoint robustness check, Appendix[K](https://arxiv.org/html/2608.08224#A11)the second\-model \(Llama\-3\.2\-3B\) transfer results of Section[6\.5](https://arxiv.org/html/2608.08224#S6.SS5), and Appendix[L](https://arxiv.org/html/2608.08224#A12)the measured cost of the regularizer\.
## Appendix AFull Specification of the Control Coefficient
##### Notation and function space\.
The vocabulary𝒱\\mathcal\{V\}is finite, with𝒱⋆=⋃T≥0𝒱T\\mathcal\{V\}^\{\\star\}=\\bigcup\_\{T\\geq 0\}\\mathcal\{V\}^\{T\}; a prompt isx∈𝒳x\\in\\mathcal\{X\}and a response isy∈𝒱⋆y\\in\\mathcal\{V\}^\{\\star\}\. The policy factorizes autoregressively asπθ\(y∣x\)=∏tπθ\(yt∣x,y<t\)\\pi\_\{\\theta\}\(y\\mid x\)=\\prod\_\{t\}\\pi\_\{\\theta\}\(y\_\{t\}\\mid x,y\_\{<t\}\), with parametersθ∈Θ⊆ℝP\\theta\\in\\Theta\\subseteq\\mathbb\{R\}^\{P\}\(an open set\) and reference policyπ0=πθ0\\pi\_\{0\}=\\pi\_\{\\theta\_\{0\}\}\. The residual\-stream dimension isdd; the state at layerℓ\\ellishℓ∈ℝdh^\{\\ell\}\\in\\mathbb\{R\}^\{d\}; circuit unitkkresides at layerℓ\(k\)\\ell\(k\)with mapfk\(⋅;θ\):ℝd→ℝdf\_\{k\}\(\\cdot;\\theta\):\\mathbb\{R\}^\{d\}\\to\\mathbb\{R\}^\{d\}\. The gradient of the target log\-likelihood with respect to the final residual is thereadout directionwm:=∇hL∑tlogπθ\(yt⋆∣x,y<t⋆\)∈ℝdw\_\{m\}:=\\nabla\_\{h^\{L\}\}\\sum\_\{t\}\\log\\pi\_\{\\theta\}\(y^\{\\star\}\_\{t\}\\mid x,y^\{\\star\}\_\{<t\}\)\\in\\mathbb\{R\}^\{d\}\(the linear readout direction, at the final layer, of the metric used by attribution patching\); the update Jacobian of layerℓ′\\ell^\{\\prime\}isJℓ′:=∂\(∑k:ℓ\(k\)=ℓ′fk\)/∂hℓ′J\_\{\\ell^\{\\prime\}\}:=\\partial\\big\(\\sum\_\{k:\\ell\(k\)=\\ell^\{\\prime\}\}f\_\{k\}\\big\)/\\partial h^\{\\ell^\{\\prime\}\}\. The control matrix hasM=MfamM=M\_\{\\mathrm\{fam\}\}rows \(one per task family, each averaged over itsnnprobe samples\) andKKcolumns for the gated components\.
###### Assumption A\.1\(smoothness and integrability\)\.
Writeπθ\(y∣x;g\)\\pi\_\{\\theta\}\(y\\mid x;g\)for the gated policy of Eq\. \([3](https://arxiv.org/html/2608.08224#S3.E3)\)\. For almost all\(x,y⋆\)\(x,y^\{\\star\}\)the map\(θ,g\)↦logπθ\(y⋆∣x;g\)\(\\theta,g\)\\mapsto\\log\\pi\_\{\\theta\}\(y^\{\\star\}\\mid x;g\)is smooth \(C∞C^\{\\infty\}\); for eachmm,\|logπθ\(y⋆∣x;g\)\|\\big\|\\log\\pi\_\{\\theta\}\(y^\{\\star\}\\mid x;g\)\\big\|and its partial derivatives inθ\\thetaand inggare dominated by𝒟m\\mathcal\{D\}\_\{m\}\-integrable functions, so that expectation and differentiation commute\. Henceℓm\(θ;g\)=𝔼𝒟m\[logπθ\(y⋆∣x;g\)\]\\ell\_\{m\}\(\\theta;g\)=\\mathbb\{E\}\_\{\\mathcal\{D\}\_\{m\}\}\[\\log\\pi\_\{\\theta\}\(y^\{\\star\}\\mid x;g\)\]is smooth in both arguments, and so isJm\(θ;g\)J\_\{m\}\(\\theta;g\), which differs fromℓm\(θ;g\)\\ell\_\{m\}\(\\theta;g\)by the gate\-independent constant𝔼𝒟m\[logπ0\(y⋆∣x\)\]\\mathbb\{E\}\_\{\\mathcal\{D\}\_\{m\}\}\[\\log\\pi\_\{0\}\(y^\{\\star\}\\mid x\)\]: every statement below may therefore be read for either object\.
###### Assumption A\.2\(nominal identity\)\.
πθ\(⋅∣⋅;𝟏\)=πθ\\pi\_\{\\theta\}\(\\cdot\\mid\\cdot\\,;\\mathbf\{1\}\)=\\pi\_\{\\theta\}: at the nominal gate value the gated network is the original model\. The gating is thus an analysis probe, and all control quantities are differentiated in a neighborhood ofg=𝟏g=\\mathbf\{1\}\.
###### Assumption A\.3\(normalizer non\-degeneracy\)\.
There existsϵ0\>0\\epsilon\_\{0\}\>0\(set to10−410^\{\-4\}in the implementation\) such that every family admitted for analysis satisfies\|ℓm\(θ\)\|≥ϵ0\|\\ell\_\{m\}\(\\theta\)\|\\geq\\epsilon\_\{0\}; otherwise Eq\. \([4](https://arxiv.org/html/2608.08224#S3.E4)\) is unidentifiable and the family is discarded\.
###### Proposition A\.1\(first\-order expansion; main text Eq\. \([5](https://arxiv.org/html/2608.08224#S3.E5)\)\)\.
Under Assumptions[A\.1](https://arxiv.org/html/2608.08224#A1.Thmassumption1)–[A\.3](https://arxiv.org/html/2608.08224#A1.Thmassumption3), withu=loggu=\\log g,
ℓm\(θ;eu\)−ℓm\(θ\)ℓm\(θ\)=∑kCm,kuk\+O\(∥u∥22\)\.\\frac\{\\ell\_\{m\}\(\\theta;e^\{u\}\)\-\\ell\_\{m\}\(\\theta\)\}\{\\ell\_\{m\}\(\\theta\)\}=\\sum\_\{k\}C\_\{m,k\}\\,u\_\{k\}\+O\(\\lVert u\\rVert\_\{2\}^\{2\}\)\.
###### Proof\.
Letℓ~m\(u\)=ℓm\(θ;eu\)\\widetilde\{\\ell\}\_\{m\}\(u\)=\\ell\_\{m\}\(\\theta;e^\{u\}\), which is smooth by Assumption[A\.1](https://arxiv.org/html/2608.08224#A1.Thmassumption1)\. The chain rule gives∂ukℓ~m\|0=∂gkℓm\|g=𝟏⋅euk\|0=∂gkℓm\|g=𝟏\\partial\_\{u\_\{k\}\}\\widetilde\{\\ell\}\_\{m\}\|\_\{0\}=\\partial\_\{g\_\{k\}\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}\\cdot e^\{u\_\{k\}\}\|\_\{0\}=\\partial\_\{g\_\{k\}\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}\. A first\-order Taylor expansion with Peano remainder, divided byℓm\(θ\)≠0\\ell\_\{m\}\(\\theta\)\\neq 0\(Assumption[A\.3](https://arxiv.org/html/2608.08224#A1.Thmassumption3)\), and substitutingCm,k=1ℓm∂gkℓm\|g=𝟏C\_\{m,k\}=\\tfrac\{1\}\{\\ell\_\{m\}\}\\partial\_\{g\_\{k\}\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}, yields the claim\. ∎
###### Corollary A\.1\(logarithmic form\)\.
Sinceℓm<0\\ell\_\{m\}<0, writingz=δℓm/ℓmz=\\delta\\ell\_\{m\}/\\ell\_\{m\}andlog\(1\+z\)=z\+O\(z2\)\\log\(1\+z\)=z\+O\(z^\{2\}\)givesδlog\|ℓm\|=∑kCm,kδloggk\+O\(∥δlogg∥22\)\\delta\\log\|\\ell\_\{m\}\|=\\sum\_\{k\}C\_\{m,k\}\\,\\delta\\log g\_\{k\}\+O\(\\lVert\\delta\\log g\\rVert\_\{2\}^\{2\}\), the sign\-robust analogue of the classical MCA log–log derivative\.
###### Lemma A\.1\(activation axis and control axis are not interchangeable\)\.
There exist networks and task families for which no scalarcm\>0c\_\{m\}\>0satisfiesCm,⋅=cmFm,⋅C\_\{m,\\cdot\}=c\_\{m\}F\_\{m,\\cdot\}\.
###### Proof\.
JmJ\_\{m\}is determined by the projection ofhLh^\{L\}ontowmw\_\{m\}\. The first\-order effect of gategkg\_\{k\}onJmJ\_\{m\}propagates through Eq\. \([3](https://arxiv.org/html/2608.08224#S3.E3)\) tohLh^\{L\}and is read alongwmw\_\{m\}\. Choose a circuitkkwhose writeoutk⟂wm\\mathrm\{out\}\_\{k\}\\perp w\_\{m\}and whose downstream Jacobian does not rotate it intowmw\_\{m\}; then∂gkJm\|g=𝟏=0\\partial\_\{g\_\{k\}\}J\_\{m\}\|\_\{g=\\mathbf\{1\}\}=0, henceCm,k≈0C\_\{m,k\}\\approx 0, whileFm,k=𝔼∥outk∥2F\_\{m,k\}=\\mathbb\{E\}\\lVert\\mathrm\{out\}\_\{k\}\\rVert\_\{2\}can be made independently large\. Dually, a circuit with small write aligned withwmw\_\{m\}has low usage and high control\. The coexistence makes the large\-component supports ofCm,⋅C\_\{m,\\cdot\}andFm,⋅F\_\{m,\\cdot\}differ, excluding any positive proportionality\. ∎
## Appendix BFull Specification of the Shared Control Bottleneck
###### Proposition B\.1\(failure of the summation theorem\)\.
In general,∑kCm,k≠1\\sum\_\{k\}C\_\{m,k\}\\neq 1\.
###### Proof\.
By Eq\. \([4](https://arxiv.org/html/2608.08224#S3.E4)\),∑kgk∂gkℓm\|g=𝟏=ℓm∑kCm,k\\sum\_\{k\}g\_\{k\}\\partial\_\{g\_\{k\}\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}=\\ell\_\{m\}\\sum\_\{k\}C\_\{m,k\}\. Ifg↦ℓm\(θ;g\)g\\mapsto\\ell\_\{m\}\(\\theta;g\)were first\-order homogeneous atg=𝟏g=\\mathbf\{1\}, Euler’s theorem would give the left\-hand side=ℓm=\\ell\_\{m\}, whence∑kCm,k=1\\sum\_\{k\}C\_\{m,k\}=1\. But the residual update in Eq\. \([3](https://arxiv.org/html/2608.08224#S3.E3)\) contains the identity termhℓh^\{\\ell\}: under uniform scalingg↦αgg\\mapsto\\alpha g,hℓh^\{\\ell\}does not scale withα\\alpha, and layer normalization is not homogeneous, soℓm\(θ;αg\)\\ell\_\{m\}\(\\theta;\\alpha g\)is not first\-order homogeneous inα\\alpha, the left\-hand side≠ℓm\\neq\\ell\_\{m\}, and therefore∑kCm,k≠1\\sum\_\{k\}C\_\{m,k\}\\neq 1\. ∎
###### Theorem B\.1\(no single\-gate monopoly; main text Theorem[1](https://arxiv.org/html/2608.08224#Thmtheorem1)\)\.
Under generic weights,∂gkJm\|g=𝟏≠0\\partial\_\{g\_\{k\}\}J\_\{m\}\|\_\{g=\\mathbf\{1\}\}\\neq 0for everykk\.
###### Proof\.
Differentiating fork∉Sk\\notin S, gategkg\_\{k\}writesfk\(hℓ\(k\);θ\)f\_\{k\}\(h^\{\\ell\(k\)\};\\theta\)at its resident layer and propagates through the subsequent layer Jacobians tohLh^\{L\}:
∂Jm∂gk\|g=𝟏=wm⊤\(∏ℓ′=ℓ\(k\)\+1L−1\(I\+Jℓ′\)\)fk\(hℓ\(k\);θ\),\\frac\{\\partial J\_\{m\}\}\{\\partial g\_\{k\}\}\\bigg\|\_\{g=\\mathbf\{1\}\}=w\_\{m\}^\{\\top\}\\Big\(\\textstyle\\prod\_\{\\ell^\{\\prime\}=\\ell\(k\)\+1\}^\{L\-1\}\(I\+J\_\{\\ell^\{\\prime\}\}\)\\Big\)f\_\{k\}\(h^\{\\ell\(k\)\};\\theta\),withJℓ′J\_\{\\ell^\{\\prime\}\}the layer update Jacobian defined above\. Vanishing at a givenkkrequires the propagated write to lie inwm⟂w\_\{m\}^\{\\perp\}; for eachkk, with real\-analytic activations this is the zero set of a not\-identically\-zero real\-analytic function ofθ\\theta, hence a measure\-zero set in parameter space, and a finite intersection remains measure zero\. Hence no gate derivative vanishes under generic weights, and in particular the control support is not a single gate\. ∎
###### Proposition B\.2\(bounds, saturation, invariance ofBsharedB\_\{\\mathrm\{shared\}\}\)\.
LetGXG\_\{X\}have eigenvaluesλ1≥⋯≥λM≥0\\lambda\_\{1\}\\geq\\dots\\geq\\lambda\_\{M\}\\geq 0\. Then \(i\)1/M≤Bshared≤11/M\\leq B\_\{\\mathrm\{shared\}\}\\leq 1; \(ii\)=1⇔rank\(X\)=1=1\\iff\\operatorname\{rank\}\(X\)=1\(collinear rows\); \(iii\)=1/M⇔λ1=⋯=λM=1/M\\iff\\lambda\_\{1\}=\\dots=\\lambda\_\{M\}\(orthogonal, equal\-norm rows\); \(iv\) invariant underX↦cXX\\mapsto cX\(c≠0c\\neq 0\) andX↦XQX\\mapsto XQ\(QQ⊤=IQQ^\{\\top\}=I\)\.
###### Proof\.
\(i\)–\(iii\):Bshared=λ1/∑iλiB\_\{\\mathrm\{shared\}\}=\\lambda\_\{1\}/\\sum\_\{i\}\\lambda\_\{i\};∑iλi≤Mλ1\\sum\_\{i\}\\lambda\_\{i\}\\leq M\\lambda\_\{1\}gives the lower bound \(equality⇔\\iffall eigenvalues equal\), andλ1≤∑iλi\\lambda\_\{1\}\\leq\\sum\_\{i\}\\lambda\_\{i\}the upper bound \(equality⇔λ≥2=0⇔\\iff\\lambda\_\{\\geq 2\}=0\\iffrank one⇔\\iffcollinear rows\)\. \(iv\):GcX=c2GXG\_\{cX\}=c^\{2\}G\_\{X\}scales all eigenvalues byc2c^\{2\}, leaving the ratio unchanged, whileGXQ=XQQ⊤X⊤=GXG\_\{XQ\}=XQQ^\{\\top\}X^\{\\top\}=G\_\{X\}\. ∎
###### Proposition B\.3\(gap range and characterization; main text Proposition[3](https://arxiv.org/html/2608.08224#Thmproposition3)\)\.
AsBshared\(F\)→1B\_\{\\mathrm\{shared\}\}\(F\)\\to 1: \(i\)ACG∈\[0,Bshared\(F\)−1/M\]→\[0,1−1/M\]\\mathrm\{ACG\}\\in\[0,B\_\{\\mathrm\{shared\}\}\(F\)\-1/M\]\\to\[0,1\-1/M\]; \(ii\)ACG→0⇔Bshared\(C\)→Bshared\(F\)\\mathrm\{ACG\}\\to 0\\iff B\_\{\\mathrm\{shared\}\}\(C\)\\to B\_\{\\mathrm\{shared\}\}\(F\); \(iii\)ACG\\mathrm\{ACG\}is large⇔Bshared\(C\)\\iff B\_\{\\mathrm\{shared\}\}\(C\)lies far below the ceiling\.
###### Proof\.
The range follows from Proposition[B\.2](https://arxiv.org/html/2608.08224#A2.Thmproposition2)\(i\),Bshared\(C\)∈\[1/M,1\]B\_\{\\mathrm\{shared\}\}\(C\)\\in\[1/M,1\], with Eq\. \([8](https://arxiv.org/html/2608.08224#S4.E8)\); parts \(ii\)–\(iii\) are read from Eq\. \([8](https://arxiv.org/html/2608.08224#S4.E8)\) and given mechanistic meaning via Proposition[B\.2](https://arxiv.org/html/2608.08224#A2.Thmproposition2)\(ii\)–\(iii\)\. ∎
###### Proposition B\.4\(perturbation variance\)\.
When the gate logarithms are perturbed byξ∼𝒩\(0,Σ\)\\xi\\sim\\mathcal\{N\}\(0,\\Sigma\), the variance of the relative change is
Var\[δℓmℓm\]≈Cm,⋅⊤ΣCm,⋅→Σ=σ2Iσ2∥Cm,⋅∥22\.\\operatorname\{Var\}\\\!\\left\[\\frac\{\\delta\\ell\_\{m\}\}\{\\ell\_\{m\}\}\\right\]\\approx C\_\{m,\\cdot\}^\{\\top\}\\Sigma\\,C\_\{m,\\cdot\}\\ \\xrightarrow\{\\ \\Sigma=\\sigma^\{2\}I\\ \}\\ \\sigma^\{2\}\\lVert C\_\{m,\\cdot\}\\rVert\_\{2\}^\{2\}\.
###### Proof\.
By Eq\. \([5](https://arxiv.org/html/2608.08224#S3.E5)\),δℓm/ℓm≈Cm,⋅⊤ξ\\delta\\ell\_\{m\}/\\ell\_\{m\}\\approx C\_\{m,\\cdot\}^\{\\top\}\\xi; the variance of a zero\-mean Gaussian linear form isCm,⋅⊤ΣCm,⋅C\_\{m,\\cdot\}^\{\\top\}\\Sigma C\_\{m,\\cdot\}, which in the isotropic case equalsσ2∥Cm,⋅∥22\\sigma^\{2\}\\lVert C\_\{m,\\cdot\}\\rVert\_\{2\}^\{2\}\. ∎
###### Corollary B\.1\(concentration implies fragility\)\.
At fixed total control energytr\(GC\)\\operatorname\{tr\}\(G\_\{C\}\), the largerBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\), the more energy concentrates along the leading direction and the larger the perturbation variance of Proposition[B\.4](https://arxiv.org/html/2608.08224#A2.Thmproposition4)for aligned families\.
###### Proof\.
Immediate from Proposition[B\.4](https://arxiv.org/html/2608.08224#A2.Thmproposition4): at fixedtr\(GC\)\\operatorname\{tr\}\(G\_\{C\}\), raisingBshared\(C\)=λmax/tr\(GC\)B\_\{\\mathrm\{shared\}\}\(C\)=\\lambda\_\{\\max\}/\\operatorname\{tr\}\(G\_\{C\}\)raisesλmax\\lambda\_\{\\max\}, which is the variance coefficient along the leading direction\. ∎
This is a statement about the geometry ofCCalone; its connection to real deployment \(quantization, pruning, out\-of\-distribution behavior\) is left to future work\.
## Appendix CFull Specification of the Control\-Diverse Regularizer
###### Proposition C\.1\(non\-smoothness of the spectral extremum at degeneracy\)\.
The spectral extremum has eigenvalue gradient and eigenvector sensitivity
∂λ1∂C\\displaystyle\\frac\{\\partial\\lambda\_\{1\}\}\{\\partial C\}=2v1v1⊤C,\\displaystyle=2v\_\{1\}v\_\{1\}^\{\\top\}C,∂v1∂C\\displaystyle\\frac\{\\partial v\_\{1\}\}\{\\partial C\}∝∑i≥2vivi⊤λ1−λi\(Magnus and Neudecker[2019](https://arxiv.org/html/2608.08224#bib.bib29)\);\\displaystyle\\propto\\sum\_\{i\\geq 2\}\\frac\{v\_\{i\}v\_\{i\}^\{\\top\}\}\{\\lambda\_\{1\}\-\\lambda\_\{i\}\}\\quad\\text\{\\cite\[citep\]\{\(\\@@bibref\{AuthorsPhrase1Year\}\{magnus2019matrix\}\{\\@@citephrase\{ \}\}\{\}\)\}\};the eigenvalue gradient2v1v1⊤C2v\_\{1\}v\_\{1\}^\{\\top\}Cstays bounded, but the eigenvector sensitivity diverges at top\-eigenvalue degeneracyλ1=λ2\\lambda\_\{1\}=\\lambda\_\{2\}asv1v\_\{1\}becomes non\-unique, so backpropagation through the eigendecomposition is unstable there\.
###### Proposition C\.2\(legitimacy of the spectral\-moment\-ratio proxy\)\.
The spectral\-moment ratio
ℛ\(C\)=tr\(GC2\)tr\(GC\)2=∑iλ^i2\\mathcal\{R\}\(C\)=\\frac\{\\operatorname\{tr\}\(G\_\{C\}^\{2\}\)\}\{\\operatorname\{tr\}\(G\_\{C\}\)^\{2\}\}=\\sum\_\{i\}\\hat\{\\lambda\}\_\{i\}^\{2\}is \(i\) purely polynomial \(matvec, no eigendecomposition\); \(ii\) co\-monotone withBshared=maxiλ^iB\_\{\\mathrm\{shared\}\}=\\max\_\{i\}\\hat\{\\lambda\}\_\{i\}with respect to the majorization order, sharing the extremal values\{1,1/M\}\\\{1,1/M\\\}; \(iii\)C∞C^\{\\infty\}atC≠0C\\neq 0, free of any1/\(λ1−λi\)1/\(\\lambda\_\{1\}\-\\lambda\_\{i\}\)singularity\.
###### Proof\.
\(ii\): both∑iλ^i2\\sum\_\{i\}\\hat\{\\lambda\}\_\{i\}^\{2\}andmaxiλ^i\\max\_\{i\}\\hat\{\\lambda\}\_\{i\}are Schur\-convex\(Marshallet al\.[2011](https://arxiv.org/html/2608.08224#bib.bib30)\)and co\-monotone with respect to the spectral majorization order, attaining their lower bound at the uniform spectrum and upper bound at the single\-point spectrum; since majorization is only a partial order, spectra incomparable under it may rank differently, so we track empirical agreement of the two during training\. \(i\) and \(iii\) are immediate\. ∎
###### Lemma C\.1\(closed\-form sensitivity; main text Section[5\.3](https://arxiv.org/html/2608.08224#S5.SS3)\)\.
W=∂ℛ/∂C\|C¯=4tr\(G¯\)2\(G¯C¯−ℛ¯tr\(G¯\)C¯\)W=\\partial\\mathcal\{R\}/\\partial C\|\_\{\\bar\{C\}\}=\\tfrac\{4\}\{\\operatorname\{tr\}\(\\bar\{G\}\)^\{2\}\}\\big\(\\bar\{G\}\\bar\{C\}\-\\bar\{\\mathcal\{R\}\}\\operatorname\{tr\}\(\\bar\{G\}\)\\bar\{C\}\\big\)\.
###### Proof\.
With∂tr\(G2\)/∂C=4GC\\partial\\operatorname\{tr\}\(G^\{2\}\)/\\partial C=4GCand∂tr\(G\)/∂C=2C\\partial\\operatorname\{tr\}\(G\)/\\partial C=2C, the quotient rule gives
∂ℛ∂C\\displaystyle\\frac\{\\partial\\mathcal\{R\}\}\{\\partial C\}=4GCtr\(G\)2−2tr\(G2\)tr\(G\)2Ctr\(G\)4\\displaystyle=\\frac\{4GC\\operatorname\{tr\}\(G\)^\{2\}\-2\\operatorname\{tr\}\(G^\{2\}\)\\operatorname\{tr\}\(G\)\\,2C\}\{\\operatorname\{tr\}\(G\)^\{4\}\}=4tr\(G\)2\(GC−ℛtr\(G\)C\),\\displaystyle=\\frac\{4\}\{\\operatorname\{tr\}\(G\)^\{2\}\}\\big\(GC\-\\mathcal\{R\}\\operatorname\{tr\}\(G\)C\\big\),which, evaluated atC¯\\bar\{C\}, yields the claim\. ∎
###### Theorem C\.1\(irreplaceability of the control axis\)\.
ReplacingC\(θ\)C\(\\theta\)by the activation usageAk=𝔼∥outk∥2A\_\{k\}=\\mathbb\{E\}\\lVert\\mathrm\{out\}\_\{k\}\\rVert\_\{2\}fails on two structural grounds: \(i\) it violatesF≠CF\\neq C\(Lemma[A\.1](https://arxiv.org/html/2608.08224#A1.Thmlemma1)\); \(ii\) the per\-gate scalarwk=∑mrelu\(Wm,k\)w\_\{k\}=\\sum\_\{m\}\\mathrm\{relu\}\(W\_\{m,k\}\)sums out the family dimension and cannot express per\-task decoupling\.
###### Proof\.
\(i\) follows from Lemma[A\.1](https://arxiv.org/html/2608.08224#A1.Thmlemma1); \(ii\)AkA\_\{k\}is family\-independent and∑m\\sum\_\{m\}discards the per\-task directional information inWW\. ∎
###### Theorem C\.2\(single\-backpropagation equivalence via central differences; main text Theorem[3](https://arxiv.org/html/2608.08224#Thmtheorem3)\)\.
Equation \([13](https://arxiv.org/html/2608.08224#S5.E13)\) holds: the proxy gradient is obtained by a single backpropagation and differs from∇θℛ\\nabla\_\{\\theta\}\\mathcal\{R\}byO\(ϵ2\)O\(\\epsilon^\{2\}\)\.
###### Proof\.
Decomposing the Frobenius inner product by rows and substitutingCm,⋅=1ℓm∂gℓmC\_\{m,\\cdot\}=\\tfrac\{1\}\{\\ell\_\{m\}\}\\partial\_\{g\}\\ell\_\{m\}gives⟨W,C⟩=∑m1ℓmDWm,⋅ℓm\\langle W,C\\rangle=\\sum\_\{m\}\\tfrac\{1\}\{\\ell\_\{m\}\}D\_\{W\_\{m,\\cdot\}\}\\ell\_\{m\}\. Expandingϕ\(t\)=ℓm\(θ;𝟏\+tW^m,⋅\)\\phi\(t\)=\\ell\_\{m\}\(\\theta;\\mathbf\{1\}\+t\\hat\{W\}\_\{m,\\cdot\}\)at0, we haveϕ\(ϵ\)−ϕ\(−ϵ\)=2ϵϕ′\(0\)\+O\(ϵ3\)\\phi\(\\epsilon\)\-\\phi\(\-\\epsilon\)=2\\epsilon\\phi^\{\\prime\}\(0\)\+O\(\\epsilon^\{3\}\)\(ϕ\\phismooth by Assumption[A\.1](https://arxiv.org/html/2608.08224#A1.Thmassumption1)\) withϕ′\(0\)=DW^m,⋅ℓm\\phi^\{\\prime\}\(0\)=D\_\{\\hat\{W\}\_\{m,\\cdot\}\}\\ell\_\{m\}; multiplying back by∥Wm,⋅∥2\\lVert W\_\{m,\\cdot\}\\rVert\_\{2\}restores the term with errorO\(ϵ2\)O\(\\epsilon^\{2\}\)\. The gate scalings𝟏±ϵW^\\mathbf\{1\}\\pm\\epsilon\\hat\{W\}are constant vectors independent ofθ\\theta, so the two forward\-pass scalars are ordinarily differentiable inθ\\thetaand a single backpropagation yields their gradient; the composition equals∇θ⟨W,C\(θ\)⟩=∇θℛ\\nabla\_\{\\theta\}\\langle W,C\(\\theta\)\\rangle=\\nabla\_\{\\theta\}\\mathcal\{R\}, with difference the central\-difference remainderO\(ϵ2\)O\(\\epsilon^\{2\}\)\. ∎
###### Corollary C\.1\(directional consistency\)\.
The cosine between the proxy gradient and∇θℛ\\nabla\_\{\\theta\}\\mathcal\{R\}tends to11asϵ→0\\epsilon\\to 0\.
###### Proof\.
By Theorem[C\.2](https://arxiv.org/html/2608.08224#A3.Thmtheorem2)the two differ byO\(ϵ2\)O\(\\epsilon^\{2\}\)while∇θℛ\\nabla\_\{\\theta\}\\mathcal\{R\}is fixed, so the angle between them vanishes withϵ\\epsilon\. ∎
At toy scale the measured cosine is≈1\.0\\approx 1\.0\.
Two safeguards follow: skip a family when∥Wm,⋅∥2<10−12\\lVert W\_\{m,\\cdot\}\\rVert\_\{2\}<10^\{\-12\}, and when\|ℓm\|<ϵ0\|\\ell\_\{m\}\|<\\epsilon\_\{0\}stabilize1/ℓm1/\\ell\_\{m\}by1/\(ℓm±ϵ0\)1/\(\\ell\_\{m\}\\pm\\epsilon\_\{0\}\)\. The working point isϵ=0\.05\\epsilon=0\.05\.
###### Definition C\.1\(inner\-loop projection\)\.
Takeτ∈\[1/M,1\]\\tau\\in\[1/M,1\], step sizeη\\eta, and capKmaxK\_\{\\max\}\. After each main update, iterate on the probe along−∇θℒproxy\-\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}untilBshared\(probe\)≤τB\_\{\\mathrm\{shared\}\}\(\\text\{probe\}\)\\leq\\tauorKmaxK\_\{\\max\}is reached\. Whereasλ\\lambdagoverns a soft trade\-off,\(τ,Kmax,η\)\(\\tau,K\_\{\\max\},\\eta\)turn it into a near\-hard constraint that curbs the transient collinearity spike during warm\-up\. The working point isτ=70\\tau=70,Kmax=12K\_\{\\max\}=12,η=4×10−3\\eta=4\\times 10^\{\-3\}\.
###### Proposition C\.3\(small\-probe sufficiency\)\.
Write each family row as a mean overnnprobe samples,C∈ℝMfam×KC\\in\\mathbb\{R\}^\{M\_\{\\mathrm\{fam\}\}\\times K\}, under a rank\-one signal plus per\-row noiseC=s1u⊤\+NC=s\\,\\mathbf\{1\}u^\{\\top\}\+N\(∥u∥2=1\\lVert u\\rVert\_\{2\}=1\)\. Then \(i\) the sample complexity of locating the shared spikeuuis∼1/gap2\\sim 1/\\mathrm\{gap\}^\{2\}, independent ofKK\(Davis and Kahan[1970](https://arxiv.org/html/2608.08224#bib.bib28)\), so a few samples per family stabilize the leading direction; \(ii\) the concentration ratioBshared\(C\)=λmax/tr\(GC\)B\_\{\\mathrm\{shared\}\}\(C\)=\\lambda\_\{\\max\}/\\operatorname\{tr\}\(G\_\{C\}\)is a biased estimator: per\-row noise adds totr\(GC\)=∥C∥F2\\operatorname\{tr\}\(G\_\{C\}\)=\\lVert C\\rVert\_\{F\}^\{2\}and shrinks as1/n1/n, so a small probe raises the denominator and under\-reads concentration, making the estimatenn\-dependent\.
###### Proof\.
\(i\) is Davis–Kahan applied to the rank\-one signal\. \(ii\) WithGC=s2Mfam𝟏¯𝟏¯⊤\+\(zero\-mean cross terms\)\+NN⊤G\_\{C\}=s^\{2\}M\_\{\\mathrm\{fam\}\}\\,\\bar\{\\mathbf\{1\}\}\\bar\{\\mathbf\{1\}\}^\{\\top\}\+\(\\text\{zero\-mean cross terms\}\)\+NN^\{\\top\}, the trace is additive,tr\(GC\)=s2Mfam\+tr\(NN⊤\)\\operatorname\{tr\}\(G\_\{C\}\)=s^\{2\}M\_\{\\mathrm\{fam\}\}\+\\operatorname\{tr\}\(NN^\{\\top\}\)with𝔼tr\(NN⊤\)∝1/n\\mathbb\{E\}\\operatorname\{tr\}\(NN^\{\\top\}\)\\propto 1/nafter averagingnnsamples per row; sinceλmax\\lambda\_\{\\max\}is signal\-dominated, this extra denominator mass lowers the ratio at smallnn\. ∎
## Appendix DCD\-RFT Training Step
Algorithm[1](https://arxiv.org/html/2608.08224#alg1)gives one training step of CD\-RFT on top of a GRPO/PPO backbone, using the single\-backward\-pass first\-order proxy of Section[5\.4](https://arxiv.org/html/2608.08224#S5.SS4)\. Steps 2’s two gated forward passes carry constant gate scalings independent ofθ\\theta, so the entire regularizer gradient is obtained by one backward pass and differs from∇θℛ\\nabla\_\{\\theta\}\\mathcal\{R\}byO\(ϵ2\)O\(\\epsilon^\{2\}\)\(Theorem[C\.2](https://arxiv.org/html/2608.08224#A3.Thmtheorem2)\); the stopping test in Step 4 uses a cheap first\-order readout ofBsharedB\_\{\\mathrm\{shared\}\}on the same probe\.
Algorithm 1One CD\-RFT training step1:policy
θ\\theta; probe set
Π\\Pi; weight
λ\\lambda; step
ϵ\\epsilon; target
τ\\tau; step sizes
η,ηproj\\eta,\\eta\_\{\\mathrm\{proj\}\}; cap
KmaxK\_\{\\max\}
2:Backbone update:sample rollouts, compute
ℒpost\\mathcal\{L\}\_\{\\mathrm\{post\}\}and
∇θℒpost\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{post\}\}
3:set sublayer gates
g←𝟏g\\\!\\leftarrow\\\!\\mathbf\{1\}on
Π\\Pi
4:foreach family
mmdo⊳\\trianglerightone ordinary backward, stop\-gradθ\\theta
5:
ℓm←𝔼Πm\[logπθ\(y⋆\)\]\\ell\_\{m\}\\\!\\leftarrow\\\!\\mathbb\{E\}\_\{\\Pi\_\{m\}\}\[\\log\\pi\_\{\\theta\}\(y^\{\\star\}\)\]; skip if
\|ℓm\|<ϵ0\|\\ell\_\{m\}\|<\\epsilon\_\{0\}
6:
C¯m,⋅←\(1/ℓm\)∂gℓm\|g=𝟏\\bar\{C\}\_\{m,\\cdot\}\\\!\\leftarrow\\\!\(1/\\ell\_\{m\}\)\\,\\partial\_\{g\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}
7:endfor
8:
G¯←C¯C¯⊤\\bar\{G\}\\\!\\leftarrow\\\!\\bar\{C\}\\bar\{C\}^\{\\top\};
ℛ¯←tr\(G¯2\)/tr\(G¯\)2\\bar\{\\mathcal\{R\}\}\\\!\\leftarrow\\\!\\operatorname\{tr\}\(\\bar\{G\}^\{2\}\)/\\operatorname\{tr\}\(\\bar\{G\}\)^\{2\}
9:
W←4tr\(G¯\)2\(G¯C¯−ℛ¯tr\(G¯\)C¯\)W\\\!\\leftarrow\\\!\\tfrac\{4\}\{\\operatorname\{tr\}\(\\bar\{G\}\)^\{2\}\}\(\\bar\{G\}\\bar\{C\}\-\\bar\{\\mathcal\{R\}\}\\operatorname\{tr\}\(\\bar\{G\}\)\\bar\{C\}\)⊳\\trianglerightclosed form, Lemma[C\.1](https://arxiv.org/html/2608.08224#A3.Thmlemma1)
10:
ℒproxy←0\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}\\\!\\leftarrow\\\!0
11:foreach family
mmdo
12:
W^m←Wm,⋅/∥Wm,⋅∥\\hat\{W\}\_\{m\}\\\!\\leftarrow\\\!W\_\{m,\\cdot\}/\\lVert W\_\{m,\\cdot\}\\rVert⊳\\trianglerightunit dir\., Rmk\.[C\.2](https://arxiv.org/html/2608.08224#A3.Thmremark2)
13:
ℓ\+←ℓm\(θ;𝟏\+ϵW^m\)\\ell^\{\+\}\\\!\\leftarrow\\\!\\ell\_\{m\}\(\\theta;\\mathbf\{1\}\\\!\+\\\!\\epsilon\\hat\{W\}\_\{m\}\);
ℓ−←ℓm\(θ;𝟏−ϵW^m\)\\ell^\{\-\}\\\!\\leftarrow\\\!\\ell\_\{m\}\(\\theta;\\mathbf\{1\}\\\!\-\\\!\\epsilon\\hat\{W\}\_\{m\}\)
14:
ℒproxy←ℒproxy\+∥Wm,⋅∥ℓm⋅ℓ\+−ℓ−2ϵ\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}\\\!\\leftarrow\\\!\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}\+\\tfrac\{\\lVert W\_\{m,\\cdot\}\\rVert\}\{\\ell\_\{m\}\}\\cdot\\tfrac\{\\ell^\{\+\}\-\\ell^\{\-\}\}\{2\\epsilon\}
15:endfor
16:
∇θℒproxy←Backward\(λℒproxy\)\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}\\\!\\leftarrow\\\!\\textsc\{Backward\}\(\\lambda\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}\)⊳\\trianglerightsingle pass, flash\-safe
17:
θ←θ−η\(∇θℒpost\+∇θℒproxy\)\\theta\\\!\\leftarrow\\\!\\theta\-\\eta\(\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{post\}\}\+\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{proxy\}\}\)
18:
k←0k\\\!\\leftarrow\\\!0⊳\\trianglerightinner\-loop projection, Def\.[C\.1](https://arxiv.org/html/2608.08224#A3.Thmdefinition1)
19:while
Bshared\(Π;θ\)\>τB\_\{\\mathrm\{shared\}\}\(\\Pi;\\theta\)\>\\tauand
k<Kmaxk<K\_\{\\max\}do
20:recompute
C¯,W,ℒproxy\\bar\{C\},W,\\mathcal\{L\}\_\{\\mathrm\{proxy\}\};
θ←θ−ηproj∇θℒproxy\\theta\\\!\\leftarrow\\\!\\theta\-\\eta\_\{\\mathrm\{proj\}\}\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{proxy\}\};
k←k\+1k\\\!\\leftarrow\\\!k\+1
21:endwhile
22:return
θ\\theta
## Appendix EFull Experimental Settings
##### Training data\.
Three families of verifiable tasks in a balanced mixture, each resampled to 3,000 prompts \(9,000 total\), Table[E\.1](https://arxiv.org/html/2608.08224#A5.T1)\.
Table E\.1:Training data\.
##### Evaluation protocol\.
Nine held\-out benchmarks, fixed temperature applied identically to all methods, audited for leakage below; max generation length 4096 for mathematics and 2048 for code/logic; hard set = AIME24/AIME25/Minerva \(Table[E\.2](https://arxiv.org/html/2608.08224#A5.T2)\)\. The summarizedpass@k\\textup\{pass@\}kof Table[2](https://arxiv.org/html/2608.08224#S6.T2)therefore uses a per\-benchmarkkk: pass@256 for the math anchors, pass@16 for MATH500 and ordering, pass@64 for code and Logic\-graph\. The value is the largestkkeach set supports at its sample count, subject to the set still discriminating between methods: on the easier benchmarks a largekksaturates every arm near the ceiling, which hides both the ordering between methods and the large\-kkcoverage degradation of the reward\-maximizing baselines that Section[6\.3](https://arxiv.org/html/2608.08224#S6.SS3)examines\. AMC23 illustrates the saturated regime \(pass@256≈100\\approx 100for every method, Table[G\.1](https://arxiv.org/html/2608.08224#A7.T1)\); the hard set retains headroom at the samekkand is where the coverage comparison is read\. Paired arms share initialization, data order, and schedule, so each capability comparison in Section[6\.3](https://arxiv.org/html/2608.08224#S6.SS3)isolates the regularizer; the temperature sweep \(Appendix[H](https://arxiv.org/html/2608.08224#A8)\) and the checkpoint sweep \(Appendix[J](https://arxiv.org/html/2608.08224#A10)\) further test that reading\.
Table E\.2:Evaluation protocol \(primary experiment, Qwen2\.5\-7B\)\. Benchmarks: MATH500\(Hendryckset al\.[2021](https://arxiv.org/html/2608.08224#bib.bib32)\), Minerva\(Lewkowyczet al\.[2022](https://arxiv.org/html/2608.08224#bib.bib33)\), HumanEval\+\(Chenet al\.[2021](https://arxiv.org/html/2608.08224#bib.bib31); Liuet al\.[2023](https://arxiv.org/html/2608.08224#bib.bib35)\), MBPP\(Austinet al\.[2021](https://arxiv.org/html/2608.08224#bib.bib34)\), ordering\-puzzle/Logic\-graph\(Chenget al\.[2025](https://arxiv.org/html/2608.08224#bib.bib10)\)\. The Llama\-3\.2\-3B evaluation \(Appendix[K](https://arxiv.org/html/2608.08224#A11)\) uses the same benchmarks at the training rollout temperature\.
##### Train/evaluation decontamination\.
We audit every held\-out benchmark against every training pool at three levels: raw string identity; a normalized hash \(lowercased, all non\-alphanumeric characters stripped\); and near\-duplication by word88\-gram Jaccard over an inverted index, reported at thresholds\.5\.5and\.8\.8\. For the code family we additionally match the executable content, comparing normalized assertion sets of the evaluation test suites against the training functional tests\. Table[E\.3](https://arxiv.org/html/2608.08224#A5.T3)gives the audit\.
Eight of the nine benchmarks are clean at every level, with no exact, normalized, or near\-duplicate hit and no shared assertion\. The exception is MATH500, of which77items \(1\.4%1\.4\\%\) are normalized\-exact matches of items in the DeepScaleR pool\. Because each family is resampled to3,0003\{,\}000prompts from a much larger pool, pool membership overstates exposure: replaying the training shuffle shows that exactly11of those77items \(0\.2%0\.2\\%of MATH500\) was drawn into the data the runs actually saw\. On that item the trained arms do not exceed the untrained base \(9/169/16correct for base,8/168/16for GRPO,9/169/16for CD\-RFT\), and on all77the base and GRPO \(β=0\\beta\{=\}0\)pass@1\\textup\{pass@\}1are identical \(17\.917\.9\), so there is no memorization signature\.
We nonetheless re\-score MATH500 with all77items removed, identically for every arm, as a conservative upper bound\. Every arm gains between\+0\.1\+0\.1and\+0\.7\+0\.7points \(the removed items are harder than average\), and the paired CD\-RFT−\-GRPO difference keeps its sign atk=1,2,4,8k\{=\}1,2,4,8on both axes; atk=16k\{=\}16, where MATH500 is near saturation \(≈90\\approx 90for all arms\), the difference moves within±0\.4\\pm 0\.4points of zero on both pairs\. No conclusion in Section[6\.3](https://arxiv.org/html/2608.08224#S6.SS3)depends on the contaminated items\. The logic row deserves a separate note: the logic training pool and the logic benchmarks are drawn from the same GURU generator family, so their88\-gram overlap \(maxJ=\.35\\max J=\.35on ordering\-puzzle\) measures shared templates, not shared instances; exact and normalized instance matches are zero, which is what leakage would require\.
Table E\.3:Train/evaluation contamination audit\. “near” = word88\-gram Jaccard≥\.8\\geq\.8; “assert” = shared normalized test assertions \(code only\)\.maxJ\\max Jis the largest Jaccard against any training item\.
##### Hyperparameters \(primary experiment, Qwen2\.5\-7B\)\.
On the primary model all methods share the same recipe, differing only in the regularizer/KL \(Table[E\.4](https://arxiv.org/html/2608.08224#A5.T4)\); the second\-model Llama\-3\.2\-3B setup and its differences are given in Appendix[K](https://arxiv.org/html/2608.08224#A11)\.
Table E\.4:Hyperparameters \(primary experiment, Qwen2\.5\-7B\)\. All arms share seed 0, so each paired comparison holds initialization and data order fixed\.
##### Software and hardware\.
Sampling and evaluation use vLLM 0\.23\.0 with the rollout GPU\-memory utilization set to 0\.6; training, probing, and evaluation are each performed on a single GPU, on NVIDIA A100 \(80GB\) and NVIDIA RTX PRO 6000 \(Blackwell\) GPUs\. Table[E\.5](https://arxiv.org/html/2608.08224#A5.T5)lists the software stack\.
Table E\.5:Software stack\.
## Appendix FAdditional Mechanistic Results
### F\.1Activation\-Level Metrics
Section[6\.2](https://arxiv.org/html/2608.08224#S6.SS2)reports that two of the three activation\-level metrics ofZhanget al\.\([2026](https://arxiv.org/html/2608.08224#bib.bib22)\)reproduce while the third does not; Tables[F\.1](https://arxiv.org/html/2608.08224#A6.T1)and[F\.2](https://arxiv.org/html/2608.08224#A6.T2)give the underlying numbers\. All three are computed from EAP edge attributions\(Syedet al\.[2024](https://arxiv.org/html/2608.08224#bib.bib12)\)on the both\-correct subset of the base model and the method under test, the standard attribution\-difference protocol\. Because each pairing induces a slightly different both\-correct subset, and the Base row shown is the subset baseline of theβ=0\\beta\{=\}0pairing, the tables are read within a column\.
Table F\.1:The two activation\-level signatures ofZhanget al\.\([2026](https://arxiv.org/html/2608.08224#bib.bib22)\)that reproduce robustly: activation intensity rises and distribution kurtosis falls, for every method on every domain \(12/12 each\), relative to the base model\. graph = Logic\-graph\.Activation intensity rises and kurtosis falls for every method on every domain \(Table[F\.1](https://arxiv.org/html/2608.08224#A6.T1)\)\. The signature that reinforcement learning raises circuit usage and flattens its magnitude distribution therefore reproduces not only for vanilla GRPO but for all four trained models, and forms the starting point on which the analysis builds\.
Table F\.2:Information complexity, the direction\-unstable third metric\. Bootstrap standard deviations \(B=2000B\{=\}2000\) on the both\-correct subset of each pairing \(n=73n\{=\}73–100100on MATH500/MBPP,n=19n\{=\}19–2323on Logic\-graph\)\. graph = Logic\-graph, which carries the logic\-domain evidence\.Information complexity \(Table[F\.2](https://arxiv.org/html/2608.08224#A6.T2)\) does not behave this way\. On MATH500 and MBPP it stays high under CD\-RFT \(β=0\\beta\{=\}0\) while collapsing to near zero under the matched GRPO \(β=0\\beta\{=\}0\), \.219±\\pm\.048 against \.008±\\pm\.030 and \.270±\\pm\.019 against \.015±\\pm\.004, gaps of several bootstrap standard deviations; the same ordering reappears on Logic\-graph \(\.568±\\pm\.037 against \.496±\\pm\.057\)\. Yet the metric does not track capability: GRPO \(β=0\.01\\beta\{=\}0\.01\) is stronger than the base on every domain of Table[2](https://arxiv.org/html/2608.08224#S6.T2), while its information complexity on MBPP falls from \.255 to \.020, a drop of more than90%90\\%\. A metric that splits this way between two equally healthy models cannot on its own say what reinforcement learning changed, which is what motivates moving to the control axis\. The well\-powered form of this reading is theβ=0\\beta\{=\}0pairing, where both arms sit high enough on mathematics and code to separate\.
### F\.2Random Reference forBsharedB\_\{\\mathrm\{shared\}\}
The lower bound1/M1/Mof Proposition[2](https://arxiv.org/html/2608.08224#Thmproposition2)is attained only when the rows are exactly orthogonal with equal norms, which random rows in finite dimension never are, so1/M1/Mis a bound and not the value a random control matrix would produce\. We therefore estimate the random reference by simulation at the shape used throughout \(M=3M\{=\}3families,K=56K\{=\}56sublayer gates\): drawing i\.i\.d\. Gaussian rows over50,00050\{,\}000trials gives
Bsharedrand=42\.8±3\.4,p95=49\.0,p99=51\.9,B\_\{\\mathrm\{shared\}\}^\{\\mathrm\{rand\}\}=42\.8\\pm 3\.4,\\qquad p\_\{95\}=49\.0,\\quad p\_\{99\}=51\.9,against the bound1/M=33\.31/M=33\.3\. This reference depends only on\(M,K\)\(M,K\), hence is shared by every arm and every model in Figure[F\.3](https://arxiv.org/html/2608.08224#A6.F3)\. All five base models of Figure[F\.3](https://arxiv.org/html/2608.08224#A6.F3)exceed thep99p\_\{99\}of the random reference: Qwen2\.5\-3B59\.259\.2, Qwen2\.5\-1\.5B64\.364\.3, Llama\-3\.1\-8B71\.371\.3, Llama\-3\.2\-3B80\.580\.5, Qwen2\.5\-7B83\.483\.4, i\.e\. between\+16\+16and\+41\+41above it\. These are the exactBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)values of that figure, on the same protocol as the Base row of Table[2](https://arxiv.org/html/2608.08224#S6.T2)\. The cross\-model reading of Section[6\.2](https://arxiv.org/html/2608.08224#S6.SS2)therefore holds against a simulated random reference and not merely against the bound\.
BsharedB\_\{\\mathrm\{shared\}\}also responds to row\-norm heterogeneity, not to direction alone\. Proposition[2](https://arxiv.org/html/2608.08224#Thmproposition2)\(iv\) gives invariance under global scalingC↦cCC\\mapsto cCand under orthogonal gate reparameterizationC↦CQC\\mapsto CQ, but not under*per\-row*rescaling, and each row ofCCcarries its own normalizer1/ℓm1/\\ell\_\{m\}, whose scale differs across families with target length\. This is why every comparison we report is made at a fixed protocol and, for the method claim, as a same\-seed same\-probe paired difference \(Figure[F\.2](https://arxiv.org/html/2608.08224#A6.F2)\), both of which hold the row\-norm structure of the estimator fixed; the cross\-model panel, whose arms do not share a protocol, is read only by ordering\.
### F\.3Direction Versus Row Norm
Holding the protocol fixed still leaves a question the paired difference cannot answer on its own, so we separate the two contributions directly\. Writingdm=∥Cm,⋅∥22d\_\{m\}=\\lVert C\_\{m,\\cdot\}\\rVert\_\{2\}^\{2\},D=diag\(d\)D=\\operatorname\{diag\}\(d\)andRRfor the matrix of row cosines, the Gram matrix factors asGC=D1/2RD1/2G\_\{C\}=D^\{1/2\}RD^\{1/2\}, and we report alongsideBsharedB\_\{\\mathrm\{shared\}\}two quantities computed from the sameCC:
Bdir\(C\)=λmax\(R\)M,Bnorm\(C\)=maxmdm∑m′dm′\.B\_\{\\mathrm\{dir\}\}\(C\)=\\frac\{\\lambda\_\{\\max\}\(R\)\}\{M\},\\qquad B\_\{\\mathrm\{norm\}\}\(C\)=\\max\_\{m\}\\frac\{d\_\{m\}\}\{\\sum\_\{m^\{\\prime\}\}d\_\{m^\{\\prime\}\}\}\.\(17\)BdirB\_\{\\mathrm\{dir\}\}normalizes each row to unit length first, so it is invariant to per\-row rescaling and measures direction sharing alone; it keeps the range\[1/M,1\]\[1/M,1\], attaining1/M1/Mexactly when the rows are pairwise orthogonal and11when they are collinear\.BnormB\_\{\\mathrm\{norm\}\}is whatBsharedB\_\{\\mathrm\{shared\}\}would equal if the rows were exactly orthogonal, i\.e\. the row\-norm contribution alone\. The two are not additive, so we plot them side by side rather than as a decomposition ofBsharedB\_\{\\mathrm\{shared\}\}\.
Figure F\.1:CD\-RFT lowers direction sharing\.\(a\)All three measures at theβ=0\\beta\{=\}0arms;BdirB\_\{\\mathrm\{dir\}\}sits far belowBsharedB\_\{\\mathrm\{shared\}\}at every arm, soBsharedB\_\{\\mathrm\{shared\}\}overstates how collinear the control directions are\.\(b\)Paired per seed against the matched GRPO:ΔBdir=−7\.6±2\.1\\Delta B\_\{\\mathrm\{dir\}\}=\-7\.6\\pm 2\.1is negative on7/77/7seeds and its95%95\\%interval excludes zero, while the interval onΔBshared\\Delta B\_\{\\mathrm\{shared\}\}does not\.\(c\)The seed\-to\-seed variation ofΔBshared\\Delta B\_\{\\mathrm\{shared\}\}lies on the diagonal againstΔBnorm\\Delta B\_\{\\mathrm\{norm\}\}\(r=0\.99r\{=\}0\.99\): the row\-norm term is what injects it, and the one seed that flipsΔBshared\\Delta B\_\{\\mathrm\{shared\}\}positive \(s6\) still hasΔBdir<0\\Delta B\_\{\\mathrm\{dir\}\}<0\.\(d\)Pairwise direction alignment in magnitude; the code–logic cosine falls from0\.480\.48to0\.160\.16, and CD\-RFT is the lowest on all three pairs\.Figure[F\.1](https://arxiv.org/html/2608.08224#A6.F1)reports both on theβ=0\\beta\{=\}0pair, over the seven probe seeds of Table[2](https://arxiv.org/html/2608.08224#S6.T2)\. Direction sharing falls under CD\-RFT: the pairedΔBdir\\Delta B\_\{\\mathrm\{dir\}\}is−7\.6±2\.1\-7\.6\\pm 2\.1points and negative on every seed, against6/76/7forΔBshared\\Delta B\_\{\\mathrm\{shared\}\}, and its95%95\\%interval \(Studenttt, six degrees of freedom\) excludes zero where the interval onΔBshared\\Delta B\_\{\\mathrm\{shared\}\}does not\. Removing the row norms therefore sharpens the effect rather than dissolving it, because the row\-norm term carries most of the seed\-to\-seed variance:ΔBshared\\Delta B\_\{\\mathrm\{shared\}\}tracksΔBnorm\\Delta B\_\{\\mathrm\{norm\}\}across seeds atr=0\.99r=0\.99, and the single seed on whichΔBshared\\Delta B\_\{\\mathrm\{shared\}\}turns positive is one whereΔBdir\\Delta B\_\{\\mathrm\{dir\}\}remains negative\. The row cosines move with it: in magnitude CD\-RFT is the lowest of the three arms on every family pair, the code–logic pair dropping from0\.480\.48to0\.160\.16\. Magnitudes are the relevant summary because a signed average over seeds cancels, and anti\-alignment raisesλmax\(R\)\\lambda\_\{\\max\}\(R\)just as alignment does\.
Because∥Cm,⋅∥2=∥∂gℓm\|g=𝟏∥2/\|ℓm\|\\lVert C\_\{m,\\cdot\}\\rVert\_\{2\}=\\lVert\\partial\_\{g\}\\ell\_\{m\}\|\_\{g=\\mathbf\{1\}\}\\rVert\_\{2\}/\\lvert\\ell\_\{m\}\\rvert, a regularizer could in principle lowerBsharedB\_\{\\mathrm\{shared\}\}by inflating\|ℓm\|\\lvert\\ell\_\{m\}\\rverton the family that dominates the energy budget, which would work against the reward flux rather than with it\. This is not what happens\. On the family holding the largest share at baseline,\|ℓm\|\\lvert\\ell\_\{m\}\\rvertrises by12\.5%12\.5\\%from the matched GRPO while its row norm falls by63%63\\%, so the implied gate sensitivity∥∂gℓm∥2\\lVert\\partial\_\{g\}\\ell\_\{m\}\\rVert\_\{2\}falls by59%59\\%: the normalizer accounts for about a fifth of the change and the rest is a genuine reduction in how strongly that family’s reward flux responds to the gates\.
### F\.4Paired Verification and Cross\-Model Generality
AbsoluteBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)depends on which probe samples are drawn, so Figure[F\.2](https://arxiv.org/html/2608.08224#A6.F2)re\-reads the measurement of Table[2](https://arxiv.org/html/2608.08224#S6.T2)as a same\-seed, same\-probe paired difference, which cancels the shared probe noise\. Figure[F\.3](https://arxiv.org/html/2608.08224#A6.F3)then checks that the bottleneck is not an artifact of the single model we intervene on, and Figure[F\.4](https://arxiv.org/html/2608.08224#A6.F4)repeats the distributedness check of Figure[2](https://arxiv.org/html/2608.08224#S5.F2)\(b\) at the finer EAP edge granularity\. The participation ratio of a control row isPR\(Cm,⋅\)=\(∑k\|Cm,k\|\)2/∑kCm,k2\\mathrm\{PR\}\(C\_\{m,\\cdot\}\)=\\big\(\\sum\_\{k\}\|C\_\{m,k\}\|\\big\)^\{2\}\\big/\\sum\_\{k\}C\_\{m,k\}^\{2\}, the effective number of gates it spreads over:PR=1\\mathrm\{PR\}=1when a single gate carries the row andPR=K\\mathrm\{PR\}=Kwhen all gates carry it equally\.
Figure F\.2:Forest plot of the paired differenceΔ=Bshared\(C\)CD\-RFT−Bshared\(C\)GRPO\\Delta=B\_\{\\mathrm\{shared\}\}\(C\)^\{\\text\{CD\-RFT\}\}\-B\_\{\\mathrm\{shared\}\}\(C\)^\{\\text\{GRPO\}\}under the same seed and probes\. CD\-RFT \(β=0\\beta\{=\}0\) shows a robust negative shift \(Δ=−12\.3±3\.1\\Delta=\-12\.3\\pm 3\.1, negative in 6 of 7 seeds\); CD\-RFT \(β=0\.01\\beta\{=\}0\.01\) is consistent in direction but smaller\.Figure F\.3:Bshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)for five base models under the same multi\-task probe and exact protocol\. All lie above the9999th percentile of the simulated random reference of Appendix[F\.2](https://arxiv.org/html/2608.08224#A6.SS2); the two models we train CD\-RFT on \(Qwen2\.5\-7B and Llama\-3\.2\-3B, in blue\) have the deepest bottlenecks, and hence the most decoupling headroom\.Figure F\.4:Distributedness of Figure[2](https://arxiv.org/html/2608.08224#S5.F2)\(a\) at the finer EAP edge granularity \(K≈3221K\\approx 3221\) on the offline Part\-I control vectors\.\(a\)Participation ratio and\(b\)cumulative\|C\|\|C\|mass give the same no\-monopoly conclusion; the ratio is larger here only because the edge decomposition is finer, and aggregation into sublayers is non\-orthogonal, so the sublayer figure reports the on\-granularity value\.
## Appendix GAdditional Capability Results
### G\.1Full Per\-Benchmark pass@kkLadder
Table[G\.1](https://arxiv.org/html/2608.08224#A7.T1)gives the per\-benchmark results behind the domain summary of Table[2](https://arxiv.org/html/2608.08224#S6.T2)\(step\-1000, five matched methods\)\. The coverage column reports the highestkkof each set: pass@256 for the math anchors, pass@16 for MATH500 and ordering, pass@64 for code and Logic\-graph\. AMC23 \(40 problems\) is near\-saturated there \(pass@256≈100\\approx 100for every method\), which treats all methods alike; the hard\-set coverage advantage is carried by AIME24/AIME25/Minerva\.
Table G\.1:Per\-benchmarkpass@k\\textup\{pass@\}kladder underlying Table[2](https://arxiv.org/html/2608.08224#S6.T2)\. All entries are percentages\.
### G\.2Per\-Benchmark Paired Gains and Coverage
Figure[G\.1](https://arxiv.org/html/2608.08224#A7.F1)resolves the domain summary of Table[2](https://arxiv.org/html/2608.08224#S6.T2)into the per\-benchmark paired gain of CD\-RFT over its matched GRPO baseline, and Figure[G\.2](https://arxiv.org/html/2608.08224#A7.F2)traces how the coverage advantage of the with\-KL variant grows withkkon the hard sets\.
Figure G\.1:Per\-benchmark paired gains, CD\-RFT minus its matched GRPO baseline, across all nine benchmarks in four panels \(β=0\\beta\{=\}0/β=0\.01\\beta\{=\}0\.01crossed withpass@1\\textup\{pass@\}1/pass@k\\textup\{pass@\}k\)\. Theβ=0\\beta\{=\}0axis leads onpass@1\\textup\{pass@\}1\(8/9\); theβ=0\.01\\beta\{=\}0\.01axis leads onpass@k\\textup\{pass@\}kcoverage \(8/9\)\.Figure G\.2:Hard\-set \(AIME24/AIME25/Minerva\) meanpass@k\\textup\{pass@\}kas a function ofkk\(log scale\)\. CD\-RFT \(β=0\.01\\beta\{=\}0\.01\) widens its coverage advantage over GRPO \(β=0\.01\\beta\{=\}0\.01\) at largekk, reaching \+8\.1pp at pass@256\.
## Appendix HFull Temperature\-Sweep Results
We repeat the evaluation at four temperatures \(the sweep of each domain including its main\-table working temperature\) and count the fraction of paired×\\timestemperature cells \(8 in total\) in which “CD\-RFT≥\\geqmatched RL” holds \(Table[H\.1](https://arxiv.org/html/2608.08224#A8.T1); the trend is in Figure[H\.1](https://arxiv.org/html/2608.08224#A8.F1)\)\.
Table H\.1:Temperature robustness: paired\-ordering win rate \(CD\-RFT≥\\geqmatched RL, out of 8\)\.On the math hard set and on code, thepass@1\\textup\{pass@\}1advantage of CD\-RFT holds across all four temperatures \(8/8 each\) andpass@k\\textup\{pass@\}kholds at 7/8\. Lowering the temperature raisespass@1\\textup\{pass@\}1and raising it raisespass@k\\textup\{pass@\}k, yet the method ordering does not flip\. The logic domain shows lower ordering consistency \(pass@1\\textup\{pass@\}13/8\), in line with the mild gain margin of this domain in the capability comparison, while the featured configuration CD\-RFT \(β=0\\beta\{=\}0\) shows no collapse on any domain\.
Figure H\.1:Sampling\-temperature sweep\. Lowering the temperature raisespass@1\\textup\{pass@\}1and raising it raisespass@k\\textup\{pass@\}k; the method ordering on the math hard set and code does not flip across the four temperatures\.
## Appendix IProxy\-Target Sensitivity
We sweep the target concentrationτ∈\{55,65,70\}\\tau\\in\\\{55,65,70\\\}with every other setting fixed at the main recipe \(Qwen2\.5\-7B, step\-1000, matched protocol\), soτ\\tauis the sole variable\. Table[I\.1](https://arxiv.org/html/2608.08224#A9.T1)givespass@1\\textup\{pass@\}1and the coveragepass@k\\textup\{pass@\}kfor each benchmark\.
The method is insensitive toτ\\tauacross a broad band: the two looser settings,7070and6565, lie within a small margin of each other on every benchmark and simply exchange which axis they favour\. The main working point7070takespass@1\\textup\{pass@\}1on the discriminative benchmarks \(MATH500, AMC23, code, Logic\-graph\), while the slightly tighter6565trades that small deficit for large\-kkcoverage on the harder or more open\-ended sets \(MATH500, Minerva, code, ordering\)\. Only over\-compression,τ=55\\tau\{=\}55, breaks the pattern: it trails on both axes across essentially all discriminative benchmarks, marking a floor below which tightening the constraint costs single\-shot accuracy without buying coverage\.τ\\taushould therefore be kept in the loose band, its exact value chosen by whether the deployment weights single\-shot accuracy \(7070\) or coverage \(6565\)\.
Table I\.1:Proxy\-target sweep on Qwen2\.5\-7B \(pass@1\\textup\{pass@\}1/ coveragepass@k\\textup\{pass@\}k, forτ=55/65/70\\tau\{=\}55/65/70\)\. Coveragekkis the highest sampled per set \(MATH500/ordering p@16 and p@64 respectively, math anchors p@256, code/Logic\-graph p@64\)\. Bold marks the best of the three\. Logic\-graph uses the mid\-difficulty split, matching the main table\. All entries are percentages\.
## Appendix JCross\-Checkpoint Robustness
The main comparison reports one checkpoint \(step\-1000\), fixed under the evaluation compute budget before any benchmark was scored and applied identically to all five arms; it is not the final step of training\. Figure[J\.1](https://arxiv.org/html/2608.08224#A10.F1)sweeps theβ=0\\beta\{=\}0pair across checkpoints: CD\-RFT \(β=0\\beta\{=\}0\) sits at or above its matched GRPO at every checkpoint on both axes\.
Table[J\.1](https://arxiv.org/html/2608.08224#A10.T1)gives the per\-benchmark ladder at the*final*step\-1500 checkpoint under the full main\-table protocol, the same benchmarks, sample counts, and temperatures as Table[2](https://arxiv.org/html/2608.08224#S6.T2), applied identically to both arms, with the untrained base as an anchor\.
Table J\.1:Per\-benchmark ladder at the final step\-1500 checkpoint,β=0\\beta\{=\}0pair, full main\-table protocol \(cf\. Table[G\.1](https://arxiv.org/html/2608.08224#A7.T1)at step\-1000\)\. G = GRPO, CD = CD\-RFT; Base is untrained and step\-independent\. Bold marks the row\-wise best\. All entries are percentages\.Figure J\.1:Cross\-checkpoint robustness for theβ=0\\beta\{=\}0pair: overallpass@1\\textup\{pass@\}1andpass@k\\textup\{pass@\}kversus training step\. CD\-RFT \(β=0\\beta\{=\}0\) stays at or above GRPO \(β=0\\beta\{=\}0\) at every checkpoint\. The mathematics domain is scored here on subsampled MATH500 and Minerva, so the figure is read by ordering; Table[J\.1](https://arxiv.org/html/2608.08224#A10.T1)gives the final checkpoint under the full main\-table protocol\.
## Appendix KSecond Model: Llama\-3\.2\-3B
This appendix details the second\-model transfer experiment of Section[6\.5](https://arxiv.org/html/2608.08224#S6.SS5): the setup differences from the primary Qwen2\.5\-7B run, the full per\-benchmarkpass@k\\textup\{pass@\}kladders, and the sampling\-temperature behaviour\.
##### Setup\.
The method, the training data \(the balanced three\-family mixture of Table[E\.1](https://arxiv.org/html/2608.08224#A5.T1)\), and the benchmarks are shared with the primary experiment\. The recipe is adapted for the weaker base model, and every adaptation is applied identically to the two trained arms, so the GRPO→\\toCD\-RFT comparison stays controlled; Table[K\.1](https://arxiv.org/html/2608.08224#A11.T1)lists the differences\.
Evaluation protocol\.The trained arms are evaluated few\-shot \(K=2K\{=\}2,K=4K\{=\}4for code\) with the same demonstrations, because the plain Llama\-3\.2\-3B template otherwise holds them below the ability they have already reached\. The*Base*column is instead the*zero\-shot*untrained model, which sits at the reward floor and degenerates into format\-violating and repeated generations—the starting point from which reinforcement learning departs\. Since the two protocols differ,the pairedΔ\\Deltais computed only between GRPO and CD\-RFT; Base serves as an untrained\-floor reference and is not differenced against them\.
On the targetτ\\tau\.The valueτ=50\\tau\{=\}50was fixed before this line was evaluated, neither transferred from the primary model nor tuned against benchmark scores, and it enters only the CD\-RFT arm; by the proxy\-target sweep of Section[6\.5](https://arxiv.org/html/2608.08224#S6.SS5)the comparison does not rest on its exact value\.
ItemQwen2\.5\-7BLlama\-3\.2\-3BKLβ\\beta0 and 0\.010 onlyPrompt templateplain \(\\boxed\)R1 think/answerFew\-shot02 \(code 4\)Rollout temperature0\.60\.5Max generation length5121024Target concentrationτ\\tau\(%\)7050Reported checkpointstep\-1000step\-1122Table K\.1:Llama\-3\.2\-3B setup, as differences from the primary Qwen2\.5\-7B recipe \(Table[E\.4](https://arxiv.org/html/2608.08224#A5.T4)\)\. All other knobs \(λ=1\\lambda\{=\}1,ϵ=0\.05\\epsilon\{=\}0\.05, inner\-loop cap1212, learning rate/schedule\) are unchanged and shared across arms\.
##### Mechanism: control\-sharing bottleneck\.
Table[K\.2](https://arxiv.org/html/2608.08224#A11.T2)reports the exactBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)for the three arms under the same protocol as the primary model \(Section[6\.2](https://arxiv.org/html/2608.08224#S6.SS2); npf=3=3, second\-order double\-backward, five probe seeds paired to cancel shared\-probe noise\), using the Llama\-3\.2\-3B tokenizer to build the probe\. The signature reproduces: base\>\>GRPO\>\>CD\-RFT, i\.e\. reinforcement learning lowers control sharing and the regularizer lowers it further, with the same\-seed pairedCD−RFT−CD\-RFT\{\-\}GRPO=−5\.5±1\.9=\-5\.5\\pm 1\.9\(2\.9σ2\.9\\sigma,4/54/5seeds negative\)\.
Table K\.2:Llama\-3\.2\-3B exactBshared\(C\)B\_\{\\mathrm\{shared\}\}\(C\)\(step\-1122, five probe seeds, mean±\\pmSE\), same protocol as Table[2](https://arxiv.org/html/2608.08224#S6.T2)\. Same\-seed paired differences at right\.
##### Fullpass@k\\textup\{pass@\}kladder\.
Table[K\.3](https://arxiv.org/html/2608.08224#A11.T3)gives the complete ladder at the training rollout temperature, with all three arms evaluated few\-shot under the fully matched protocol \(so*Base*here is the few\-shot untrained model, complementing the zero\-shot floor anchor of Table[3](https://arxiv.org/html/2608.08224#S6.T3)\)\. The paired CD\-RFT−\-GRPO gap is positive at nearly everykkand widens withkkin every domain, the same signature as on Qwen2\.5\-7B; atk=64k\{=\}64on ordering\-puzzle, GRPO drops below the few\-shot base \(37\.0<41\.037\.0<41\.0, the coverage collapse ofYueet al\.\([2025](https://arxiv.org/html/2608.08224#bib.bib6)\)\) while CD\-RFT restores it above \(46\.046\.0\)\.
Table K\.3:Llama\-3\.2\-3B full per\-benchmarkpass@k\\textup\{pass@\}kladder at the training rollout temperature; all three arms \(Base//GRPO//CD\-RFT\) evaluated few\-shot under the matched protocol\. Bold marks the best of the three\. GSM8K is held out from the training mixture of both models\. All entries are percentages\.
## Appendix LComputational Overhead of the Control Regularizer
Table[L\.1](https://arxiv.org/html/2608.08224#A12.T1)measures what the control regularizer costs per step\. The comparison is fully controlled: the same RTX PRO 6000 GPU, the sameβ=0\\beta\{=\}0backbone, the same generation length, and the same first 1000 steps, with the finite\-difference regularizer as the only variable toggled on or off\.
Table L\.1:Per\-step cost of the control regularizer \(Qwen2\.5\-7B, single RTX PRO 6000,β=0\\beta\{=\}0, first 1000 steps, only the regularizer toggled\)\. This is the worst case: atβ=0\\beta\{=\}0the inner projection is active at essentially every step\.The measurement is the worst case for our method by construction\. Atβ=0\\beta\{=\}0nothing else holds the concentration down, so the inner projection engages at essentially every step and pays its full cost; under theβ=0\.01\\beta\{=\}0\.01backbone the KL penalty already suppressesBsharedB\_\{\\mathrm\{shared\}\}, the projection fires on only 16 of 1500 steps, and the overhead is correspondingly smaller\. We report only the fully\-engaged configuration\.
Two details fix the reading\. The reported value is the mean step time, and both arms include a few steps slowed by node contention, so on the median step the gap is somewhat wider \(22\.8→25\.522\.8\\rightarrow 25\.5s,\+11\.8%\+11\.8\\%\)\. Wall\-clock time is also not comparable across GPUs, because among the four main arms the CD\-RFT models were trained on RTX PRO 6000 and the GRPO models on A100, and since the former is the faster GPU a direct wall\-clock comparison would favour the proposed method\. The overhead is therefore quoted only from the same\-GPU, same\-backbone measurement above, which isolates the regularizer itself\. In either case the conclusion is the one that Remark[C\.4](https://arxiv.org/html/2608.08224#A3.Thmremark4)anticipates, that rollout generation dominates the step, the central\-difference probe is a small addition on top of it, and the regularizer remains a low\-overhead component\.相似文章
论SFT的泛化:强化学习视角与奖励修正
本文从强化学习的视角分析了标准监督微调(SFT)的局限性,并提出了一种简单的梯度重新缩放方法——动态微调(DFT),该方法提高了LLM的泛化能力,并与离线RL性能相匹配。
当RL在SFT后失效:恢复模型可塑性以实现稳健的SFT到RL交接
本文研究了在大型语言模型的先SFT后RL流程中,过度监督微调(SFT)后模型可塑性的丧失问题,并提出了一种名为Rejuvenation的方法,该方法通过基于基线的模型融合和定向神经元重置来恢复可塑性,从而持续提升RL性能。
互惠协同训练(RCT):通过强化学习耦合基于梯度与不可微模型
# 互惠协同训练(RCT):通过强化学习耦合基于梯度与不可微模型 来源:[https://arxiv.org/html/2604.16378](https://arxiv.org/html/2604.16378) Yunshuo Tian¹, Akayou Kitessa¹, Tanuja Chitnis², 和 Yijun Zhao¹ 1 纽约市福特汉姆大学计算机与信息科学系 2 马萨诸塞州波士顿市Mass General Brigham医院神经科 ###### 摘要 大型语言模型 \(LLMs\) 与经典机器学习方法提供互补...
CellRFT: 用于单细胞扰动建模的强化微调
CellRFT引入了一个强化微调框架,该框架使用生物评估作为直接训练反馈,以改进单细胞扰动模型,解决了代理训练损失与生物评估标准之间的不匹配问题。
为什么多步骤工具使用强化学习会崩溃以及监督信号如何修复它
本文研究了为什么多步骤工具使用强化学习(RL)常常崩溃或收益有限,并将控制令牌中的概率尖峰识别为关键原因。研究表明,将监督微调与RL交替进行可以提高稳定性,并探索了各种监督信号以指导稳健训练。