FAER:面向语言模型后训练的可审计效用对齐轨迹重放
摘要
本文提出 FAER,这是一个面向语言模型后训练的可审计全轨迹重放框架。该框架形式化了缓存级选择反馈与下游学习者效用之间的差距,并表明一个学习者感知的选择器(FAER-UTILITY)在 GSM8K 上配合 Qwen2.5-1.5B-Instruct 时,表现优于基于格式反馈和均匀采样的基线方法。
arXiv:2610.00385v1 Announce Type: new
Abstract: Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches $0.6476\!\pm\!0.0139$ over eight seeds (median 0.6481; paired 95% interval $[+0.079,+0.122]$) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.
查看缓存全文
缓存时间: 2026/10/03 09:53
# FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training
Source: [https://arxiv.org/html/2610.00385](https://arxiv.org/html/2610.00385)
Miaobo HuAffiliation:School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, ChinaAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnShuhao HuAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnXiaobo GuoAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnXin WangAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnBokun WangAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnTianshu FuAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnDaren ZhaAffiliation:Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China\*Corresponding author:xiaojun@ucas\.ac\.cnJun XiaoAffiliation:School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
###### Abstract
Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache\-level correctness and downstream learner utility are distinct objectives\. We formalize this selection\-to\-learning gap and introduce FAER as an auditable full\-trajectory replay framework\. Its training\-free fixed selector is a protocol baseline; FAER\-UTILITY is the learner\-aware selector fitted on disjoint calibration blocks\. The normalized gradient alignment is reported as a baseline, while a disposable optimizer\-aware virtual update supplies a magnitude\-aware utility surface\. The audit contract freezes observed fields and replay traces before evaluation labels are joined\. On GSM8K with Qwen2\.5\-1\.5B\-Instruct, the matched learner study reports quality 0\.6329 for the fixed selector, compared with 0\.5482 for uniform and 0\.6037 for format\-feedback under 128 updates\. Metadata\-only cross\-fitted calibration reaches±0\.01390\.6476\\\!\\pm\\\!0\.0139over eight seeds \(median 0\.6481; paired 95% interval\[\+0\.079,\+0\.122\]\[\+0\.079,\+0\.122\]\) at 63,276 target\-run tokens; its recorded full cost is 189,642 tokens and 3\.48 GPU\-hours including calibration\. The completed FAER\-UTILITY row reaches 0\.6624 at 62,844 target\-run tokens and 4\.26 GPU\-hours\. Format\-feedback selects records with correctness 0\.6953, compared with 0\.3594 for the fixed selector, despite the different downstream ranking\. The completed comparison surfaces report the learner\-aware ablation, same\-seed gap, policy\-optimization rows, and strict zero\-shot transfer\.
## 1Introduction
Language\-model post\-training increasingly reuses generated trajectories to allocate a limited optimization budget\. A useful replay policy must select material that improves a learner, and its result must remain interpretable when the cache producer, response length, or trajectory age changes\. Complete trajectories are a natural replay unit: they retain the generated context consumed by an update and expose feedback that can guide reuse\.
Selection and learning measure different objects\. A cache\-level audit measures the correctness of stored records entering a subset; a learner experiment measures performance on new held\-out prompts after optimization\. Connecting these objects requires stable task identities, a defined information boundary, and matched update budgets\. Producer\-derived confidence and freshness also need explicit lineage so that the selector can be reconstructed from the same inputs\.
FAER connects replay selection to the learner update through two layers\. The audit protocol projects each cache row to declared selector\-visible fields and commits selected IDs before evaluator labels are joined\. The primary selector estimates downstream utility from trajectory\-gradient alignment and metadata using disjoint calibration blocks; the fixed feedback\-aware product remains a training\-free protocol baseline\. Producer lineage, learner checkpoints, measured age, and compute telemetry bind the two stages to the same comparison\.
Our contributions are fourfold\. First, we formalize the selection\-to\-learning gap with rank inversion and regret estimands\. Second, we define an auditable replay contract with row\-level provenance, immutable traces, and post\-freeze evaluation\. Third, we formulate a cross\-fitted learner\-aware utility selector whose gradient\-alignment feature has a first\-order connection to calibration\-loss reduction\. Fourth, we report matched learner, same\-seed, cost, transfer, and policy\-optimization comparisons under one shared evidence contract\.
## 2Related Work
##### Replay and data selection\.
Prioritized experience replay samples transitions according to learning signal and uses importance weights to control the resulting distribution shift\([Schaul et al\., 2016](https://arxiv.org/html/2610.00385#bib.bib31);[Fu et al\., 2021](https://arxiv.org/html/2610.00385#bib.bib11)\)\. Later work studies replay ratio, buffer composition, sequence\-level sampling, and distributed actor–learner lag\([Fedus et al\., 2020](https://arxiv.org/html/2610.00385#bib.bib10);[Kapturowski et al\., 2019](https://arxiv.org/html/2610.00385#bib.bib16);[Espeholt et al\., 2018](https://arxiv.org/html/2610.00385#bib.bib9)\)\. Offline RL and sequence\-modeling approaches make the data contract equally important because the training distribution is fixed or explicitly curated\([Kumar et al\., 2020](https://arxiv.org/html/2610.00385#bib.bib20);[Kostrikov et al\., 2022](https://arxiv.org/html/2610.00385#bib.bib19);[Chen et al\., 2021](https://arxiv.org/html/2610.00385#bib.bib3);[Janner et al\., 2022](https://arxiv.org/html/2610.00385#bib.bib15)\)\. FAER connects complete\-trajectory selection to a matched learner objective and records the information available at each stage\.
##### Learning\-aware data valuation\.
Influence functions and gradient\-based data valuation relate example\-level gradients to changes in a held\-out objective\([Koh & Liang, 2017](https://arxiv.org/html/2610.00385#bib.bib17)\)\. FAER uses this local connection to define a replay priority for complete trajectories, then measures its multi\-step utility under a matched learner rather than treating the first\-order score as an end\-to\-end guarantee\.
##### Language\-model post\-training\.
Instruction tuning and preference optimization establish the update stage at which replay choices ultimately matter\([Ouyang et al\., 2022](https://arxiv.org/html/2610.00385#bib.bib25);[Rafailov et al\., 2023](https://arxiv.org/html/2610.00385#bib.bib28);[Schulman et al\., 2017](https://arxiv.org/html/2610.00385#bib.bib32)\)\. For reasoning models, verifier\-based supervision and group\-relative optimization make correctness, confidence, and response length natural selection signals\([Cobbe et al\., 2021](https://arxiv.org/html/2610.00385#bib.bib4);[Lightman et al\., 2024](https://arxiv.org/html/2610.00385#bib.bib22);[Shao et al\., 2024](https://arxiv.org/html/2610.00385#bib.bib33);[DeepSeek\-AI, 2025](https://arxiv.org/html/2610.00385#bib.bib6);[Zheng et al\., 2024](https://arxiv.org/html/2610.00385#bib.bib39)\)\. Recent work connects replay directly to LLM post\-training: FreshPER applies age decay to experience priorities[\(FreshPER\)](https://arxiv.org/abs/2604.16918), Prioritized Replay for RL Post\-training uses problem\-level statistics[\(RL post\-training PER\)](https://arxiv.org/abs/2601.02648), RLEP mixes verified trajectories with new rollouts[\(RLEP\)](https://arxiv.org/abs/2507.07451), and M2PO studies stale\-data off\-policy updates[\(M2PO\)](https://proceedings.iclr.cc/paper_files/paper/2026/hash/8590b946e4f0c2b1262cb5b047924653-Abstract-Conference.html)\. Reliability\-Adjusted Prioritized Experience Replay integrates reliability in an actual replay learner[\(reliability\-adjusted PER\)](https://proceedings.iclr.cc/paper_files/paper/2026/hash/60e3d7caf8cfdaf6598467f9b0ebf36a-Abstract-Conference.html)\. These papers motivate measuring the full path from stored trajectory to learner update; FAER links selector diagnostics on frozen caches to matched learner outcomes with a shared provenance contract\.
##### Agent memory and feedback\.
Reflexion and Self\-Refine reuse textual feedback across attempts, while Voyager maintains an expanding library of executable skills\([Shinn et al\., 2023](https://arxiv.org/html/2610.00385#bib.bib34);[Madaan et al\., 2023](https://arxiv.org/html/2610.00385#bib.bib23);[Wang et al\., 2023a](https://arxiv.org/html/2610.00385#bib.bib35)\)\. ReAct interleaves reasoning with actions, showing why a trajectory can include structure beyond a final answer\([Yao et al\., 2023](https://arxiv.org/html/2610.00385#bib.bib38)\)\. These settings raise related questions about identity, feedback provenance, and replay unit\. Our unit is a complete model\-cache trajectory; the protocol records the selected ID and full\-trace digest before adding the offline task label\.
##### Reasoning evaluation and reporting\.
Chain\-of\-thought, self\-consistency, and zero\-shot reasoning studies establish the importance of the generated trace and the evaluation contract\([Wei et al\., 2022](https://arxiv.org/html/2610.00385#bib.bib37);[Wang et al\., 2023b](https://arxiv.org/html/2610.00385#bib.bib36);[Kojima et al\., 2022](https://arxiv.org/html/2610.00385#bib.bib18)\)\. Model reports for Qwen2\.5\-Math and DeepSeek\-V3 make the model, tokenizer, and data identity part of a reproducible result record\([Qwen Team, 2024](https://arxiv.org/html/2610.00385#bib.bib27);[DeepSeek\-AI, 2024](https://arxiv.org/html/2610.00385#bib.bib5)\)\. Offline RL surveys and causal estimands further motivate explicit sampling units, treatment assignments, and outcome joins\([Levine et al\., 2020](https://arxiv.org/html/2610.00385#bib.bib21);[Rubin, 1974](https://arxiv.org/html/2610.00385#bib.bib29)\)\. FAER brings these reporting requirements to trajectory selection by recording the observed projection, random draw, trace digest, and post\-freeze evaluator join as one auditable contract\.
##### Freshness, provenance, and evaluation\.
AgentTrails treats trace lineage as an artifact\-level object[\(AgentTrails\)](https://arxiv.org/abs/2607.18816), while datasheets and model cards specify provenance fields for datasets and models\([Gebru et al\., 2021](https://arxiv.org/html/2610.00385#bib.bib12);[Mitchell et al\., 2019](https://arxiv.org/html/2610.00385#bib.bib24)\)\. Reproducible RL work emphasizes seed variation and implementation details because small protocol changes can alter conclusions\([Henderson et al\., 2018](https://arxiv.org/html/2610.00385#bib.bib13);[Agarwal et al\., 2021](https://arxiv.org/html/2610.00385#bib.bib1);[Pineau et al\., 2021](https://arxiv.org/html/2610.00385#bib.bib26)\)\. Task and data documentation preserve the labeling and population context of these records\([Bender & Friedman, 2018](https://arxiv.org/html/2610.00385#bib.bib2)\)\. Our paired\-sampling analysis follows finite\-population estimation and randomization\-based uncertainty principles\([Horvitz & Thompson, 1952](https://arxiv.org/html/2610.00385#bib.bib14);[Sarndal et al\., 1992](https://arxiv.org/html/2610.00385#bib.bib30);[Dodge et al\., 2019](https://arxiv.org/html/2610.00385#bib.bib7);[Dror et al\., 2018](https://arxiv.org/html/2610.00385#bib.bib8)\)\. The distinction is operational: recorded row order supports an index\-age stress test, whereas measured freshness requires timestamps and version provenance for every trajectory\.
## 3Method
### 3\.1Problem and Information Boundary
Let a frozen cache be𝒞=\{xi\}i=1N\\mathcal\{C\}=\\\{x\_\{i\}\\\}\_\{i=1\}^\{N\}, where a record contains an ID, full generated trajectory, binary format\-status feedbackrir\_\{i\}, verifier confidenceρi\\rho\_\{i\}, model confidencecic\_\{i\}, token countLiL\_\{i\}, and producer/join metadata\. The selector receives the declared observed\-field projectionϕ\(xi\)=\(ri,ρi,ci,Li,idi\)\\phi\(x\_\{i\}\)=\(r\_\{i\},\\rho\_\{i\},c\_\{i\},L\_\{i\},\\mathrm\{id\}\_\{i\}\)\. The task reference and post\-freeze correctness labelyiy\_\{i\}are held by the evaluator\. Selection is therefore a two\-view computation: a priority policy consumesϕ\(xi\)\\phi\(x\_\{i\}\), while the evaluator joinsyiy\_\{i\}only after the selected IDs and trace digest have been committed\.
This separation yields two different estimands\. For a realized drawSmS\_\{m\}under policymm, the selected\-record rate isq^m=\|Sm\|−1∑i∈Smyi\\widehat\{q\}\_\{m\}=\|S\_\{m\}\|^\{\-1\}\\sum\_\{i\\in S\_\{m\}\}y\_\{i\}\. It describes that draw from that cache\. Across counter\-seeds,q¯m\\bar\{q\}\_\{m\}and the sample SD summarize selection variation conditional on the frozen rows and policy\. Held\-out task quality after a learner update is a separate estimand with its own reporting surface\. The procedure, row unit, denominator, and conditioning set are fixed in this definition, so all reported numbers have the same interpretation\.
The selector serialization recursively rejects target and correctness keys\. The audit scans all 64 training and 128 test rows, substitutes taint values for private target fields, and compares both the serialized selector view and each policy priority\. The downstream evaluation join reads a separately stored label table after trace freezing\. Producer manifests bind the cache to source run, model/data digests, and record count\. This establishes selector\-side exclusion and byte\-level cache identity; semantic provenance for derived fields such as verifier confidence is tracked separately in the producer contract in Appendix[B\.4](https://arxiv.org/html/2610.00385#A2.SS4)\.
To make the mismatch measurable, letqmq\_\{m\}denote the selected\-record correctness estimand and letQmQ\_\{m\}denote the held\-out learner quality for policymm\. We summarize the alignment with
τSL=τ\(\{qm\}m∈ℳ,\{Qm\}m∈ℳ\),ISL=∑m<n\[\(qm−qn\)\(Qm−Qn\)<0\]\.\\tau\_\{\\mathrm\{SL\}\}=\\tau\\\!\\left\(\\\{q\_\{m\}\\\}\_\{m\\in\\mathcal\{M\}\},\\\{Q\_\{m\}\\\}\_\{m\\in\\mathcal\{M\}\}\\right\),\\qquad I\_\{\\mathrm\{SL\}\}=\\sum\_\{m<n\}\\mathbf\{1\}\\\!\\left\[\(q\_\{m\}\-q\_\{n\}\)\(Q\_\{m\}\-Q\_\{n\}\)<0\\right\]\.\(1\)Hereτ\\tauis Kendall rank correlation andISLI\_\{\\mathrm\{SL\}\}counts policy\-order inversions among arms with both endpoints\. Spearman correlation is reported alongside Kendall correlation because the policy set is small and the estimands have different units\. We additionally report the normalized inversion rate and the regret incurred by selecting the policy with highest cached correctness:
ISLnorm=ISL\(\|ℳ\|2\),RSL=maxm∈ℳQm−Qmq,mq=argmaxm∈ℳqm\.I\_\{\\mathrm\{SL\}\}^\{\\mathrm\{norm\}\}=\\frac\{I\_\{\\mathrm\{SL\}\}\}\{\\binom\{\|\\mathcal\{M\}\|\}\{2\}\},\\qquad R\_\{\\mathrm\{SL\}\}=\\max\_\{m\\in\\mathcal\{M\}\}Q\_\{m\}\-Q\_\{m\_\{q\}\},\\qquad m\_\{q\}=\\arg\\max\_\{m\\in\\mathcal\{M\}\}q\_\{m\}\.\(2\)Ties inmqm\_\{q\}follow the fixed policy order in the run manifest\. For replicated comparisons these quantities are computed within each producer–learner block\(qm,s,Qm,s\)\(q\_\{m,s\},Q\_\{m,s\}\)and summarized by block bootstrap\.
Figure 1:FAER algorithm and audit contract\. The online selector sees the frozen trajectory cache and observed fields, scores and samples 32 records without replacement, and commits an ordered trace digest\. The evaluator and three audit lanes join hidden labels and provenance only after the freeze boundary\.
### 3\.2Priority and Audit Sampling
For recordii, letri∈\{0,1\}r\_\{i\}\\in\\\{0,1\\\}be observed format\-status feedback,ρi∈\[0,1\]\\rho\_\{i\}\\in\[0,1\]verifier confidence,si=1−cis\_\{i\}=1\-c\_\{i\}surprise,ui=1−ρiu\_\{i\}=1\-\\rho\_\{i\}uncertainty,cic\_\{i\}model confidence, andLiL\_\{i\}generated tokens\. The experiment compares seven saved priority rules:
pi\(m\)=\{1m=uniform,r~im=format\-feedback,ρ~im=reliability,s~im=surprise,u~im=uncertainty,ρ~i\(0\.5\+r~i\)\(0\.5\+s~i\)\(0\.5\+u~i\)m=feedback\-aware,c~i/Lim=token\-efficiency,p\_\{i\}^\{\(m\)\}=\\begin\{cases\}1&m=\\mathrm\{uniform\},\\\\ \\tilde\{r\}\_\{i\}&m=\\mathrm\{format\\mbox\{\-\}feedback\},\\\\ \\tilde\{\\rho\}\_\{i\}&m=\\mathrm\{reliability\},\\\\ \\tilde\{s\}\_\{i\}&m=\\mathrm\{surprise\},\\\\ \\tilde\{u\}\_\{i\}&m=\\mathrm\{uncertainty\},\\\\ \\tilde\{\\rho\}\_\{i\}\(0\.5\+\\tilde\{r\}\_\{i\}\)\(0\.5\+\\tilde\{s\}\_\{i\}\)\(0\.5\+\\tilde\{u\}\_\{i\}\)&m=\\mathrm\{feedback\\mbox\{\-\}aware\},\\\\ \\tilde\{c\}\_\{i\}/\\sqrt\{L\_\{i\}\}&m=\\mathrm\{token\\mbox\{\-\}efficiency\},\\end\{cases\}\(3\)wherex~=max\(ϵ,x\)\\tilde\{x\}=\\max\(\\epsilon,x\)andϵ=10−6\\epsilon=10^\{\-6\}provides support for every row\. At each step, a remaining row is sampled with probability proportional to its priority and removed from the pool\. Uniform is the equal\-weight control; format\-feedback, reliability, surprise, and uncertainty isolate individual cached signals; token\-efficiency normalizes confidence by the square root of response length; feedback\-aware combines the four factors with fixed offsets\. The exact policy aliases match the retained run manifest\.
Holding the other factors fixed, the verifier\-dependent factor isg\(ρ\)=ρ\(0\.5\+1−ρ\)=ρ\(1\.5−ρ\)g\(\\rho\)=\\rho\(0\.5\+1\-\\rho\)=\\rho\(1\.5\-\\rho\)\. Its derivative isg′\(ρ\)=1\.5−2ρg^\{\\prime\}\(\\rho\)=1\.5\-2\\rhoand its second derivative is−2\-2, so the maximum occurs atρ=0\.75\\rho=0\.75with value0\.56250\.5625\. The equal valuesg\(0\.5\)=g\(1\)=0\.5g\(0\.5\)=g\(1\)=0\.5make the rule non\-monotone in confidence\. We therefore treat this multiplicative rule asFAER\-Fixed, a training\-free instance of the audit contract rather than a claimed optimal utility estimator\. The separate age stress usesexp\[−λai\(b\)\]\\exp\[\-\\lambda a\_\{i\}\(b\)\]withai\(b\)a\_\{i\}\(b\)defined by the declared original\-row index; the cache scan records 0/192 measured timestamp fields\.
The seed is part of the trace manifest\. Seed 13 supplies the per\-policy slice; seeds 13, 17, 23, 29, 31, 37, 41, and 43 supply eight draws from the same cache\. We report selected\-record rate, sample SD, min–max, and pairwise Jaccard\. The existing ledger records policy\-wise weighted draws; a common\-random\-number paired design with inclusion probabilities is reported as a separate experiment in Appendix[D\.5](https://arxiv.org/html/2610.00385#A4.SS5)\.
### 3\.3Cross\-Fitted Learner\-Aware Utility Replay
The fixed product is a training\-free protocol baseline\. The primary utility selector adds a gradient feature computed at a frozen reference checkpoint using disjoint calibration data\. For candidate trajectoryxix\_\{i\}in target blockbb, define
gi\(−b\)=∇θℓ\(θ0,xi\),gcal\(−b\)=∇θℒcal\(−b\)\(θ0\),ai\(−b\)=⟨gi\(−b\),gcal\(−b\)⟩‖gi\(−b\)‖2‖gcal\(−b\)‖2\+ϵ\.g\_\{i\}^\{\(\-b\)\}=\\nabla\_\{\\theta\}\\ell\(\\theta\_\{0\};x\_\{i\}\),\\qquad g\_\{\\mathrm\{cal\}\}^\{\(\-b\)\}=\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\mathrm\{cal\}\}^\{\(\-b\)\}\(\\theta\_\{0\}\),\\qquad a\_\{i\}^\{\(\-b\)\}=\\frac\{\\langle g\_\{i\}^\{\(\-b\)\},g\_\{\\mathrm\{cal\}\}^\{\(\-b\)\}\\rangle\}\{\\\|g\_\{i\}^\{\(\-b\)\}\\\|\_\{2\}\\\|g\_\{\\mathrm\{cal\}\}^\{\(\-b\)\}\\\|\_\{2\}\+\\epsilon\}\.\(4\)The feature uses no target\-block correctness label\. Under aβ\\beta\-smooth calibration loss, one gradient updateθ′=θ−ηgi\\theta^\{\\prime\}=\\theta\-\\eta g\_\{i\}satisfies
ℒcal\(θ′\)≤ℒcal\(θ\)−η⟨gcal,gi⟩\+βη22‖gi‖22\.\\mathcal\{L\}\_\{\\mathrm\{cal\}\}\(\\theta^\{\\prime\}\)\\leq\\mathcal\{L\}\_\{\\mathrm\{cal\}\}\(\\theta\)\-\\eta\\langle g\_\{\\mathrm\{cal\}\},g\_\{i\}\\rangle\+\\frac\{\\beta\\eta^\{2\}\}\{2\}\\\|g\_\{i\}\\\|\_\{2\}^\{2\}\.\(5\)The bound is expressed in the unnormalized inner product and candidate\-gradient norm; cosine alignment is therefore a normalized comparison baseline rather than an identity implied by the bound\. To make optimizer geometry explicit, we also define a disposable virtual update using the same frozen learner operator and optimizer state:
Δθi\(−b\)=𝒰1\(θ0,s0,xi\)−θ0,vi,FO\(−b\)=−⟨gcal\(−b\),Δθi\(−b\)⟩\.\\Delta\\theta\_\{i\}^\{\(\-b\)\}=\\mathcal\{U\}\_\{1\}\(\\theta\_\{0\},s\_\{0\};x\_\{i\}\)\-\\theta\_\{0\},\\qquad v\_\{i,\\mathrm\{FO\}\}^\{\(\-b\)\}=\-\\langle g\_\{\\mathrm\{cal\}\}^\{\(\-b\)\},\\Delta\\theta\_\{i\}^\{\(\-b\)\}\\rangle\.\(6\)The virtual state is discarded immediately\. A smoothness\-compatible score adds a quadratic correction,
vi\(−b\)=vi,FO\(−b\)−λ\(−b\)‖Δθi\(−b\)‖22,v\_\{i\}^\{\(\-b\)\}=v\_\{i,\\mathrm\{FO\}\}^\{\(\-b\)\}\-\\lambda^\{\(\-b\)\}\\\|\\Delta\\theta\_\{i\}^\{\(\-b\)\}\\\|\_\{2\}^\{2\},\(7\)withλ\(−b\)\\lambda^\{\(\-b\)\}selected only on non\-target calibration blocks\. The optimizer\-aware fidelity and ablation surfaces are reported in Appendix[D\.22](https://arxiv.org/html/2610.00385#A4.SS22); the completed learner rows use the normalized alignment treatment on the reported calibration surface\.
The exponential replay rule also has a variational interpretation\. For a fixed reference distributionqi\>0q\_\{i\}\>0, the distribution
p∗=argmaxp∈ΔN\{∑ipiei−τDKL\(p∥q\)\}⟹pi∗=qiexp\(ei/τ\)∑jqjexp\(ej/τ\)p^\{\*\}=\\arg\\max\_\{p\\in\\Delta\_\{N\}\}\\left\\\{\\sum\_\{i\}p\_\{i\}e\_\{i\}\-\\tau D\_\{\\mathrm\{KL\}\}\(p\\\|q\)\\right\\\}\\quad\\Longrightarrow\\quad p\_\{i\}^\{\*\}=\\frac\{q\_\{i\}\\exp\(e\_\{i\}/\\tau\)\}\{\\sum\_\{j\}q\_\{j\}\\exp\(e\_\{j\}/\\tau\)\}\(8\)is the unique KL\-regularized utility maximizer\. FAER samples this distribution sequentially without replacement, renormalizing over the remaining rows\. Uniformqqis the default; the reference choice and temperature are part of the calibration contract\.
The metadata score uses nonredundant observed fields, withsi=1−cis\_\{i\}=1\-c\_\{i\}:
ei\(𝐰,γ\)=wrri\+wρρi\+wssi−wLlog\(1\+Li\)\+γai\(−b\),piutility=exp\(ei\(𝐰,γ\)/τ\)\.e\_\{i\}\(\\mathbf\{w\},\\gamma\)=w\_\{r\}r\_\{i\}\+w\_\{\\rho\}\\rho\_\{i\}\+w\_\{s\}s\_\{i\}\-w\_\{L\}\\log\(1\+L\_\{i\}\)\+\\gamma a\_\{i\}^\{\(\-b\)\},\\qquad p\_\{i\}^\{\\mathrm\{utility\}\}=\\exp\(e\_\{i\}\(\\mathbf\{w\},\\gamma\)/\\tau\)\.\(9\)We omit a separate coefficient forui=1−ρiu\_\{i\}=1\-\\rho\_\{i\}because it is linearly redundant withρi\\rho\_\{i\}under an intercept\-equivalent normalization\. Sequential sampling remains without replacement\. The metadata\-only ablation setsγ=0\\gamma=0; alignment\-only replay sets the metadata coefficients to zero except for the shared length support term\.
For target blockbb,\(𝐰\(−b\),γ\(−b\),τ\(−b\)\)\(\\mathbf\{w\}^\{\(\-b\)\},\\gamma^\{\(\-b\)\},\\tau^\{\(\-b\)\}\)are selected using only the remaining calibration blocks to maximize their matched learner quality\. The parameters, calibration\-data digest, checkpoint, gradient implementation, and optimizer configuration are frozen before the target draw\. Target\-block correctness and held\-out quality remain inaccessible until the selected trace and learner checkpoint are committed\.
##### Audit boundary for learner\-aware features\.
The audit binds the reference checkpoint, disjoint calibration IDs, gradient implementation, fitted parameters, ordered selected IDs, and replay digest\. Mutating target\-block labels must leave the alignment feature, observed projection, priorities, and trace unchanged; the completed matched mutation count and denominator are 1,920/1,920\. The optimizer\-aware extension additionally records checkpoint and optimizer\-state digests, and requires the disposable virtual step to leave both states unchanged\.
### 3\.4From Selection to a Matched Learner
For armmmand replication seedss, the replay interface supplies an ordered trajectory batchSm,s,tS\_\{m,s,t\}to a common learner at updatett\. Denote that shared update operator by𝒰\\mathcal\{U\}and its fixed configuration byη\\eta:
θm,s,t\+1=𝒰\(θm,s,t,Sm,s,t,η\),Qm,s=1\|𝒟test\|∑x∈𝒟testcorrect\(decode\(θm,s,T,x\),yx\)\.\\theta\_\{m,s,t\+1\}=\\mathcal\{U\}\(\\theta\_\{m,s,t\},S\_\{m,s,t\};\\eta\),\\qquad Q\_\{m,s\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{\\rm test\}\|\}\\sum\_\{x\\in\\mathcal\{D\}\_\{\\rm test\}\}\\operatorname\{correct\}\(\\operatorname\{decode\}\(\\theta\_\{m,s,T\},x\),y\_\{x\}\)\.\(10\)The selector determines the replay input; the learner configuration, stopping rule, and sealed evaluation task IDs are common across arms\. The run record carries the checkpoint step and digest alongsideQm,sQ\_\{m,s\}, generated tokens, verifier calls, and latency\. This separates a change in replay membership from a change in optimization or evaluation\. On the six arms with both endpoints, the cache ranking is format\-feedback\>\>token\-efficiency\>\>reliability\>\>feedback\-aware\>\>uniform\>\>uncertainty, whereas the learner ranking is feedback\-aware\>\>token\-efficiency\>\>format\-feedback\>\>reliability\>\>uncertainty\>\>uniform\. Their measured alignment isτSL=0\.33\\tau\_\{\\mathrm\{SL\}\}=0\.33, SpearmanρSL=0\.54\\rho\_\{\\mathrm\{SL\}\}=0\.54, andISL=5I\_\{\\mathrm\{SL\}\}=5of 15 pairs; these six\-arm statistics are descriptive, while the same\-seed block surface is the primary replicated analysis\.
The matched learner fixes the model, tokenizer, AdamW/bf16 optimizer, update horizon, token and verifier caps, initialization, stopping rule, and sealed task IDs across arms\. The learner endpoint is the masked autoregressive objective
ℒ\(θ;Bt\)=−1\|Bt\|∑x∈Bt∑j∈ℳ\(x\)logpθ\(xj∣x<j\),\\mathcal\{L\}\(\\theta;B\_\{t\}\)=\-\\frac\{1\}\{\|B\_\{t\}\|\}\\sum\_\{x\\in B\_\{t\}\}\\sum\_\{j\\in\\mathcal\{M\}\(x\)\}\\log p\_\{\\theta\}\(x\_\{j\}\\mid x\_\{<j\}\),\(11\)whereℳ\(x\)\\mathcal\{M\}\(x\)covers assistant\-response tokens in a selected full trajectory\. The supplied run uses equal token loss weights, greedy decoding, SHA\-256\-bound checkpoints, and seeded cyclic replay of 32 records for 128 updates\. The age input, sequential sampling complexity, edge cases, and trace commitment are specified in Appendices[A](https://arxiv.org/html/2610.00385#A1)and[C](https://arxiv.org/html/2610.00385#A3); the learner uses no importance correction\.
### 3\.5Paired Outcomes and Replication
The paired selection estimand and inclusion\-probability design are defined in Appendix[A\.3](https://arxiv.org/html/2610.00385#A1.SS3)\. The learner comparison uses the replication seed as its block, binding producer cache, initialization, trace, checkpoint, and sealed evaluation IDs; full pairing and randomization details appear in Appendix[D\.5](https://arxiv.org/html/2610.00385#A4.SS5)\.
## 4Experiments
### 4\.1Experimental Setup
We evaluate complete Qwen\-generated trajectories on GSM8K using three linked stages: fixed\-cache selection, measured\-age selection, and a matched learner\. Table[1](https://arxiv.org/html/2610.00385#S4.T1)records their populations\. The reference family\-A cache contains 64 train and 128 held\-out trajectories, and the family\-B cache contains 256 and 1,319\. Each policy draws 32 unique records\. Selection sensitivity uses seeds 13, 17, 23, 29, 31, 37, 41, and 43; the core independent learner/cache/evaluation replication uses seeds 13, 17, and 23, with the supplied replication surface extending the learner comparison to five\- and eight\-seed aggregates\.
Table 1:Experimental populations and replication units\. Each stage has its own cache binding and denominator\.The matched learner contract uses Qwen2\.5\-1\.5B\-Instruct, AdamW, bf16, learning rate2×10−52\\times 10^\{\-5\}, 128 updates, a 65,536 generated\-token cap, and a 4,096 verifier\-call cap\. Stable task\-ID splits contain 256/128/256 train/calibration/test tasks, with zero overlaps\. The independent producers supply 640 rows per seed and completeage\_secondsmetadata for 1,920 rows\. Six arms are specified: uniform, format\-feedback, reliability, feedback\-aware, token\-efficiency, and uncertainty\-PER\. The main table reports the four primary arms; Reliability and Uncertainty\-PER are added with their measured quality, token, latency, and paired\-interval outcomes in Appendix[C](https://arxiv.org/html/2610.00385#A3)\.
Selected\-record correctness is the number of correct cached answers among 32 selected rows\. Held\-out learner quality evaluates generated answers on the sealed task split after optimization\. Cached token means describe stored response lengths; learner tokens count generated tokens during the matched run\. We report p50/p95 synchronized end\-to\-end latency in seconds\. Eight\-draw summaries use sample SD; paired contrasts use the supplied 95% intervals and the inference design in Appendix[D\.5](https://arxiv.org/html/2610.00385#A4.SS5)\. Evaluator validation and unresolved\-output accounting appear in Appendix[B\.6](https://arxiv.org/html/2610.00385#A2.SS6)\.
##### Cost accounting and run inclusion\.
We distinguish target execution costCtargetC\_\{\\mathrm\{target\}\}, selector\-fitting costCcalibrationC\_\{\\mathrm\{calibration\}\}, and total costCtotal=Ctarget\+CcalibrationC\_\{\\mathrm\{total\}\}=C\_\{\\mathrm\{target\}\}\+C\_\{\\mathrm\{calibration\}\}\. Fixed selectors have zero calibration cost\. All preregistered runs remain in the analysis regardless of whether they satisfy the analysis success criterion; only predeclared execution failures, corrupted artifact identity, incomplete checkpoint serialization, or matched\-budget violations permit exclusion\.
### 4\.2Main Learner Results
Feedback\-aware replay reaches held\-out quality 0\.6329, compared with 0\.5482 for uniform and 0\.6037 for format\-feedback \(Table[2](https://arxiv.org/html/2610.00385#S4.T2)\)\. Reliability reaches 0\.5948 and Uncertainty\-PER reaches 0\.5817 under the same contract \(Appendix[C](https://arxiv.org/html/2610.00385#A3)\)\. The cross\-fitted metadata\-only calibration reaches0\.6476±0\.01390\.6476\\pm 0\.0139over eight independent seeds, with median 0\.6481 and paired interval\[\+0\.079,\+0\.122\]\[\+0\.079,\+0\.122\]\. Its target execution uses 63,276 generated tokens; the full recorded calibration cost is 189,642 tokens and 3\.48 GPU\-hours\. The optimizer\-aware FAER\-UTILITY row is 0\.6624 quality at 62,844 target tokens, 189,210 total tokens, and 4\.26 GPU\-hours, with paired interval\[\+0\.094,\+0\.134\]\[\+0\.094,\+0\.134\]\. The paired feedback\-aware\-minus\-uniform gain is\+0\.0847\+0\.0847, with 95% interval\[\+0\.058,\+0\.112\]\[\+0\.058,\+0\.112\]\. All six core arms share the model, split, update count, and budget cap, while learned selector parameters are frozen before each target block\.
Table 2:Matched learner on the sealed GSM8K split\. Target tokens count only the target replay run; total tokens and GPU\-hours include selector calibration\. Quality is higher\-is\-better and cost is lower\-is\-better\. The core rows use three independent seeds; the metadata\-only calibrated row is the eight\-seed aggregate\.FAER\-Fixed uses 63,771 generated tokens, with p50/p95 latency of 0\.829/1\.289 seconds\. Token\-efficiency reaches 0\.6186 quality and the lowest token count \(63,428\) and latency \(0\.806/1\.251 seconds\)\. Figure[2](https://arxiv.org/html/2610.00385#S4.F2)shows the quality–cost trade\-off and the paired intervals\. Format\-feedback and token\-efficiency also improve over uniform, with intervals\[\+0\.031,\+0\.080\]\[\+0\.031,\+0\.080\]and\[\+0\.043,\+0\.098\]\[\+0\.043,\+0\.098\], respectively\. The optimizer\-aware row is highest on this completed surface; its selector cost is reported separately from target execution\. The extra arms place the composite above reliability and uncertainty\-PER in the same three\-seed comparison, while the frozen\-cache correctness ranking remains different from the learner ranking\.
Figure 2:Matched learner outcomes\. Left: held\-out quality versus target\-run generated tokens; right: paired quality differences from uniform with 95% intervals\. The interval bars belong to the differences, and the quality points show reported means\.
### 4\.3Same\-Seed Selection\-to\-Learning Gap
The core gap analysis is repeated within matched producer–learner blocks\. For each seed and replay arm,qm,sq\_\{m,s\}is computed from the exact trajectory set consumed by the learner andQm,sQ\_\{m,s\}is evaluated from its resulting checkpoint on the same sealed split\. The table establishes the estimand and reporting surface for the matched blocks\.
Table 3:Same\-seed selection\-to\-learning analysis\. Each block uses one producer cache, one learner initialization, and the same sealed evaluation IDs across arms\. Rank statistics and regret are summarized over independent blocks\.The paired analysis uses the seed as the blocking unit and recomputes rank statistics inside each bootstrap resample\. This directly tests whether the reported disagreement between cached correctness and learner quality persists when cache and checkpoint outcomes arise from one replay intervention\.
### 4\.4Learner\-Aware Utility Ablation
We isolate the gradient\-alignment contribution with metadata\-only, alignment\-only, full, and negative\-control selectors\. Shuffling alignment values preserves their marginal distribution while breaking trajectory identity; the random\-gradient control uses a norm\-matched direction\. The table reports the complete learner\-aware comparison\.
Table 4:Learner\-aware utility ablation under the common replay and update budget\. Quality is higher\-is\-better; target tokens count the target replay run\.The full selector is compared with both structured metadata calibration and controls that preserve either the alignment marginal or its scale\. This comparison separates learner\-aware trajectory identity from a change in priority dispersion\.
### 4\.5Selection Composition and Priority Analysis
Format\-feedback reaches 0\.6953 selected\-record correctness versus 0\.3594 forFAER\-Fixed; the complete selection table, scatter, and audit packages are in Appendix[D\.12](https://arxiv.org/html/2610.00385#A4.SS12)and Appendix[D](https://arxiv.org/html/2610.00385#A4)\.
## 5Limitations
The evidence covers GSM8K/Qwen2\.5\-1\.5B\-Instruct, two cache families, matched three\-seed runs, and five\- and eight\-seed aggregates at a 128\-update horizon\.FAER\-Fixedis a training\-free protocol baseline;FAER\-Utilityis learner\-aware\. Gradient alignment is local, and the audit contract records information flow and artifact identity while producer\-field semantics remain an empirical property of the evaluated packages\.
## 6Conclusion
FAER treats replay selection as a learner\-utility estimation problem with an explicit audit boundary\. The matched study shows thatFAER\-Fixedreaches 0\.6329 held\-out quality versus 0\.5482 for uniform and 0\.6037 for format\-feedback, while format\-feedback reaches frozen\-cache correctness 0\.6953 versus 0\.3594 for the fixed selector\. The completed FAER\-UTILITY row reaches 0\.6624 under the target contract and the policy\-optimization row reaches 0\.6575 final quality\. The six\-arm descriptive gap has KendallτSL=0\.33\\tau\_\{\\mathrm\{SL\}\}=0\.33, SpearmanρSL=0\.54\\rho\_\{\\mathrm\{SL\}\}=0\.54, and five pairwise inversions\. Metadata\-only calibration reaches0\.6476±0\.01390\.6476\\pm 0\.0139with full cost reported separately\. Same\-seed, zero\-shot transfer, policy\-optimization, and optimizer\-aware diagnostic surfaces remain tied to their declared learner, trace, and evaluation boundaries\.
## AI\-use
We used generative AI tools for language polishing and for summarizing references\. We have not used generative AI tools to generate experimental results, create synthetic datasets, formulate mathematical claims, provide proofs, or make decisions regarding research conclusions\. The design of the methodology, experimental setup, analysis, and interpretation of results were conducted and verified by the authors\. Other required disclosure tasks not mentioned above are not applicable to this work\. We take full responsibility for the final content of this work, including all text, claims, analyses, and artifacts produced with the assistance of generative AI tools\.
## ETHICS STATEMENT
This work uses model\-generated trajectories and public mathematical reasoning benchmarks\. It involves no human subjects, personal data, user study, or deployment affecting individuals\. The audit contract separates selector\-visible metadata from post\-freeze evaluation labels, and the reported experiments use fixed task splits, bounded compute, and reproducible offline evaluation\. The supplementary release excludes model weights and any source artifacts whose redistribution is restricted by their licenses\.
## REPRODUCIBILITY STATEMENT
The code for trajectory selection, audit checks, sampling, learner evaluation, statistical summaries, and figure generation is provided in the Supplementary Material\. The manuscript specifies the model and tokenizer identity, task splits, seeds, update horizon, token and verifier budgets, selector contracts, evaluator rules, and aggregation units\. Run manifests, selected\-ID traces, checkpoint identifiers, result mappings, and digest\-bound artifact records are included so that each reported table and figure can be reconstructed from the corresponding package\.
## References
- Agarwal et al\. \(2021\)Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G\. Bellemare\.Deep reinforcement learning at the edge of the statistical precipice\.In*Advances in Neural Information Processing Systems*, volume 34, 2021\.
- Bender & Friedman \(2018\)Emily M\. Bender and Batya Friedman\.Data statements for natural language processing: Toward mitigating system bias and enabling better science\.*Transactions of the Association for Computational Linguistics*, 6:587–604, 2018\.
- Chen et al\. \(2021\)Lili Chen et al\.Decision transformer: Reinforcement learning via sequence modeling\.In*Advances in Neural Information Processing Systems*, 2021\.
- Cobbe et al\. \(2021\)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\.Training verifiers to solve math word problems\.In*arXiv preprint arXiv:2110\.14168*, 2021\.
- DeepSeek\-AI \(2024\)DeepSeek\-AI\.Deepseek\-v3 technical report\.*arXiv preprint arXiv:2412\.19437*, 2024\.
- DeepSeek\-AI \(2025\)DeepSeek\-AI\.Deepseek\-r1: Incentivizing reasoning capability in llms via reinforcement learning\.*arXiv preprint arXiv:2501\.12948*, 2025\.
- Dodge et al\. \(2019\)Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A\. Smith\.Show your work: Improved reporting of experimental results\.In*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing*, pp\. 2185–2194, 2019\.
- Dror et al\. \(2018\)Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart\.The hitchhiker’s guide to testing statistical significance in natural language processing\.In*Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics*, pp\. 1383–1392, 2018\.
- Espeholt et al\. \(2018\)Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu\.Impala: Scalable distributed deep\-rl with importance weighted actor\-learner architectures\.In*Proceedings of the 35th International Conference on Machine Learning*, pp\. 1407–1416, 2018\.
- Fedus et al\. \(2020\)William Fedus, Carles Gelada, Yoshua Bengio, Marc G\. Bellemare, and Hugo Larochelle\.Revisiting fundamentals of experience replay\.In*Proceedings of the 37th International Conference on Machine Learning*, pp\. 3061–3071, 2020\.
- Fu et al\. \(2021\)Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine\.D4rl: Datasets for deep data\-driven reinforcement learning\.In*International Conference on Learning Representations*, 2021\.
- Gebru et al\. \(2021\)Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daume, and Kate Crawford\.Datasheets for datasets\.*Communications of the ACM*, 64\(12\):86–92, 2021\.
- Henderson et al\. \(2018\)Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger\.Deep reinforcement learning that matters\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 32, pp\. 3207–3214, 2018\.
- Horvitz & Thompson \(1952\)Daniel G\. Horvitz and Donovan J\. Thompson\.A generalization of sampling without replacement from a finite universe\.*Journal of the American Statistical Association*, 47\(260\):663–685, 1952\.
- Janner et al\. \(2022\)Michael Janner, Yilun Du, Joshua B\. Tenenbaum, and Sergey Levine\.Planning with diffusion for flexible behavior synthesis\.In*International Conference on Machine Learning*, 2022\.
- Kapturowski et al\. \(2019\)Steven Kapturowski, Georg Ostrovski, Will Dabney, John Quan, and Remi Munos\.Recurrent experience replay in distributed reinforcement learning\.In*International Conference on Learning Representations*, 2019\.
- Koh & Liang \(2017\)Pang Wei Koh and Percy Liang\.Understanding black\-box predictions via influence functions\.In*Proceedings of the 34th International Conference on Machine Learning*, 2017\.
- Kojima et al\. \(2022\)Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa\.Large language models are zero\-shot reasoners\.In*Advances in Neural Information Processing Systems*, volume 35, 2022\.
- Kostrikov et al\. \(2022\)Ilya Kostrikov, Ashvin Nair, and Sergey Levine\.Offline reinforcement learning with implicit q\-learning\.In*International Conference on Learning Representations*, 2022\.
- Kumar et al\. \(2020\)Aviral Kumar, Rishabh Zhou, George Tucker, and Sergey Levine\.Conservative q\-learning for offline reinforcement learning\.In*Advances in Neural Information Processing Systems*, 2020\.
- Levine et al\. \(2020\)Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu\.Offline reinforcement learning: Tutorial, review, and perspectives on open problems\.In*arXiv preprint arXiv:2005\.01643*, 2020\.
- Lightman et al\. \(2024\)Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In*International Conference on Learning Representations*, 2024\.
- Madaan et al\. \(2023\)Aman Madaan et al\.Self\-refine: Iterative refinement with self\-feedback\.In*Advances in Neural Information Processing Systems*, 2023\.
- Mitchell et al\. \(2019\)Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru\.Model cards for model reporting\.In*Proceedings of the Conference on Fairness, Accountability, and Transparency*, pp\. 220–229, 2019\.
- Ouyang et al\. \(2022\)Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L\. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe\.Training language models to follow instructions with human feedback\.In*Advances in Neural Information Processing Systems*, 2022\.
- Pineau et al\. \(2021\)Joelle Pineau, Philippe Vincent\-Lamarre, Koustuv Sinha, Vincent Lariviere, Alina Beygelzimer, Florence d’Alche Buc, Emily Fox, and Hugo Larochelle\.Improving reproducibility in machine learning research \(a report from the neurips 2019 reproducibility program\)\.*Journal of Machine Learning Research*, 22\(164\):1–20, 2021\.
- Qwen Team \(2024\)Qwen Team\.Qwen2\.5\-math technical report: Toward mathematical expert model via self\-improvement\.*arXiv preprint arXiv:2409\.12122*, 2024\.
- Rafailov et al\. \(2023\)Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D\. Manning, and Chelsea Finn\.Direct preference optimization: Your language model is secretly a reward model\.In*Advances in Neural Information Processing Systems*, 2023\.
- Rubin \(1974\)Donald B\. Rubin\.Estimating causal effects of treatments in randomized and nonrandomized studies\.In*Journal of Educational Psychology*, volume 66, pp\. 688–701, 1974\.
- Sarndal et al\. \(1992\)Carl\-Erik Sarndal, Bengt Swensson, and Jan Wretman\.*Model Assisted Survey Sampling*\.Springer, 1992\.
- Schaul et al\. \(2016\)Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver\.Prioritized experience replay\.In*International Conference on Learning Representations*, 2016\.
- Schulman et al\. \(2017\)John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov\.Proximal policy optimization algorithms\.In*arXiv preprint arXiv:1707\.06347*, 2017\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y\. K\. Li, Y\. Wu, and Daya Guo\.Deepseekmath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Shinn et al\. \(2023\)Noah Shinn, Federico Cassano, Ashwin Berman, Ashay Gopinath, Karthik Narasimhan, and Shunyu Yao\.Reflexion: Language agents with verbal reinforcement learning\.In*Advances in Neural Information Processing Systems*, 2023\.
- Wang et al\. \(2023a\)Guanzhi Wang et al\.Voyager: An open\-ended embodied agent with large language models\.In*Advances in Neural Information Processing Systems*, 2023a\.
- Wang et al\. \(2023b\)Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V\. Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou\.Self\-consistency improves chain of thought reasoning in language models\.In*International Conference on Learning Representations*, 2023b\.
- Wei et al\. \(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V\. Le, and Denny Zhou\.Chain\-of\-thought prompting elicits reasoning in large language models\.In*Advances in Neural Information Processing Systems*, volume 35, 2022\.
- Yao et al\. \(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao\.React: Synergizing reasoning and acting in language models\.In*International Conference on Learning Representations*, 2023\.
- Zheng et al\. \(2024\)Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin, and Qwen Team\.Processbench: Identifying process errors in mathematical reasoning\.*arXiv preprint arXiv:2412\.06559*, 2024\.
## Appendix ATask, Sampling, and Method Details
### A\.1Selection unit and field contract
The sampling unit is a complete trajectory record rather than a token, answer span, or intermediate reasoning step\. A record is identified by a stable task/cache ID and contains the generated text plus producer\-side fields used by the priority rules\. For the primary caches, the declared selector fields are format\-status feedback, verifier confidence, model confidence, token count, and ID\. The target answer and oracle correctness label are stored in a separate evaluation view\. The split manifest binds each cache to its source run, model and data digests, number of rows, and split name\.
This unit keeps the audit question aligned with replay practice: a selected record is the object that would be supplied to a learner\. If a trajectory contains multiple textual segments, the stored sequence is retained as one object and the reported token count refers to the generated response represented by the cache row\. Selection operates on full records, with duplicate\-ID rejection and train/test disjointness validation before priorities are computed\.
### A\.2Sampling sequence
For a cache ofNNrows and a policy priority vectorp\(m\)p^\{\(m\)\}, the selection procedure drawsk=32k=32unique indices sequentially\. At drawtt, letRtR\_\{t\}be the set of rows not yet selected\. The next row is sampled from the categorical distributionpi\(m\)/∑j∈Rtpj\(m\)p\_\{i\}^\{\(m\)\}/\\sum\_\{j\\in R\_\{t\}\}p\_\{j\}^\{\(m\)\}fori∈Rti\\in R\_\{t\}\. The chosen row is removed, and the normalized distribution is recomputed\. The run seed and ordered draw list are serialized with a trace digest\. Task labels enter at the post\-freeze evaluation join after thekkdraws are committed\.
The operational sequence is:
1. 1\.Load the split manifest and row records; validate required field types, unique IDs, row count, and source digest\.
2. 2\.Project each record to its declared observed fields and recursively reject target/correctness keys\.
3. 3\.Compute the named policy priority for every row, apply the10−610^\{\-6\}support floor, and sample without replacement using the run seed\.
4. 4\.Serialize the ordered selected IDs, per\-draw policy state, seed, and replay digest\.
5. 5\.Freeze the selection trace, then join correctness labels by task ID and compute the selected\-record count over the fixed denominator of 32\.
The run also replays the stored action trace and checks deterministic digest agreement\. A trace match tests that the same rows are selected from the same manifest; it is distinct from the post\-freeze correctness outcome\.
### A\.3Policy interpretation
The seven policy names refer to the retained executable contracts\. Uniform assigns equal weights\. Format\-feedback uses the observed binary format\-status signal; reliability uses verifier confidence; surprise uses one minus model confidence; uncertainty uses one minus verifier confidence\. Feedback\-aware multiplies verifier confidence by offset terms for format\-feedback, surprise, and uncertainty\. Token\-efficiency divides confidence by the square root of generated length\. These policies consume observed cached fields; the separate target label enters the post\-freeze evaluator join\.
For the composite, the four factors serve different operational roles\. The verifier factor favors high verifier confidence up to its turning point; format\-feedback gives extra mass to format\-valid traces; surprise emphasizes lower model confidence; and uncertainty emphasizes lower verifier confidence\. The last two factors therefore encode exploration\-like preferences rather than a calibrated probability of correctness\. Multiplication makes the components interact, so a component’s effect depends on the values of the other three fields\. The analytic derivative in Section[A\.5](https://arxiv.org/html/2610.00385#A1.SS5)isolates the verifier factor while holding those fields fixed\.
### A\.4Sampling variation and finite\-population estimands
The primary repeated draws are weighted without replacement\. LetIi\(m,s\)I\_\{i\}^\{\(m,s\)\}indicate whether rowiiis selected by policymmunder seedss\. The observed selected\-record rate isq^m,s=k−1∑iIi\(m,s\)yi\\widehat\{q\}\_\{m,s\}=k^\{\-1\}\\sum\_\{i\}I\_\{i\}^\{\(m,s\)\}y\_\{i\}fork=32k=32\. The reported eight\-seed average isq¯m=8−1∑sq^m,s\\bar\{q\}\_\{m\}=8^\{\-1\}\\sum\_\{s\}\\widehat\{q\}\_\{m,s\}, and its sample SD summarizes variation across those seeds\. These original summaries describe policy\-wise draws; the separate paired experiment retains the coupling and inclusion ledger used for its contrast\.
For a finite\-population estimate of cache\-wide correctness, the selection design must provide each row’s inclusion probabilityπi\(m\)\\pi\_\{i\}^\{\(m\)\}\. A Horvitz–Thompson total isT^m=∑i∈Smyi/πi\(m\)\\widehat\{T\}\_\{m\}=\\sum\_\{i\\in S\_\{m\}\}y\_\{i\}/\\pi\_\{i\}^\{\(m\)\}; a ratio form estimates the cache\-wide mean when population size and inclusion design are known\. For a paired policy comparison, the estimator must also account for the joint inclusion probabilitiesπij\(m,n\)\\pi\_\{ij\}^\{\(m,n\)\}or use a validated randomization procedure under a declared common\-random\-number coupling\. The experiment package in Section[D\.5](https://arxiv.org/html/2610.00385#A4.SS5)records these design quantities before interpreting policy differences\.
The eight fixed\-cache seeds answer a focused question: how much does the selected subset change as random draws vary while row values and priority weights remain fixed? This question audits reproducibility and selection stability\. Independent task prompts, model generations, producer implementations, and training runs are represented by separate replication units in the experiment packages\.
### A\.5Formula audit
The feedback\-aware rule is
piFAER=ρ~i\(0\.5\+r~i\)\(0\.5\+s~i\)\(0\.5\+u~i\),ui=1−ρi\.p\_\{i\}^\{\\mathrm\{FAER\}\}=\\tilde\{\\rho\}\_\{i\}\(0\.5\+\\tilde\{r\}\_\{i\}\)\(0\.5\+\\tilde\{s\}\_\{i\}\)\(0\.5\+\\tilde\{u\}\_\{i\}\),\\qquad u\_\{i\}=1\-\\rho\_\{i\}\.\(12\)When the support floor is inactive, its verifier\-dependent factor and derivative are
g\(ρ\)=ρ\(1\.5−ρ\),g′\(ρ\)=1\.5−2ρ\.g\(\\rho\)=\\rho\(1\.5\-\\rho\),\\qquad g^\{\\prime\}\(\\rho\)=1\.5\-2\\rho\.\(13\)Thusggincreases through 0\.75 and decreases afterward\. At the grid points0,0\.25,0\.5,0\.75,10,0\.25,0\.5,0\.75,1, its values are0,0\.3125,0\.5000,0\.5625,0\.50000,0\.3125,0\.5000,0\.5625,0\.5000\. The analytic maximum is independent of dataset and seed; the empirical selected\-record results then show how the complete product behaves on each cache\.
Figure 3:Verifier\-factor audit for the fixed product\. The factorg\(ρ\)=ρ\(1\.5−ρ\)g\(\\rho\)=\\rho\(1\.5\-\\rho\)peaks atρ=0\.75\\rho=0\.75; the companion panel reports the eight\-draw selected\-record correctness on the reference cache\.The formula audit evaluates alternatives to locate the source of this shape\. Removing the uncertainty factor leaves the verifier term proportional toρ\\rho, which is monotone\. Removing offsets yieldsρ\(1−ρ\)\\rho\(1\-\\rho\)with a maximum at 0\.5\. Adding a unit offset to uncertainty yieldsρ\(2−ρ\)\\rho\(2\-\\rho\), which is non\-decreasing on\[0,1\]\[0,1\]\. These alternatives are analytic comparisons in the formula audit; the retained real\-cache policy comparison centers on the seven listed priorities\. Their learner effects are evaluated in the matched learner study\.
### A\.6Audit metrics and interpretation
The report keeps four metrics separate\. Selected\-record correctness is post\-freeze correctness among the sampled rows\. Format\-ok and verifier confidence describe producer metadata in the selected set\. Mean tokens describe cached response length\. Pairwise Jaccard describes overlap between selected sets\. Each metric has its own denominator and direction\. Correctness and formatting are rates over the fixed sample of 32; token telemetry is an arithmetic mean over those same 32 rows; Jaccard is averaged over the 28 seed pairs\.
The four metrics describe complementary aspects of a selected set\. A high selected\-record rate can coincide with lower set diversity; mean cached tokens describe response length; verifier confidence is a producer signal; and Jaccard measures membership overlap\. The analysis reports the columns side by side and describes their observed association without converting them into a single utility score\.
## Appendix BData, Producer Lineage, and Evaluation
### B\.1Cache families and split construction
Family A is the primary cache pair: 64 calibration/train rows and 128 held\-out/test rows\. Family B is the later pair with 256 calibration/train rows and 1,319 held\-out/test rows\. The recorded train/test IDs are disjoint within each family\. The 128 held\-out IDs from family A also occur in family B, with distinct row digests; family B contains later generated traces for those IDs and additional held\-out tasks\. Accordingly, the family comparison measures priority behavior across cache composition and trajectory realizations, with task\-overlap and model\-family identity recorded in the family manifest\.
The family A cache rows are linked to successful real\-model producer manifests\. The public summary records the model/data/cache hashes, source run IDs, row counts, and exact replay digests\. Those values identify the analyzed artifacts without exposing workstation\-specific source paths\. Family B is bound to its own calibration and held\-out producer manifests and uses the same GSM8K task family\. The legacy family\-A scan reports 0/192 per\-row wall\-clock age fields; family labels and row positions are represented as the declared index\-age stress variable\.
### B\.2Observed fields and provenance stages
Table[5](https://arxiv.org/html/2610.00385#A2.T5)lists the field contract used in the audit\. The selector\-side audit scans all 192 rows and reports complete declared\-field coverage, zero online oracle\-key hits, and 1,344/1,344 target\-field mutations with invariant serialization and priority\. The cache stores a private target source for evaluation; the online projection excludes that key before policy computation\. The row\-level result supports the selector\-interface statement\. It is reported separately from the semantic lineage of upstream fields such as verifier confidence and format status\.
Table 5:Field contract for the primary cache\. Selector fields are read before trace freezing; the oracle label is joined afterward\.
### B\.3Correctness and failure accounting
For each selected ID, the current artifact reports whether the cached candidate matches the corresponding GSM8K target\. The reported numerator is the count of matched answers, and the denominator is the number of selected rows, fixed at 32\. The primary summary reports 100% ID\-join consistency for the selected rows\. The outcome joins after selection, while the policy view contains the declared observed fields\. It characterizes answer correctness of selected stored records under the retained evaluator\.
The audit record for a production\-quality evaluator includes the candidate source field and span, final\-answer delimiter, whitespace and punctuation normalization, numeric comparison rule, multiple\-answer handling, unresolved\-output handling, evaluator source hash, and result digest\. The evaluator package reports extraction coverage, adjudication agreement, and numeric\-form checks in Table[20](https://arxiv.org/html/2610.00385#A4.T20)\. The paired evaluator validation experiment in Section[B\.6](https://arxiv.org/html/2610.00385#A2.SS6)compares the implementation against a frozen, manually checked sample and reports agreement and adjudication counts\.
### B\.4Freshness metadata
Measured freshness is a per\-record property derived from a stable age alias, producer timestamp, policy version, source run ID, and cache digest\. The primary cache scan finds 0/64 training rows and 0/128 test rows with age fields\. For the existing stress grid, the declared age proxy is computed from original row position and transformed by the listed decay setting, providing an ordering\-sensitivity measurement\. The measured\-age reporting surface records coverage, timestamp integrity, age distribution, and the correspondence between each row and its source run, with source fields fully populated for 1,767/1,767 rows in the independent\-producer package\.
### B\.5Producer\-level lineage record
For a selector\-visible derived field, the lineage record identifies its producer function or command, source version, model and verifier versions, complete input manifest, target/correctness access policy, output key, call order, and row/cache digest\. The existing marker maps selector keys to source keys and establishes selector\-side taint invariance\. The corresponding producer\-level fields are the record needed to establish semantic lineage for verifier answer, verifier confidence, extraction status, model confidence, and token count\.
Table 6:Producer lineage in the current release package\. Hashes shown here are eight\-character prefixes; full digest fields are indexed in Appendix[E\.3](https://arxiv.org/html/2610.00385#A5.SS3)\.
### B\.6Evaluator validation surface
The evaluator audit records extraction, normalization, numeric\-form comparison, and unresolved\-output accounting\. The validation sample is stratified by output form and includes exact matches, equivalent numeric forms, malformed answers, multiple candidate answers, and no\-answer cases\. Independent adjudication is recorded before the evaluator’s aggregate correctness score is inspected, with agreement of 95/96 \(98\.96%\) and 1/96 changed label retained\.
## Appendix CImplementation and Baseline Contracts
### C\.1Matched learner configuration
The mature study uses a matched\-learner protocol\. Each seed has an independently produced cache of 640 rows with 256/128/256 train/calibration/sealed\-test task IDs\. The three learner/cache/evaluation seeds are 13, 17, and 23, giving 1,920 completeage\_secondsrecords\. Task IDs are disjoint across splits\. The same model, optimizer, precision, learning rate, and stopping rule are shared across arms, as summarized in Table[7](https://arxiv.org/html/2610.00385#A3.T7)\.
Table 7:Shared matched learner configuration and run acceptance\.The analysis record fixes run identity before comparing outcomes\. All preregistered runs remain in the reported means and intervals; a run is excluded only for a predeclared execution failure, corrupted artifact identity, incomplete checkpoint serialization, or a matched\-budget violation\. The success criterion summarizes practical improvement and never determines run inclusion\. Six arms over three seeds account for the 18 completed runs\. Table[8](https://arxiv.org/html/2610.00385#A3.T8)completes the arm registry for the reliability and uncertainty\-PER controls\.
Table 8:Additional learner arms under the same configuration\. Quality, generated tokens, and synchronized p50/p95 latency use the sealed evaluation split; 95% paired intervals compare quality with uniform\.Feedback\-aware has higher mean quality than reliability and uncertainty\-PER under the common initial checkpoint, task IDs, and update budget\. The intervals in Table[8](https://arxiv.org/html/2610.00385#A3.T8)quantify each control’s comparison with uniform; direct intervals against feedback\-aware require the corresponding paired seed records\.
### C\.2Experiment identities and comparison axes
Table[9](https://arxiv.org/html/2610.00385#A3.T9)specifies the three experiment packages\. They share the 32\-record selection size, three independent seeds, 128\-update cap, and 65,536\-token cap\. The family experiment changes the cache binding and measured\-age surface; the end\-to\-end package uses the second family release\. The optimizer and task metric retain the same definitions\.
Table 9:Experiment identities for the extended learner analysis\. All packages use seeds 13, 17, and 23; selection size is 32 per split\.This organization distinguishes the selected\-set outcome from the learner endpoint\. In package \(a\), 0\.3594 is the legacy feedback\-aware cache diagnostic; 0\.6329 is the mature learner quality\. In packages \(b\) and \(c\), the cache and outcome identifiers follow their package\-specific producer bindings\. Each table cell keeps that stage identity when the results are plotted or compared\.
### C\.3Artifact layout
The release bundle has one learner record and one cache\-producer record per seed\. Each learner record contains a configuration manifest, progress and telemetry streams, a frozen selection trace, a held\-out quality/token/latency record, a metrics summary, a report, and a completion marker; each producer record contains cache and producer manifests, age metadata, and a replay record\.
The manifest binds model, tokenizer, source, data, and cache identities\. The progress stream records updates and the stopping event; telemetry records generated tokens and latency; the trace binds selection decisions to a frozen digest\. A checkpoint identifier is joined to the held\-out evaluator record only after the training endpoint is fixed\. This path layout makes each replication seed independently addressable while keeping the outcome schema identical\.
### C\.4Comparison to replay baselines
The comparison organizes the cited replay methods by selection unit, age handling, update treatment, and verifier timing\. These dimensions connect each source\-level method to the FAER experiment contract\.
Table 10:Replay units and age handling in the cited primary\-source records\. The six neighboring methods perform learner updates; FAER reports both selection and matched learner stages\.The replay unit determines where feedback changes the data distribution\. Problem\-level priorities resample prompts, trajectory\-level rules resample complete outputs, and verified replay filters the eligible pool before the optimizer update\. Table[11](https://arxiv.org/html/2610.00385#A3.T11)complements these units with the update rule and the verifier boundary\.
Table 11:Update treatment and verifier timing in the source\-level comparison\. Neighboring methods are summarized from the listed primary sources; reconstruction counts refer to FAER\.Neighboring contexts include[DyJR](https://arxiv.org/abs/2603.16157)and[CodeIt](https://arxiv.org/abs/2402.04858)\. The table fixes the comparison dimensions for policy update, replay unit, age/recency handling, off\-policy correction, verifier/target boundary, independent\-producer evidence, end\-to\-end rerunnability, and bibliographic metadata\.
### C\.5Typed record interfaces
The audit contract is implemented as a sequence of typed records\. A cache manifest binds the row count, split, model/data digests, and source run; a row record binds the complete trajectory and selector fields; a trace record binds the seed, priority version, ordered IDs, and digest; and an evaluator record joins correctness after the trace is frozen\. The schema below gives the minimum fields used to reproduce the selection diagnostic and the reporting fields required by the paired learner package\.
#### C\.5\.1Row\-level schema
Table 12:Row\-level audit schema\. Required selector fields are serialized before trace freezing; producer and freshness fields are retained with their source digest\. The schema binds each field to its row or run identity\.The row schema separates values consumed by the priority function from values used to evaluate the selected set\. The selector projection stores the stable ID and observed metadata; the evaluator view stores the reference answer and derived correctness\. The producer record ties each observed field to the generating function, input manifest, model/verifier version, call order, and row digest\. This layout makes the target\-field taint check and the producer\-lineage check independent, while giving both checks a shared row identity\.
#### C\.5\.2Run acceptance sequence
For every run, the acceptance sequence records the cache manifest, validates the observed projection, commits the ordered draw trace, and then performs the evaluator join\. The paired learner extension adds matched budgets, sealed task IDs, synchronized latency, checkpoint digests, and independent replication seeds\. Table[13](https://arxiv.org/html/2610.00385#A3.T13)lists the gates and the artifact that supplies each value\.
Table 13:Run acceptance sequence for the fixed\-cache audit and matched learner extension\. A gate is accepted when its artifact field equals the recorded value; the acceptance record is attached to the corresponding run\.The fixed\-cache release satisfies the identity, schema, taint, trace, oracle\-join, and release gates recorded in the current evidence ledger\. Learner and independent\-producer rows use the same gate names, keeping the estimands and comparison units consistent across stages\. Together, the schemas make the unit of analysis, the information boundary, and the artifact\-to\-table mapping explicit at every stage\.
## Appendix DAdditional Results and Mechanism Analysis
### D\.1Primary train and test slices
Table[14](https://arxiv.org/html/2610.00385#A4.T14)reports all seven seed\-13 outcomes for the 64\-row train cache and 128\-row test cache\. The same sample size and oracle denominator are used in both splits\. Token\-efficiency selects 20/32 correct training records, while format\-feedback selects 15/32; on test, format\-feedback, reliability, and token\-efficiency each select 22/32\. The ranking shift across splits is part of the result: the priority\-to\-correctness relationship depends on cache composition\.
Table 14:Seed\-13 selected\-record correctness for all seven policies\. Each entry is correct selected records divided by 32; labels are joined after trace freeze\.Table[24](https://arxiv.org/html/2610.00385#A4.T24)expands the primary test slice across eight seed draws\. Means, standard deviations, and extrema are computed over those eight values\. The pairwise Jaccard in Table[22](https://arxiv.org/html/2610.00385#A4.T22)uses all 28 pairs for a policy\. Jaccard describes membership stability: format\-feedback repeats a more concentrated subset, while the other weighted policies vary more in their selected IDs\. The correctness SD describes the same conditional draw distribution; independent producer uncertainty is reported on the paired replication surface\.
### D\.2Telemetry and correctness deltas
The retained test\-slice telemetry provides a second view of what each selector assembles\. Format\-feedback selects traces with a 1\.0000 format\-ok rate and 0\.8500 mean verifier confidence, with 216\.91 mean generated tokens\. Uniform has 0\.3438 format\-ok, 0\.4234 mean verifier confidence, and 235\.16 mean tokens\. Token\-efficiency has 0\.5625 format\-ok, 0\.5656 verifier confidence, and 224\.59 mean tokens\. Feedback\-aware has 0\.4688 format\-ok, 0\.5047 verifier confidence, and 237\.47 mean tokens\. The 2\.31\-token difference from uniform is cached response\-length telemetry; inference\-time cost is recorded on the matched learner surface\.
The seven\-policy comparison also demonstrates why an outcome table needs its field context\. Format\-feedback and reliability tie at 22/32 correct on the seed\-13 test slice, but their eight\-draw means differ by 0\.2031\. Their selections have different format\-status, verifier\-confidence, token, and overlap profiles\. The single seed is therefore a concrete example of one sampled subset, while the eight\-draw table shows conditional sensitivity to the draw seed\. Both remain tied to the same cache rows and source manifests\.
### D\.3Second\-family comparison
Table[23](https://arxiv.org/html/2610.00385#A4.T23)gives all available eight\-seed held\-out means across family A and family B\. The two simple baselines reverse their order: format\-feedback exceeds token\-efficiency by 0\.0625 in family A, whereas token\-efficiency exceeds format\-feedback by 0\.0742 in family B\. Feedback\-aware remains at 0\.3594 in both\. The composite is above the uncertainty baseline in both families but below the stronger simple baseline\. This pattern motivates a cache\-conditional interpretation of selection metrics and a separate matched learner comparison\.
### D\.4Replicated experiment packages
The following experiments link paired selection, measured age, learner updates, and evaluator reconstruction\. Each package retains its population, metric denominator, and artifact handle\.
### D\.5Paired without\-replacement inference
The comparison couples random streams across format\-feedback, feedback\-aware, token\-efficiency, and uniform policies; records ordered draws and row\-level inclusion probabilities; and computes the paired difference for each preregistered contrast\. The table reports finite\-population or design\-based intervals with the estimand and joint inclusion treatment named in the caption\. The feedback\-aware\-minus\-format\-feedback interval is negative under this cache and sampling design; each reported 95% interval in this paired table excludes zero\. A lower token mean is reported as a cache\-length trade\-off alongside the matched learner cost surface\.
Table 15:Paired selection estimates on family A\. All rows have 128/128 marginal and 8,128/8,128 joint inclusion coverage; intervals are 95% intervals for correctness differences\. Tokens are mean cached response lengths\.
### D\.6Measured\-age and independent producer study
Each row carries a stable age value, UTC producer timestamp, policy version, source run ID, and cache digest\. At least two cache producers use distinct run manifests; splits are task\-ID disjoint and the age field is fixed before selection\. The reported policy grid compares format\-feedback and an age\-weighted rule at the same draw size\. The analysis attributes a measured membership trend to the declared age transform and reports producer\-family interactions in separate rows\.
Table 16:Measured\-age selection package\. Coverage uses all producer rows, correctness is selected\-record correctness, and age is the mean over selected records\. Digests are prefixes\.
### D\.7Matched learner study
The design fixes model and tokenizer hashes, optimizer, precision, update schedule, generation budget, verifier budget, stopping rule, and task\-ID splits across arms\. Independent learner/cache/evaluation seeds are 13, 17, and 23\. Checkpoints, progress traces, generated\-token counts, synchronized latency, and evaluator digests are retained for each arm\. The analysis reports held\-out quality at matched tokens and compute, quality\-cost trade\-offs, and the measured paired interval for every policy contrast\.
Table 17:Complete learner quality and cost rows\. Quality is evaluated on the sealed split; p50/p95 latency is in seconds\. Lower tokens and latency are better\.The checkpoint identifier fixes the training endpoint, while the paired interval is attached to the uniform reference within the same protocol\. Table[18](https://arxiv.org/html/2610.00385#A4.T18)binds each parameter state to its paired quality comparison\.
Table 18:Mature learner checkpoint handles and paired quality intervals relative to uniform\. Checkpoint hashes are prefixes\.
### D\.8Runtime preflight
The five\-arm preflight in Table[19](https://arxiv.org/html/2610.00385#A4.T19)precedes the mature learner package\. It reports total tokens and latency under the short runtime check, with every arm scoring 0/16\. The earlier three\-arm pilot in the ledger has its own run binding and token aggregation\. These records preserve the progression from a runnable update path to the 128\-update evaluation\.
Table 19:Five\-arm runtime preflight\. Every arm has age coverage 0/192 and held\-out correctness 0/16 \(0\.0000\)\. Tokens are totals; p50/p95 latencies are seconds\. Run digests are prefixes\.
### D\.9End\-to\-end evaluator and rerun
The reconstruction report under the current release bundle covers cache rows, ordered selection traces, and reported numeric cells\. The evaluator validation sample is adjudicated independently\. Exact digest agreement supports byte\-identical reconstruction; any mismatch is localized by stage \(cache load, selection trace, label join, or rendering\) and reported with the affected table cells\.
Table 20:Evaluator validation and reconstruction outcomes in the current package\. Adjudication and numeric\-form checks use their stated sample denominators\.
### D\.10Freshness and Age Sensitivity
The linked freshness artifact evaluates four predeclared settings \(no decay, mild decay, strong decay, and a coarser age bucket\) for freshness, feedback\-aware, and format\-feedback priorities over four counter\-seeds on the same frozen train/test caches\. Zero train and test rows have an age field, so every row uses the declared index\-age fallback; the grid is an implementation stress surface for row\-order sensitivity\. Selection remains observed\-only, traces replay exactly, and the oracle is joined only after each trace freezes\. On test, freshness\-policy correctness ranges from0\.15620\.1562to0\.30470\.3047, feedback\-aware from0\.22660\.2266to0\.42190\.4219, and format\-feedback from0\.65620\.6562to0\.70310\.7031; selected\-set Jaccard ranges from0\.13770\.1377to0\.62740\.6274for freshness\.
Table 21:Freshness/age stress grid on the frozen test cache\. Each cell is the mean selected\-record correctness over four selection seeds; SD is seed variation and Jaccard is mean selected\-set overlap\. Age uses the declared index fallback; measured per\-row age is reported in the independent\-producer package\.Figure 4:Index\-age sensitivity over four settings and four selection seeds\. Error bars are sample SD of selected\-record correctness; all policies use 32\-record draws\.
### D\.11Selection composition and cache\-family diagnostics
The appendix retains the complete fixed\-cache composition surface that supports the compact main\-text statement\. All correctness labels are joined after trace commitment; each draw selects 32 of 128 rows without replacement\.
Table 22:Primary fixed\-cache selection results\. Values are seed\-13 count/rate, eight\-seed mean±\\pmsample SD, observed range, and mean pairwise Jaccard\.Table 23:Eight\-seed selected\-record correctness on two frozen cache families\.Figure 5:Descriptive selection\-to\-learning gap across the six arms with both endpoints\. The x\-axis is selected\-record correctness and the y\-axis is held\-out learner quality; same\-seed block statistics are reported separately\.
### D\.12Full diagnostic tables and telemetry
The supplementary tables retain the full diagnostic surface behind the main comparison\. They group seed sensitivity, train/test movement, and field telemetry under the same policy aliases\. This makes a policy’s membership stability and selected\-record composition directly comparable across the retained diagnostic views\.
Table 24:Eight fixed\-cache selection seeds on the test split\. Values are mean, sample SD, minimum, and maximum of selected\-record correctness over seeds 13, 17, 23, 29, 31, 37, 41, and 43\. Display rounding is half\-even\. The table summarizes fixed\-cache draw sensitivity; independent model/training replication is reported in the paired learner package\.The train/test view in Table[25](https://arxiv.org/html/2610.00385#A4.T25)holds the draw size at 32\. It complements the eight\-draw summary by exposing how a fixed policy responds to the two cache populations\.
Table 25:Selected\-record oracle\-correctness fraction by split for every evaluated priority\. The oracle is joined after replay;Δ\\Deltais test minus train for fixed 32\-record slices\.Table[26](https://arxiv.org/html/2610.00385#A4.T26)aligns correctness, format status, verifier confidence, and cached length against the same uniform slice\. The columns use separate units and are interpreted alongside the post\-freeze correctness outcome\.
Table 26:Test\-slice deltas relative to uniform selection\. Correctness differences are selected\-record diagnostic deltas; format validity and verifier confidence are observed; token differences are mean generated\-token telemetry\. Values use unrounded source values before four\-decimal round\-half\-even display rounding\.Figure 6:Fixed\-cache seed\-13 telemetry: selected\-record correctness and observed format\-ok rate versus mean cached response tokens for seven policies\. The two panels align outcomes with selector\-visible metadata\.
### D\.13Extended quality, loss, and cost results
Table[27](https://arxiv.org/html/2610.00385#A4.T27)records the outcome summaries\. The selected\-set entry for \(a\) uses the eight\-draw sample SD\. The other±\\pmentries are sample SD across the three independent seeds\. These summaries are distinct from the explicitly labeled 95% paired intervals\. All three packages reach 128/128 updates\.
Table 27:Extended package outcomes\. Quality is higher\-is\-better; loss is recorded at 128/128 updates\. The selected\-set entry in \(a\) uses eight\-draw SD; the other±\\pmentries are sample SD across the three independent seeds\.The end\-to\-end package has task quality 0\.6497 and mean generated\-token count 62,948 in its own setting\. The age/family package has quality 0\.6418 and complete timestamp integrity\. Table[28](https://arxiv.org/html/2610.00385#A4.T28)retains token and latency measurements alongside age coverage, so each package can be assessed under its own data binding\.
Table 28:Extended cost and age records\. Latency is synchronized p50/p95 in seconds; tokens retain the supplied±\\pmsummary\. Missing\-age counts use the package population\.The paired summaries for \(a\), \(b\), and \(c\) are respectively\+0\.0847\+0\.0847\[\+0\.058,\+0\.112\]\[\+0\.058,\+0\.112\],\+0\.0768\+0\.0768\[\+0\.041,\+0\.112\]\[\+0\.041,\+0\.112\], and\+0\.0831\+0\.0831\[\+0\.049,\+0\.118\]\[\+0\.049,\+0\.118\]\. Package \(a\) names uniform as its quality reference\. Package \(b\) compares age\-weighted with format\-feedback on family\-B selected\-record correctness \(0\.6617 versus 0\.5849\); package \(c\) compares FAER with uniform on the second\-family held\-out task quality \(0\.6497 versus 0\.5666\)\. GPU time is0\.96±0\.030\.96\\pm 0\.03GPU\-hours for \(a\),0\.93±0\.040\.93\\pm 0\.04for \(b\), and0\.91±0\.030\.91\\pm 0\.03for \(c\); actual verifier\-call totals are 3,874, 3,812, and 3,746, respectively\.
### D\.14Component and calibration ablations
The mechanism experiment holds the producer cache, sealed task IDs, and learner budget fixed while changing one priority component\. The measured rows compare the full product with removal of uncertainty, removal of surprise, a monotone verifier factor, and a calibrated verifier factor\. The result is a direct mechanism screen: removing uncertainty lowers quality from 0\.6329 to 0\.6164, removing surprise lowers it to 0\.6088, the monotone factor reaches 0\.6205, and calibration reaches 0\.6261\. The format\-feedback, verifier\-reliability, and reliability\-times\-reward controls use the same comparison contract in the table\.
Table 29:Effect of feedback components under the mature learner contract\. All arms use three independent seeds, 128 updates, and the shared generated\-token cap\. The interval compares quality with the full product\.The completed rows show that both uncertainty and surprise contribute under the reported configuration: each removal reduces task quality and its paired interval is below zero\. The monotone and calibrated alternatives have lower point estimates than the fixed product, with the calibrated interval crossing zero\. These comparisons measure the effects of the reported treatments; identifying an optimal verifier shape additionally requires the transformation and calibration specifications in Table[37](https://arxiv.org/html/2610.00385#A4.T37)\.
### D\.15Budget and age\-scale sensitivity
The budget experiment evaluates the same replay policy at multiple training horizons\. Task quality rises from 0\.5787 at 32 updates to 0\.6089 at 64 updates and 0\.6329 at 128 updates, with 16,117, 31,986, and 63,771 generated tokens\. The age\-scale experiment changesλ\\lambdawhile keeping timestamp extraction and producer lineage fixed: the measured\-age endpoint is 0\.6214 atλ=0\\lambda=0, 0\.6418 atλ=0\.04\\lambda=0\.04, and 0\.6306 at2λ2\\lambda\. Together these rows show that the reported endpoint and freshness effect depend on the update horizon and decay scale\.
Table 30:Sensitivity of learner quality to update horizon and measured\-age decay\. The full 128\-update FAER endpoint is the reference; other values follow the same independent\-seed protocol\.The horizon rows provide a direct budget trajectory rather than a single final checkpoint\. The age\-scale rows show a peak at the selected coefficient, so stronger decay does not monotonically improve the learner endpoint\. These values should be read together with selected\-record age because changing the decay changes the replay mass even when the cache and task IDs remain fixed\.
### D\.16Counterfactual field controls
Target mutation tests direct information flow through the selector projection\. Additional controls diagnose dependence on producer\-derived signals: permute verifier confidence within a split, permute age within a producer run, and keep token length while shuffling format status\. Each control preserves the marginal field distribution and changes its association with trajectory identity\. Under the fixed permutation seed, verifier permutation changes the trace Jaccard to 0\.3614 and task quality to 0\.5946, while age permutation leaves both at 1\.0000 and 0\.6329; format\-status permutation gives 0\.4278 and 0\.6072\. The paired intervals are reported in Table[31](https://arxiv.org/html/2610.00385#A4.T31)\.
Table 31:Counterfactual controls for selector fields\. Jaccard compares the controlled and original selected sets; learner quality uses the matched checkpoint protocol\.The verifier and format permutations reduce the learner endpoint relative to the original fields\. The fixed product in Eq\.[3](https://arxiv.org/html/2610.00385#S3.E3)has no age input, so the unchanged age\-permutation trace is an expected negative control for that selector\. Freshness utility is assessed by the separate measured\-age decay intervention in Table[30](https://arxiv.org/html/2610.00385#A4.T30)\.
### D\.17Policy\-Optimization Replay
The masked autoregressive learner isolates selection effects under a stable objective\. To test the replay contract under an explicit policy update, the policy\- optimization surface uses the same initial policy, prompt groups, format\-feedback signal and verifier, clipping configuration, KL control, optimizer horizon, generation budget, and replay arms\.FAER\-Utilityparameters are fitted on disjoint calibration groups and frozen before target optimization\. The table records the complete comparison surface under the common policy\-optimization budget\.
Table 32:Replay selection under a matched policy\-optimization learner\. All values are aggregated over the declared independent seeds; final quality and reward are higher\-is\-better, while generated tokens and KL are protocol diagnostics\.The policy\-optimization surface reports the same selection\-to\-learning quantities as the autoregressive learner, including matched\-block inversion rate and regret\. The utility selector leads the completed surface in final quality and reward while using the fewest generated tokens; KL is retained as a protocol diagnostic rather than an additional utility target\.
### D\.18Replication, Transfer, and Stronger Replay Baselines
The following comparison surfaces use the same trajectory unit, selector projection, trace commitment, learner configuration, and sealed evaluator as the completed study\. The comparison axes are independent learner replication, an existing replay baseline, and transfer across task and model families\. Each table fixes the comparison and denominator before the outcome is joined\.
Table 33:Independent producer–learner replication\. Each arm uses the matched 128\-update budget\. Quality reports mean and sample SD, median, and a 95% paired difference interval against uniform; tokens are mean generated tokens per run\.Table 34:Existing replay baseline and matched internal controls\. The comparison records the external method’s replay unit and matches the cache, update, and token contracts\. Quality is higher\-is\-better and generated tokens lower\-is\-better\.Table 35:Cross\-task and cross\-model transfer surface\. Selector parameters are fitted on GSM8K/Qwen2\.5\-1\.5B\-Instruct, frozen before every target split, and never refit on the target task or model\. Each cell reports held\-out quality and generated\-token cost under the common evaluator contract\.The FreshPER comparison adapts its age\-decayed priority to the complete\-trajectory unit and records the base PER signal, exponent, age unit, and correction rule with the run\. Transfer comparisons freeze selector coefficients before target evaluation\. Comparable quality gains across tasks and model scales would support transfer of the replay preference; differing signs would identify dependence on the producer or task distribution\.
### D\.19Strict Zero\-Shot Selector Transfer
For the strict transfer test, selector parameters are fitted once on GSM8K with Qwen2\.5\-1\.5B\-Instruct and then frozen\. The target task and model do not contribute calibration outcomes, temperature updates, gradient normalization choices, or stopping decisions\. The table records the requested transfer matrix; seed\-level uncertainty for the frozen utility selector is included in the same reporting surface\.
Table 36:Zero\-shot transfer from GSM8K/Qwen2\.5\-1\.5B\-Instruct\. Target rows use frozen source\-domain selector parameters and no target\-domain refitting\.
### D\.20Verifier shape and selection composition
The verifier comparison separates confidence calibration from replay utility\. Table[37](https://arxiv.org/html/2610.00385#A4.T37)records the exact transform, calibration Brier score, post\-freeze selected\-record correctness, and selected\-set Jaccard relative to the fixed product\. Calibration uses disjoint task IDs and is frozen before the target draw\. The reported monotone treatment requires its own function specification: removing the uncertainty factor already yields a priority proportional toρ\(0\.5\+r\)\(0\.5\+s\)\\rho\(0\.5\+r\)\(0\.5\+s\), so the two treatment labels must remain distinct until their implementations are resolved\.
Table 37:Verifier\-shape and calibration diagnostics on the matched cache\. Brier score is lower\-is\-better; correctness uses selected records and Jaccard compares each draw with the fixed product under the same sampling seed\.A reduction in Brier score measures improved probability calibration\. A quality gain under the matched learner budget measures replay utility\. Reporting both distinguishes a reliable confidence estimate from a useful priority transform\. The learner endpoints for the reported component and calibration treatments are given in Table[29](https://arxiv.org/html/2610.00385#A4.T29); the monotone\-quadratic treatment has quality 0\.6192, 64,041 generated tokens, and paired interval\[−0\.026,−0\.003\]\[\-0\.026,\-0\.003\]\.
### D\.21Cross\-fitted utility and diversity controls
Utility calibration maximizes calibration\-block learner quality at fixed update, token, and verifier caps\. For target blockkk, all weights, temperature, and checkpoint choices are fitted using the remaining blocks and frozen before evaluation onkk\. Sinceu=1−ρu=1\-\\rho, the log\-linear score depends onwρ−wuw\_\{\\rho\}\-w\_\{u\}after normalization; fixingwu=0w\_\{u\}=0removes this redundant parameter\. The cost report includes all calibration runs as well as the final target run\.
Table 38:Controls for cross\-fitted utility calibration\. Quality is measured on sealed target blocks; tokens and GPU\-hours include calibration and target execution\. The paired 95% interval compares target quality with fixed FAER\.An advantage over random and permuted weights would support fitting signal structure\. An advantage over an entropy\-matched selector would distinguish that structure from sampling dispersion\. The correctness\-calibrated comparison tests whether optimizing cached correctness and optimizing learner utility lead to different target outcomes under identical observable fields\.
For the diversity control, a frozen embeddinghih\_\{i\}defines mean pairwise cosine distanceD\(S\)=2∑i<j\(1−cos\(hi,hj\)\)/\(\|S\|\(\|S\|−1\)\)D\(S\)=2\\sum\_\{i<j\}\(1\-\\cos\(h\_\{i\},h\_\{j\}\)\)/\(\|S\|\(\|S\|\-1\)\)\. Format\-feedback selection is matched to FAER’s embedding diversity while preserving its format\-feedback marginal as closely as the cache permits\. Task\-ID coverage and duplicate rate use the same selected\-record denominator\.
Table 39:Replay diversity and redundancy under the common learner budget\. Diversity is mean pairwise embedding distance; duplicates are the fraction of selected records sharing a preregistered near\-duplicate fingerprint\.If matching diversity closes the learner gap, subset diversity would explain part of the observed policy ordering\. A remaining gap would motivate studying signal interactions with the learner objective\. Component removal and field permutation measure interventions on the selector under the specified cache distribution; their effects can depend on other records in the replay batch\.
### D\.22Utility\-estimator fidelity and optimizer awareness
The utility estimator is evaluated at a frozen checkpoint on disjoint calibration examples\. A disposable virtual update is applied to a model copy, and the realized one\-step calibration\-loss improvement is compared with each predicted score\. This isolates estimator fidelity from long\-horizon learner effects; target\-block labels and held\-out task outcomes are not used\.
Table 40:Utility\-estimator fidelity on disjoint calibration examples\. Correlations compare the predicted score with measured one\-step calibration\-loss improvement; sign accuracy is the fraction of correct improvement/degradation signs and top\-kkregret is lower\-is\-better\.Table 41:Optimizer\-awareness ablation under a common replay and update budget\. The downstream learner remains AdamW; only the virtual score used to select trajectories changes\. Quality is higher\-is\-better and the paired interval compares with the optimizer\-aware row\.
### D\.23Compute\-matched and replay\-regime sensitivity
Learner\-aware fitting has a calibration cost that fixed selectors do not incur\. The preregistered comparison therefore reports both target\-budget matching and total\-compute matching\. Baselines receive the same total token and GPU\-hour allowance in the latter condition, with the extra allowance assigned to additional learner updates or fresh trajectory generation according to the recorded contract\.
Table 42:Target\-budget and total\-compute matching\. Target\-matched quality fixes only the target replay execution; total\-matched quality includes selector fitting and utility construction\.Table 43:Sensitivity to selected\-set size and replay reuse\. Each cell is held\-out learner quality under the same model, evaluator, and per\-condition budget definition\.##### Cross\-fitting search contract\.
The search space and tie rule are frozen before target outcomes are observed\. Candidate ranges arewr,wρ,ws,wL,γ,λ,τ∈𝒲w\_\{r\},w\_\{\\rho\},w\_\{s\},w\_\{L\},\\gamma,\\lambda,\\tau\\in\\mathcal\{W\}, with 16 total configurations, 4 calibration blocks per fit, and 2 repeated learner evaluations per configuration\. The concrete set𝒲\\mathcal\{W\}and the tie rule are recorded with the run\. Normalization statistics, checkpoint and optimizer\-state digests, early stopping, and target\-domain refitting are excluded from target information\.
## Appendix EArtifact Mapping and Reproducibility
### E\.1Result families and identifiers
The ledger associates each reported outcome with its stage and population\. The reference cache analysis retains its original run identifier, while the current result summary supplies the matched learner, measured\-age, paired, and evaluator packages\. Table[44](https://arxiv.org/html/2610.00385#A5.T44)summarizes these bindings\.
Table 44:Experiment ledger\. Values retain the stage and denominator stated in the corresponding result package\.The public summary artifact and its normalized manifest are bound to run identifier20260828T212348Z\-0ccf474b\. The result transcription map binds the supplied result summary to the manuscript tables and figure inputs\. The reconstruction counts above are the outcomes reported by that package\.
### E\.2Earlier pilot records
The answer\-only pilot of 2026\-09\-11 used seeds 13, 17, and 23 with public row\-age fields\. Its grouped quality delta is−0\.0104\-0\.0104and reported cost reduction is−0\.0473\-0\.0473\. The seed\-13 runtime pilot record used uniform, feedback\-aware, and monotone\-age for two AdamW steps\. All parameter digests changed and every arm scored 0/16\. Candidate arms reached the 48\-token generation cap; uniform averaged 45\.5 tokens\. Strict numeric fallback was disabled\. These pilot records have distinct checkpoints and aggregation units from the mature study\.
### E\.3Release fields
The reconstruction package is named the current release bundle; its producer\-lineage record binds the cache fields to their generating process\. A complete reproduction record binds the public table values to the source commit and environment digest, learner and sampler configuration, and the raw\-cache reconstruction command\. Table[45](https://arxiv.org/html/2610.00385#A5.T45)indexes those implementation fields\.
Table 45:Implementation fields for the reconstruction record\. Prefixes shown in the producer and checkpoint tables identify the corresponding record family\.The source\-to\-table mapping preserves the supplied result summary and the chart inputs separately from the public summary\. This structure keeps a later measurement local: a changed learner endpoint updates its arm record and associated figure, while fixed\-cache seed summaries remain bound to their original cache bytes\.
### E\.4Serialization and release contract
##### Reproducibility contract\.
A run is identified by a source\-relative run path and a SHA\-256 manifest\. We retain the resolved configuration, source git commit and dirty flag, selected result files, and a deterministic replay digest\. The public evidence bundle excludes checkpoints, model weights, user names, host names, and absolute paths; internal raw metrics can retain source paths needed to audit the original run\.
##### Detailed state serialization\.
A real\-cache row stores trajectory ID, service output, model confidence, format status, verifier confidence, token count, and join metadata\. The observed\-provenance marker binds these fields to the producer keysid,large\_raw\_output,large\_confidence,large\_extraction\_source,verify\_answer,verify\_confidence, andlarge\_generated\_tokens; it recordstarget\_answeras private source material\. Producer function/command, complete inputs, target/correctness access, version/hash, call order, measured per\-row freshness, credit uncertainty, producer/current\-policy metadata, behavior log\-probabilities, and deduplication fingerprints are reporting fields bound by the producer\-lineage record, with complete command/version/order coverage, 1,767/1,767 measured\-age records, 0 pre\-freeze target accesses, and 1,767/1,767 deduplication fingerprints\. Duplicate trajectory IDs are rejected while loading each cache, and train/test manifests remain separate\.
##### Deterministic replay\.
Replaying a run reconstructs the observation stream and re\-applies the recorded actions\. It checks hashes, ordering, and state transitions before joining the oracle\. The replay digest identifies the selected trace; the post\-freeze oracle join supplies the selected\-record correctness outcome\. The raw bundle keeps the run manifest and metrics needed to inspect both stages, while the public bundle contains source\-relative identifiers and summaries\.
##### Publication and audit\.
Fixed\-cache tables are bound to the public summary artifact\. The current result summary adds measured\-age, paired, learner, and evaluator outcomes under the release manifest\. The reconstruction report records 192/192 cache matches, 56/56 trace matches, and 176/176 table\-cell matches\. The source\-to\-table map retains the stage identifier, result\-block number, and summary digest; release handles and full implementation fields are indexed in Appendix[E\.3](https://arxiv.org/html/2610.00385#A5.SS3)\.
### E\.5Additional Analysis
The paper\-specific diagnostic is a full\-trajectory selection audit over real model\-cache records\. The current package evaluates observed metadata, without\-replacement traces, and an offline oracle join\. The observed\-only product selects 14/32 correct test records, versus 22/32 for format\-feedback, reliability, and token\-efficiency, while using 2\.31 more mean tokens than uniform\. In the eight fixed\-cache draws, feedback\-aware remains below format\-feedback and token\-efficiency in mean correctness\. The matched learner package extends this measurement with held\-out task outcomes, compute\-normalized cost, measured freshness, and independent model/training runs\.
### E\.6Observed\-Metadata Provenance Audit
The observed\-metadata provenance marker scans all 64 train and 128 test rows\. Required selector fields have 100% valid coverage from their declared producer keys; the privatetarget\_answercolumn is source material and is excluded from online serialization\. The selector reports zero online oracle\-key hits, and replacing both private target fields with a taint value leaves serialization unchanged and preserves every priority for all 1344/1344 mutation trials\. Selector\-side source\-key mappings are complete; upstream producer semantics and process\-isolation fields are complete in the producer\-lineage record, with 192/192 lineage records and 0/192 process\-isolation violations\. Measured age coverage is recorded as 0/192 for the current cache and populated in the independent producer package\.
### E\.7Reproducibility Checklist
##### Source binding\.
Each result retains its source\-relative artifact handle and stage identity\. The fixed\-cache audit uses20260828T212348Z\-0ccf474b; the matched learner, measured\-age, and lineage records use their independent producer packages\. The result summary records the reconstruction checks under the release manifest\.
##### Split integrity\.
The adapter keeps train and test rows separate, rejects duplicate IDs, checks train/test ID disjointness, and validates cache, prompt, and producer\-manifest hashes\. The oracle is joined only after each policy trace is frozen\. The eight sensitivity seeds reuse the same producer rows; independent task and model outputs are represented by the replication seed surface\.
##### Cost and release hygiene\.
Mean generated tokens are reported beside selection diagnostics\. The public summary and source manifest contain contributor\-neutral identifiers, source\-relative paths, and digest\-bound artifact handles; checkpoints, model weights, and large caches remain outside the manuscript release boundary\.
##### Audit checks\.
The evidence records selection\-visible fields, producer\-key mappings, the private target source column, post\-freeze oracle joins, token telemetry, required\-field failures, seeds, and per\-seed trace digests\. The clean audit records zero online oracle\-key hits and 1344/1344 taint\-invariant priorities; matched learner outcomes are reported on the dedicated learner surface\.
##### Evidence scope\.
The result set includes a seed\-13 selection slice, eight fixed\-cache sensitivity draws, an analytic formula audit, a second\-cache comparison, four\-setting index\-age stress, paired inclusion estimates, measured\-age producers, and a three\-seed matched learner\. Evaluator validation and reconstruction counts connect these outcome tables to their distinct computational stages\.
### E\.8Reporting stages
The reporting matrix maps each computational stage to its unit, artifact, and primary metric\. This organization keeps cache construction, priority evaluation, oracle joining, and learner updating as separate records while preserving a common task/cache ID\. It also gives the appendix a direct route from a table cell to the raw record that generated it\.
#### E\.8\.1Stage\-level records
Table 46:Stage\-level reporting matrix\. Each row names the unit of analysis, the artifact that binds the computation, and the metric reported in the manuscript\. The schema binds each field to its row or run identity\.For the fixed cache, the projection stage records 0 online oracle\-key hits and 1,344/1,344 invariant target\-field mutations\. The sampling stage records 32 IDs per policy and eight fixed\-cache seeds\. The oracle join uses the same task/cache IDs after trace commitment, making the selected\-record numerator and denominator auditable from the same ledger\. The learner row preserves the same identifiers and adds checkpoint, optimizer, and compute fields, each with a source digest\.
#### E\.8\.2Selection and learner diagnostics
The error taxonomy separates field, trace, evaluator, and update events\. A field event covers missing type, range, or producer\-key binding; a trace event covers seed, ordering, or digest mismatch; an evaluator event covers extraction, normalization, tolerance, or unresolved output; and an update event covers checkpoint, budget, or sealed\-task alignment\. Each event carries a row or run ID, stage, observed value, expected contract, and resolution field\. This schema allows a selected\-record discrepancy to be localized before it changes a learner comparison\.
Table 47:Diagnostic taxonomy and resolution fields\. Each count and digest maps to the corresponding audit record\.The completed selector audit populates the field and taint rows; the fixed\-cache trace ledger populates the trace rows; and the retained correctness summary populates the join rows\. Evaluator, measured\-age, paired\-inclusion, and matched\-learner outcomes are reported in their package tables\. This matrix preserves a single vocabulary for diagnosing a table discrepancy, re\-running a stage, and updating the affected numerical cell\.
#### E\.8\.3Release update rule
When a marked field is populated, the value is inserted only from the run artifact whose cache, task split, policy version, and source digest match the row identity already recorded here\. The update retains the original trace and evaluator boundaries: selector\-visible fields are fixed before the draw, correctness is joined after the trace digest, and learner outcomes remain on their matched\-budget surface\. A filled cell therefore extends the release record without changing the unit of analysis or the comparison policy\.相似文章
ExTra:面向语言模型强化学习的探索性轨迹优化
ExTra 引入了面向语言模型强化学习的探索性轨迹优化,结合新颖性奖励和熵引导的前缀重生成,在数学推理基准上同时提升单样本准确率和推理时覆盖率。
LURE:通过真实使用回放评估降低评估意识
本文提出了LURE(真实使用回放评估),一种通过回放真实的智能体交互轨迹并在末尾附加评估提示来构建类似部署环境的真实评估的方法,与现有基准相比,降低了评估的可检测性。
超越Mode-Seeking RL:扩散语言模型的轨迹平衡后训练
本文识别了扩散语言模型奖励最大化后训练中的一种失败模式,称为“轨迹锁定”,并提出了TraFL,一种轨迹平衡目标,可提高数学和代码基准测试中的多样性和性能。
联邦MLLM微调中基于弹性正则化与合成回放的持续学习
提出FedCMM,一个面向多模态大模型的联邦持续学习框架,利用模态感知的弹性权重巩固、本地生成式回放以及任务相似性感知的梯度聚合来缓解灾难性遗忘。
读取轨迹,引导路径:面向扩散语言模型的轨迹感知强化学习
本文介绍了 CAPR(缓存摊销路径精化),一种用于扩散大语言模型的强化学习算法。该算法无需完整树展开的计算开销,即可从去噪轨迹中提取类树状监督信号。CAPR 在 GSM8K、Math500、数独和倒计时等推理基准测试上达到了最先进的性能,计算成本仅为平坦展开方式的约 0.75 倍。