EvoSteer:基于参考锚定信用分配的在线自进化图编排

arXiv cs.AI 论文

摘要

EvoSteer 为基于 LLM 的多智能体系统提出了一种在线自进化图编排范式,利用参考锚定的流匹配损失(AnchorTB)以及经过验证的技能准入机制,在执行过程中修复失败的步骤,在涵盖问答、数学、代码与决策等十二个数据集上均优于各类基线方法。

arXiv:2609.38661v1 Announce Type: new Abstract: In recent years, LLM-based multi-agent systems have been widely applied to orchestrate tool-using agents into executable communication graphs. However, existing self-evolving orchestration still faces key challenges, including post-hoc evolution that revises the team only after the trajectory ends, credit diffusion that gives every action the same terminal advantage under confounded baselines, and skill admission that is uncalibrated and never retired. To address these challenges, we propose EvoSteer, a new paradigm of Online Self-Evolving Graph Orchestration -- the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate. To support this paradigm, we introduce Anchored Trajectory Balance (AnchorTB), a regression-style flow-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference. Built on the learned flow, we further propose Validated Skill Admission, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget. Moreover, AnchorTB combines measured task-level reference reward statistics with prefix-dependent corrections. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making. Our code is available at https://github.com/beita6969/evosteer.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:41

# Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment
Source: [https://arxiv.org/html/2609.38661](https://arxiv.org/html/2609.38661)
Hanwen ZhangAffiliation:The Chinese University of Hong Kong, Shenzhen, China Dalian University of Technology, ChinaQiang HuangZijia WangPengfei GuoAffiliation:Fudan University, China University of Oxford, UK North China Electric Power University, ChinaYuchen ZhangAffiliation:The University of Texas Health Science Center at Houston, USAJionghao ZhuXiaoying Tang

###### Abstract

In recent years, LLM\-based multi\-agent systems have been widely applied to orchestrate tool\-using agents into executable communication graphs\. However, existing self\-evolving orchestration still faces key challenges, including*post\-hoc evolution*that revises the team only after the trajectory ends,*credit diffusion*that gives every action the same terminal advantage under confounded baselines, and*skill admission*that is uncalibrated and never retired\. To address these challenges, we propose EvoSteer, a new paradigm of Online Self\-Evolving Graph Orchestration—the orchestrator builds a running team and repairs its plausible but failing steps from execution features and a learned value estimate\. To support this paradigm, we introduce*Anchored Trajectory Balance*\(AnchorTB\), a regression\-style flow\-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference\. Built on the learned flow, we further propose*Validated Skill Admission*, in which a candidate skill is tried before promotion and promoted only if paired evidence passes a sequential test under a shared nominal testing budget\. Moreover, AnchorTB combines measured task\-level reference reward statistics with prefix\-dependent corrections\. Experimental results on twelve datasets show that EvoSteer significantly outperforms baselines across question answering, mathematical reasoning, code generation, and interactive decision making\. Our code is available at[https://github\.com/beita6969/evosteer](https://github.com/beita6969/evosteer)\.

## 1Introduction

In recent years, a variety of powerful LLM\-based multi\-agent systems have been applied to solve a wide range of complex tasks\([Yao et al\., 2022b](https://arxiv.org/html/2609.38661#bib.bib57);[Hong et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib13);[Zhuge et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib69)\), gradually moving beyond a single model call toward teams of tool\-using agents that complete tasks end\-to\-end\.

Figure 1:Outputs that look plausible can still be headed for failure\. EvoSteer reads a low continuation estimate off the execution record and edits the running team, here routing the checker’s report back to the solver; dashed: the same team left unedited\.In this process, graph orchestration has become a key bridge from task goals to reproducible execution: by assigning each node a role and bound skills and each edge a communication protocol on an agent communication graph \(Fig\.[1](https://arxiv.org/html/2609.38661#S1.F1), left\), agents can complete complex tasks with improved controllability, compositionality, and reusability\([Zhang et al\., 2025b](https://arxiv.org/html/2609.38661#bib.bib62);[Zhang et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib63)\)\. However, in practice, orchestration still evolves only between trajectories, by revising skills, prompts, or topology after a batch of runs has ended\([Li & Ramakrishnan, 2026](https://arxiv.org/html/2609.38661#bib.bib21);[Pan et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib33)\), making a misconfigured team costly to detect, as its outputs still look plausible, and impossible to repair while the workflow is still being executed\.

Figure 2:Three lines of work and ours\.\(a\)Post\-hoc evolution revises the team only after the run\.\(b\)Outcome\-driven optimization spreads one terminal reward over every step\.\(c\)Skill evolution admits by judges or replay\.\(d\)EvoSteer edits the team inside the same run, deciding from measured execution features and a reference value estimate, and admits a skill only after a sequential test\.To address these issues, three main lines of work have emerged, as shown in Figure[2](https://arxiv.org/html/2609.38661#S1.F2)\. First,*self\-evolving orchestration*adapts topologies, role prompts, and skill libraries from the outcomes of completed trajectories\([Li & Ramakrishnan, 2026](https://arxiv.org/html/2609.38661#bib.bib21);[Pan et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib33);[Zhang et al\., 2026d](https://arxiv.org/html/2609.38661#bib.bib66)\), treating the team as an object of learning\. Second,*credit assignment for multi\-agent LLMs*converts a shared terminal reward into per\-agent or per\-message signals through learned critics, leave\-one\-out baselines, counterfactual replay, and flow\-matching objectives\([Chen et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib5);[Li et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib22);[Shah, 2026](https://arxiv.org/html/2609.38661#bib.bib37);[Zhang et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib64)\), telling which step was responsible for the outcome\. Third,*skill libraries*let capability grow across tasks by extracting reusable procedures from trajectories and curating them with judges, replay checks, or posteriors\([Wang et al\., 2023](https://arxiv.org/html/2609.38661#bib.bib47);[Yang et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib52);[Wu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib50)\)\.

However, these methods still face three challenges\.\(i\) Post\-hoc evolution\.Self\-evolving orchestration revises the team after the trajectory ends, by rounds or batches\([Pan et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib33);[Li & Ramakrishnan, 2026](https://arxiv.org/html/2609.38661#bib.bib21)\); during execution the orchestrator can only keep adding agents, and the decisive error is locally indistinguishable from a recoverable step\([Zhang et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib58)\), so a wrong agent is neither rerun nor removed until the task has failed\.\(ii\) Credit diffusion\.Under a shared terminal reward, trajectory\-level methods give every action the same coefficient of a single terminal advantage\([Chen et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib5);[Zhang, 2026](https://arxiv.org/html/2609.38661#bib.bib59)\), and learned state values confound how often a task is solved, execution noise, and structural choice, so the orchestrator cannot tell whether a low score came from a task the reference rarely solves or from a badly organized team; counterfactual replay localizes the failing step\([Shah, 2026](https://arxiv.org/html/2609.38661#bib.bib37);[Bonagiri et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib1)\)but does not provide a training signal for every decision\.\(iii\) Uncalibrated skill admission\.Skills are admitted by LLM judgement or by success counts treated as reliable belief\([Yang et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib52);[Wu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib50)\); such point estimates cannot separate insufficient evidence from a confirmed effect, so an accidental strategy from one trajectory is consolidated\([Chen et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib3)\), and since no test tracks whether an admitted skill still helps on later tasks, it stays in the library and is never retired\.

To address these challenges, we proposeEvoSteer\(Fig\.[2](https://arxiv.org/html/2609.38661#S1.F2)d\), a new paradigm of*Online Self\-Evolving Graph Orchestration*, in which the orchestrator dynamically builds and repairs a running team from execution features and a learned value estimate—replacing the run\-then\-revise loop\. The orchestrator issues one atomic edit per turn, rerun and drop included; each is executed at once, and its feedback and the value estimate enter the state, so a plausible but failing result can be redone mid\-run\. To support this paradigm, we introduce*Anchored Trajectory Balance*\(AnchorTB\), a regression\-style flow\-matching loss that assigns each orchestration action a coefficient by balancing subtrajectories against a frozen reference\. It combines measured task\-level reference reward statistics with prefix\-dependent corrections\. Built on the learned flow, we further propose*Validated Skill Admission*: a skill proposer distils candidates from scored experience, the orchestrator may try a candidate before promotion, and paired rollouts differing only in the candidate are judged by a sequential sign test under a shared nominal testing budget, which promotes, retires, or defers\. Trained end\-to\-end on natural and paired trajectories with the reference frozen throughout, EvoSteer turns construction, repair, and skill growth into one set of learned decisions under a single training objective\.

We evaluate on benchmarks across question answering, mathematical reasoning, interactive decision making, and code generation\. Results show EvoSteer outperforms workflow search, flow\- and RL\-based orchestration training, and skill\-evolution methods on every benchmark and improves all seven executors—a foundation for orchestration that evolves while it runs\.

## 2Related Work

#### Agent Task Orchestration\.

Orchestration graphs are searched under execution feedback\([Zhang et al\., 2025b](https://arxiv.org/html/2609.38661#bib.bib62);[Zhang et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib63)\)\. Later work evolves the team between runs\([Dang et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib6);[Hao et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib10)\), mutates topology between test\-time rounds\([Xu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib51)\), repairs finished trajectories\([Luan et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib29);[Lu & Zhang, 2026](https://arxiv.org/html/2609.38661#bib.bib28);[Zhao et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib67)\), or audits unfolding prefixes\([Wang et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib46);[Jiang et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib15);[Li et al\., 2026d](https://arxiv.org/html/2609.38661#bib.bib24)\), but keeps the running team fixed\. Credit comes from counterfactual replay\([Chen et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib4);[Deshmukh et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib7)\), group\-relative objectives\([Mishra et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib32)\), or \(sub\)trajectory balance\([Malkin et al\., 2022](https://arxiv.org/html/2609.38661#bib.bib31);[Madan et al\., 2023](https://arxiv.org/html/2609.38661#bib.bib30);[Venkatraman et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib45)\)and its extensions\([Fawkes & Hartford, 2026](https://arxiv.org/html/2609.38661#bib.bib8);[Liu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib27);[Wang et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib48);[Liang et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib25)\)\. EvoSteer reruns and drops agents in the running team and measures flows from a frozen reference\.

#### Skill Self\-Evolution\.

Skill libraries check candidates by replay\([He & Yang, 2026](https://arxiv.org/html/2609.38661#bib.bib11)\), paired trajectories\([Gao et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib9)\), leave\-one\-out attribution\([Shen et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib42);[Hu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib14);[Wang et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib49)\), or success\-count posteriors\([Wu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib50)\)\. Others couple per\-skill credit to library edits\([Yao et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib55)\), gate contaminated candidates\([Zhang & Li, 2026](https://arxiv.org/html/2609.38661#bib.bib65)\), or gate self\-modification with sequential or paired tests\([Shawn, 2026](https://arxiv.org/html/2609.38661#bib.bib41);[Sengupta, 2026](https://arxiv.org/html/2609.38661#bib.bib36);[Shang & Yang, 2026](https://arxiv.org/html/2609.38661#bib.bib38)\)\. EvoSteer gates skills in a live episode under one testing budget per run, and candidates stay selectable while tested\.

## 3Preliminaries

Definition 1: Typed Agent Communication Graph\.A typed agent communication graph \(henceforth communication graph\) is a directed graphG=\(V,E,o\)G=\(V,E,o\)over agent nodes, where each nodevi=\(ci,𝒦i\)v\_\{i\}=\(c\_\{i\},\\mathcal\{K\}\_\{i\}\)carries a rolecic\_\{i\}from a role catalogue and a set of bound skills𝒦i\\mathcal\{K\}\_\{i\}drawn from the visible active skill set, each edge\(vi,vj,πi​j\)\(v\_\{i\},v\_\{j\},\\pi\_\{ij\}\)carries a communication protocolπi​j\\pi\_\{ij\}that decides whether and how the source’s result revises the target, ando∈Vo\\in Vtogether with an output rule fixes the final answer \(our typed edges and nodes extend the untyped topologies of[Zhang et al\. \(2024\)](https://arxiv.org/html/2609.38661#bib.bib60);[Zhang et al\. \(2025a\)](https://arxiv.org/html/2609.38661#bib.bib61)and the computational graphs of[Zhuge et al\. \(2024\)](https://arxiv.org/html/2609.38661#bib.bib69)\)\.

Definition 2: Orchestration Trajectory\.A complete sequence ofTTorchestration actions from the empty team to a legal stop defines an orchestration trajectory, in which every action is executed as soon as it is issued, before the next action is chosen:

τ=\{\(at,otexec\)\}t=0T−1⇒r⁡\(x\),st\+1=st⊕\(at,otexec\),\\tau=\\bigl\\\{\(a\_\{t\},\\ o\_\{t\}^\{\\mathrm\{exec\}\}\)\\bigr\\\}\_\{t=0\}^\{T\-1\}\\ \\Rightarrow\\ r\(x\),\\qquad s\_\{t\+1\}=s\_\{t\}\\oplus\(a\_\{t\},\\ o\_\{t\}^\{\\mathrm\{exec\}\}\),\(1\)whereata\_\{t\}is an atomic edit with type in\{\\\{add\_agent,add\_edge,bind\_skill,set\_output,rerun\_agent,drop\_agent,stop\}\\\},otexeco\_\{t\}^\{\\mathrm\{exec\}\}is the execution feedback returned by running the agent, communication, repair, or output that the action triggers, andxxis the completed historyx=sTx=s\_\{T\}with task rewardr⁡\(x\)∈\[0,1\]r\(x\)\\in\[0,1\]andT≥1T\\geq 1\. Actionata\_\{t\}is chosen atsts\_\{t\}and producesst\+1s\_\{t\+1\}, fort=0,…,T−1t=0,\\ldots,T\-1\. Because each state retains the full history, every non\-root state has a unique parent\. The deterministic feature record is suppressed here and made explicit in Eq\. \([3](https://arxiv.org/html/2609.38661#S4.E3)\)\.

Problem Statement\.Given a taskqq, a frozen executor with tools, a frozen reference policyρ\\rho, and a skill library𝒮\\mathcal\{S\}, we seek an orchestratorπθ\\pi\_\{\\theta\}trained toward an ideal reward\-proportional history distribution relative toρ\\rho\([Venkatraman et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib45)\):

P∗​\(x∣C\)=ρ⁡\(x∣C\)​Rβ​\(x\)Z⁡\(C\),Rβ​\(x\)=1\+\(eβ−1\)​r​\(x\),P^\{\*\}\(x\\mid C\)=\\frac\{\\rho\(x\\mid C\)\\,R\_\{\\beta\}\(x\)\}\{Z\(C\)\},\\qquad R\_\{\\beta\}\(x\)=1\+\(e^\{\\beta\}\-1\)\\,r\(x\),\(2\)where0≤β<∞0\\leq\\beta<\\infty, andCCcollects the task, tools, roles, and the skill and value\-head snapshots of the current batch,ρ⁡\(x∣C\)\\rho\(x\\mid C\)is the history measure induced by the reference policy and the executor’s stochastic transitions, andRβR\_\{\\beta\}keeps a positive mass for failures while weighting a full scoreeβe^\{\\beta\}times a zero score\. Equation \([2](https://arxiv.org/html/2609.38661#S3.E2)\) is the training target; any deterministic executor, and more generally one meeting the condition of Proposition[A\.10](https://arxiv.org/html/2609.38661#A1.Thmproposition10), lets action choices alone realize it\.

## 4Methodology: EvoSteer

Figure 3:EvoSteer architecture\.Top:πθ\\pi\_\{\\theta\}builds and runs the team from execution featuresftf\_\{t\}and reference estimatev^k\\hat\{v\}\_\{k\}; rerun and new edges are ordinary actions\.Bottom left:AnchorTB balances subtrajectories againstρ\\rho;u~,δ\\tilde\{u\},\\deltaabbreviate trajectory\-specific flow estimates and residuals\.Bottom right:paired rollouts and a sign test under a shared nominalα\\alphabudget promote or retire candidateσ\\sigma\.As illustrated in Figure[3](https://arxiv.org/html/2609.38661#S4.F3), this section introduces the EvoSteer framework, including online self\-evolving graph orchestration \(Section 4\.1\), anchored trajectory balance \(Section 4\.2\), and validated skill admission \(Section 4\.3\), all built on the frozen referenceρ\\rho\.

### 4\.1Online Self\-Evolving Graph Orchestration

As shown in Figure[3](https://arxiv.org/html/2609.38661#S4.F3)\(top\), EvoSteer follows an Orchestrator–Executor paradigm\([Yao et al\., 2022b](https://arxiv.org/html/2609.38661#bib.bib57);[Zhu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib68);[Zhang et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib63)\): a trainable orchestratorπθ\\pi\_\{\\theta\}organizes tool\-using agents, each on a frozen executor, inside a structured environment\.

Structured Environment\.The environmentℰ=\(𝒫,𝒮\(k\),ℛexec,v^k\)\\mathcal\{E\}=\(\\mathcal\{P\},\\mathcal\{S\}^\{\(k\)\},\\mathcal\{R\}\_\{\\mathrm\{exec\}\},\\hat\{v\}\_\{k\}\)maintains the role catalogue𝒫\\mathcal\{P\}, i\.e\., the roles with their tool and session permissions; the active skill set𝒮\(k\)\\mathcal\{S\}^\{\(k\)\}visible in training stepkk\(slot counts in Appendix C\); the executor runtimeℛexec\\mathcal\{R\}\_\{\\mathrm\{exec\}\}with a shared resource budget; and the reference value headv^k\\hat\{v\}\_\{k\}of Eq\. \([5](https://arxiv.org/html/2609.38661#S4.E5)\)\. The two snapshots𝒮\(k\)\\mathcal\{S\}^\{\(k\)\}andv^k\\hat\{v\}\_\{k\}are fixed for a whole rollout batch, while the legal mask𝒜⁡\(s\)\\mathcal\{A\}\(s\)remains state\-dependent\.

Interleaved Execution\.Each orchestrator action is validated, applied to the graph, and executed at once; its feedback and a fixed\-length execution feature vector are appended to the state:

st\+1=st⊕\(at,otexec,ft\+1\),ft\+1=f⁡\(st,at,otexec\)∈ℝ30,s\_\{t\+1\}=s\_\{t\}\\oplus\\bigl\(a\_\{t\},\\ o\_\{t\}^\{\\mathrm\{exec\}\},\\ f\_\{t\+1\}\\bigr\),\\qquad f\_\{t\+1\}=f\\bigl\(s\_\{t\},a\_\{t\},o\_\{t\}^\{\\mathrm\{exec\}\}\\bigr\)\\in\\mathbb\{R\}^\{30\},\(3\)whereotexeco\_\{t\}^\{\\mathrm\{exec\}\}is the execution feedback of Definition 2 andffrecords graph size, action type, repair outcome, whether agents answered and agreed, communication revision, environment phase, and budget use\. These features need no extra model call and cover the execution\-side failure modes catalogued for multi\-agent systems\([Cemri et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib2)\);f⁡\(s\)f\(s\)denotes the latest vector inss\.

Learnable Repair\.Repair shares one action space and one policyat∼πθ\(⋅∣st\)a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\}\)with construction:

𝒜⁡\(st\)⊆𝒜con∪𝒜rep∪\{stop\},𝒜rep=\{rerun\_agent,drop\_agent\},\\mathcal\{A\}\(s\_\{t\}\)\\subseteq\\mathcal\{A\}\_\{\\mathrm\{con\}\}\\cup\\mathcal\{A\}\_\{\\mathrm\{rep\}\}\\cup\\\{\\textsc\{stop\}\\\},\\qquad\\mathcal\{A\}\_\{\\mathrm\{rep\}\}=\\\{\\textsc\{rerun\\\_agent\},\\ \\textsc\{drop\\\_agent\}\\\},\(4\)where𝒜con=\{add\_agent,add\_edge,bind\_skill,set\_output\}\\mathcal\{A\}\_\{\\mathrm\{con\}\}=\\\{\\textsc\{add\\\_agent\},\\textsc\{add\\\_edge\},\\textsc\{bind\\\_skill\},\\textsc\{set\\\_output\}\\\}holds the construction edits of Definition 2, and the mask𝒜⁡\(st\)\\mathcal\{A\}\(s\_\{t\}\)\([Li et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib20)\)follows from the current graph, roles, skill slots, output modes, and budget, admitting repairs only when valid\. One objective trains all three kinds of decision: how many agents a task needs, when a result is worth redoing, and when the answer in hand suffices\. Team size and repair count follow from the policy under the shared budget, and the terminal rewardr⁡\(x\)r\(x\)is computed once, after the stop, on the output node’s answer\.

In\-Loop Reference Value Estimate\.A lightweight head uses execution features and task type to estimate reference continuation reward by squared\-error regression:

v^ω​\(s\)=σ⁡\(MLPω​\[f⁡\(s\);onehot⁡\(tasktype⁡\(q\)\)\]\),ℒv=𝔼\(s,r\)∼𝒟ρ​\(v^ω​\(s\)−r\)2,\\hat\{v\}\_\{\\omega\}\(s\)=\\sigma\\bigl\(\\mathrm\{MLP\}\_\{\\omega\}\[f\(s\);\\,\\mathrm\{onehot\}\(\\mathrm\{tasktype\}\(q\)\)\]\\bigr\),\\qquad\\mathcal\{L\}\_\{v\}=\\mathbb\{E\}\_\{\(s,r\)\\sim\\mathcal\{D\}\_\{\\rho\}\}\\bigl\(\\hat\{v\}\_\{\\omega\}\(s\)\-r\\bigr\)^\{2\},\(5\)whereMLPω\\mathrm\{MLP\}\_\{\\omega\}is a two\-layer head on30\+630\{\+\}6inputs \(Appendix B\),𝒟ρ\\mathcal\{D\}\_\{\\rho\}holds states of legal reference continuations with their terminal rewards, andv^k\\hat\{v\}\_\{k\}inℰ\\mathcal\{E\}is the snapshot ofv^ω\\hat\{v\}\_\{\\omega\}frozen for batchkk\. The estimate and its one\-step change are shown to the orchestrator as feedback; both are deterministic readouts of the retained history under the frozen head and carry cross\-task experience into in\-loop decisions, including when to stop\. Its population MSE optimum is the feature\-conditional mean of the terminal reward under the training data law, and this optimum is calibrated \(Proposition[B\.8](https://arxiv.org/html/2609.38661#A2.Thmproposition8)\)\.

Proposition 1\.*Executing every action as it is issued and placing rerun and drop in the action space lets the orchestrator edit and retry a running team through the same action interface\.**Proof\.*Appendix[C\.6](https://arxiv.org/html/2609.38661#A3.SS6), via the history tree and legal actions of Lemma[A\.1](https://arxiv.org/html/2609.38661#A1.Thmproposition1)\.

### 4\.2Anchored Trajectory Balance

As shown in Figure[3](https://arxiv.org/html/2609.38661#S4.F3)\(bottom left\) and step by step in Figure[4](https://arxiv.org/html/2609.38661#S4.F4), we trainπθ\\pi\_\{\\theta\}on the history tree of Section 3 toward the ideal reward\-proportional target of Eq\. \([2](https://arxiv.org/html/2609.38661#S3.E2)\), relative to the frozen reference\.

① Score each actionEq\. \(6\)② Anchor each stateEq\. \(8\)③ Balance every spanEq\. \(7\)④ Sum into creditProp\. 2s0s\_\{0\}s1s\_\{1\}s2s\_\{2\}s3s\_\{3\}a0a\_\{0\}Δ​ℓ0\\Delta\\ell\_\{0\}a1a\_\{1\}Δ​ℓ1\\Delta\\ell\_\{1\}a2a\_\{2\}Δ​ℓ2\\Delta\\ell\_\{2\}Δ​ℓt=ℓθ​\(at\)−ℓρ​\(at\)\\Delta\\ell\_\{t\}=\\ell\_\{\\theta\}\(a\_\{t\}\)\-\\ell\_\{\\rho\}\(a\_\{t\}\)ℓ\\ell: log\-probability summed overthe executed tool\-call tokensρ\\rho: same backbone, frozen;both scored at the samests\_\{t\}s0s\_\{0\}s1s\_\{1\}s2s\_\{2\}s3s\_\{3\}u~0\\tilde\{u\}\_\{0\}u~1\\tilde\{u\}\_\{1\}u~2\\tilde\{u\}\_\{2\}log⁡Rβ\\log R\_\{\\beta\}measured, no gradientsg⁡\[clip\[0,β\]​\(uq,x\+c⁡\(g\)\)\]\\operatorname\{sg\}\\bigl\[\\mathrm\{clip\}\_\{\[0,\\beta\]\}\\bigl\(u\_\{q,x\}\+c\(g\)\\bigr\)\\bigr\]\+\+learned residualbψ​\(hρ​\(s\),f⁡\(s\)\)b\_\{\\psi\}\\bigl\(h\_\{\\rho\}\(s\),f\(s\)\\bigr\)uq,xu\_\{q,x\}: task reference anchorc⁡\(g\)c\(g\): structure corrections0s\_\{0\}s1s\_\{1\}s2s\_\{2\}s3s\_\{3\}δ0:1\\delta\_\{0:1\}δ1:2\\delta\_\{1:2\}δ2:3\\delta\_\{2:3\}δ0:2\\delta\_\{0:2\}δ1:3\\delta\_\{1:3\}δ0:3\\delta\_\{0:3\}δi:j=u~i\+∑t=ij−1Δℓt−u~j\\delta\_\{i:j\}=\\tilde\{u\}\_\{i\}\+\\sum\_\{t=i\}^\{j\-1\}\\Delta\\ell\_\{t\}\-\\tilde\{u\}\_\{j\}ℒx=16∑i<jδi:j2\\mathcal\{L\}\_\{x\}=\\tfrac\{1\}\{6\}\\sum\_\{i<j\}\\delta\_\{i:j\}^\{2\}δ0:1\\delta\_\{0:1\}δ1:2\\delta\_\{1:2\}δ2:3\\delta\_\{2:3\}δ0:2\\delta\_\{0:2\}δ1:3\\delta\_\{1:3\}δ0:3\\delta\_\{0:3\}a0a\_\{0\}a1a\_\{1\}a2a\_\{2\}∂ℒx∂ℓθ​\(at\)=2K3∑i≤t<jδi:j\\dfrac\{\\partial\\mathcal\{L\}\_\{x\}\}\{\\partial\\ell\_\{\\theta\}\(a\_\{t\}\)\}=\\tfrac\{2\}\{K\_\{3\}\}\\sum\_\{i\\leq t<j\}\\delta\_\{i:j\}rows use 3, 4, and 3 spansTB keeps onlyδ0:3\\delta\_\{0:3\}\(boxed\):one signal for all actions

Figure 4:AnchorTB on a three\-action history: ① score each action against the frozen reference; ② anchor each state with a measured, stop\-gradient anchor \(blue\) and a learned residual \(pink\); ③ form allK3=6K\_\{3\}=6span residuals; ④ sum each action’s spans into its coefficient\.Action Log\-Ratio\.An orchestrator action is a tool call ofKtK\_\{t\}tokens sampled under a grammar mask; its log\-probability and its log\-ratio to the reference are computed on the executed tokens:

ℓθ​\(at∣st\)=∑j=1Ktlog⁡Pθ​\(at,j∣st,at,<j\),Δ​ℓt=ℓθ​\(at∣st\)−ℓρ​\(at∣st\),\\ell\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)=\\sum\_\{j=1\}^\{K\_\{t\}\}\\log P\_\{\\theta\}\(a\_\{t,j\}\\mid s\_\{t\},a\_\{t,<j\}\),\\qquad\\Delta\\ell\_\{t\}=\\ell\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)\-\\ell\_\{\\rho\}\(a\_\{t\}\\mid s\_\{t\}\),\(6\)whereρ\\rhois the same backbone without the trainable adapter, scored under the same context, tokens, and mask\. Positions with a single legal token contribute zero, so only real choices enter the ratio\.

Subtrajectory Residual and Loss\.Every non\-root state has a unique parent, so the backward policy is identically one\. Motivated by reference\-relative subtrajectory balance\([Madan et al\., 2023](https://arxiv.org/html/2609.38661#bib.bib30);[Venkatraman et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib45)\), we regress each residual toward zero:

δi:j\(x\)=u~x\(si\)\+∑t=ij−1Δℓt−u~x\(sj\),u~x\(sT\)=logRβ\(x\),ℒx=1KT∑0≤i<j≤T\(δi:j\(x\)\)2,\\delta^\{\(x\)\}\_\{i:j\}=\\tilde\{u\}\_\{x\}\(s\_\{i\}\)\+\\sum\_\{t=i\}^\{j\-1\}\\Delta\\ell\_\{t\}\-\\tilde\{u\}\_\{x\}\(s\_\{j\}\),\\qquad\\tilde\{u\}\_\{x\}\(s\_\{T\}\)=\\log R\_\{\\beta\}\(x\),\\qquad\\mathcal\{L\}\_\{x\}=\\frac\{1\}\{K\_\{T\}\}\\sum\_\{0\\leq i<j\\leq T\}\\bigl\(\\delta^\{\(x\)\}\_\{i:j\}\\bigr\)^\{2\},\(7\)where nonterminal flows use Eq\. \([8](https://arxiv.org/html/2609.38661#S4.E8)\) with the same trajectory indexxxthroughout \(batch index suppressed\), the terminal value overrides the learned parameterization, andKT=T⁡\(T\+1\)/2K\_\{T\}=T\(T\+1\)/2counts the equally weighted subtrajectories forT≥1T\\geq 1, so each action’s coefficient sums the residuals of the spans containing it \(Figure[4](https://arxiv.org/html/2609.38661#S4.F4), ④\)\.

Measured State Flows\.The ideal relative log flowu∗\(s\)=log𝔼ρ\[Rβ\(x\)∣s,C\]u^\{\*\}\(s\)=\\log\\mathbb\{E\}\_\{\\rho\}\[R\_\{\\beta\}\(x\)\\mid s,C\]satisfiesu∗​\(s0\)=log⁡Z⁡\(C\)u^\{\*\}\(s\_\{0\}\)=\\log Z\(C\)\. A shared flow with zero residuals recoversu∗u^\{\*\}andP∗P^\{\*\}\(Proposition[A\.8](https://arxiv.org/html/2609.38661#A1.Thmproposition8)\); this motivates reference anchors with a learned residual:

u~x​\(s\)=sg⁡\[clip\[0,β\]⁡\(uq,x\+c⁡\(g⁡\(s\)\)\)\]\+bψ​\(hρ​\(s\),f⁡\(s\)\),uq,x=log⁡\[1\+\(eβ−1\)​p^q\(−x\)\],\\tilde\{u\}\_\{x\}\(s\)=\\operatorname\{sg\}\\\!\\Bigl\[\\clip\_\{\[0,\\beta\]\}\\bigl\(u\_\{q,x\}\+c\(g\(s\)\)\\bigr\)\\Bigr\]\+b\_\{\\psi\}\\bigl\(h\_\{\\rho\}\(s\),f\(s\)\\bigr\),\\qquad u\_\{q,x\}=\\log\\bigl\[1\+\(e^\{\\beta\}\-1\)\\,\\hat\{p\}\_\{q\}^\{\(\-x\)\}\\bigr\],\(8\)For nonterminalss,p^q\(−x\)\\hat\{p\}\_\{q\}^\{\(\-x\)\}is a task\-level partial leave\-one\-out shrinkage estimate of the reference mean reward on the same task, computed from reference rollout outcomes \(Appendix B\);c⁡\(g\)c\(g\)is a correction read one batch behind from a bucket of canonical prefix structuresgg,hρ​\(s\)h\_\{\\rho\}\(s\)the frozen reference encoding, andbψb\_\{\\psi\}a zero\-initialized MLP\. Under leave\-one\-out, one prefix can differ across scored trajectories\. The correction averages reward ratios over prefix visits before the logarithm,

yx=Rβ​\(x\)euq,x−1,c⁡\(g\)=clip\[−cmax,cmax\]⁡log⁡\(1\+∑\(x,t\)∈𝒱gyxng\+n0\),y\_\{x\}=\\frac\{R\_\{\\beta\}\(x\)\}\{e^\{u\_\{q,x\}\}\}\-1,\\qquad c\(g\)=\\clip\_\{\[\-c\_\{\\max\},\\,c\_\{\\max\}\]\}\\log\\Bigl\(1\+\\frac\{\\sum\_\{\(x,t\)\\in\\mathcal\{V\}\_\{g\}\}y\_\{x\}\}\{n\_\{g\}\+n\_\{0\}\}\\Bigr\),\(9\)where𝒱g\\mathcal\{V\}\_\{g\}is the multiset of eligible reference prefix visits to structuregg, paired prefixes included \(Appendix C\),ng=\|𝒱g\|n\_\{g\}=\|\\mathcal\{V\}\_\{g\}\|, andn0≥0n\_\{0\}\\geq 0withng\+n0\>0n\_\{g\}\+n\_\{0\}\>0\. The pseudo\-count shrinks rare structures toward zero correction andcmaxc\_\{\\max\}bounds it \(values in Appendix C\), and the logarithm’s argument stays positive \(Lemma[B\.6](https://arxiv.org/html/2609.38661#A2.Thmproposition6)\)\. The measured terms are held fixed, so onlyπθ\\pi\_\{\\theta\}andbψb\_\{\\psi\}are updated, withbψb\_\{\\psi\}carrying the remainder\.

Proposition 2\.*AnchorTB assigns each orchestration action a regression coefficient from all subtrajectories containing it, and combines a task\-level reference anchor with prefix\-dependent corrections\.**Proof\.*Appendix[C\.6](https://arxiv.org/html/2609.38661#A3.SS6), via the coefficient of Proposition[A\.14](https://arxiv.org/html/2609.38661#A1.Thmproposition14)and the corrections of Appendix B\.

### 4\.3Validated Skill Admission

Prior work admits a skill from a single trajectory, by a judge, by counting successes, or by a trust hierarchy over verifiers\([Yang et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib52);[Chen et al\., 2026a](https://arxiv.org/html/2609.38661#bib.bib3);[Wu et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib50);[Shang et al\., 2026](https://arxiv.org/html/2609.38661#bib.bib39)\), leaving underspecified*whether*a candidate helps and*when*it may be promoted\. EvoSteer answers both inside the loop: a candidate is an ordinary entry of the action set in Eq\. \([4](https://arxiv.org/html/2609.38661#S4.E4)\) that the orchestrator may bind while its evidence accumulates, and that evidence is measured under the frozen reference of Section 4\.2 \(Figure[3](https://arxiv.org/html/2609.38661#S4.F3), bottom right\), so credit and admission share one reference\.

Skill Proposal and Action\-Set Integration\.In an author window, an independent author model distils at most one candidateσ\\sigma, a structured procedure, from scored, side\-effect\-free trajectories \(proposal rules and skill format in Appendix C\)\. A candidate occupies a candidate slot of𝒮\(k\)\\mathcal\{S\}^\{\(k\)\}, so the action mask of Eq\. \([4](https://arxiv.org/html/2609.38661#S4.E4)\) already exposes it toadd\_agentandbind\_skillbefore it is validated; the full skill text is rendered only when the bound agent executes\.

Paired Counterfactual Rollouts\.For every batch task in a family with a candidate, two extra trajectories are launched from the same record and the same forced first role, one bindingσ\\sigmaand one not, and both are continued by the reference policy:

x\+∼ρ\(⋅∣C\+,add\_agent\(c1,σ\)\),x−∼ρ\(⋅∣C−,add\_agent\(c1,∅\)\),x^\{\+\}\\sim\\rho\\bigl\(\\cdot\\mid C^\{\+\},\\ \\textsc\{add\\\_agent\}\(c\_\{1\},\\ \\sigma\)\\bigr\),\\qquad x^\{\-\}\\sim\\rho\\bigl\(\\cdot\\mid C^\{\-\},\\ \\textsc\{add\\\_agent\}\(c\_\{1\},\\ \\varnothing\)\\bigr\),\(10\)wherec1c\_\{1\}is the forced first role\. The arm contextsC±C^\{\\pm\}share the initial task, runtime, and value\-head snapshot;C−C^\{\-\}masksσ\\sigmathroughout\. Pairs are scored by the terminal reward of Eq\. \([2](https://arxiv.org/html/2609.38661#S3.E2)\); withb±=𝕀\[r\(x±\)≥12\]b^\{\\pm\}=\\mathbb\{I\}\[r\(x^\{\\pm\}\)\\geq\\tfrac\{1\}\{2\}\], the comparisons accumulate as winsW=∑𝕀\[b\+\>b−\]W=\\sum\\mathbb\{I\}\[b^\{\+\}\>b^\{\-\}\], lossesL=∑𝕀\[b\+<b−\]L=\\sum\\mathbb\{I\}\[b^\{\+\}<b^\{\-\}\], and tiesT0=∑𝕀\[b\+=b−\]T\_\{0\}=\\sum\\mathbb\{I\}\[b^\{\+\}=b^\{\-\}\]over complete, side\-effect\-free pairs, which also enter AnchorTB as off\-policy paths, re\-scored under each arm’s own context and legal mask \(Appendix C\)\.

Sequential Validation\.At validation rounds between batches, each candidate with new wins or losses gets one binomial sign\-test tail per direction, exact under the fixed\-look sign model \(Appendix C\):

p\+=ℙ\{Bin\(D,12\)≥W\},p−=ℙ\{Bin\(D,12\)≥L\},D=W\+L,p\_\{\+\}=\\mathbb\{P\}\\bigl\\\{\\mathrm\{Bin\}\(D,\\tfrac\{1\}\{2\}\)\\geq W\\bigr\\\},\\qquad p\_\{\-\}=\\mathbb\{P\}\\bigl\\\{\\mathrm\{Bin\}\(D,\\tfrac\{1\}\{2\}\)\\geq L\\bigr\\\},\\qquad D=W\+L,\(11\)withp\+=p−=1p\_\{\+\}=p\_\{\-\}=1whenD=0D=0\. A candidate whosep\+p\_\{\+\}clears its boundary is promoted to validated, one whosep−p\_\{\-\}clears it is retired, and otherwise it keeps accumulating\. One budget covers the whole run on two levels, and the same counts give the reported effect:

αj,ℓ=αj⁡\(j\+1\)​ℓ​\(ℓ\+1\),Δ^=W−LW\+L\+T0,\\alpha\_\{j,\\ell\}=\\frac\{\\alpha\}\{j\(j\+1\)\\,\\ell\(\\ell\+1\)\},\\qquad\\hat\{\\Delta\}=\\frac\{W\-L\}\{W\+L\+T\_\{0\}\},\(12\)whereΔ^\\hat\{\\Delta\}is reported only whenW\+L\+T0\>0W\+L\+T\_\{0\}\>0and measures the paired threshold\-pass\-rate difference\. Herejjglobally indexes the registered comparison andℓ\\ellits observation; each direction receivesαj,ℓ/2\\alpha\_\{j,\\ell\}/2, and the budget is spent only when new evidence arrives, so the total nominal level is at mostα\\alpha\(Appendix C\)\. By a union bound, family\-wise error is at mostα\\alphawhenever each directional p\-value is valid for the selection and sampling rules in use, even for dependent tests\.

Proposition 3\.*Validated Skill Admission lets a candidate be tried before it is promoted, and promotes or retires it only when paired evidence passes a sequential test under a shared alpha\-spending budget\.**Proof\.*Appendix[C\.6](https://arxiv.org/html/2609.38661#A3.SS6), via the candidate\-slot and status\-transition rules, the nominal budget accounting of Proposition[C\.5](https://arxiv.org/html/2609.38661#A3.Thmproposition5), and the conditional FWER result of Proposition[C\.6](https://arxiv.org/html/2609.38661#A3.Thmproposition6)\.

## 5Experiments

We evaluate EvoSteer through the following research questions \(RQs\)\.RQ1: How does EvoSteer compare with the baselines in distribution?RQ2: Does it generalize to held\-out benchmarks?RQ3: Does it transfer across executor backbones?RQ4: What does each component contribute, and how do fixed orchestration paradigms compare?RQ5: How does AnchorTB compare with other objectives, and do the measured anchor, validated admission, and in\-run edits each behave as designed?

BaselineSFTGRPO†AFlowAgent\+RLSkill evolutionOursDatasetMetricQwen3\.5v4\-flashQwen3\.5Qwen3\.5Qwen3\.5AgentFlowFlowSteerSkillFlowSkillOptEvoSteer \(Δ↑\\Delta\\uparrow\)\(a\) In\-Distribution \(IID\) benchmarksHotpotQAAns EM57\.66±1\.2870\.94±0\.3563\.59±1\.1869\.22±0\.4387\.34±0\.8686\.09±0\.8689\.22±0\.6589\.06±0\.0088\.28±0\.0092\.34±0\.35\(\+34\.7\)Ans F173\.75±1\.2482\.33±0\.4777\.70±1\.2381\.64±0\.5088\.97±0\.8588\.42±0\.8890\.30±0\.6692\.00±0\.1190\.34±0\.4492\.84±0\.35\(\+19\.1\)NQ\-OpenAns EM23\.59±0\.8638\.59±1\.3125\.78±0\.0024\.22±1\.1076\.56±0\.7881\.41±1\.2879\.53±1\.2882\.66±0\.8682\.34±1\.3185\.00±0\.65\(\+61\.4\)Ans F133\.04±1\.2249\.60±1\.7837\.69±0\.1432\.59±1\.4580\.77±0\.8285\.46±1\.7584\.31±1\.3686\.56±0\.8086\.28±1\.2788\.92±0\.69\(\+55\.9\)MedQAAcc\.71\.41±0\.4384\.53±0\.3573\.13±0\.4377\.66±0\.4389\.22±0\.6588\.44±0\.8690\.31±1\.1891\.41±0\.7891\.72±0\.7093\.13±0\.86\(\+21\.7\)AIME 2026Acc\.48\.67±2\.9850\.67±1\.4932\.67±2\.7932\.00±1\.8353\.33±0\.0060\.67±2\.7964\.67±2\.9863\.33±2\.3666\.00±1\.4974\.67±1\.83\(\+26\.0\)MBPP\+Pass@178\.91±1\.1084\.53±0\.8678\.44±0\.7081\.88±0\.3585\.00±0\.6586\.09±0\.6587\.34±0\.8688\.59±0\.7090\.16±0\.4392\.19±1\.10\(\+13\.3\)ALFWorldSR46\.88±1\.2460\.47±0\.4340\.31±0\.7050\.47±0\.7073\.28±0\.6576\.09±0\.4381\.56±0\.7083\.28±0\.7085\.78±0\.8689\.69±1\.02\(\+42\.8\)Avg\. \(IID\)Ans EM40\.63±0\.7754\.77±0\.6844\.69±0\.5946\.72±0\.5981\.95±0\.5883\.75±0\.7784\.38±0\.7285\.86±0\.4385\.31±0\.6588\.67±0\.37\(\+48\.0\)Ans F153\.39±0\.8765\.97±0\.9257\.70±0\.6257\.12±0\.7784\.87±0\.5986\.94±0\.9887\.31±0\.7689\.28±0\.4088\.31±0\.6790\.88±0\.39\(\+37\.5\)Acc\.61\.46±0\.8670\.05±0\.4556\.14±0\.7560\.50±0\.5175\.21±0\.2877\.82±0\.7680\.97±0\.8581\.65±0\.6783\.41±0\.4887\.42±0\.63\(\+26\.0\)\(b\) Out\-of\-Distribution \(OOD\) benchmarksTriviaQAAns EM44\.22±0\.4370\.31±0\.0046\.09±0\.5547\.34±1\.1890\.94±1\.0589\.06±0\.0091\.25±0\.3590\.16±1\.0592\.34±1\.2895\.63±1\.18\(\+51\.4\)Ans F153\.46±0\.5280\.81±0\.2158\.77±0\.7058\.50±1\.4592\.34±1\.0690\.44±0\.0492\.71±0\.3591\.79±1\.0793\.41±1\.3096\.81±1\.23\(\+43\.4\)MuSiQueAns EM39\.22±0\.8650\.00±0\.7839\.06±0\.5540\.63±0\.5577\.66±0\.7079\.06±1\.1680\.00±1\.1880\.94±0\.4381\.09±0\.6584\.53±0\.86\(\+45\.3\)Ans F149\.12±1\.1758\.73±0\.7250\.79±0\.9550\.09±0\.6882\.70±0\.7484\.36±1\.2485\.74±1\.4784\.82±0\.4586\.09±0\.6987\.84±0\.89\(\+38\.7\)GPQAAcc\.61\.72±1\.1073\.91±0\.4368\.91±0\.8667\.50±0\.8975\.63±1\.0279\.69±0\.9683\.75±0\.6582\.19±0\.8681\.88±0\.3585\.63±0\.70\(\+23\.9\)MATH\-HardAcc\.89\.06±0\.7892\.19±0\.5588\.59±1\.3187\.50±0\.7891\.09±1\.0592\.66±0\.7093\.44±0\.7092\.34±1\.0294\.69±0\.6595\.47±0\.86\(\+6\.4\)SWE\-BenchResolved15\.94±0\.4336\.25±0\.7015\.00±0\.8616\.88±1\.0528\.28±0\.6538\.75±0\.7040\.31±0\.7039\.84±0\.7841\.56±0\.6542\.50±0\.70\(\+26\.6\)WebShopSR32\.66±0\.8665\.47±0\.8632\.81±0\.7836\.25±0\.7059\.53±0\.8676\.09±0\.4378\.28±0\.8682\.81±1\.1085\.94±0\.7890\.63±0\.00\(\+58\.0\)Avg\. \(OOD\)Ans EM41\.72±0\.4860\.16±0\.3942\.58±0\.3943\.98±0\.6584\.30±0\.6384\.06±0\.5885\.63±0\.6285\.55±0\.5786\.72±0\.7290\.08±0\.73\(\+48\.4\)Ans F151\.29±0\.6469\.77±0\.3854\.78±0\.5954\.29±0\.8087\.52±0\.6587\.40±0\.6289\.22±0\.7688\.31±0\.5889\.75±0\.7492\.32±0\.76\(\+41\.0\)Acc\.49\.84±0\.4166\.95±0\.3351\.33±0\.4952\.03±0\.4363\.63±0\.4571\.80±0\.3673\.95±0\.3774\.30±0\.4776\.02±0\.3178\.55±0\.33\(\+28\.7\)

Table 1:Main results \(five\-run mean±\\pmstd\)\. All methods run on the Qwen3\.5\-9B executor except v4\-flash \(DeepSeek\-V4\-Flash\); GRPO†trains the backbone;Δ↑\\Delta\\uparrow: gain over Qwen3\.5\-9B\.### 5\.1Experimental Setup

Datasets\.Six in\-distribution \(IID\) benchmarks supply the training tasks and IID tests: HotpotQA\([Yang et al\., 2018](https://arxiv.org/html/2609.38661#bib.bib54)\), NQ\-Open\([Kwiatkowski et al\., 2019](https://arxiv.org/html/2609.38661#bib.bib19)\), MedQA\([Jin et al\., 2021](https://arxiv.org/html/2609.38661#bib.bib17)\), AIME 2026, MBPP\+\([Liu et al\., 2023](https://arxiv.org/html/2609.38661#bib.bib26)\), and ALFWorld\([Shridhar et al\., 2020](https://arxiv.org/html/2609.38661#bib.bib43)\)\. Six out\-of\-distribution \(OOD\) benchmarks are held out: TriviaQA\([Joshi et al\., 2017](https://arxiv.org/html/2609.38661#bib.bib18)\), MuSiQue\([Trivedi et al\., 2022](https://arxiv.org/html/2609.38661#bib.bib44)\), GPQA\([Rein et al\., 2023](https://arxiv.org/html/2609.38661#bib.bib34)\), MATH\-Hard\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.38661#bib.bib12)\), SWE\-Bench Verified\([Jimenez et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib16)\), and WebShop\([Yao et al\., 2022a](https://arxiv.org/html/2609.38661#bib.bib56)\)\. Each test set has 128 items \(AIME 2026: 30\) disjoint from training\.

Baselines\.We compare with direct prompting \(Qwen3\.5\-9B, DeepSeek\-V4\-Flash\), backbone training \(SFT, GRPO†\([Shao et al\., 2024](https://arxiv.org/html/2609.38661#bib.bib40)\)\), workflow search \(AFlow\([Zhang et al\., 2025b](https://arxiv.org/html/2609.38661#bib.bib62)\)\), RL\-trained orchestration \(AgentFlow\([Li et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib23)\), FlowSteer\([Zhang et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib63)\)\), and skill evolution \(SkillFlow\([Zhang et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib64)\), SkillOpt\([Yang et al\., 2026b](https://arxiv.org/html/2609.38661#bib.bib53)\)\)\. All use Qwen3\.5\-9B as executor and trained model, the same training tasks, and their authors’ best configurations; DeepSeek\-V4\-Flash only writes EvoSteer’s candidate skills, and other authors move the averages by under 2 points\. EvoSteer is tested with learned parts frozen, and every test score is a five\-run mean \(Appendix[D](https://arxiv.org/html/2609.38661#A4)\)\.

VariantIIDOODHotpotQANQ\-OpenMedQAAIMEMBPP\+ALFWorldTriviaQAMuSiQueGPQAMATHSWEWebShopAns EMAns EMAcc\.Acc\.Pass@1SRAns EMAns EMAcc\.Acc\.ResolvedSRQwen3\.5\-9B \(frozen\)57\.6623\.5971\.4148\.6778\.9146\.8844\.2239\.2261\.7289\.0615\.9432\.66EvoSteer arch\., untrained82\.3473\.2882\.9753\.3387\.6673\.5990\.4771\.8875\.3190\.0029\.5369\.84Fixed orchestration paradigmsSingle agent with tools75\.9467\.1980\.7848\.6785\.6370\.1686\.2561\.4170\.3189\.8427\.0366\.09Fixed multi\-agent template82\.0374\.8485\.4756\.6783\.2866\.7289\.8470\.3173\.7590\.4724\.0654\.84Plan\-then\-execute80\.4772\.5081\.8851\.3384\.8465\.6390\.7868\.5971\.5689\.5325\.9457\.81Searched workflow87\.1976\.4188\.9154\.6785\.1673\.1391\.5677\.5075\.6390\.7828\.2859\.06Online orchestration \(Sec\. 4\.1\)−\-Interleaved execution89\.3883\.7589\.5358\.6790\.6380\.9495\.0079\.8481\.4193\.2831\.2575\.47−\-Execution features88\.2883\.7590\.3166\.6791\.4181\.8893\.9181\.2582\.3493\.2837\.3482\.81−\-Learnable repair90\.3184\.0689\.8467\.3390\.4783\.5995\.3180\.1682\.8192\.5031\.8887\.03−\-Reference value head83\.2884\.2291\.0968\.0089\.8480\.1694\.2279\.2281\.0992\.8139\.2282\.66AnchorTB \(Sec\. 4\.2\)−\-Measured flows90\.1683\.5989\.2256\.6790\.0077\.8194\.3876\.2581\.8891\.0933\.7580\.63−\-Flow corrections85\.9482\.6689\.8464\.0090\.3182\.5095\.0083\.7582\.8192\.1938\.1385\.16Skill admission \(Sec\. 4\.3\)−\-Skill evolution88\.2880\.9490\.1668\.6784\.5382\.8194\.3880\.6383\.2892\.8138\.5985\.16−\-Sequential validation90\.4781\.7289\.2266\.0090\.0078\.5994\.8481\.2583\.2894\.5335\.0086\.09EvoSteer \(Full\)92\.3485\.0093\.1374\.6792\.1989\.6995\.6384\.5385\.6395\.4742\.5090\.63

Table 2:Component ablation and paradigm comparison \(untrained: initialπθ\\pi\_\{\\theta\}\)\. Paradigms \(same executor and budget\) fix the team before execution: a ReAct\-style agent with all tools, a hand\-designed template, a planner that writes the graph once, and a workflow searched offline on training tasks\. Ablations:−\-interleaved execution builds the graph first;−\-execution features dropsfffrom the state ofπθ\\pi\_\{\\theta\};−\-learnable repair masksrerunanddrop;−\-reference value head withholdsv^k\\hat\{v\}\_\{k\}andΔ​v^\\Delta\\hat\{v\};−\-measured flows learnsuqu\_\{q\}as a free scalar;−\-flow corrections turns offc⁡\(g\)c\(g\)andbψb\_\{\\psi\};−\-skill evolution runs without any skills;−\-sequential validation admits afterk=5k=5successes\.\(a\) IID benchmark profile on each frozen backbone

\(b\) Training dynamics

\(c\) OOD scores aggregated by task domain

Figure 5:Backbone transfer and training dynamics\.\(a\)IID scores of six other frozen executors \(dashed\) and with the same trained orchestrator \(solid\); the radius is linear over 20–80 on the inner half and 80–100 on the outer half\.\(b\)Accuracy ofπθ\\pi\_\{\\theta\}and of the frozenρ\\rho, and loss against the anchor\-only level, over 240 steps\.\(c\)OOD scores per task domain\.ObjectiveTriviaMuSiQueGPQAMATHSWEWebShopGPU msGPU sTokensEMEMAcc\.Acc\.Res\.SR/ episode/ step/ problemNo training90\.4771\.8875\.3190\.0029\.5369\.84–––PPO93\.5979\.2280\.6392\.8135\.9482\.972,002384\.417,834\.3GRPO92\.1977\.3477\.5091\.5632\.6678\.441,174225\.417,886\.0Trajectory balance94\.8477\.0381\.2592\.6633\.4477\.34911174\.99,730\.1Tempered TB95\.6381\.0982\.5093\.7540\.0087\.811,737333\.519,373\.8AnchorTB \(ours\)95\.6384\.5385\.6395\.4742\.5090\.631,254240\.819,287\.4

\(a\) Objective comparison and training cost \(RQ5\)
\(b\) Pairwise IID differences
\(c\) Anchor AUC
\(d\) Admission gain
\(e\) Edit outcomes
\(f\) Replanning gain

Figure 6:Objective comparison and mechanism analysis \(definitions in Appendix[D](https://arxiv.org/html/2609.38661#A4)\)\.\(a\)OOD scores and cost per training objective\.\(b\)Row minus column, IID mean \(pp\); TTB: tempered TB\.\(c\)AUC for predicting reference\-rollout correctness\.\(d\)Paired skill gain, admitting all or only validated candidates\.\(e\)In\-run edits by type and outcome\.\(f\)Gain of value\-guided replanning in training\.
### 5\.2Main Results and Backbone Transfer \(RQ1–RQ3\)

Main results\.EvoSteer is best on all twelve benchmarks \(Table[1](https://arxiv.org/html/2609.38661#S5.T1)\), ahead of the strongest baseline by 2\.81 EM and 4\.01 accuracy points on IID averages and by 3\.36 and 2\.53 on OOD, with standard deviations at most 1\.83\. The lead is largest on AIME 2026 \(8\.67\) and smallest on MATH\-Hard \(0\.78\), where the backbone is near its ceiling\. Training the backbone alone \(SFT, GRPO†\) adds at most 6\.09 points to the IID EM average, so the gain comes from orchestration over the same frozen executor, and with a 9B executor EvoSteer beats the larger DeepSeek\-V4\-Flash on every benchmark\.

Backbone transfer\.The trained orchestrator transfers unchanged to six other executors and improves all 72 backbone–benchmark scores \(Figure[5](https://arxiv.org/html/2609.38661#S5.F5)a,c\), so one trained policy serves every executor\. Weaker executors gain more, and the IID spread across backbones shrinks from 17\.5 to 4\.3 points; gains are largest on interactive tasks and smallest on math, where frozen scores are already near the ceiling\.

### 5\.3Component Analysis \(RQ4, RQ5\)

Online orchestration \(Section 4\.1\)\.Each of its mechanisms contributes \(Table[2](https://arxiv.org/html/2609.38661#S5.T2)\): building the graph before execution, dropping the execution featuresff, masking rerun and drop, and withholding the value estimate cost 5\.69, 4\.12, 3\.57, and 5\.07 IID points\. Keeping interleaving, text feedback as in FlowSteer, and the value estimate but notffcosts 3\.91 OOD points, soffcarries information the text does not\. Fixed paradigms help decomposable questions but lag where results must be redone, and the best trails EvoSteer by 10\.26 IID points; even untrained, the architecture leads all four on the OOD average\. In\-run edits mostly help: 84\.2% raise the graded answer score and 6\.6% lower it \(Figure[6](https://arxiv.org/html/2609.38661#S5.F6)e\)\. Repairs, the most frequent edit, raise the score in 80\.7% of cases, and value\-guided replanning helps both executors on every IID dataset, most on the interactive ALFWorld \(Figure[6](https://arxiv.org/html/2609.38661#S5.F6)f\)\.

AnchorTB \(Section 4\.2\)\.Measured flows are the largest single contribution on IID: a free\-scalaruqu\_\{q\}costs 6\.59 IID points, and removingc⁡\(g\)c\(g\)andbψb\_\{\\psi\}costs 5\.29\. The anchor is informative on its own, predicting reference outcomes before a rollout starts with AUC 0\.812 against 0\.506 for the task\-family mean \(Figure[6](https://arxiv.org/html/2609.38661#S5.F6)c\), and the learned part lowers the loss below the anchor\-only level \(Figure[5](https://arxiv.org/html/2609.38661#S5.F5)b\)\. With only the loss changed, AnchorTB beats tempered TB\([Zhang et al\., 2026c](https://arxiv.org/html/2609.38661#bib.bib64)\), TB\([Malkin et al\., 2022](https://arxiv.org/html/2609.38661#bib.bib31)\), PPO\([Schulman et al\., 2017](https://arxiv.org/html/2609.38661#bib.bib35)\), and GRPO by 3\.17–8\.81 IID points \(Figure[6](https://arxiv.org/html/2609.38661#S5.F6)a,b\); TB lags most on MuSiQue, SWE\-Bench, and WebShop\. AnchorTB needs 28% less update compute than tempered TB, at a rollout cost within 9% of PPO and GRPO, so its gains come at comparable compute\.

Validated admission \(Section 4\.3\)\.Removing skill evolution costs 5\.27 IID and 3\.26 OOD points, most on MBPP\+ \(7\.66\)\. A fixed success count in place of the sequential test costs 5\.17 IID points, nearly the same for any count from 1 to 10 \(Appendix[D](https://arxiv.org/html/2609.38661#A4)\), and even trails the no\-skill variant on MedQA, AIME 2026, and ALFWorld\. Accepting every candidate helps on average \(\+2\.76\) but hurts MedQA and AIME 2026, whereas validated admission helps all six IID datasets \(\+5\.95; Figure[6](https://arxiv.org/html/2609.38661#S5.F6)d\)\.

Architecture and training\.The untrained architecture adds 21\.01 IID and 24\.04 OOD points to the frozen executor, and AnchorTB training adds 12\.31 and 11\.22, so each supplies much of the gain\.

## 6Conclusion

EvoSteer unifies construction, repair, and skill growth in one learned loop: each action executes as it is issued, AnchorTB credits it against a frozen reference, and paired sequential tests decide which skills are kept\. It beats all baselines on twelve benchmarks and improves every executor backbone\.

## AI Use Statement

In this work, generative AI tools were used to edit and polish the text for clarity and readability, assist withLaTeXtypesetting and page layout, and adjust the placement and sizing of existing figures and tables\. They were also used for retrieval and discovery, to help search for and identify related work; all bibliographic entries were taken from Google Scholar\. The authors manually reviewed and verified all AI\-assisted text, figures, numerical values, citations, and formatting against the underlying implementation and experimental records\. The authors take full responsibility for the final content of the paper and all artifacts produced with the assistance of generative AI\.

## Ethics Statement

This work follows the ICLR Code of Ethics\. We aim to conduct and report our research with scientific integrity, transparency, and reproducibility\. Experimental results are reported without fabrication, falsification, or intentional misrepresentation, and sufficient implementation and evaluation details are provided to facilitate verification and reproduction\. We acknowledge prior work and use existing benchmarks, models, and associated resources in accordance with their intended research purposes and applicable licenses\. This study involves no human participants or personal information\.

## Reproducibility Statement

We support reproducibility by documenting the orchestration environment, action space, and value estimate in Section 4\.1, the AnchorTB objective in Section 4\.2, and the validated skill admission procedure in Section 4\.3\. Section 5\.1 specifies the benchmarks, baselines, and evaluation setting\. The appendix provides the assumptions and complete proofs of all theoretical results \(Appendices[A](https://arxiv.org/html/2609.38661#A1)–[C](https://arxiv.org/html/2609.38661#A3)\), the training algorithm and all hyperparameters \(Appendix[C](https://arxiv.org/html/2609.38661#A3), Table[3](https://arxiv.org/html/2609.38661#A3.T3)\), and the benchmark splits, metrics, baseline settings, evaluation protocol, and compute \(Appendix[D](https://arxiv.org/html/2609.38661#A4)\)\. All test scores are five\-run means, and Table[1](https://arxiv.org/html/2609.38661#S5.T1)also reports their standard deviations\.Code availability\.Our code, the exact training and test splits, and the full configuration are publicly available at[https://github\.com/beita6969/evosteer](https://github.com/beita6969/evosteer)\.

## References

- Bonagiri et al\. \(2026\)Akash Bonagiri, Devang Borkar, Gerard Janno Anderias, Setareh Rafatirad, and Houman Homayoun\.Causalflow: Causal attribution and counterfactual repair for llm agent failures\.*arXiv preprint arXiv:2605\.25338*, 2026\.
- Cemri et al\. \(2026\)Mert Cemri, Melissa Z Pan, Shuyi Yang, Lakshya A Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, et al\.Why do multi\-agent llm systems fail?*Advances in Neural Information Processing Systems*, 38, 2026\.
- Chen et al\. \(2026a\)Kunfeng Chen, Qihuang Zhong, Juhua Liu, and Bo Du\.Skillcat: Contrastive assessment and topology\-aware skill self\-evolution for llm agents\.*arXiv preprint arXiv:2606\.13317*, 2026a\.
- Chen et al\. \(2026b\)Xudong Chen, Yixin Liu, Hua Wei, and Kaize Ding\.Lemon: Learning executable multi\-agent orchestration via counterfactual reinforcement learning\.*arXiv preprint arXiv:2605\.14483*, 2026b\.
- Chen et al\. \(2026c\)Yanjun Chen, Yirong Sun, Hanlin Wang, Jinghan Wang, Xinming Zhang, Xiaoyu Shen, Wenjie Li, and Wei Zhang\.Exact is easier: Credit assignment for cooperative llm agents\.*arXiv preprint arXiv:2603\.06859*, 2026c\.
- Dang et al\. \(2026\)Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan, Zihao Xie, Ruijie Shi, Weize Chen, Cheng Yang, Xiaoyin Che, Ye Tian, et al\.Multi\-agent collaboration via evolving orchestration\.*Advances in neural information processing systems*, 38:165025–165059, 2026\.
- Deshmukh et al\. \(2026\)Shripad Deshmukh, Jayakumar Subramanian, Raghavendra Addanki, and Nikos Vlassis\.Cosac: Counterfactual credit assignment in sequential cooperative teams\.*arXiv preprint arXiv:2604\.17693*, 2026\.
- Fawkes & Hartford \(2026\)Jake Fawkes and Jason Hartford\.ff\-trajectory balance: A loss family for tuning gflownets, generative models, and llms with off\-and on\-policy data\.*arXiv preprint arXiv:2605\.15417*, 2026\.
- Gao et al\. \(2026\)Haowen Gao, Haoran Chen, Can Wang, Shasha Guo, Liang Pang, Zhaoyang Liu, Huawei Shen, and Xueqi Cheng\.Skillaudit: Ground\-truth\-free skill evolution via paired trajectory auditing\.*arXiv preprint arXiv:2606\.14239*, 2026\.
- Hao et al\. \(2026\)Zhezheng Hao, Tianfu Wang, Huanshuo Dong, Ziyan Liu, Hong Wang, Xiankun Lin, Qiang Lin, Can Wang, Hande Dong, and Jiawei Chen\.Evolve as a team: Collaborative self\-evolution for llm\-based multi\-agent systems\.*arXiv preprint arXiv:2605\.29790*, 2026\.
- He & Yang \(2026\)Yu He and Weikai Yang\.Skillcommit: Evolving agent skills through behaviorally validated scope expansion\.*arXiv preprint arXiv:2608\.15165*, 2026\.
- Hendrycks et al\. \(2021\)Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the math dataset\.*arXiv preprint arXiv:2103\.03874*, 2021\.
- Hong et al\. \(2024\)Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Steven Yau, Zijuan Lin, Liyang Zhou, et al\.Metagpt: Meta programming for a multi\-agent collaborative framework\.In*International Conference on Learning Representations*, volume 2024, pp\. 23247–23275, 2024\.
- Hu et al\. \(2026\)Wentao Hu, Zhendong Chu, Yiming Zhang, Junda Wu, Ming Jin, Xiangyu Zhao, Yilei Shao, Yanfeng Wang, and Qingsong Wen\.Skillbrew: Multi\-objective curation of skill banks for llm agents\.*arXiv preprint arXiv:2605\.29440*, 2026\.
- Jiang et al\. \(2026\)Yanze Jiang, Mingxuan Li, Yuhao Wang, Shengfang Zhai, and Jiaheng Zhang\.Don’t solve, just compare: Tiny advisors for runtime intervention in llm agents\.*arXiv preprint arXiv:2608\.21027*, 2026\.
- Jimenez et al\. \(2024\)Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan\.Swe\-bench: Can language models resolve real\-world github issues?In*International Conference on Learning Representations*, volume 2024, pp\. 54107–54157, 2024\.
- Jin et al\. \(2021\)Di Jin, Eileen Pan, Nassim Oufattole, Wei\-Hung Weng, Hanyi Fang, and Peter Szolovits\.What disease does this patient have? a large\-scale open domain question answering dataset from medical exams\.*Applied Sciences*, 11\(14\):6421, 2021\.
- Joshi et al\. \(2017\)Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer\.Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension\.In*Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 1601–1611, 2017\.
- Kwiatkowski et al\. \(2019\)Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al\.Natural questions: a benchmark for question answering research\.*Transactions of the Association for Computational Linguistics*, 7:453–466, 2019\.
- Li et al\. \(2026a\)Haoran Li, Shulun Chen, Shaoyuan Sun, and Hanchen Wang\.Multi\-agent coordination adaptation via structure\-guided orchestration\.*arXiv preprint arXiv:2605\.25746*, 2026a\.
- Li & Ramakrishnan \(2026\)Sha Li and Naren Ramakrishnan\.Experience as a compass: Multi\-agent rag with evolving orchestration and agent prompts\.*arXiv preprint arXiv:2604\.00901*, 2026\.
- Li et al\. \(2026b\)Zhongyi Li, Wan Tian, Jinju Chen, Huiming Zhang, Yang Liu, Yikun Ban, and Fuzhen Zhuang\.Counterfactual credit policy optimization for multi\-agent collaboration\.*arXiv preprint arXiv:2603\.21563*, 2026b\.
- Li et al\. \(2026c\)Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Y Zou, and Pan Lu\.In\-the\-flow agentic system optimization for effective planning and tool use\.In*International Conference on Learning Representations*, volume 2026, pp\. 50524–50570, 2026c\.
- Li et al\. \(2026d\)Zongyue Li, Chengyue Yu, Lei Zang, Chenyi Zhuang, Linjian Mo, and Leilei Gan\.Last step matters: Early uncertainty cannot predict failure in long\-horizon agents\.*arXiv preprint arXiv:2608\.29685*, 2026d\.
- Liang et al\. \(2026\)Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang, Yingzong Min, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, et al\.Granularity\-adaptive credit assignment for long\-horizon llm agent reinforcement learning\.*arXiv preprint arXiv:2609\.12424*, 2026\.
- Liu et al\. \(2023\)Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang\.Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation\.*Advances in neural information processing systems*, 36:21558–21572, 2023\.
- Liu et al\. \(2026\)Xiaodong Liu, Michael Xu, Jack W Stokes, Paul Smolensky, Doug Burger, and Jianfeng Gao\.Gflowrl: Scaling distribution\-matching rl to large language models\.*arXiv preprint arXiv:2607\.13394*, 2026\.
- Lu & Zhang \(2026\)Hanxiao Lu and Tianyi Zhang\.Autonomous repair for multi\-agent systems via monte\-carlo tree search\.*arXiv preprint arXiv:2607\.29055*, 2026\.
- Luan et al\. \(2026\)Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, and Xiaohong Chen\.Repair or resample? rethinking failure debugging in llm multi\-agent systems\.*arXiv preprint arXiv:2608\.25920*, 2026\.
- Madan et al\. \(2023\)Kanika Madan, Jarrid Rector\-Brooks, Maksym Korablyov, Emmanuel Bengio, Moksh Jain, Andrei Cristian Nica, Tom Bosc, Yoshua Bengio, and Nikolay Malkin\.Learning gflownets from partial episodes for improved convergence and stability\.In*International Conference on Machine Learning*, pp\. 23467–23483\. PMLR, 2023\.
- Malkin et al\. \(2022\)Nikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun, and Yoshua Bengio\.Trajectory balance: Improved credit assignment in gflownets\.*Advances in Neural Information Processing Systems*, 35:5955–5967, 2022\.
- Mishra et al\. \(2026\)Amritansh Mishra, Supriyo Chakraborty, and Berkcan Kapusuzoglu\.On the policy gradient foundations of group relative policy optimization: Credit assignment, gradient sparsity, and rank collapse\.*arXiv preprint arXiv:2606\.29238*, 2026\.
- Pan et al\. \(2026\)Shuai Pan, Yixiang Liu, Jiaye Gao, Te Gao, Weiwen Liu, Jianghao Lin, Zhihui Fu, Jun Wang, Weinan Zhang, and Yong Yu\.Skillmas: Skill co\-evolution with llm\-based multi\-agent system\.*arXiv preprint arXiv:2605\.09341*, 2026\.
- Rein et al\. \(2023\)David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman\.Gpqa: A graduate\-level google\-proof q&a benchmark\.*arXiv preprint arXiv:2311\.12022*, 2023\.
- Schulman et al\. \(2017\)John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov\.Proximal policy optimization algorithms\.*arXiv preprint arXiv:1707\.06347*, 2017\.
- Sengupta \(2026\)Biswa Sengupta\.Self\-evolving agents with anytime\-valid certificates\.*arXiv preprint arXiv:2607\.00871*, 2026\.
- Shah \(2026\)Jaineet Shah\.Causal agent replay: Counterfactual attribution for llm\-agent failures\.*arXiv preprint arXiv:2606\.08275*, 2026\.
- Shang & Yang \(2026\)Fangxin Shang and Yehui Yang\.Hypothesis\-driven skill optimization for llm agents\.*arXiv preprint arXiv:2606\.22330*, 2026\.
- Shang et al\. \(2026\)Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, and Ning Zheng\.When self\-evolution backfires: Pre\-commit gating against skill contamination in llm agents\.*arXiv preprint arXiv:2608\.05810*, 2026\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al\.Deepseekmath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Shawn \(2026\)Zayx Shawn\.Pace: Anytime\-valid acceptance tests for self\-evolving agents\.*arXiv preprint arXiv:2606\.08106*, 2026\.
- Shen et al\. \(2026\)Junhao Shen, Teng Zhang, Xiaoyan Zhao, and Hong Cheng\.Dynamic skill lifecycle management for agentic reinforcement learning\.*arXiv preprint arXiv:2605\.10923*, 2026\.
- Shridhar et al\. \(2020\)Mohit Shridhar, Xingdi Yuan, Marc\-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht\.Alfworld: Aligning text and embodied environments for interactive learning\.*arXiv preprint arXiv:2010\.03768*, 2020\.
- Trivedi et al\. \(2022\)Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal\.♪ musique: Multihop questions via single\-hop question composition\.*Transactions of the Association for Computational Linguistics*, 10:539–554, 2022\.
- Venkatraman et al\. \(2024\)Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, et al\.Amortizing intractable inference in diffusion models for vision, language, and control\.*Advances in neural information processing systems*, 37:76080–76114, 2024\.
- Wang et al\. \(2026a\)Chenyu Wang, Yunbo Lyu, Junda He, Zhou Yang, Chenxing Zhong, Yaniv Harel, and David Lo\.Fail\-fast, restart\-smart: Early failure prediction and restart for swe agentic tasks\.*arXiv preprint arXiv:2608\.03222*, 2026a\.
- Wang et al\. \(2023\)Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar\.Voyager: An open\-ended embodied agent with large language models\.*arXiv preprint arXiv:2305\.16291*, 2023\.
- Wang et al\. \(2026b\)Tao Wang, Suhang Zheng, and Xiaoxiao Xu\.Rtmc: Step\-level credit assignment via rollout trees\.*arXiv preprint arXiv:2604\.11037*, 2026b\.
- Wang et al\. \(2026c\)Yixuan Wang, Yiyang Zhou, Yiming Liang, Congyu Zhang, Fuxiao Liu, Jiawei Zhou, and Huaxiu Yao\.Not all skills help: Measuring and repairing agent knowledge\.*arXiv preprint arXiv:2606\.15390*, 2026c\.
- Wu et al\. \(2026\)Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Wenjie Zhang, Zhichao Shi, Xuhui Jiang, Chengjin Xu, Jia Li, and Jian Guo\.Bayesian\-agent: Posterior\-guided skill evolution for llm agent harnesses\.*arXiv preprint arXiv:2606\.08348*, 2026\.
- Xu et al\. \(2026\)Chen Xu, Yicheng Hu, Ruizi Wang, Xinyu Lin, Wenjie Wang, Dongrui Liu, and Fuli Feng\.Tacomas: Test\-time co\-evolution of topology and capability in llm\-based multi\-agent systems\.*arXiv preprint arXiv:2605\.09539*, 2026\.
- Yang et al\. \(2026a\)Shidong Yang, Ziyu Ma, Tongwen Huang, Xucong Wang, Renda Li, Yiming Hu, Yong Wang, and Xiangxiang Chu\.Skillforge: Evolving verifiable skills for reinforcement learning agents\.*arXiv preprint arXiv:2608\.24747*, 2026a\.
- Yang et al\. \(2026b\)Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, Yan Li, Xuemei Gao, Qi Dai, Bei Liu, et al\.Skillopt: Executive strategy for self\-evolving agent skills\.*arXiv preprint arXiv:2605\.23904*, 2026b\.
- Yang et al\. \(2018\)Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning\.Hotpotqa: A dataset for diverse, explainable multi\-hop question answering\.In*Proceedings of the 2018 conference on empirical methods in natural language processing*, pp\. 2369–2380, 2018\.
- Yao et al\. \(2026\)Huaiyuan Yao, Xiaoou Liu, Charles Fleming, Tianlong Chen, and Hua Wei\.Maskills: Continual skills optimization for multi\-agent llm systems\.*arXiv preprint arXiv:2609\.02094*, 2026\.
- Yao et al\. \(2022a\)Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan\.Webshop: Towards scalable real\-world web interaction with grounded language agents\.*Advances in Neural Information Processing Systems*, 35:20744–20757, 2022a\.
- Yao et al\. \(2022b\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao\.React: Synergizing reasoning and acting in language models\.*arXiv preprint arXiv:2210\.03629*, 2022b\.
- Zhang et al\. \(2026a\)Boxuan Zhang, Jianing Zhu, Zeru Shi, Dongfang Liu, and Ruixiang Tang\.Agentforesight: Online auditing for early failure prediction in multi\-agent systems\.*arXiv preprint arXiv:2605\.08715*, 2026a\.
- Zhang \(2026\)Chenchen Zhang\.Reinforcement learning for llm\-based multi\-agent systems through orchestration traces\.*arXiv preprint arXiv:2605\.02801*, 2026\.
- Zhang et al\. \(2024\)Guibin Zhang, Yanwei Yue, Xiangguo Sun, Guancheng Wan, Miao Yu, Junfeng Fang, Kun Wang, Tianlong Chen, and Dawei Cheng\.G\-designer: Architecting multi\-agent communication topologies via graph neural networks\.*arXiv preprint arXiv:2410\.11782*, 2024\.
- Zhang et al\. \(2025a\)Guibin Zhang, Yanwei Yue, Zhixun Li, Sukwon Yun, Guancheng Wan, Kun Wang, Dawei Cheng, Jeffrey Yu, and Tianlong Chen\.Cut the crap: An economical communication pipeline for llm\-based multi\-agent systems\.In*International Conference on Learning Representations*, volume 2025, pp\. 75389–75428, 2025a\.
- Zhang et al\. \(2025b\)Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, Jinlin Wang, et al\.Aflow: Automating agentic workflow generation\.In*International Conference on Learning Representations*, volume 2025, pp\. 34040–34077, 2025b\.
- Zhang et al\. \(2026b\)Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Qika Lin, Rui Mao, Erik Cambria, Xiaoying Tang, and Haoran Luo\.Flowsteer: Towards agents designing agentic workflows via reinforced progressive canvas editing\.*arXiv preprint arXiv:2602\.01664*, 2026b\.
- Zhang et al\. \(2026c\)Mingda Zhang, Tiesunlong Shen, Haoran Luo, Wenjin Liu, Zikai Xiao, Erik Cambria, and Xiaoying Tang\.Skillflow: Flow\-driven recursive skill evolution for agentic orchestration\.*arXiv preprint arXiv:2605\.14089*, 2026c\.
- Zhang & Li \(2026\)Yan Zhang and Shibo Li\.Consistencygate: Preventing memory contamination in llm agents via self\-consistency admission control\.*arXiv preprint arXiv:2607\.22962*, 2026\.
- Zhang et al\. \(2026d\)Yaolun Zhang, Tianyi Xu, Shengyu Dai, Zhenwen Shao, Qingyun Wu, and Huazheng Wang\.Evochamber: Test\-time co\-evolution of multi\-agent system at individual, team, and population scales\.*arXiv preprint arXiv:2605\.11136*, 2026d\.
- Zhao et al\. \(2026\)Chenyu Zhao, Shenglin Zhang, Wenwei Gu, Yongqian Sun, Dan Pei, Chetan Bansal, Saravan Rajmohan, and Minghua Ma\.Agenttether: Graph\-guided diagnosis and runtime intervention for reliable llm agent operation\.*arXiv preprint arXiv:2607\.06273*, 2026\.
- Zhu et al\. \(2026\)Junze Zhu, Weihao Chen, Xuanwang Zhang, Zhen Wu, and Xinyu Dai\.Recognize your orchestrator: An entropy dynamics perspective for llm multi\-agent systems\.*arXiv preprint arXiv:2606\.01351*, 2026\.
- Zhuge et al\. \(2024\)Mingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, and Jürgen Schmidhuber\.Language agents as optimizable graphs\.*arXiv preprint arXiv:2402\.16823*, 2024\.

## Appendix AAnchored Trajectory Balance

This appendix treats three objects in turn: the reward\-tilted history law, the regression loss evaluated on recorded trajectories, and the statistical procedure for skill validation\. We first establish structural and distributional properties conditional on a frozen environment context, then analyze the implemented trajectory\-indexed anchors and paired tests\. Algebraic identities hold on every recorded trajectory, and each distributional or statistical result holds under the assumptions it states\.

### A\.1Formal Setting, Histories, and Legal Actions

###### Definition A\.1\(Context and complete\-history representation\)\.

Fix a task and a rollout\-batch contextCC\. The context includes the task, tools, role catalogue, executor configuration, skill\-menu and value\-head snapshots, and rules that generate legal masks\. A statests\_\{t\}retains the complete ordered history, including the current communication graph, execution records, features, and budget information\. Withs0s\_\{0\}the initial state,

st\+1=st⊕\(at,otexec,ft\+1\),at:st⟶st\+1,0≤t<T,s\_\{t\+1\}=s\_\{t\}\\oplus\(a\_\{t\},o\_\{t\}^\{\\mathrm\{exec\}\},f\_\{t\+1\}\),\\qquad a\_\{t\}:s\_\{t\}\\longrightarrow s\_\{t\+1\},\\qquad 0\\leq t<T,\(A\.1\)whereft\+1=f⁡\(st,at,otexec\)f\_\{t\+1\}=f\(s\_\{t\},a\_\{t\},o\_\{t\}^\{\\mathrm\{exec\}\}\)is the feature record produced by actionata\_\{t\}\. Thusf⁡\(st\+1\)=ft\+1f\(s\_\{t\+1\}\)=f\_\{t\+1\}denotes the latest feature record; the initial state has its interface\-specified initial features\. The displayed value estimate and its change are deterministic functions of this retained history and the frozen head inCC; they add no transition randomness\. A complete history isx=\(s0,a0,s1,…,aT−1,sT\)x=\(s\_\{0\},a\_\{0\},s\_\{1\},\\ldots,a\_\{T\-1\},s\_\{T\}\), identified with its terminal statesTs\_\{T\}, which retains the whole record\. The finalstoptransition is included inT≥1T\\geq 1\. Writex⪰sx\\succeq swhenxxextends prefixss, and let\|s\|\|s\|count appended action records\.

###### Definition A\.2\(Legal actions and fixed executor\)\.

Let𝒜C​\(s\)\\mathcal\{A\}\_\{C\}\(s\)be the nonempty legal action set at a nonterminal history\. It is a state\-dependent subset of the batch\-frozen action vocabulary\. Let𝖪C​\(s′∣s,a\)\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)be the normalized kernel of execution, tool observations, and history updates after a legal action\. For a normalized history\-dependent action policypp, its joint successor kernel isp⁡\(a∣s,C\)​𝖪C​\(s′∣s,a\)p\(a\\mid s,C\)\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)\. The executor is fixed: this conditional kernel is the same for everypp, and its outcomes may be stochastic and history\-dependent\.

###### Lemma A\.1\(Strict growth and unique parentage\)\.

Under Definition[A\.1](https://arxiv.org/html/2609.38661#A1.Thmdefinition1), the history graph is a rooted tree: each edge increases\|s\|\|s\|by one and every non\-root state has a unique parent\. Its backward policy on parents is therefore identically one, so no backward model needs to be learned\.

Equation \([A\.1](https://arxiv.org/html/2609.38661#A1.E1)\) appends exactly one action record, so\|st\+1\|=\|st\|\+1\|s\_\{t\+1\}\|=\|s\_\{t\}\|\+1\. A directed cycle would strictly increase this integer and return to its original value, a contradiction\. Deleting the last action, observation, and feature record from a non\-root history recovers its preceding prefix\. Any parent must produce precisely this last appended record and the same preceding ordered history, so no second parent exists\. A normalized distribution on this singleton parent set has probability one\. Returning to an earlier topology keeps the intervening records, so the history still grows\. ∎

###### Assumption A\.1\(Support, normalization, and termination\)\.

The serialized history space is finite or countable\. At every reference\-reachable nonterminal prefix,ρ\(⋅∣s,C\)\\rho\(\\cdot\\mid s,C\)is normalized and positive on𝒜C​\(s\)\\mathcal\{A\}\_\{C\}\(s\)\. A policy compared to it in a log\-ratio has the same positive action support and uses the same𝖪C\\mathsf\{K\}\_\{C\}\. From every such prefix, both continuation processes reach a terminal history almost surely, and every terminal record has a rewardr⁡\(x\)∈\[0,1\]r\(x\)\\in\[0,1\], with0≤β<∞0\\leq\\beta<\\infty\.

###### Proposition A\.2\(Normalized terminal and prefix laws\)\.

Under Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1), terminal histories form a probability\-one partition forp∈\{ρ,πθ\}p\\in\\\{\\rho,\\pi\_\{\\theta\}\\\}\. In particular,

∑x∈𝒳CPp​\(x∣C\)=1,Pp​\(s∣C\)=∑x⪰sPp​\(x∣C\),Pp​\(s0∣C\)=1,\\sum\_\{x\\in\\mathcal\{X\}\_\{C\}\}P\_\{p\}\(x\\mid C\)=1,\\qquad P\_\{p\}\(s\\mid C\)=\\sum\_\{x\\succeq s\}P\_\{p\}\(x\\mid C\),\\qquad P\_\{p\}\(s\_\{0\}\\mid C\)=1,\(A\.2\)wherePpP\_\{p\}denotes the history law induced byppand the executor\.

Normalized local kernels define the successive action and observation probabilities\. Different terminal histories describe disjoint events: a terminal record cannot be a proper prefix of a later executed history\. Almost\-sure termination makes their union a probability\-one event\. The same reasoning, restricted to the event of reachingss, partitions that event by its terminal descendants\. Countable additivity proves both sums; the root is reached with probability one\. ∎

### A\.2Path Laws and Positive Affine Tilting

###### Lemma A\.3\(Path factorization and kernel cancellation\)\.

Under Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1),

Pp​\(x∣C\)=∏t=0T−1p⁡\(at∣st,C\)​𝖪C​\(st\+1∣st,at\)\.P\_\{p\}\(x\\mid C\)=\\prod\_\{t=0\}^\{T\-1\}p\(a\_\{t\}\\mid s\_\{t\},C\)\\mathsf\{K\}\_\{C\}\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\}\)\.\(A\.3\)For a supported continuation ofs=sms=s\_\{m\}, defineΔ​ℓt=log⁡πθ​\(at∣st,C\)−log⁡ρ⁡\(at∣st,C\)\\Delta\\ell\_\{t\}=\\log\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\},C\)\-\\log\\rho\(a\_\{t\}\\mid s\_\{t\},C\)\. Then

Pθ​\(x∣s,C\)Pρ​\(x∣s,C\)=∏t=mT−1πθ​\(at∣st,C\)ρ⁡\(at∣st,C\)=exp⁡\(∑t=mT−1Δ​ℓt\)\.\\frac\{P\_\{\\theta\}\(x\\mid s,C\)\}\{P\_\{\\rho\}\(x\\mid s,C\)\}=\\prod\_\{t=m\}^\{T\-1\}\\frac\{\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\},C\)\}\{\\rho\(a\_\{t\}\\mid s\_\{t\},C\)\}=\\exp\\\!\\left\(\\sum\_\{t=m\}^\{T\-1\}\\Delta\\ell\_\{t\}\\right\)\.\(A\.4\)HerePθ=PπθP\_\{\\theta\}=P\_\{\\pi\_\{\\theta\}\}andPρ​\(x∣C\)P\_\{\\rho\}\(x\\mid C\)is the quantity writtenρ⁡\(x∣C\)\\rho\(x\\mid C\)in Eq\. \([2](https://arxiv.org/html/2609.38661#S3.E2)\)\.

Apply the probability chain rule, conditioning each action on the history before it and each observation on that history and action\. This gives Eq\. \([A\.3](https://arxiv.org/html/2609.38661#A1.E3)\); starting atsms\_\{m\}gives the corresponding suffix product\. On the common support every executor factor in numerator and denominator is equal and positive, so it cancels\. The remaining product consists of the action ratios, whose logarithm is the displayed sum\. This cancellation compares two policies on the same recorded history\. ∎

###### Lemma A\.4\(Masked token likelihoods\)\.

Suppose legal actions have unique finite token serializations, including any required action\-ending symbol, and masked token generation terminates almost surely\. Use the same legal token sets for both policies, with each policy normalized on each such set\. The sum of executed token log\-probabilities is the log\-probability of that action\. Singleton legal sets contribute zero to its log\-ratio, since both policies give them probability one\.

For actiona=\(a1,…,aK\)a=\(a\_\{1\},\\ldots,a\_\{K\}\), the autoregressive chain rule gives

p⁡\(a∣s,C\)=∏j=1Kp⁡\(aj∣s,a<j,C\),p\(a\\mid s,C\)=\\prod\_\{j=1\}^\{K\}p\(a\_\{j\}\\mid s,a\_\{<j\},C\),with any termination probability included in the serialization\. Taking logarithms gives the sum, and taking the difference for the two policies gives the action log\-ratio\. If a legal token set is a singleton, both normalized probabilities equal one\. Its two log\-probabilities are zero\. Each model’s normalization constant on the common mask remains part of its normalized probability\. ∎

###### Proposition A\.5\(Affine tilt, normalizer, and ideal reward gain\)\.

Letμ=𝔼ρ​\[r⁡\(X\)∣C\]\\mu=\\mathbb\{E\}\_\{\\rho\}\[r\(X\)\\mid C\]andaβ=eβ−1a\_\{\\beta\}=e^\{\\beta\}\-1\. The target in Eq\. \([2](https://arxiv.org/html/2609.38661#S3.E2)\) is normalized, has the same terminal support asPρP\_\{\\rho\}, and satisfies

1≤Rβ​\(x\)≤eβ,Z⁡\(C\)=1\+aβ​μ∈\[1,eβ\],\\displaystyle 1\\leq R\_\{\\beta\}\(x\)\\leq e^\{\\beta\},\\qquad Z\(C\)=1\+a\_\{\\beta\}\\mu\\in\[1,e^\{\\beta\}\],\(A\.5\)𝔼P∗​\[r⁡\(X\)∣C\]−μ=aβ​Varρ​\(r⁡\(X\)∣C\)1\+aβ​μ≥0\.\\displaystyle\\mathbb\{E\}\_\{P^\{\*\}\}\[r\(X\)\\mid C\]\-\\mu=\\frac\{a\_\{\\beta\}\\operatorname\{Var\}\_\{\\rho\}\(r\(X\)\\mid C\)\}\{1\+a\_\{\\beta\}\\mu\}\\geq 0\.\(A\.6\)Atβ=0\\beta=0,P∗=PρP^\{\*\}=P\_\{\\rho\}\. Forβ\>0\\beta\>0, strict gain in Eq\. \([A\.6](https://arxiv.org/html/2609.38661#A1.E6)\) occurs exactly when the reference reward has positive variance, which holds whenever the reference both succeeds and fails\.

The reward range follows fromr∈\[0,1\]r\\in\[0,1\]andaβ≥0a\_\{\\beta\}\\geq 0\. SummingPρ​\(x\)​\(1\+aβ​r​\(x\)\)P\_\{\\rho\}\(x\)\(1\+a\_\{\\beta\}r\(x\)\)givesZ=1\+aβ​μZ=1\+a\_\{\\beta\}\\mu; it is finite and positive, so dividing byZZnormalizes the target and keeps its support\. Moreover,

𝔼P∗​r=𝔼ρ​r\+aβ​𝔼ρ​r21\+aβ​μ\.\\mathbb\{E\}\_\{P^\{\*\}\}r=\\frac\{\\mathbb\{E\}\_\{\\rho\}r\+a\_\{\\beta\}\\mathbb\{E\}\_\{\\rho\}r^\{2\}\}\{1\+a\_\{\\beta\}\\mu\}\.Subtractingμ\\mugivesaβ​\(𝔼ρ​r2−μ2\)/\(1\+aβ​μ\)a\_\{\\beta\}\(\\mathbb\{E\}\_\{\\rho\}r^\{2\}\-\\mu^\{2\}\)/\(1\+a\_\{\\beta\}\\mu\), and the equality and strictness statements follow froma0=0a\_\{0\}=0andaβ\>0a\_\{\\beta\}\>0forβ\>0\\beta\>0\. ∎

### A\.3Conditional Relative Flows and Uniqueness

###### Definition A\.3\(Unnormalized target flow and relative flow\)\.

ForPρ​\(s∣C\)\>0P\_\{\\rho\}\(s\\mid C\)\>0, define

F∗​\(s∣C\)=∑x⪰sPρ​\(x∣C\)​Rβ​\(x\),U∗​\(s∣C\)=F∗​\(s∣C\)Pρ​\(s∣C\),F^\{\*\}\(s\\mid C\)=\\sum\_\{x\\succeq s\}P\_\{\\rho\}\(x\\mid C\)R\_\{\\beta\}\(x\),\\qquad U^\{\*\}\(s\\mid C\)=\\frac\{F^\{\*\}\(s\\mid C\)\}\{P\_\{\\rho\}\(s\\mid C\)\},\(A\.7\)andu∗​\(s∣C\)=log⁡U∗​\(s∣C\)u^\{\*\}\(s\\mid C\)=\\log U^\{\*\}\(s\\mid C\)\. ThusF∗F^\{\*\}is unnormalized\. The normalized target prefix probability isP∗​\(s∣C\)=F∗​\(s∣C\)/Z⁡\(C\)P^\{\*\}\(s\\mid C\)=F^\{\*\}\(s\\mid C\)/Z\(C\), soP∗​\(s∣C\)/Pρ​\(s∣C\)=U∗​\(s∣C\)/Z⁡\(C\)P^\{\*\}\(s\\mid C\)/P\_\{\\rho\}\(s\\mid C\)=U^\{\*\}\(s\\mid C\)/Z\(C\)\.

###### Lemma A\.6\(Prefix partition and conditional recursion\)\.

Under Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1), at every supported prefix,

U∗\(s∣C\)=𝔼ρ\[Rβ\(X\)∣s,C\]=1\+aβ𝔼ρ\[r\(X\)∣s,C\]∈\[1,eβ\]\.U^\{\*\}\(s\\mid C\)=\\mathbb\{E\}\_\{\\rho\}\[R\_\{\\beta\}\(X\)\\mid s,C\]=1\+a\_\{\\beta\}\\mathbb\{E\}\_\{\\rho\}\[r\(X\)\\mid s,C\]\\in\[1,e^\{\\beta\}\]\.\(A\.8\)For nonterminalss, writingCh⁡\(s\)\\operatorname\{Ch\}\(s\)for its supported children,

F∗​\(s\)\\displaystyle F^\{\*\}\(s\)=∑s′∈Ch⁡\(s\)F∗​\(s′\),\\displaystyle=\\sum\_\{s^\{\\prime\}\\in\\operatorname\{Ch\}\(s\)\}F^\{\*\}\(s^\{\\prime\}\),\(A\.9\)U∗​\(s\)\\displaystyle U^\{\*\}\(s\)=∑a∈𝒜C​\(s\)ρ⁡\(a∣s,C\)​∑s′𝖪C​\(s′∣s,a\)​U∗​\(s′\)\.\\displaystyle=\\sum\_\{a\\in\\mathcal\{A\}\_\{C\}\(s\)\}\\rho\(a\\mid s,C\)\\sum\_\{s^\{\\prime\}\}\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)U^\{\*\}\(s^\{\\prime\}\)\.\(A\.10\)The boundaries areU∗​\(x\)=Rβ​\(x\)U^\{\*\}\(x\)=R\_\{\\beta\}\(x\)andU∗​\(s0\)=Z⁡\(C\)U^\{\*\}\(s\_\{0\}\)=Z\(C\)\.

Divide the descendant sum definingF∗​\(s\)F^\{\*\}\(s\)by the positive prefix probability\. By Eq\. \([A\.2](https://arxiv.org/html/2609.38661#A1.E2)\), the resulting weights are the normalized reference continuation law, yielding Eq\. \([A\.8](https://arxiv.org/html/2609.38661#A1.E8)\)\. Descendants of a nonterminalsspartition by their unique first child, proving Eq\. \([A\.9](https://arxiv.org/html/2609.38661#A1.E9)\)\. Each child retains its incoming action, soPρ​\(s′\)=Pρ​\(s\)​ρ​\(a∣s,C\)​𝖪C​\(s′∣s,a\)P\_\{\\rho\}\(s^\{\\prime\}\)=P\_\{\\rho\}\(s\)\\rho\(a\\mid s,C\)\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)\. SubstituteF∗​\(s′\)=Pρ​\(s′\)​U∗​\(s′\)F^\{\*\}\(s^\{\\prime\}\)=P\_\{\\rho\}\(s^\{\\prime\}\)U^\{\*\}\(s^\{\\prime\}\)in the child sum and divide byPρ​\(s\)P\_\{\\rho\}\(s\)to obtain Eq\. \([A\.10](https://arxiv.org/html/2609.38661#A1.E10)\)\. No additional transition weight belongs in the sum ofF∗​\(s′\)F^\{\*\}\(s^\{\\prime\}\), since it already includes the probability of reaching the child\. A terminal has only itself as a terminal descendant; the root descendant sum isZZ\. ∎

###### Proposition A\.7\(Uniqueness with an appropriate horizon condition\)\.

On a history tree with a uniform finite action horizonHH,U∗U^\{\*\}is the unique finite\-valued solution of Eq\. \([A\.10](https://arxiv.org/html/2609.38661#A1.E10)\) with terminal boundaryV​\(x\)=Rβ​\(x\)V\(x\)=R\_\{\\beta\}\(x\)\. On a possibly infinite\-depth tree with almost\-sure reference termination from every supported prefix,U∗U^\{\*\}is the unique*bounded*solution with this boundary\.

For the finite\-horizon statement, terminal values coincide\. SupposeV=U∗V=U^\{\*\}at all supported successors of a nonterminalss\. Applying the same recursion to both functions gives equality atss\. Induct backward through the at mostHHremaining transitions\. This proves equality at every supported history, also for countably many children since their already identified bounded values have well\-defined weighted sums, so the backward induction goes through unchanged\.

For the infinite\-depth statement, start a reference continuation atss, and letτ\\taube its remaining termination time\. Extend the state process mathematically by retaining its terminal state afterτ\\tau\. The recursion and iterated conditional expectation giveV\(s\)=𝔼ρ\[V\(Sn∧τ\)∣s,C\]V\(s\)=\\mathbb\{E\}\_\{\\rho\}\[V\(S\_\{n\\wedge\\tau\}\)\\mid s,C\]for every integernn\. Almost\-sure termination impliesV⁡\(Sn∧τ\)→Rβ​\(X\)V\(S\_\{n\\wedge\\tau\}\)\\to R\_\{\\beta\}\(X\)\. Boundedness permits passage to the expectation, soV\(s\)=𝔼ρ\[Rβ\(X\)∣s,C\]=U∗\(s\)V\(s\)=\\mathbb\{E\}\_\{\\rho\}\[R\_\{\\beta\}\(X\)\\mid s,C\]=U^\{\*\}\(s\)\. The same argument holds at each supported prefix\. Equation \([A\.8](https://arxiv.org/html/2609.38661#A1.E8)\) already ensures thatU∗U^\{\*\}itself is bounded\. ∎

### A\.4Ideal Zero\-Residual Consistency

###### Definition A\.4\(Shared\-flow residual\)\.

A shared flow is a real\-valued functionu⁡\(s\)u\(s\)assigning one value to a prefix independently of which later continuation is scored\. Its terminal boundary isu⁡\(sT\)=log⁡Rβ​\(x\)u\(s\_\{T\}\)=\\log R\_\{\\beta\}\(x\)\. For0≤i<j≤T0\\leq i<j\\leq T, set

δi:ju\(x\)=u\(si\)\+∑t=ij−1Δℓt−u\(sj\),ℒxu=1KT∑i<j\(δi:ju\(x\)\)2,KT=T⁡\(T\+1\)2\.\\delta^\{u\}\_\{i:j\}\(x\)=u\(s\_\{i\}\)\+\\sum\_\{t=i\}^\{j\-1\}\\Delta\\ell\_\{t\}\-u\(s\_\{j\}\),\\qquad\\mathcal\{L\}\_\{x\}^\{u\}=\\frac\{1\}\{K\_\{T\}\}\\sum\_\{i<j\}\(\\delta^\{u\}\_\{i:j\}\(x\)\)^\{2\},\\quad K\_\{T\}=\\frac\{T\(T\+1\)\}\{2\}\.\(A\.11\)This is Eq\. \([7](https://arxiv.org/html/2609.38661#S4.E7)\) with a common state function; Appendix[B](https://arxiv.org/html/2609.38661#A2)treats the trajectory\-indexed estimate\.

###### Proposition A\.8\(Ideal reference\-relative consistency\)\.

Under Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1), suppose a shareduusatisfies the terminal boundary andδt:t\+1u\(x\)=0\\delta^\{u\}\_\{t:t\+1\}\(x\)=0on every supported complete history and every action in that history\. Then

eu⁡\(s\)=U∗​\(s∣C\),u⁡\(s0\)=log⁡Z⁡\(C\),Pθ​\(x∣C\)=P∗​\(x∣C\)e^\{u\(s\)\}=U^\{\*\}\(s\\mid C\),\\qquad u\(s\_\{0\}\)=\\log Z\(C\),\\qquad P\_\{\\theta\}\(x\\mid C\)=P^\{\*\}\(x\\mid C\)\(A\.12\)on the reference support\. Requiring zero residual for all subtrajectories is equivalent here to requiring it for all singleton intervals, so checking one\-step residuals suffices\.

Fix a supported prefixs=sms=s\_\{m\}\. Adding the singleton equalities along any supported complete continuation gives

∑t=mT−1Δ​ℓt=∑t=mT−1\(u⁡\(st\+1\)−u⁡\(st\)\)=log⁡Rβ​\(x\)−u⁡\(s\)\.\\sum\_\{t=m\}^\{T\-1\}\\Delta\\ell\_\{t\}=\\sum\_\{t=m\}^\{T\-1\}\(u\(s\_\{t\+1\}\)\-u\(s\_\{t\}\)\)=\\log R\_\{\\beta\}\(x\)\-u\(s\)\.By Lemma[A\.3](https://arxiv.org/html/2609.38661#A1.Thmproposition3),eu⁡\(s\)​Pθ​\(x∣s,C\)=Pρ​\(x∣s,C\)​Rβ​\(x\)e^\{u\(s\)\}P\_\{\\theta\}\(x\\mid s,C\)=P\_\{\\rho\}\(x\\mid s,C\)R\_\{\\beta\}\(x\)\. Sum over all terminal descendants\. The shared valueeu⁡\(s\)e^\{u\(s\)\}can be taken outside the sum, and the policy continuation probabilities sum to one by termination\. Hence

eu⁡\(s\)=∑x⪰sPρ​\(x∣s,C\)​Rβ​\(x\)=U∗​\(s∣C\)\.e^\{u\(s\)\}=\\sum\_\{x\\succeq s\}P\_\{\\rho\}\(x\\mid s,C\)R\_\{\\beta\}\(x\)=U^\{\*\}\(s\\mid C\)\.At terminalss, this equality follows directly from the boundary; at the root, it yields the normalizeru⁡\(s0\)=log⁡Z⁡\(C\)u\(s\_\{0\}\)=\\log Z\(C\), and substituting this into the root\-to\-terminal identity givesPθ=P∗P\_\{\\theta\}=P^\{\*\}\. Finally, the same finite sum of singleton equalities telescopes on any intervali:ji:j\. The reverse implication follows because singleton intervals are included among all subtrajectories\. ∎

###### Corollary A\.9\(Population zero loss under full coverage\)\.

Suppose the preceding shared\-flow and history\-law assumptions hold and a fixed sampling lawν\\nuassigns positive mass to every reference\-supported terminal history\. If𝔼ν​\[ℒXu\]=0\\mathbb\{E\}\_\{\\nu\}\[\\mathcal\{L\}\_\{X\}^\{u\}\]=0, then Proposition[A\.8](https://arxiv.org/html/2609.38661#A1.Thmproposition8)applies\.

Every loss is nonnegative\. A positive loss at a history of positiveν\\nu\-mass would give a positive expectation, so each supported history has zero loss\. Its finite sum of squared residuals then has every term zero, in particular each singleton term\. The consistency proposition therefore applies\. ∎

### A\.5Fixed\-Executor Realizability and Its Obstruction

###### Proposition A\.10\(Action\-only realizability criterion\)\.

Under the reference\-law and reward premises of Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1), define, on its support,

M∗​\(s,a\)\\displaystyle M^\{\*\}\(s,a\)=∑s′𝖪C​\(s′∣s,a\)​U∗​\(s′\),\\displaystyle=\\sum\_\{s^\{\\prime\}\}\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)U^\{\*\}\(s^\{\\prime\}\),\(A\.13\)Q∗​\(a,s′∣s\)\\displaystyle Q^\{\*\}\(a,s^\{\\prime\}\\mid s\)=ρ⁡\(a∣s,C\)​𝖪C​\(s′∣s,a\)​U∗​\(s′\)U∗​\(s\),\\displaystyle=\\rho\(a\\mid s,C\)\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)\\frac\{U^\{\*\}\(s^\{\\prime\}\)\}\{U^\{\*\}\(s\)\},\(A\.14\)π∗​\(a∣s\)\\displaystyle\\pi^\{\*\}\(a\\mid s\)=ρ⁡\(a∣s,C\)​M∗​\(s,a\)U∗​\(s\),\\displaystyle=\\rho\(a\\mid s,C\)\\frac\{M^\{\*\}\(s,a\)\}\{U^\{\*\}\(s\)\},𝖪∗​\(s′∣s,a\)\\displaystyle\\mathsf\{K\}^\{\*\}\(s^\{\\prime\}\\mid s,a\)=𝖪C​\(s′∣s,a\)​U∗​\(s′\)M∗​\(s,a\)\.\\displaystyle=\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)\\frac\{U^\{\*\}\(s^\{\\prime\}\)\}\{M^\{\*\}\(s,a\)\}\.\(A\.15\)HereQ∗Q^\{\*\}is the target joint successor law andQ∗=π∗​𝖪∗Q^\{\*\}=\\pi^\{\*\}\\mathsf\{K\}^\{\*\}\. Some arbitrary complete\-history action policy using the*fixed*executor realizesP∗P^\{\*\}if and only if

U∗\(s′\)=M∗\(s,a\)𝖪C\(⋅∣s,a\)\-almost surelyU^\{\*\}\(s^\{\\prime\}\)=M^\{\*\}\(s,a\)\\quad\\mathsf\{K\}\_\{C\}\(\\cdot\\mid s,a\)\\text\{\-almost surely\}\(A\.16\)for every supported\(s,a\)\(s,a\)\. If realizable, that action policy is uniquelyπ∗\\pi^\{\*\}on the reference support\. A deterministic executor satisfies Eq\. \([A\.16](https://arxiv.org/html/2609.38661#A1.E16)\) automatically\.

The target probability of a supported child divided by the target probability of its parent isF∗​\(s′\)/F∗​\(s\)F^\{\*\}\(s^\{\\prime\}\)/F^\{\*\}\(s\)\. Substitute the reference prefix factorization from Lemma[A\.6](https://arxiv.org/html/2609.38661#A1.Thmproposition6); since the child records its action, this gives Eq\. \([A\.14](https://arxiv.org/html/2609.38661#A1.E14)\)\. Summing over children for a fixed action givesπ∗\\pi^\{\*\}, and dividing by this positive action marginal gives𝖪∗\\mathsf\{K\}^\{\*\}\. Both normalize byM∗M^\{\*\}and the flow recursion\.

For necessity, a policyπ\\pirealizing the complete target law must also realize its prefix and joint successor probabilities\. Its joint kernel isπ⁡\(a∣s\)​𝖪C​\(s′∣s,a\)\\pi\(a\\mid s\)\\mathsf\{K\}\_\{C\}\(s^\{\\prime\}\\mid s,a\)\. Equating it toQ∗Q^\{\*\}and canceling a positive executor factor yieldsπ⁡\(a∣s\)=ρ⁡\(a∣s\)​U∗​\(s′\)/U∗​\(s\)\\pi\(a\\mid s\)=\\rho\(a\\mid s\)U^\{\*\}\(s^\{\\prime\}\)/U^\{\*\}\(s\)\. The left side does not vary withs′s^\{\\prime\}, soU∗​\(s′\)U^\{\*\}\(s^\{\\prime\}\)is constant on this successor support\. Averaging that constant givesM∗​\(s,a\)M^\{\*\}\(s,a\), proving Eq\. \([A\.16](https://arxiv.org/html/2609.38661#A1.E16)\) andπ=π∗\\pi=\\pi^\{\*\}\.

For sufficiency, under this condition𝖪∗=𝖪C\\mathsf\{K\}^\{\*\}=\\mathsf\{K\}\_\{C\}andπ∗​𝖪C=Q∗\\pi^\{\*\}\\mathsf\{K\}\_\{C\}=Q^\{\*\}\. Multiplying these joint transitions on a complete history telescopes theU∗U^\{\*\}ratios, yieldingPρ​\(x\)​U∗​\(x\)/U∗​\(s0\)=P∗​\(x\)P\_\{\\rho\}\(x\)U^\{\*\}\(x\)/U^\{\*\}\(s\_\{0\}\)=P^\{\*\}\(x\)\. These terminal probabilities sum to one, so the constructed process terminates almost surely and realizes the target\. Uniqueness follows from its already determined action marginals\. A deterministic kernel has only one supported successor, which proves the last statement\. ∎

###### Proposition A\.11\(Irreducible executor contribution to relative entropy\)\.

Under the reference\-law and reward premises of Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1), suppose the supported tree has a uniform finite horizonHH\. LetΠC\\Pi\_\{C\}contain all normalized complete\-history action policies supported on the reference legal actions, including policies with some zero action probabilities\. Forπ∈ΠC\\pi\\in\\Pi\_\{C\}, whereDKL\(p∥q\)=∑zp\(z\)log\[p\(z\)/q\(z\)\]D\_\{\\mathrm\{KL\}\}\(p\\\|q\)=\\sum\_\{z\}p\(z\)\\log\[p\(z\)/q\(z\)\], zero\-ppterms contribute zero, and a positive\-pp, zero\-qqterm gives infinite divergence,

DKL\(P∗∥Pπ\)\\displaystyle D\_\{\\mathrm\{KL\}\}\(P^\{\*\}\\\|P\_\{\\pi\}\)=𝔼P∗\[∑t=0T−1DKL\(π∗\(⋅∣St\)∥π\(⋅∣St\)\)\]\+ℐC,\\displaystyle=\\mathbb\{E\}\_\{P^\{\*\}\}\\\!\\left\[\\sum\_\{t=0\}^\{T\-1\}D\_\{\\mathrm\{KL\}\}\(\\pi^\{\*\}\(\\cdot\\mid S\_\{t\}\)\\\|\\pi\(\\cdot\\mid S\_\{t\}\)\)\\right\]\+\\mathcal\{I\}\_\{C\},\(A\.17\)ℐC\\displaystyle\\mathcal\{I\}\_\{C\}=𝔼P∗\[∑t=0T−1DKL\(𝖪∗\(⋅∣St,At\)∥𝖪C\(⋅∣St,At\)\)\],\\displaystyle=\\mathbb\{E\}\_\{P^\{\*\}\}\\\!\\left\[\\sum\_\{t=0\}^\{T\-1\}D\_\{\\mathrm\{KL\}\}\(\\mathsf\{K\}^\{\*\}\(\\cdot\\mid S\_\{t\},A\_\{t\}\)\\\|\\mathsf\{K\}\_\{C\}\(\\cdot\\mid S\_\{t\},A\_\{t\}\)\)\\right\],\(A\.18\)infπ∈ΠCDKL\(P∗∥Pπ\)\\displaystyle\\inf\_\{\\pi\\in\\Pi\_\{C\}\}D\_\{\\mathrm\{KL\}\}\(P^\{\*\}\\\|P\_\{\\pi\}\)=ℐC\.\\displaystyle=\\mathcal\{I\}\_\{C\}\.\(A\.19\)The infimum is attained by using the action marginalπ∗\\pi^\{\*\}with the fixed executor\. Moreover,ℐC=0\\mathcal\{I\}\_\{C\}=0exactly when Eq\. \([A\.16](https://arxiv.org/html/2609.38661#A1.E16)\) holds on the reference support\.

Factor the target withπ∗​𝖪∗\\pi^\{\*\}\\mathsf\{K\}^\{\*\}and the comparison process withπ​𝖪C\\pi\\mathsf\{K\}\_\{C\}\. Their log density ratio is

∑t=0T−1log⁡π∗​\(At∣St\)π⁡\(At∣St\)\+∑t=0T−1log⁡𝖪∗​\(St\+1∣St,At\)𝖪C​\(St\+1∣St,At\)\.\\sum\_\{t=0\}^\{T\-1\}\\log\\frac\{\\pi^\{\*\}\(A\_\{t\}\\mid S\_\{t\}\)\}\{\\pi\(A\_\{t\}\\mid S\_\{t\}\)\}\+\\sum\_\{t=0\}^\{T\-1\}\\log\\frac\{\\mathsf\{K\}^\{\*\}\(S\_\{t\+1\}\\mid S\_\{t\},A\_\{t\}\)\}\{\\mathsf\{K\}\_\{C\}\(S\_\{t\+1\}\\mid S\_\{t\},A\_\{t\}\)\}\.For the first sum, condition on each nonterminalStS\_\{t\}and average over its target action law; the result is the action KL term in Eq\. \([A\.17](https://arxiv.org/html/2609.38661#A1.E17)\)\. For the second, condition on\(St,At\)\(S\_\{t\},A\_\{t\}\)and average over𝖪∗\\mathsf\{K\}^\{\*\}; the result is Eq\. \([A\.18](https://arxiv.org/html/2609.38661#A1.E18)\)\. Variable lengths cause no additional term: for this calculation only, histories may be padded toHHwith deterministic terminal coordinates of log\-ratio zero\.

These steps remain valid as extended expectations\. For probability vectorsp,qp,q, the negative part of∑p​log⁡\(p/q\)\\sum p\\log\(p/q\)is bounded by∑p<qq⁡\(p/q\)​log⁡\(q/p\)≤1/e\\sum\_\{p<q\}q\(p/q\)\\log\(q/p\)\\leq 1/e, sincez​log⁡\(1/z\)≤1/ez\\log\(1/z\)\\leq 1/eon\[0,1\]\[0,1\]\. There are at most2​H2Hconditional terms\. AlsoU∗,M∗∈\[1,eβ\]U^\{\*\},M^\{\*\}\\in\[1,e^\{\\beta\}\], so the executor log\-ratio is bounded in absolute value byβ\\beta\. No subtraction of two infinite positive expectations is required\.

Each conditional KL is nonnegative\. Takingπ=π∗\\pi=\\pi^\{\*\}makes every action term zero and proves the infimum formula; finite horizon makes this fixed\-executor policy a terminating member ofΠC\\Pi\_\{C\}\. The remaining sum is zero precisely when𝖪∗=𝖪C\\mathsf\{K\}^\{\*\}=\\mathsf\{K\}\_\{C\}at every target\-supported state\-action pair\. The target and reference have the same support, and Eq\. \([A\.15](https://arxiv.org/html/2609.38661#A1.E15)\) makes this equality equivalent to Eq\. \([A\.16](https://arxiv.org/html/2609.38661#A1.E16)\)\. ∎

### A\.6Subtrajectory Algebra and Credit Coefficients

Fix a recorded trajectoryxxof lengthT≥1T\\geq 1\. All action log\-ratios and endpoint values are finite\. Writeviv\_\{i\}for the value used atsis\_\{i\}on this record, allowingvi=u~x​\(si\)v\_\{i\}=\\widetilde\{u\}\_\{x\}\(s\_\{i\}\); reuse that same value in every interval of the record\. Putet=vt\+Δ​ℓt−vt\+1e\_\{t\}=v\_\{t\}\+\\Delta\\ell\_\{t\}\-v\_\{t\+1\},0≤t<T0\\leq t<T, andδ\(x\)i:j=vi\+∑t=ij−1Δℓt−vj\\delta^\{\(x\)\}\_\{i:j\}=v\_\{i\}\+\\sum\_\{t=i\}^\{j\-1\}\\Delta\\ell\_\{t\}\-v\_\{j\}\.

###### Lemma A\.12\(Within\-trajectory telescoping\)\.

For every interval0≤i<j≤T0\\leq i<j\\leq T,δ\(x\)i:j=∑t=ij−1et\\delta^\{\(x\)\}\_\{i:j\}=\\sum\_\{t=i\}^\{j\-1\}e\_\{t\}\.

Expand the sum ofete\_\{t\}\. Each internal valuevi\+1,…,vj−1v\_\{i\+1\},\\ldots,v\_\{j\-1\}appears once with each sign and cancels; onlyvi−vjv\_\{i\}\-v\_\{j\}and the action log\-ratios remain\. No other trajectory is used\. ∎

###### Proposition A\.13\(Matrix form, positive definiteness, and zero loss\)\.

For all0≤i<j≤T0\\leq i<j\\leq Tand0≤t<T0\\leq t<T, letAT\[\(i,j\),t\]=𝟏\{i≤t<j\}A\_\{T\}\[\(i,j\),t\]=\\mathbf\{1\}\\\{i\\leq t<j\\\}\. WithGT=AT𝖳​ATG\_\{T\}=A\_\{T\}^\{\\mathsf\{T\}\}A\_\{T\},

δ\(x\)=AT​e,ℒx=1KT​e𝖳​GT​e,\(GT\)r​s=\(min⁡\(r,s\)\+1\)​\(T−max⁡\(r,s\)\),\\delta^\{\(x\)\}=A\_\{T\}e,\\qquad\\mathcal\{L\}\_\{x\}=\\frac\{1\}\{K\_\{T\}\}e^\{\\mathsf\{T\}\}G\_\{T\}e,\\qquad\(G\_\{T\}\)\_\{rs\}=\(\\min\(r,s\)\+1\)\(T\-\\max\(r,s\)\),\(A\.20\)where0≤r,s<T0\\leq r,s<T\. The matrixGTG\_\{T\}is positive definite and

ℒx≥1KT∑t=0T−1et2,ℒx=0⟺e=0⟺δ\(x\)i:j=0for everyi<j\.\\mathcal\{L\}\_\{x\}\\geq\\frac\{1\}\{K\_\{T\}\}\\sum\_\{t=0\}^\{T\-1\}e\_\{t\}^\{2\},\\qquad\\mathcal\{L\}\_\{x\}=0\\ \\Longleftrightarrow\\ e=0\\ \\Longleftrightarrow\\ \\delta^\{\(x\)\}\_\{i:j\}=0\\text\{ for every \}i<j\.\(A\.21\)

Lemma[A\.12](https://arxiv.org/html/2609.38661#A1.Thmproposition12)gives the matrix multiplication and hence the quadratic form\. Entry\(r,s\)\(r,s\)counts intervals containing both actions: there aremin⁡\(r,s\)\+1\\min\(r,s\)\+1possible left endpoints andT−max⁡\(r,s\)T\-\\max\(r,s\)possible right endpoints\. This proves the entry formula\. The rows indexed by\(t,t\+1\)\(t,t\+1\)form an identity matrix, soATA\_\{T\}has full column rank\. For any nonzeroee,‖AT​e‖22\>0\\\|A\_\{T\}e\\\|\_\{2\}^\{2\}\>0, proving positive definiteness\. Retaining only the singleton squares proves the lower bound and forcese=0e=0if the loss is zero\. Conversely,e=0e=0gives every interval residual zero and therefore zero loss\. ∎

###### Proposition A\.14\(Action coefficients and cumulative\-sum identity\)\.

Holding the endpoint values fixed when differentiating with respect to action log\-ratios, the coefficient for actionttis

Γt\(x\):=∂ℒx∂Δ​ℓt=2KT∑i≤t<jδi:j\(x\)=2KT\(GTe\)t\.\\Gamma\_\{t\}\(x\):=\\frac\{\\partial\\mathcal\{L\}\_\{x\}\}\{\\partial\\Delta\\ell\_\{t\}\}=\\frac\{2\}\{K\_\{T\}\}\\sum\_\{i\\leq t<j\}\\delta^\{\(x\)\}\_\{i:j\}=\\frac\{2\}\{K\_\{T\}\}\(G\_\{T\}e\)\_\{t\}\.\(A\.22\)There are\(t\+1\)​\(T−t\)\(t\+1\)\(T\-t\)intervals in this sum\. Ifh0=0h\_\{0\}=0andhj=∑t=0j−1eth\_\{j\}=\\sum\_\{t=0\}^\{j\-1\}e\_\{t\}, then

KT​ℒx=\(T\+1\)​∑j=0Thj2−\(∑j=0Thj\)2\.K\_\{T\}\\mathcal\{L\}\_\{x\}=\(T\+1\)\\sum\_\{j=0\}^\{T\}h\_\{j\}^\{2\}\-\\left\(\\sum\_\{j=0\}^\{T\}h\_\{j\}\\right\)^\{2\}\.\(A\.23\)

The partial derivative ofδ\(x\)i:j\\delta^\{\(x\)\}\_\{i:j\}with respect toΔ​ℓt\\Delta\\ell\_\{t\}is𝟏\{i≤t<j\}\\mathbf\{1\}\\\{i\\leq t<j\\\}\. Differentiate the finite sum of squared residuals to obtain the first coefficient formula;AT𝖳​δ\(x\)=GT​eA\_\{T\}^\{\\mathsf\{T\}\}\\delta^\{\(x\)\}=G\_\{T\}egives the second\. An interval containingtthast\+1t\+1possible starts andT−tT\-tpossible ends\. For the cumulative identity, telescoping givesδ\(x\)i:j=hj−hi\\delta^\{\(x\)\}\_\{i:j\}=h\_\{j\}\-h\_\{i\}\. Therefore

∑i<j\(hj−hi\)2=12​∑i=0T∑j=0T\(hj−hi\)2=\(T\+1\)​∑j=0Thj2−\(∑j=0Thj\)2,\\sum\_\{i<j\}\(h\_\{j\}\-h\_\{i\}\)^\{2\}=\\frac\{1\}\{2\}\\sum\_\{i=0\}^\{T\}\\sum\_\{j=0\}^\{T\}\(h\_\{j\}\-h\_\{i\}\)^\{2\}=\(T\+1\)\\sum\_\{j=0\}^\{T\}h\_\{j\}^\{2\}\-\\left\(\\sum\_\{j=0\}^\{T\}h\_\{j\}\\right\)^\{2\},where the last equality expands the two square terms and the cross term\. ∎

### A\.7Approximate Consistency Under Uniform Residual Control

###### Proposition A\.15\(Root, log\-density, and total\-variation bounds\)\.

Under Assumption[A\.1](https://arxiv.org/html/2609.38661#A1.Thmassumption1), letuube shared with the exact terminal boundary\. If\|δ0:T⁡\(x\)u\(x\)\|≤ε\|\\delta^\{u\}\_\{0:T\(x\)\}\(x\)\|\\leq\\varepsilonfor every reference\-supported complete history, whereε≥0\\varepsilon\\geq 0, then

\|u⁡\(s0\)−log⁡Z⁡\(C\)\|\\displaystyle\|u\(s\_\{0\}\)\-\\log Z\(C\)\|≤ε,\\displaystyle\\leq\\varepsilon,\(A\.24\)\|log⁡Pθ​\(x∣C\)P∗​\(x∣C\)\|\\displaystyle\\left\|\\log\\frac\{P\_\{\\theta\}\(x\\mid C\)\}\{P^\{\*\}\(x\\mid C\)\}\\right\|≤2​ε,\\displaystyle\\leq 2\\varepsilon,\(A\.25\)TV⁡\(Pθ,P∗\)\\displaystyle\\operatorname\{TV\}\(P\_\{\\theta\},P^\{\*\}\)≤tanh⁡\(ε\),\\displaystyle\\leq\\tanh\(\\varepsilon\),\(A\.26\)whereTV⁡\(P,Q\)=12​∑x\|P⁡\(x\)−Q⁡\(x\)\|\\operatorname\{TV\}\(P,Q\)=\\tfrac\{1\}\{2\}\\sum\_\{x\}\|P\(x\)\-Q\(x\)\|\.

Writed\(x\)=δ0:T⁡\(x\)u\(x\)d\(x\)=\\delta^\{u\}\_\{0:T\(x\)\}\(x\)\. Equation \([A\.4](https://arxiv.org/html/2609.38661#A1.E4)\) and the terminal boundary give

Pθ​\(x\)/P∗​\(x\)=exp⁡\(d⁡\(x\)\+log⁡Z−u⁡\(s0\)\)\.P\_\{\\theta\}\(x\)/P^\{\*\}\(x\)=\\exp\\bigl\(d\(x\)\+\\log Z\-u\(s\_\{0\}\)\\bigr\)\.Summing againstP∗P^\{\*\}and using normalization shows

u⁡\(s0\)−log⁡Z=log⁡𝔼P∗​ed⁡\(X\)\.u\(s\_\{0\}\)\-\\log Z=\\log\\mathbb\{E\}\_\{P^\{\*\}\}e^\{d\(X\)\}\.Becausee−ε≤ed⁡\(X\)≤eεe^\{\-\\varepsilon\}\\leq e^\{d\(X\)\}\\leq e^\{\\varepsilon\}, this logarithm lies in\[−ε,ε\]\[\-\\varepsilon,\\varepsilon\], proving the root bound\. Subtract it fromd⁡\(x\)d\(x\)to obtain the log\-density bound\. In particular, forg⁡\(x\)=Pθ​\(x\)/P∗​\(x\)g\(x\)=P\_\{\\theta\}\(x\)/P^\{\*\}\(x\),e−2​ε≤g⁡\(x\)≤e2​εe^\{\-2\\varepsilon\}\\leq g\(x\)\\leq e^\{2\\varepsilon\}\. For any positivezzin this interval,

\|z−1\|z\+1≤e2​ε−1e2​ε\+1=tanh⁡\(ε\)\.\\frac\{\|z\-1\|\}\{z\+1\}\\leq\\frac\{e^\{2\\varepsilon\}\-1\}\{e^\{2\\varepsilon\}\+1\}=\\tanh\(\\varepsilon\)\.The fraction increases forz≥1z\\geq 1and decreases forz≤1z\\leq 1, so its maximum is at an interval endpoint\. Integrating\|g−1\|≤tanh⁡\(ε\)​\(g\+1\)\|g\-1\|\\leq\\tanh\(\\varepsilon\)\(g\+1\)underP∗P^\{\*\}with𝔼P∗​g=1\\mathbb\{E\}\_\{P^\{\*\}\}g=1gives Eq\. \([A\.26](https://arxiv.org/html/2609.38661#A1.E26)\)\. ∎

###### Corollary A\.16\(A uniformly controlled all\-subtrajectory loss\)\.

If additionallyη≥0\\eta\\geq 0,T⁡\(x\)≤HT\(x\)\\leq H, andℒxu≤η2\\mathcal\{L\}\_\{x\}^\{u\}\\leq\\eta^\{2\}for every supported complete history, the preceding bounds hold withε=H⁡\(H\+1\)/2​η\\varepsilon=\\sqrt\{H\(H\+1\)/2\}\\,\\eta\.

The full\-path square is one nonnegative summand ofKT​ℒxuK\_\{T\}\\mathcal\{L\}\_\{x\}^\{u\}, so Proposition[A\.15](https://arxiv.org/html/2609.38661#A1.Thmproposition15)applies with

\|δu0:T\|≤KTη≤H⁡\(H\+1\)/2η=ε\.∎\|\\delta^\{u\}\_\{0:T\}\|\\leq\\sqrt\{K\_\{T\}\}\\,\\eta\\leq\\sqrt\{H\(H\+1\)/2\}\\,\\eta=\\varepsilon\.\\qed

### A\.8Mixed\-Behavior Regression and Gradient Bounds

###### Lemma A\.17\(Record\-conditioned gradient routing\)\.

Condition on a recorded batch, its contexts, and detached statistics\. Assumebψ​\(hρ​\(s\),f⁡\(s\)\)b\_\{\\psi\}\(h\_\{\\rho\}\(s\),f\(s\)\)has noθ\\theta\-dependence\. The AnchorTB sample gradients are

∇θℒx\\displaystyle\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{x\}=∑t=0T−1Γt​\(x\)​∇θ​log⁡πθ​\(at∣st,C\),\\displaystyle=\\sum\_\{t=0\}^\{T\-1\}\\Gamma\_\{t\}\(x\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\},C\),\(A\.27\)∇ψℒx\\displaystyle\\nabla\_\{\\psi\}\\mathcal\{L\}\_\{x\}=2KT∑i<jδi:j\(x\)\(∇ψu~x\(si\)−∇ψu~x\(sj\)\),\\displaystyle=\\frac\{2\}\{K\_\{T\}\}\\sum\_\{i<j\}\\delta^\{\(x\)\}\_\{i:j\}\\bigl\(\\nabla\_\{\\psi\}\\widetilde\{u\}\_\{x\}\(s\_\{i\}\)\-\\nabla\_\{\\psi\}\\widetilde\{u\}\_\{x\}\(s\_\{j\}\)\\bigr\),\(A\.28\)where∇ψu~x​\(sT\)=0\\nabla\_\{\\psi\}\\widetilde\{u\}\_\{x\}\(s\_\{T\}\)=0because the terminal reward boundary overrides the learned head\.

Differentiate each squared residual using the chain rule\. Forθ\\theta, the endpoint values and reference log\-probabilities are fixed, leaving only the current action log\-probabilities; collecting terms for the same action gives Eq\. \([A\.27](https://arxiv.org/html/2609.38661#A1.E27)\)\. Forψ\\psi, the action log\-ratios are fixed and the two endpoints have opposite signs, giving Eq\. \([A\.28](https://arxiv.org/html/2609.38661#A1.E28)\)\. Stop\-gradient sets the derivative of the measured part to zero, and the calculation conditions on the sampled history and its terminal reward\. ∎

###### Proposition A\.18\(A bounded\-Jacobian gradient second\-moment bound\)\.

Letξ\\xibe any chosen trainable parameter vector, with the record and detached statistics held fixed\. At differentiability points suppose

1KT∑i<j∥∇ξδ\(x\)i:j∥22≤G2\\frac\{1\}\{K\_\{T\}\}\\sum\_\{i<j\}\\\|\\nabla\_\{\\xi\}\\delta^\{\(x\)\}\_\{i:j\}\\\|\_\{2\}^\{2\}\\leq G^\{2\}for a constantG<∞G<\\infty\. Then

‖∇ξℒx‖22≤4​G2​ℒx\.\\\|\\nabla\_\{\\xi\}\\mathcal\{L\}\_\{x\}\\\|\_\{2\}^\{2\}\\leq 4G^\{2\}\\mathcal\{L\}\_\{x\}\.\(A\.29\)For a fixed distribution of recorded inputs satisfying this bound and𝔼​ℒX<∞\\mathbb\{E\}\\mathcal\{L\}\_\{X\}<\\infty,

tr⁡Cov⁡\(∇ξℒX\)≤𝔼​‖∇ξℒX‖22≤4​G2​𝔼​ℒX\.\\operatorname\{tr\}\\operatorname\{Cov\}\(\\nabla\_\{\\xi\}\\mathcal\{L\}\_\{X\}\)\\leq\\mathbb\{E\}\\\|\\nabla\_\{\\xi\}\\mathcal\{L\}\_\{X\}\\\|\_\{2\}^\{2\}\\leq 4G^\{2\}\\mathbb\{E\}\\mathcal\{L\}\_\{X\}\.\(A\.30\)

Apply vector Cauchy–Schwarz to the sample gradient:

‖2KT∑i<jδi:j\(x\)∇ξδi:j\(x\)‖22≤4\(1KT∑i<j\(δi:j\(x\)\)2\)\(1KT∑i<j∥∇ξδi:j\(x\)∥22\)\.\\left\\\|\\frac\{2\}\{K\_\{T\}\}\\sum\_\{i<j\}\\delta^\{\(x\)\}\_\{i:j\}\\nabla\_\{\\xi\}\\delta^\{\(x\)\}\_\{i:j\}\\right\\\|\_\{2\}^\{2\}\\leq 4\\left\(\\frac\{1\}\{K\_\{T\}\}\\sum\_\{i<j\}\(\\delta^\{\(x\)\}\_\{i:j\}\)^\{2\}\\right\)\\left\(\\frac\{1\}\{K\_\{T\}\}\\sum\_\{i<j\}\\\|\\nabla\_\{\\xi\}\\delta^\{\(x\)\}\_\{i:j\}\\\|\_\{2\}^\{2\}\\right\)\.The first factor is the loss and the second is at mostG2G^\{2\}\. Taking expectation proves the second\-moment bound\. The identitytr⁡Cov⁡\(V\)=𝔼​‖V‖22−‖𝔼​V‖22\\operatorname\{tr\}\\operatorname\{Cov\}\(V\)=\\mathbb\{E\}\\\|V\\\|\_\{2\}^\{2\}\-\\\|\\mathbb\{E\}V\\\|\_\{2\}^\{2\}for square\-integrableVVgives the rest\. ∎

#### Sampling law\.

With a fixed data lawν\\nu, AnchorTB is a regression of recorded residuals underν\\nu\. Under full coverage its exact shared zero is characterized by Corollary[A\.9](https://arxiv.org/html/2609.38661#A1.Thmproposition9), and nonzero optima depend on the sampling weights\. These sample\-gradient identities concern the AnchorTB loss; the value head and the diagnostic heads are fitted separately\.

## Appendix BShrinkage Baselines, Structural Pooling, and the Value Head

The ideal conditional mean in Eq\. \([A\.8](https://arxiv.org/html/2609.38661#A1.E8)\) is a property of a specified reference continuation law\. Measured task statistics, lagged structural buckets, and a feature\-based value head are different estimators with different data sources\. We give their algebraic properties and population conditions\.

### B\.1Implemented Trajectory\-Indexed Flows

###### Definition B\.1\(Implemented flow and terminal override\)\.

For batchkkand scored trajectoryxx, putuq,k,x=log⁡\[1\+aβ​p^q,k\(−x\)\]u\_\{q,k,x\}=\\log\[1\+a\_\{\\beta\}\\hat\{p\}\_\{q,k\}^\{\(\-x\)\}\]\. The flow estimate in Eq\. \([8](https://arxiv.org/html/2609.38661#S4.E8)\), with its terminal boundary made explicit, is

u~k,x​\(si\)=\{sg⁡\[clip\[0,β\]⁡\(uq,k,x\+ck−1​\(g⁡\(si\)\)\)\]\+bψ​\(hρ​\(si\),f⁡\(si\)\),0≤i<T,log⁡Rβ​\(x\),i=T\.\\widetilde\{u\}\_\{k,x\}\(s\_\{i\}\)=\\begin\{cases\}\\operatorname\{sg\}\\\!\\left\[\\operatorname\{clip\}\_\{\[0,\\beta\]\}\(u\_\{q,k,x\}\+c\_\{k\-1\}\(g\(s\_\{i\}\)\)\)\\right\]\+b\_\{\\psi\}\(h\_\{\\rho\}\(s\_\{i\}\),f\(s\_\{i\}\)\),&0\\leq i<T,\\\\ \\log R\_\{\\beta\}\(x\),&i=T\.\\end\{cases\}\(B\.1\)The second branch overrides the entire learned parameterization at the terminal, includingbψb\_\{\\psi\}\. We suppress the fixed batch index and writeu~x,uq,x,p^q\(−x\),c\\widetilde\{u\}\_\{x\},u\_\{q,x\},\\hat\{p\}\_\{q\}^\{\(\-x\)\},cwhen unambiguous\. The symbolg⁡\(s\)g\(s\)denotes the canonical structure bucket of the history\.

###### Proposition B\.1\(When a trajectory\-indexed family is shared\)\.

On a given collection of scored prefixes, the family\{u~x​\(s\)\}\\\{\\widetilde\{u\}\_\{x\}\(s\)\\\}defines a single state function if and only if

u~x​\(s\)=u~x′​\(s\)whenever​x⪰s​and​x′⪰s​belong to that collection\.\\widetilde\{u\}\_\{x\}\(s\)=\\widetilde\{u\}\_\{x^\{\\prime\}\}\(s\)\\quad\\text\{whenever \}x\\succeq s\\text\{ and \}x^\{\\prime\}\\succeq s\\text\{ belong to that collection\}\.\(B\.2\)Within every record, Lemma[A\.12](https://arxiv.org/html/2609.38661#A1.Thmproposition12)and Propositions[A\.13](https://arxiv.org/html/2609.38661#A1.Thmproposition13)–[A\.14](https://arxiv.org/html/2609.38661#A1.Thmproposition14)apply using Eq\. \([B\.1](https://arxiv.org/html/2609.38661#A2.E1)\)\.

If a single function exists, evaluating it atssgives the same value for every continuation, which proves necessity\. Conversely, define its value at each represented prefix by selecting any one continuation\. Equation \([B\.2](https://arxiv.org/html/2609.38661#A2.E2)\) makes the definition independent of that selection\. This proves sufficiency on the specified collection\. The within\-record identities only require that a fixed endpoint value be reused in each interval of that record, which the definition provides\. ∎

### B\.2Hierarchical Shrinkage and Its Range

###### Definition B\.2\(Hierarchical task baseline\)\.

LetSSandNNbe reward sums and trajectory counts from the healthy, non\-paired reference pool, with subscriptsg,c,qg,c,qfor global, category, and task statistics\. HeregginSg,Ng,mgS\_\{g\},N\_\{g\},m\_\{g\}means the global level, not a structure bucket\. Define

mg\\displaystyle m\_\{g\}=\{Sg/Ng,Ng\>0,1/2,Ng=0,\\displaystyle=\\begin\{cases\}S\_\{g\}/N\_\{g\},&N\_\{g\}\>0,\\\\ 1/2,&N\_\{g\}=0,\\end\{cases\}mc\\displaystyle\\qquad m\_\{c\}=Sc\+κg​mgNc\+κg,\\displaystyle=\\frac\{S\_\{c\}\+\\kappa\_\{g\}m\_\{g\}\}\{N\_\{c\}\+\\kappa\_\{g\}\},\(B\.3\)m0\\displaystyle m\_\{0\}=nd′​pd\+κc​mcnd′\+κc,\\displaystyle=\\frac\{n^\{\\prime\}\_\{d\}p\_\{d\}\+\\kappa\_\{c\}m\_\{c\}\}\{n^\{\\prime\}\_\{d\}\+\\kappa\_\{c\}\},p^q\(−x\)\\displaystyle\\qquad\\hat\{p\}\_\{q\}^\{\(\-x\)\}=Sq\(−x\)\+κq​m0Nq\(−x\)\+κq\.\\displaystyle=\\frac\{S\_\{q\}^\{\(\-x\)\}\+\\kappa\_\{q\}m\_\{0\}\}\{N\_\{q\}^\{\(\-x\)\}\+\\kappa\_\{q\}\}\.\(B\.4\)The optional record prior is\(pd,nd\)\(p\_\{d\},n\_\{d\}\), withnd′=min⁡\(nd,20\)n^\{\\prime\}\_\{d\}=\\min\(n\_\{d\},20\); absence of that prior givesnd′=0n^\{\\prime\}\_\{d\}=0\. The configured strengths are\(κg,κc,κq\)=\(2,4,0\.5\)\(\\kappa\_\{g\},\\kappa\_\{c\},\\kappa\_\{q\}\)=\(2,4,0\.5\)\. If the scored record is in the eligible task pool, it is removed once from that task sum and count; otherwise no subtraction is made\. Leave\-one\-out applies at the task level\. The estimate is clipped to\[0,1\]\[0,1\]before the reward transform\. For bounded nonbinary rewards,p^\\hat\{p\}estimates a mean score, and for binary rewards a success probability\.

###### Lemma B\.2\(Range of the hierarchical estimate\)\.

Assume nonnegative counts,0≤S≤N0\\leq S\\leq Nat every used level, a valid task\-level deletion,pd∈\[0,1\]p\_\{d\}\\in\[0,1\],nd′≥0n^\{\\prime\}\_\{d\}\\geq 0, and positive shrinkage strengths\. Thenmg,mc,m0,p^q\(−x\)∈\[0,1\]m\_\{g\},m\_\{c\},m\_\{0\},\\hat\{p\}\_\{q\}^\{\(\-x\)\}\\in\[0,1\]and0≤uq,x≤β0\\leq u\_\{q,x\}\\leq\\beta\.

A nonempty empirical meanS/NS/Nlies in\[0,1\]\[0,1\], and the empty global convention does too\. ForNc\>0N\_\{c\}\>0,mcm\_\{c\}is the convex combination ofSc/NcS\_\{c\}/N\_\{c\}andmgm\_\{g\}with weightsNc/\(Nc\+κg\)N\_\{c\}/\(N\_\{c\}\+\\kappa\_\{g\}\)andκg/\(Nc\+κg\)\\kappa\_\{g\}/\(N\_\{c\}\+\\kappa\_\{g\}\); ifNc=0N\_\{c\}=0it equalsmgm\_\{g\}\. The same argument applies tom0m\_\{0\}and the task estimate, including zero\-count cases because their shrinkage denominators stay positive\. Finallyp↦log⁡\(1\+aβ​p\)p\\mapsto\\log\(1\+a\_\{\\beta\}p\)is nondecreasing on\[0,1\]\[0,1\], with endpoint values00andβ\\beta\. This holds for any reward in\[0,1\]\[0,1\]\. ∎

###### Proposition B\.3\(Fixed\-prior shrinkage moments and sensitivity\)\.

For fixedn≥0n\\geq 0,m∈\[0,1\]m\\in\[0,1\],κ\>0\\kappa\>0, and rewardsYi∈\[0,1\]Y\_\{i\}\\in\[0,1\], letp^=\(∑i=1nYi\+κ​m\)/\(n\+κ\)\\hat\{p\}=\(\\sum\_\{i=1\}^\{n\}Y\_\{i\}\+\\kappa m\)/\(n\+\\kappa\)\. If theYiY\_\{i\}are i\.i\.d\. with meanμ\\muand varianceσ2\\sigma^\{2\}, then

𝔼​p^−μ\\displaystyle\\mathbb\{E\}\\hat\{p\}\-\\mu=κ⁡\(m−μ\)n\+κ,\\displaystyle=\\frac\{\\kappa\(m\-\\mu\)\}\{n\+\\kappa\},Var⁡\(p^\)\\displaystyle\\qquad\\operatorname\{Var\}\(\\hat\{p\}\)=n​σ2\(n\+κ\)2,\\displaystyle=\\frac\{n\\sigma^\{2\}\}\{\(n\+\\kappa\)^\{2\}\},\(B\.5\)𝔼​\(p^−μ\)2\\displaystyle\\mathbb\{E\}\(\\hat\{p\}\-\\mu\)^\{2\}=n​σ2\+κ2​\(m−μ\)2\(n\+κ\)2\.\\displaystyle=\\frac\{n\\sigma^\{2\}\+\\kappa^\{2\}\(m\-\\mu\)^\{2\}\}\{\(n\+\\kappa\)^\{2\}\}\.\(B\.6\)For a fixed record set, changing one bounded reward changesp^\\hat\{p\}by at most1/\(n\+κ\)1/\(n\+\\kappa\)\. Changing onlymmtom′m^\{\\prime\}changes it by exactlyκ​\|m−m′\|/\(n\+κ\)\\kappa\|m\-m^\{\\prime\}\|/\(n\+\\kappa\)\. Whenn≥1n\\geq 1, deleting recordiiwhile holding the prior fixed givesp^−p^−i=\(Yi−p^−i\)/\(n\+κ\)\\hat\{p\}\-\\hat\{p\}\_\{\-i\}=\(Y\_\{i\}\-\\hat\{p\}\_\{\-i\}\)/\(n\+\\kappa\), again of magnitude at most1/\(n\+κ\)1/\(n\+\\kappa\)\.

Linearity gives𝔼​p^=\(n​μ\+κ​m\)/\(n\+κ\)\\mathbb\{E\}\\hat\{p\}=\(n\\mu\+\\kappa m\)/\(n\+\\kappa\)\. Independence makes the variance of the sumn​σ2n\\sigma^\{2\}; division by the squared denominator gives its variance\. Squared bias plus variance proves the MSE expression, also atn=0n=0\. The replacement and prior sensitivities follow by subtracting the two numerators over their common denominator\. For deletion, write∑h≠iYh\+κ​m=\(n−1\+κ\)​p^−i\\sum\_\{h\\neq i\}Y\_\{h\}\+\\kappa m=\(n\-1\+\\kappa\)\\hat\{p\}\_\{\-i\}and substitute inp^\\hat\{p\}; subtractingp^−i\\hat\{p\}\_\{\-i\}gives the identity, bounded asYi,p^−i∈\[0,1\]Y\_\{i\},\\hat\{p\}\_\{\-i\}\\in\[0,1\]\. ∎

#### Pooled estimand\.

The moment formulas above condition on a fixed prior and sample size\. If records come from contextsChC\_\{h\}and pass a health eventAhA\_\{h\}, their conditional means are𝔼\[r∣Ch,Ah\]\\mathbb\{E\}\[r\\mid C\_\{h\},A\_\{h\}\], and for fixed mixture weightswhw\_\{h\}the pooled population mean is∑hwh𝔼\[r∣Ch,Ah\]\\sum\_\{h\}w\_\{h\}\\mathbb\{E\}\[r\\mid C\_\{h\},A\_\{h\}\]\.

### B\.3Anchor Sensitivity and Residual Stability

###### Lemma B\.4\(Reward\-transform sensitivity and clipping\)\.

Forp,p′∈\[0,1\]p,p^\{\\prime\}\\in\[0,1\], letfβ​\(p\)=log⁡\(1\+aβ​p\)f\_\{\\beta\}\(p\)=\\log\(1\+a\_\{\\beta\}p\)andA⁡\(p,c\)=clip\[0,β\]⁡\(fβ​\(p\)\+c\)A\(p,c\)=\\operatorname\{clip\}\_\{\[0,\\beta\]\}\(f\_\{\\beta\}\(p\)\+c\)\. Then

\|fβ​\(p\)−fβ​\(p′\)\|≤aβ​\|p−p′\|,\|A⁡\(p,c\)−A⁡\(p′,c′\)\|≤aβ​\|p−p′\|\+\|c−c′\|\.\|f\_\{\\beta\}\(p\)\-f\_\{\\beta\}\(p^\{\\prime\}\)\|\\leq a\_\{\\beta\}\|p\-p^\{\\prime\}\|,\\qquad\|A\(p,c\)\-A\(p^\{\\prime\},c^\{\\prime\}\)\|\\leq a\_\{\\beta\}\|p\-p^\{\\prime\}\|\+\|c\-c^\{\\prime\}\|\.\(B\.7\)

Forβ\>0\\beta\>0,fβ′​\(p\)=aβ/\(1\+aβ​p\)≤aβf^\{\\prime\}\_\{\\beta\}\(p\)=a\_\{\\beta\}/\(1\+a\_\{\\beta\}p\)\\leq a\_\{\\beta\}, so integration betweenppandp′p^\{\\prime\}proves the first inequality\. Atβ=0\\beta=0, both sides of that inequality are zero\. Projection onto a closed interval is nonexpansive: ordering the two inputs shows that clipping can only shorten their distance\. This holds for\[0,0\]\[0,0\]too\. This and the triangle inequality give the second bound\. ∎

###### Proposition B\.5\(Endpoint perturbations, loss, and credit stability\)\.

Compare two endpoint arraysvi,v¯iv\_\{i\},\\bar\{v\}\_\{i\}on one fixed trajectory with the same action log\-ratios, and supposemax0≤i≤T⁡\|v¯i−vi\|≤ε\\max\_\{0\\leq i\\leq T\}\|\\bar\{v\}\_\{i\}\-v\_\{i\}\|\\leq\\varepsilon\. Then

\|δ¯i:j−δi:j\|\\displaystyle\|\\bar\{\\delta\}\_\{i:j\}\-\\delta\_\{i:j\}\|≤2​ε,\\displaystyle\\leq 2\\varepsilon,\|ℒ¯x−ℒx\|\\displaystyle\\qquad\|\\sqrt\{\\bar\{\\mathcal\{L\}\}\_\{x\}\}\-\\sqrt\{\\mathcal\{L\}\_\{x\}\}\|≤2​ε,\\displaystyle\\leq 2\\varepsilon,\(B\.8\)\|Γ¯t−Γt\|\\displaystyle\|\\bar\{\\Gamma\}\_\{t\}\-\\Gamma\_\{t\}\|≤4​εKT​\(t\+1\)​\(T−t\),\\displaystyle\\leq\\frac\{4\\varepsilon\}\{K\_\{T\}\}\(t\+1\)\(T\-t\),0≤t<T\.\\displaystyle 0\\leq t<T\.\(B\.9\)If only the measured anchor changes by inputs satisfying\|p−p′\|≤εp\|p\-p^\{\\prime\}\|\\leq\\varepsilon\_\{p\},\|c−c′\|≤εc\|c\-c^\{\\prime\}\|\\leq\\varepsilon\_\{c\}, whilebψb\_\{\\psi\}and terminal rewards are fixed, these bounds hold withε=aβ​εp\+εc\\varepsilon=a\_\{\\beta\}\\varepsilon\_\{p\}\+\\varepsilon\_\{c\}\.

The difference of interval residuals is\(v¯i−vi\)−\(v¯j−vj\)\(\\bar\{v\}\_\{i\}\-v\_\{i\}\)\-\(\\bar\{v\}\_\{j\}\-v\_\{j\}\), bounded by2​ε2\\varepsilon\. There areKTK\_\{T\}residuals, so the Euclidean distance between their vectors is at most2​ε​KT2\\varepsilon\\sqrt\{K\_\{T\}\}\. Divide the reverse triangle inequality for their norms byKT\\sqrt\{K\_\{T\}\}to obtain the loss bound\. By Eq\. \([A\.22](https://arxiv.org/html/2609.38661#A1.E22)\), the coefficient difference is the sum of\(t\+1\)​\(T−t\)\(t\+1\)\(T\-t\)residual differences times2/KT2/K\_\{T\}; bounding each summand gives Eq\. \([B\.9](https://arxiv.org/html/2609.38661#A2.E9)\)\. Lemma[B\.4](https://arxiv.org/html/2609.38661#A2.Thmproposition4)gives the rest, the terminal being fixed\. ∎

### B\.4Ratio\-First Structural Pooling

###### Definition B\.3\(Visit\-multiplicity pooling\)\.

For a structural bucketgg, let𝒱g\\mathcal\{V\}\_\{g\}be the multiset of eligible prefix visits\(x,t\)\(x,t\)assigned to that bucket, and setng=\|𝒱g\|n\_\{g\}=\|\\mathcal\{V\}\_\{g\}\|\. Each included occurrence is counted in both the sum and the denominator\. PutBx=euq,xB\_\{x\}=e^\{u\_\{q,x\}\}andzx,t=Rβ​\(x\)/Bxz\_\{x,t\}=R\_\{\\beta\}\(x\)/B\_\{x\}\. Forn0≥0n\_\{0\}\\geq 0,cmax≥0c\_\{\\max\}\\geq 0, andng\+n0\>0n\_\{g\}\+n\_\{0\}\>0, define

Ag=1\+∑\(x,t\)∈𝒱g\(zx,t−1\)ng\+n0,c⁡\(g\)=clip\[−cmax,cmax\]⁡\(log⁡Ag\)\.A\_\{g\}=1\+\\frac\{\\sum\_\{\(x,t\)\\in\\mathcal\{V\}\_\{g\}\}\(z\_\{x,t\}\-1\)\}\{n\_\{g\}\+n\_\{0\}\},\\qquad c\(g\)=\\operatorname\{clip\}\_\{\[\-c\_\{\\max\},c\_\{\\max\}\]\}\(\\log A\_\{g\}\)\.\(B\.10\)These are the exact\-arithmetic quantities before the positive\-domain numerical safeguard in the implementation description\. Buckets use a canonical graph and an open/stop marker, fall back to coarse node and edge counts below eight visits, and are read one batch behind\. The configured values aren0=8n\_\{0\}=8andcmax=0\.25c\_\{\\max\}=0\.25\.

###### Lemma B\.6\(Positive denominator algebra and bounds\)\.

Under Definition[B\.3](https://arxiv.org/html/2609.38661#A2.Thmdefinition3), if eachzx,t\>0z\_\{x,t\}\>0is finite, then

Ag=n0\+∑\(x,t\)∈𝒱gzx,tng\+n0\>0\.A\_\{g\}=\\frac\{n\_\{0\}\+\\sum\_\{\(x,t\)\\in\\mathcal\{V\}\_\{g\}\}z\_\{x,t\}\}\{n\_\{g\}\+n\_\{0\}\}\>0\.\(B\.11\)IfRβ​\(x\),Bx∈\[1,eβ\]R\_\{\\beta\}\(x\),B\_\{x\}\\in\[1,e^\{\\beta\}\], thenAg∈\[e−β,eβ\]A\_\{g\}\\in\[e^\{\-\\beta\},e^\{\\beta\}\]and\|c⁡\(g\)\|≤min⁡\(β,cmax\)\|c\(g\)\|\\leq\\min\(\\beta,c\_\{\\max\}\)\. Ifng=0<n0n\_\{g\}=0<n\_\{0\}, thenAg=1A\_\{g\}=1andc⁡\(g\)=0c\(g\)=0, and the formula defines no value whenng=n0=0n\_\{g\}=n\_\{0\}=0\.

The number of subtracted ones is exactlyngn\_\{g\}, so combining terms over the common denominator gives Eq\. \([B\.11](https://arxiv.org/html/2609.38661#A2.E11)\)\. Ifng\>0n\_\{g\}\>0, its numerator contains a positive term; ifng=0n\_\{g\}=0, the positive denominator forcesn0\>0n\_\{0\}\>0\. This proves strict positivity\. The ratio is a weighted average of the valueszx,tz\_\{x,t\}and the pseudo\-observation one\. Each lies in\[e−β,eβ\]\[e^\{\-\\beta\},e^\{\\beta\}\]under the additional bounds\. The same interval contains their weighted average; logarithm and clipping give the stated bound\. Substitutingng=0n\_\{g\}=0into the ratio givesAg=1A\_\{g\}=1andc⁡\(g\)=0c\(g\)=0, which proves the empty\-bucket case\. ∎

###### Proposition B\.7\(A visit\-weighted population interpretation\)\.

For this proposition only, suppose the tuples\(Xh,Bh,Vh\)\(X\_\{h\},B\_\{h\},V\_\{h\}\)are i\.i\.d\. under a fixed law, wherezh=Rβ​\(Xh\)/Bh\>0z\_\{h\}=R\_\{\\beta\}\(X\_\{h\}\)/B\_\{h\}\>0andVhV\_\{h\}counts the eligible visits to a fixed bucket in recordhh\. Assume0<𝔼​Vh<∞0<\\mathbb\{E\}V\_\{h\}<\\inftyand𝔼⁡\[Vh​zh\]<∞\\mathbb\{E\}\[V\_\{h\}z\_\{h\}\]<\\infty\. With a fixed finiten0n\_\{0\}, pooling the firstmmrecords satisfies

n0\+∑h=1mVh​zhn0\+∑h=1mVh⟶𝔼⁡\[Vh​zh\]𝔼​Vhalmost surely\.\\frac\{n\_\{0\}\+\\sum\_\{h=1\}^\{m\}V\_\{h\}z\_\{h\}\}\{n\_\{0\}\+\\sum\_\{h=1\}^\{m\}V\_\{h\}\}\\longrightarrow\\frac\{\\mathbb\{E\}\[V\_\{h\}z\_\{h\}\]\}\{\\mathbb\{E\}V\_\{h\}\}\\quad\\text\{almost surely\}\.\(B\.12\)The log and clipped\-log quantities converge to the corresponding transforms of this positive limit\.

The strong law applied toVh​zhV\_\{h\}z\_\{h\}andVhV\_\{h\}makes their averages converge to the two expectations\. Divide numerator and denominator bymm;n0/m→0n\_\{0\}/m\\to 0, and the limiting denominator is strictly positive, so the quotient converges\. Its limiting numerator is positive becausezh\>0z\_\{h\}\>0on the positive\-probability eventVh\>0V\_\{h\}\>0\. Continuity oflog\\logat a positive limit and of clipping gives the rest\. ∎

#### Three sources of statistical discrepancy\.

First, ratio\-first pooling targets an arithmetic ratio average, and the finite\-sample Jensen gap remains: for a positive random averagez¯\\bar\{z\},𝔼​log⁡z¯≤log⁡𝔼​z¯\\mathbb\{E\}\\log\\bar\{z\}\\leq\\log\\mathbb\{E\}\\bar\{z\}\. Second, random denominators matter even before taking a logarithm:

𝔼⁡\[R/B\]=𝔼⁡\[R\]​𝔼​\[B−1\]\+Cov⁡\(R,B−1\),\\mathbb\{E\}\[R/B\]=\\mathbb\{E\}\[R\]\\mathbb\{E\}\[B^\{\-1\}\]\+\\operatorname\{Cov\}\(R,B^\{\-1\}\),\(B\.13\)when these moments exist\. This identity is the definition of covariance; withR=1R=1andBBequally likely to be one or two,𝔼⁡\[R/B\]=3/4\\mathbb\{E\}\[R/B\]=3/4while𝔼⁡\[R\]/𝔼⁡\[B\]=2/3\\mathbb\{E\}\[R\]/\\mathbb\{E\}\[B\]=2/3\. Forβ≥log⁡2\\beta\\geq\\log 2these denominators obey the preceding bounded range; shrinkage and clipping add further differences\.

Third, visit pooling weights a record byVhV\_\{h\}, so its limiting ratio in Eq\. \([B\.12](https://arxiv.org/html/2609.38661#A2.E12)\) is a visit\-weighted average\. Repeated visits share the terminal reward and can be strongly correlated\. For a fixed numbern\>0n\>0of random visit ratios with finite second moments,

Var⁡\(1n​∑v=1nzv\)=1n2​∑v=1n∑w=1nCov⁡\(zv,zw\),\\operatorname\{Var\}\\\!\\left\(\\frac\{1\}\{n\}\\sum\_\{v=1\}^\{n\}z\_\{v\}\\right\)=\\frac\{1\}\{n^\{2\}\}\\sum\_\{v=1\}^\{n\}\\sum\_\{w=1\}^\{n\}\\operatorname\{Cov\}\(z\_\{v\},z\_\{w\}\),by bilinearity of covariance, so a visit count measures visits rather than independent tasks\. Structural reads lag one batch, and buckets also include paired reference prefixes \(Section[C\.5](https://arxiv.org/html/2609.38661#A3.SS5)\)\.

### B\.5Value\-Head Projection and Calibration

###### Definition B\.4\(Feature\-level regression distribution\)\.

Let𝒟\\mathcal\{D\}be the actual distribution of scored training pairs\(S,Y\)\(S,Y\)for the value head, whereXXis the terminal continuation record fromSSandY=r⁡\(X\)∈\[0,1\]Y=r\(X\)\\in\[0,1\]its reward\. For Eq\. \([5](https://arxiv.org/html/2609.38661#S4.E5)\),𝒟\\mathcal\{D\}denotes the actual scored\-pair law represented by𝒟ρ\\mathcal\{D\}\_\{\\rho\}, including its stated data filters\. The input isz⁡\(S\)=\[f⁡\(S\);onehot⁡\(tasktype⁡\(q\)\)\]∈ℝ30\+6z\(S\)=\[f\(S\);\\operatorname\{onehot\}\(\\operatorname\{tasktype\}\(q\)\)\]\\in\\mathbb\{R\}^\{30\+6\}\. WriteZf=z⁡\(S\)Z\_\{f\}=z\(S\)for the random feature vector; it is unrelated to the normalizerZ⁡\(C\)Z\(C\)\. Definem𝒟​\(Zf\)=𝔼𝒟​\[Y∣Zf\]m\_\{\\mathcal\{D\}\}\(Z\_\{f\}\)=\\mathbb\{E\}\_\{\\mathcal\{D\}\}\[Y\\mid Z\_\{f\}\]\.

###### Proposition B\.8\(Feature\-conditional MSE projection and calibration\)\.

Over measurable square\-integrable functions ofZfZ\_\{f\}, the population risk in Eq\. \([5](https://arxiv.org/html/2609.38661#S4.E5)\) decomposes as

𝔼𝒟​\(v⁡\(Zf\)−Y\)2=𝔼𝒟​\(v⁡\(Zf\)−m𝒟​\(Zf\)\)2\+𝔼𝒟​Var⁡\(Y∣Zf\)\.\\mathbb\{E\}\_\{\\mathcal\{D\}\}\(v\(Z\_\{f\}\)\-Y\)^\{2\}=\\mathbb\{E\}\_\{\\mathcal\{D\}\}\(v\(Z\_\{f\}\)\-m\_\{\\mathcal\{D\}\}\(Z\_\{f\}\)\)^\{2\}\+\\mathbb\{E\}\_\{\\mathcal\{D\}\}\\operatorname\{Var\}\(Y\\mid Z\_\{f\}\)\.\(B\.14\)Its unique minimizer up to𝒟\\mathcal\{D\}\-null sets ism𝒟m\_\{\\mathcal\{D\}\}, and

𝔼𝒟​\[Y∣m𝒟​\(Zf\)\]=m𝒟​\(Zf\)\.\\mathbb\{E\}\_\{\\mathcal\{D\}\}\[Y\\mid m\_\{\\mathcal\{D\}\}\(Z\_\{f\}\)\]=m\_\{\\mathcal\{D\}\}\(Z\_\{f\}\)\.\(B\.15\)For any such predictorv⁡\(Zf\)v\(Z\_\{f\}\), its squared mean\-calibration error is at most its excess risk above this unrestricted optimum:

𝔼𝒟​\[\(𝔼𝒟​\[Y∣v⁡\(Zf\)\]−v⁡\(Zf\)\)2\]≤𝔼𝒟​\(m𝒟​\(Zf\)−v⁡\(Zf\)\)2\.\\mathbb\{E\}\_\{\\mathcal\{D\}\}\\\!\\bigl\[\(\\mathbb\{E\}\_\{\\mathcal\{D\}\}\[Y\\mid v\(Z\_\{f\}\)\]\-v\(Z\_\{f\}\)\)^\{2\}\\bigr\]\\leq\\mathbb\{E\}\_\{\\mathcal\{D\}\}\(m\_\{\\mathcal\{D\}\}\(Z\_\{f\}\)\-v\(Z\_\{f\}\)\)^\{2\}\.\(B\.16\)

Writev−Y=\(v−m𝒟\)\+\(m𝒟−Y\)v\-Y=\(v\-m\_\{\\mathcal\{D\}\}\)\+\(m\_\{\\mathcal\{D\}\}\-Y\)and expand its square\. The cross term has conditional expectation zero givenZfZ\_\{f\}because𝔼⁡\[Y−m𝒟∣Zf\]=0\\mathbb\{E\}\[Y\-m\_\{\\mathcal\{D\}\}\\mid Z\_\{f\}\]=0\. The remaining conditional squared error isVar⁡\(Y∣Zf\)\\operatorname\{Var\}\(Y\\mid Z\_\{f\}\), proving the risk identity\. The first term is nonnegative and vanishes exactly whenv=m𝒟v=m\_\{\\mathcal\{D\}\}almost surely, which proves the minimizer claim\. Sincem𝒟m\_\{\\mathcal\{D\}\}is measurable with respect toZfZ\_\{f\}, the tower property gives Eq\. \([B\.15](https://arxiv.org/html/2609.38661#A2.E15)\)\. For the last bound,vvis also a function ofZfZ\_\{f\}, and hence

𝔼⁡\[Y∣v\]−v=𝔼⁡\[m𝒟−v∣v\]\.\\mathbb\{E\}\[Y\\mid v\]\-v=\\mathbb\{E\}\[m\_\{\\mathcal\{D\}\}\-v\\mid v\]\.Conditional Jensen’s inequality for the square followed by expectation gives Eq\. \([B\.16](https://arxiv.org/html/2609.38661#A2.E16)\)\. ∎

#### Complete\-history values and a finite sigmoid network\.

Equality ofm𝒟​\(z​\(s\)\)m\_\{\\mathcal\{D\}\}\(z\(s\)\)with the ideal reference valueμρ\(s,C\)=𝔼ρ\[r\(X\)∣s,C\]\\mu\_\{\\rho\}\(s,C\)=\\mathbb\{E\}\_\{\\rho\}\[r\(X\)\\mid s,C\]holds when the continuation\-data law is compatible and the features are sufficient for that conditional mean; two histories that share features but have values1/41/4and3/43/4show why sufficiency is needed\. Calibration above is with respect to𝒟\\mathcal\{D\}\.

The implemented head is the sigmoid MLP of Eq\. \([5](https://arxiv.org/html/2609.38661#S4.E5)\) with30\+630\+6inputs and hidden width 32: the 30 execution features and a six\-slot task\-type one\-hot, one slot per IID benchmark\. Finite logits produce outputs strictly between zero and one; an unrestricted sigmoid predictor approaches a conditional mean of zero or one by clipping it to\[ϵ,1−ϵ\]\[\\epsilon,1\-\\epsilon\]and taking its logit, with squared excess risk at mostϵ2\\epsilon^\{2\}\. The head estimates a success probability for binary rewards and an expected score otherwise\. Its weights are exported per batch and evaluated on CPU with no extra encoder call\. The two larger outcome heads on\[hρ​\(s\);f​\(s\)\]\[h\_\{\\rho\}\(s\);f\(s\)\]serve as diagnostics\.

### B\.6The Implemented Optimum and the Ideal Flow

###### Proposition B\.9\(Squared\-log regression and the arithmetic normalizer\)\.

In the one\-action executor of Remark[A\.4](https://arxiv.org/html/2609.38661#A1.Thmremark4), take a shared scalar root valueu0u\_\{0\}and the exact terminal boundary\. Taking expectation under that reference executor law, forβ\>0\\beta\>0the population AnchorTB loss is minimized atu0=β/2u\_\{0\}=\\beta/2, not at the ideal rootlog⁡\[\(1\+eβ\)/2\]\\log\[\(1\+e^\{\\beta\}\)/2\], and its minimum isβ2/4\>0\\beta^\{2\}/4\>0\.

There is one action and both policies assign it probability one, soΔ​ℓ0=0\\Delta\\ell\_\{0\}=0,T=KT=1T=K\_\{T\}=1\. The terminal log reward is zero orβ\\beta, each with probability1/21/2\. Consequently

ℒ⁡\(u0\)=12​u02\+12​\(u0−β\)2=\(u0−β/2\)2\+β2/4\.\\mathcal\{L\}\(u\_\{0\}\)=\\tfrac\{1\}\{2\}u\_\{0\}^\{2\}\+\\tfrac\{1\}\{2\}\(u\_\{0\}\-\\beta\)^\{2\}=\(u\_\{0\}\-\\beta/2\)^\{2\}\+\\beta^\{2\}/4\.This proves the minimizer and minimum; alsoZ=\(1\+eβ\)/2Z=\(1\+e^\{\\beta\}\)/2\. Strict concavity oflog\\logon the two distinct rewards gives𝔼​log⁡Rβ=β/2<log⁡𝔼​Rβ\\mathbb\{E\}\\log R\_\{\\beta\}=\\beta/2<\\log\\mathbb\{E\}R\_\{\\beta\}forβ\>0\\beta\>0, and both vanish atβ=0\\beta=0\. ∎

The idealu∗​\(s∣C\)u^\{\*\}\(s\\mid C\), the measured task anchoruq,xu\_\{q,x\}, the implementedu~x​\(s\)\\widetilde\{u\}\_\{x\}\(s\), and a population regression minimizer are distinct objects\. The first is identified by a conditional expectation and a recursion; the next two are specified estimators; the last on its function class and training law\.

## Appendix CSequential Validation and the Training Loop

Paired validation compares two specified rollout regimes\. We identify that effect, prove fixed\-look directional tests and the error accounting of an adaptive run, and close with the training loop\.

### C\.1Paired Intervention Regimes and Effect Estimands

###### Definition C\.1\(Paired regimes, counts, and empty samples\)\.

For a registered candidate, the two arms start from the empty team with the same task record, runtime configuration, value\-head snapshot, and forced first role\. The candidate arm binds the candidate at its firstadd\_agent; the control arm does not, and keeps the candidate unavailable later\. Both arms subsequently use the reference policy under their respective histories\. They share the other validated skills, and their later decisions and observations may differ\. For complete pairii, setbi±=𝟏\{r\(xi±\)≥1/2\}b\_\{i\}^\{\\pm\}=\\mathbf\{1\}\\\{r\(x\_\{i\}^\{\\pm\}\)\\geq 1/2\\\}and define

W\\displaystyle W=∑i=1N𝟏​\{bi\+=1,bi−=0\},\\displaystyle=\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\\\{b\_\{i\}^\{\+\}=1,b\_\{i\}^\{\-\}=0\\\},L\\displaystyle\\qquad L=∑i=1N𝟏​\{bi\+=0,bi−=1\},\\displaystyle=\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\\\{b\_\{i\}^\{\+\}=0,b\_\{i\}^\{\-\}=1\\\},T0\\displaystyle T\_\{0\}=∑i=1N𝟏\{bi\+=bi−\},\\displaystyle=\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\\\{b\_\{i\}^\{\+\}=b\_\{i\}^\{\-\}\\\},D\\displaystyle\\qquad D=W\+L,N=W\+L\+T0\.\\displaystyle=W\+L,\\qquad N=W\+L\+T\_\{0\}\.\(C\.1\)WhenN\>0N\>0,Δ^=\(W−L\)/N\\hat\{\\Delta\}=\(W\-L\)/N\. ForN=0N=0the empirical effect is undefined, rather than an estimate of zero\. WithD=0D=0, both directional p\-values below are defined to be one, and there is no rejection\.

###### Lemma C\.1\(Paired effect identity\)\.

ForN\>0N\>0,

Δ^=1N​∑i=1N\(bi\+−bi−\),\|Δ^\|≤1\.\\hat\{\\Delta\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\(b\_\{i\}^\{\+\}\-b\_\{i\}^\{\-\}\),\\qquad\|\\hat\{\\Delta\}\|\\leq 1\.\(C\.2\)For a specified pair distribution, letpWp\_\{W\}andpLp\_\{L\}be its win and loss probabilities\. Its threshold\-pass effect is

Δ=ℙ⁡\(b\+=1\)−ℙ⁡\(b−=1\)=pW−pL\.\\Delta=\\mathbb\{P\}\(b^\{\+\}=1\)\-\\mathbb\{P\}\(b^\{\-\}=1\)=p\_\{W\}\-p\_\{L\}\.\(C\.3\)Neither identity requires independence of the two arms within a pair\.

A win contributes one tobi\+−bi−b\_\{i\}^\{\+\}\-b\_\{i\}^\{\-\}, a loss contributes minus one, and a tie contributes zero\. Summing proves the sample identity, and boundedness of each summand proves its range\. Taking expectations gives the population identity without factoring joint arm probabilities\. ∎

### C\.2Exact Fixed\-Look Directional Tests

###### Assumption C\.1\(Registered fixed\-look i\.i\.d\. pair model\)\.

Let𝒢j\\mathcal\{G\}\_\{j\}denote information available at registration of comparisonjj\. Conditional on it, the candidate, arm protocols, target pair distribution, and a finite nonnegative integerNNof complete pairs for this look are fixed\. The future pair outcomes are conditionally i\.i\.d\. under that distribution\. Within\-pair dependence is allowed\. No data used to choose the candidate are reused as its future validation outcomes under this assumption\. Denote the conditional win and loss probabilities bypW,pLp\_\{W\},p\_\{L\}, and the corresponding mean threshold effect byΔj=pW−pL\\Delta\_\{j\}=p\_\{W\}\-p\_\{L\}\.

###### Proposition C\.2\(Both composite directional nulls at a fixed look\)\.

Under Assumption[C\.1](https://arxiv.org/html/2609.38661#A3.Thmassumption1), define

p\+=ℙ\{BD≥W\},p−=ℙ\{BD≥L\},BD∼Bin\(D,1/2\),p\_\{\+\}=\\mathbb\{P\}\\\{B\_\{D\}\\geq W\\\},\\qquad p\_\{\-\}=\\mathbb\{P\}\\\{B\_\{D\}\\geq L\\\},\\qquad B\_\{D\}\\sim\\operatorname\{Bin\}\(D,1/2\),\(C\.4\)Here probabilities in the displayed tails are over the auxiliary binomial variable with the observedD,W,LD,W,Lheld fixed; use theD=0D=0convention of Definition[C\.1](https://arxiv.org/html/2609.38661#A3.Thmdefinition1)\. For everyu∈\[0,1\]u\\in\[0,1\],

Hj\+:Δj≤0\\displaystyle H\_\{j\}^\{\+\}:\\Delta\_\{j\}\\leq 0⟹ℙ⁡\(p\+≤u∣𝒢j\)≤u,\\displaystyle\\quad\\Longrightarrow\\quad\\mathbb\{P\}\(p\_\{\+\}\\leq u\\mid\\mathcal\{G\}\_\{j\}\)\\leq u,\(C\.5\)Hj−:Δj≥0\\displaystyle H\_\{j\}^\{\-\}:\\Delta\_\{j\}\\geq 0⟹ℙ⁡\(p−≤u∣𝒢j\)≤u\.\\displaystyle\\quad\\Longrightarrow\\quad\\mathbb\{P\}\(p\_\{\-\}\\leq u\\mid\\mathcal\{G\}\_\{j\}\)\\leq u\.\(C\.6\)Each inequality holds on its null event, so each tail is valid for its whole composite directional null\.

All probabilities in the proof condition on𝒢j\\mathcal\{G\}\_\{j\}\. PutpD=pW\+pLp\_\{D\}=p\_\{W\}\+p\_\{L\}\. IfpD=0p\_\{D\}=0, thenD=0D=0almost surely andp\+=p−=1p\_\{\+\}=p\_\{\-\}=1, which satisfies both inequalities; the same conclusion holds ifN=0N=0\. Otherwise, forddof positive probability and0≤w≤d0\\leq w\\leq d, the multinomial law of wins, losses, and ties gives

ℙ⁡\(W=w,L=d−w,D=d\)\\displaystyle\\mathbb\{P\}\(W=w,L=d\-w,D=d\)=N\!w\!​\(d−w\)\!​\(N−d\)\!​pWw​pLd−w​\(1−pD\)N−d,\\displaystyle=\\frac\{N\!\}\{w\!\(d\-w\)\!\(N\-d\)\!\}p\_\{W\}^\{w\}p\_\{L\}^\{d\-w\}\(1\-p\_\{D\}\)^\{N\-d\},ℙ⁡\(D=d\)\\displaystyle\\mathbb\{P\}\(D=d\)=\(Nd\)​pDd​\(1−pD\)N−d\.\\displaystyle=\\binom\{N\}\{d\}p\_\{D\}^\{d\}\(1\-p\_\{D\}\)^\{N\-d\}\.Dividing, with the usual limiting conventions when a cell probability is zero, proves

W∣D=d∼Bin\(d,ϑ\),ϑ=pW/pD,L∣D=d∼Bin\(d,1−ϑ\)\.W\\mid D=d\\sim\\operatorname\{Bin\}\(d,\\vartheta\),\\qquad\\vartheta=p\_\{W\}/p\_\{D\},\\qquad L\\mid D=d\\sim\\operatorname\{Bin\}\(d,1\-\\vartheta\)\.
UnderHj\+H\_\{j\}^\{\+\}we havepW≤pLp\_\{W\}\\leq p\_\{L\}, soϑ≤1/2\\vartheta\\leq 1/2\. With independent uniform variablesU1,…,UdU\_\{1\},\\ldots,U\_\{d\}, the sum∑i𝟏\{Ui≤ϑ\}\\sum\_\{i\}\\mathbf\{1\}\\\{U\_\{i\}\\leq\\vartheta\\\}is pointwise at most∑i𝟏\{Ui≤1/2\}\\sum\_\{i\}\\mathbf\{1\}\\\{U\_\{i\}\\leq 1/2\\\}\. Hence the former binomial is stochastically dominated by the latter\. Define

ku​\(d\)=min⁡\{k∈\{0,…,d\}:ℙ⁡\(Bd≥k\)≤u\},k\_\{u\}\(d\)=\\min\\\{k\\in\\\{0,\\ldots,d\\\}:\\mathbb\{P\}\(B\_\{d\}\\geq k\)\\leq u\\\},usingku​\(d\)=d\+1k\_\{u\}\(d\)=d\+1if the set is empty\. The tail is nonincreasing in its observed count, so\{p\+≤u\}=\{W≥ku\(d\)\}\\\{p\_\{\+\}\\leq u\\\}=\\\{W\\geq k\_\{u\}\(d\)\\\}conditional onD=dD=d, and stochastic domination then gives

ℙ⁡\(p\+≤u∣D=d\)≤ℙ⁡\(Bd≥ku​\(d\)\)≤u\.\\mathbb\{P\}\(p\_\{\+\}\\leq u\\mid D=d\)\\leq\\mathbb\{P\}\(B\_\{d\}\\geq k\_\{u\}\(d\)\)\\leq u\.Ford=0d=0, the assigned value one has the same validity property\. Averaging overDDproves Eq\. \([C\.5](https://arxiv.org/html/2609.38661#A3.E5)\)\.

UnderHj−H\_\{j\}^\{\-\},ϑ≥1/2\\vartheta\\geq 1/2, so1−ϑ≤1/21\-\\vartheta\\leq 1/2\. Apply the same uniform coupling to the conditional law ofLL\. The event\{p−≤u\}\\\{p\_\{\-\}\\leq u\\\}is\{L≥ku\(d\)\}\\\{L\\geq k\_\{u\}\(d\)\\\}, whose probability is at most the corresponding fair\-binomial upper tail and thus at mostuu\. Averaging overDD, includingD=0D=0, proves Eq\. \([C\.6](https://arxiv.org/html/2609.38661#A3.E6)\)\. ∎

###### Lemma C\.3\(A stronger fixed\-context conditional\-sign model\)\.

Fix a candidate and a prespecified finite collection of complete pairs\. Condition on the full vector of contexts𝐂\\mathbf\{C\}, the full discordance vector𝐃=\(𝟏\{bi\+≠bi−\}\)i\\mathbf\{D\}=\(\\mathbf\{1\}\\\{b\_\{i\}^\{\+\}\\neq b\_\{i\}^\{\-\}\\\}\)\_\{i\}, and registration information\. Suppose the retained win signsYi=𝟏​\{bi\+=1,bi−=0\}Y\_\{i\}=\\mathbf\{1\}\\\{b\_\{i\}^\{\+\}=1,b\_\{i\}^\{\-\}=0\\\}are conditionally independent\. If every retained sign has conditional probabilitypi≤1/2p\_\{i\}\\leq 1/2, thenp\+p\_\{\+\}is conditionally super\-uniform\. If everypi≥1/2p\_\{i\}\\geq 1/2, thenp−p\_\{\-\}is conditionally super\-uniform\. If allpi=1/2p\_\{i\}=1/2, thenW\|\(𝐂,𝐃,𝒢j\)W\\mid\(\\mathbf\{C\},\\mathbf\{D\},\\mathcal\{G\}\_\{j\}\)is exactlyBin⁡\(D,1/2\)\\operatorname\{Bin\}\(D,1/2\)\.

After conditioning, the retained index set and its sizeD=dD=dare fixed\. For the positive direction construct independent uniforms and represent each independent sign asYi=𝟏\{Ui≤pi\}Y\_\{i\}=\\mathbf\{1\}\\\{U\_\{i\}\\leq p\_\{i\}\\\}\. Whenpi≤1/2p\_\{i\}\\leq 1/2, each sign is bounded above by𝟏\{Ui≤1/2\}\\mathbf\{1\}\\\{U\_\{i\}\\leq 1/2\\\}, so their sum is stochastically dominated byBin⁡\(d,1/2\)\\operatorname\{Bin\}\(d,1/2\)\. Using the cutoffku​\(d\)k\_\{u\}\(d\)from the preceding proof gives the super\-uniform upper tail\. Whenpi≥1/2p\_\{i\}\\geq 1/2, the independent loss signs1−Yi1\-Y\_\{i\}instead have probabilities at most1/21/2; the same argument applies top−p\_\{\-\}\. When everypi=1/2p\_\{i\}=1/2, independence identifies the sum as precisely the fair binomial\. The zero\-discordance convention handles an empty set\. ∎

###### Corollary C\.4\(Two\-direction accounting at one look\)\.

If each valid direction receives levelb/2b/2, the probability of any false directional rejection at that look is at mostbb\. Also, for the p\-values in Eq\. \([C\.4](https://arxiv.org/html/2609.38661#A3.E4)\), both directions cannot simultaneously satisfyp±≤ap\_\{\\pm\}\\leq awhena<1/2a<1/2\.

Sum the error probabilities only over directions whose nulls are true\. There are at most two, each bounded byb/2b/2, so the union bound givesbb, without independence\. For the second claim, fair\-binomial symmetry andL=D−WL=D\-Wimplyp−=ℙ⁡\(BD≤W\)p\_\{\-\}=\\mathbb\{P\}\(B\_\{D\}\\leq W\), so

p\+\+p−=1\+ℙ⁡\(BD=W\)≥1,p\_\{\+\}\+p\_\{\-\}=1\+\\mathbb\{P\}\(B\_\{D\}=W\)\\geq 1,which no two numbers at mosta<1/2a<1/2satisfy\. IfD=0D=0, both p\-values equal one\. ∎

#### Two null models\.

The i\.i\.d\. pair proposition tests an average threshold effect under the registered pair distribution, possibly averaging over randomly sampled tasks\. The fixed\-context lemma assumes the stronger sign inequalities for every retained pair after conditioning on*all*contexts and discordance indicators; equally weighted contexts with certain wins in one and certain losses in the other show that an average zero effect can coexist with opposite deterministic conditional signs\.

### C\.3Adaptive Selection, Repeated Problems, and Repeated Looks

#### Planned looks and cumulative data\.

Suppose a fixed candidate has valid p\-values at prespecified cumulative sample sizesn1,n2,…n\_\{1\},n\_\{2\},\\ldots, with deterministic levelsa1,a2,…a\_\{1\},a\_\{2\},\\ldots\. Although the looks reuse data and their p\-values are correlated,

ℙ⁡\{some true\-null look rejects\}≤∑ℓℙ⁡\(pℓ≤aℓ\)≤∑ℓaℓ\.\\mathbb\{P\}\\\{\\text\{some true\-null look rejects\}\\\}\\leq\\sum\_\{\\ell\}\\mathbb\{P\}\(p\_\{\\ell\}\\leq a\_\{\\ell\}\)\\leq\\sum\_\{\\ell\}a\_\{\\ell\}\.Independence between looks is unnecessary, and early stopping only removes potential later rejections\. What matters is that the*actual*look indexed byℓ\\ellretains a valid null law; renumbering selected looks, choosing the sample size by favorable signs, reusing candidate\-selection outcomes, or filtering complete pairs on outcomes can change that law\. The registration\-based protocol below conditions on registration information, which is all the union bound requires\.

### C\.4Nominal Spending and Conditional Countable FWER

###### Proposition C\.5\(Alpha\-spending accounting\)\.

Let0<α<10<\\alpha<1, and assign each registered comparisonj≥1j\\geq 1and lookℓ≥1\\ell\\geq 1two directional allocations

aj,ℓ,d=α2​j​\(j\+1\)​ℓ​\(ℓ\+1\),d∈\{\+,−\}\.a\_\{j,\\ell,d\}=\\frac\{\\alpha\}\{2j\(j\+1\)\\ell\(\\ell\+1\)\},\\qquad d\\in\\\{\+,\-\\\}\.\(C\.7\)If each directional index is globally unique and used at most once, any realized run spends at mostα\\alpha\.

For every integerM≥1M\\geq 1,∑n=1M1/\[n⁡\(n\+1\)\]=∑n=1M\(1/n−1/\(n\+1\)\)=1−1/\(M\+1\)\\sum\_\{n=1\}^\{M\}1/\[n\(n\+1\)\]=\\sum\_\{n=1\}^\{M\}\(1/n\-1/\(n\+1\)\)=1\-1/\(M\+1\)\. Therefore, for finiteJ,LJ,L,

∑j=1J∑ℓ=1L∑d∈\{\+,−\}aj,ℓ,d=α⁡\(1−1J\+1\)​\(1−1L\+1\)\.\\sum\_\{j=1\}^\{J\}\\sum\_\{\\ell=1\}^\{L\}\\sum\_\{d\\in\\\{\+,\-\\\}\}a\_\{j,\\ell,d\}=\\alpha\\left\(1\-\\frac\{1\}\{J\+1\}\\right\)\\left\(1\-\\frac\{1\}\{L\+1\}\\right\)\.Monotone limits give total allocationα\\alpha, and the consumed indices, a subset, sum to at mostα\\alpha\. ∎

###### Proposition C\.6\(Conditional\-validity family\-wise error bound\)\.

For each registeredjj, letHj,dH\_\{j,d\}be the event that its directional null is true, measurable with respect to registration information𝒢j\\mathcal\{G\}\_\{j\}\. LetEj,ℓ,dE\_\{j,\\ell,d\}be the event that the directional test is actually performed and rejects a true null\. Unregistered or unperformed tests have empty rejection events\. Suppose the actual selection, sampling, and observation rules ensure

ℙ⁡\(Ej,ℓ,d∣𝒢j\)≤aj,ℓ,d​𝟏Hj,dalmost surely for every​j,ℓ,d\.\\mathbb\{P\}\(E\_\{j,\\ell,d\}\\mid\\mathcal\{G\}\_\{j\}\)\\leq a\_\{j,\\ell,d\}\\mathbf\{1\}\_\{H\_\{j,d\}\}\\quad\\text\{almost surely for every \}j,\\ell,d\.\(C\.8\)Then the probability of any false directional rejection is at mostα\\alpha\. More generally, if a pre\-run information fieldℱ0\\mathcal\{F\}\_\{0\}is contained in every𝒢j\\mathcal\{G\}\_\{j\}, the same bound holds conditional onℱ0\\mathcal\{F\}\_\{0\}\.

By the tower property, Eq\. \([C\.8](https://arxiv.org/html/2609.38661#A3.E8)\) givesℙ⁡\(Ej,ℓ,d\)≤aj,ℓ,d\\mathbb\{P\}\(E\_\{j,\\ell,d\}\)\\leq a\_\{j,\\ell,d\}\. For a finite rectangle of indices, the probability of the union is at most their summed probabilities\. Increasing these rectangles to all positive indices and using continuity of probability from below gives

ℙ⁡\(⋃j,ℓ,dEj,ℓ,d\)≤∑j,ℓ,dℙ⁡\(Ej,ℓ,d\)≤∑j,ℓ,daj,ℓ,d=α\.\\mathbb\{P\}\\\!\\left\(\\bigcup\_\{j,\\ell,d\}E\_\{j,\\ell,d\}\\right\)\\leq\\sum\_\{j,\\ell,d\}\\mathbb\{P\}\(E\_\{j,\\ell,d\}\)\\leq\\sum\_\{j,\\ell,d\}a\_\{j,\\ell,d\}=\\alpha\.For the conditional version, the tower property withℱ0⊆𝒢j\\mathcal\{F\}\_\{0\}\\subseteq\\mathcal\{G\}\_\{j\}first givesℙ⁡\(Ej,ℓ,d∣ℱ0\)≤aj,ℓ,d\\mathbb\{P\}\(E\_\{j,\\ell,d\}\\mid\\mathcal\{F\}\_\{0\}\)\\leq a\_\{j,\\ell,d\}\. Apply the finite conditional union bound and conditional monotone convergence to the same increasing sequence of unions\. This proves the bound almost surely givenℱ0\\mathcal\{F\}\_\{0\}\. The argument uses no independence between comparisons or looks, so it covers dependent tests\. ∎

###### Assumption C\.2\(A sufficient registration\-and\-future\-data protocol\)\.

For each comparisonjj, the protocol consists of the following conditions\. The candidate may be chosen adaptively from𝒢j\\mathcal\{G\}\_\{j\}, but its content, arm protocols, response law, and target pair distribution are then fixed for that comparison\. Conditional on𝒢j\\mathcal\{G\}\_\{j\}, future complete pairs form the i\.i\.d\. stream in Assumption[C\.1](https://arxiv.org/html/2609.38661#A3.Thmassumption1); past screening outcomes are not reused, and pair inclusion does not selectively retain favorable outcomes\. A nondecreasing countable sequence of finite nonnegative integer cumulative sample sizesnj,ℓn\_\{j,\\ell\}is𝒢j\\mathcal\{G\}\_\{j\}\-measurable and fixed before those future outcomes are seen\. Each potential p\-value is computed from the firstnj,ℓn\_\{j,\\ell\}pairs, and keeps its original planned index and allocation\. Whether a planned test is ultimately performed may depend on available information, but cannot change that potential test or its data law\. Unperformed planned tests consume no level\.

###### Corollary C\.7\(FWER under the sufficient protocol\)\.

Assumption[C\.2](https://arxiv.org/html/2609.38661#A3.Thmassumption2), the unique indices of Proposition[C\.5](https://arxiv.org/html/2609.38661#A3.Thmproposition5), and rejection at the allocations in Eq\. \([C\.7](https://arxiv.org/html/2609.38661#A3.E7)\) imply the bound of Proposition[C\.6](https://arxiv.org/html/2609.38661#A3.Thmproposition6)\.

Conditional on𝒢j\\mathcal\{G\}\_\{j\}, eachnj,ℓn\_\{j,\\ell\}is fixed\. Apply Proposition[C\.2](https://arxiv.org/html/2609.38661#A3.Thmproposition2)to the corresponding future\-data prefix to obtain, onHj,dH\_\{j,d\},ℙ⁡\(pj,ℓ,d≤aj,ℓ,d∣𝒢j\)≤aj,ℓ,d\\mathbb\{P\}\(p\_\{j,\\ell,d\}\\leq a\_\{j,\\ell,d\}\\mid\\mathcal\{G\}\_\{j\}\)\\leq a\_\{j,\\ell,d\}\. The actual false rejection event is a subset ofHj,d∩\{pj,ℓ,d≤aj,ℓ,d\}H\_\{j,d\}\\cap\\\{p\_\{j,\\ell,d\}\\leq a\_\{j,\\ell,d\}\\\}, even when the decision to execute a planned look uses earlier outcomes\. SinceHj,dH\_\{j,d\}is registration\-measurable, this subset relation proves Eq\. \([C\.8](https://arxiv.org/html/2609.38661#A3.E8)\)\. The preceding proposition then gives the claimed bound, also when cumulative prefixes overlap\. ∎

### C\.5Training Algorithm and Data Routing

#### Skill proposal rules\.

The author model runs only in an author window and distils at most one candidate per window from scored, side\-effect\-free trajectories, preferring task families with few skills and contrasting a high\-scoring run with a same\-task failure when one exists\. A candidate is a structured procedure with name, description, trigger, plan, pitfall, and constraint fields, and is filtered for duplicates and answer leakage\. Candidate slots are ordered by fewest paired observations\. Paired rollouts enter AnchorTB and structural buckets as off\-policy paths but not the task, category, and global baseline statistics, and the forced first action is scored at its true probability underθ\\thetaandρ\\rho\.

The following algorithm describes the frozen implementation\. It is distinct from the sufficient testing protocol in Assumption[C\.2](https://arxiv.org/html/2609.38661#A3.Thmassumption2)\. The same trajectory\-indexed endpoint values are reused across all intervals of each scored trajectory\. For paired records, “the same context” in the re\-scoring step means thatθ\\thetaandρ\\rhouse that record’s actual arm context,C\+C^\{\+\}orC−C^\{\-\}, and its legal mask\.

Algorithm 1One training step of EvoSteer1. 1\.Sample3232task records balanced over sources\.
2. 2\.Roll out2​θ\+2​ρ2\\theta\+2\\rhonatural trajectories per record, plus one candidate and one control rollout per record in a family with a candidate, under fixed skill\-menu and value\-head snapshots; execute every action on issue and append feedback and features\.
3. 3\.Gate the batch on completeness and risk; publish the current adapter to the sampling service and prefetch the next batch\.
4. 4\.Re\-score every action underθ\\thetaandρ\\rhowith the same context, tokens, and grammar mask\.
5. 5\.Fit the value head and auxiliary heads on legal reference states\.
6. 6\.Build same\-task shrinkage baselines; read structural buckets from earlier batches; accumulate this batch’s ratios\.
7. 7\.Compute all subtrajectory residuals of Eq\. \([7](https://arxiv.org/html/2609.38661#S4.E7)\) and updateθ\\thetaandbψb\_\{\\psi\}\.
8. 8\.Accumulate paired wins, losses, and ties; at each validation round, run the author window and the sequential tests of Eq\. \([11](https://arxiv.org/html/2609.38661#S4.E11)\)\.
9. 9\.Save a checkpoint with adapter, heads, optimiser, statistics, skills, and test state\.

Table 3:Configuration of the frozen implementation used for the experiments\.
#### Data routing\.

Natural healthy non\-paired reference trajectories supply the hierarchical task, category, and global statistics\. Natural and paired paths can supply the AnchorTB regression\. Structural buckets also receive paired reference prefixes, without the forced\-prefix filter of the value\-head data and with their visit multiplicities retained\. Conditional flow statements apply to a fixedCC\.

### C\.6Proofs of the Main\-Text Propositions

The first main\-text proposition is procedural: legal repairs execute before the next decision and preserve the history representation \(Lemma[A\.1](https://arxiv.org/html/2609.38661#A1.Thmproposition1)\)\. The second follows from the action coefficient in Proposition[A\.14](https://arxiv.org/html/2609.38661#A1.Thmproposition14)and the estimator components in Definition[B\.1](https://arxiv.org/html/2609.38661#A2.Thmdefinition1)\. The third combines candidate\-slot and status\-transition rules with Proposition[C\.5](https://arxiv.org/html/2609.38661#A3.Thmproposition5)’s nominal accounting; its statistical interpretation is fixed\-look validity in Proposition[C\.2](https://arxiv.org/html/2609.38661#A3.Thmproposition2)and the conditional whole\-run result of Proposition[C\.6](https://arxiv.org/html/2609.38661#A3.Thmproposition6)\.

## Appendix DExperimental Details

#### Benchmarks and splits\.

The six IID benchmarks \(HotpotQA, NQ\-Open, MedQA, AIME 2026, MBPP\+, ALFWorld\) supply the training tasks, and their test items are disjoint from those tasks; all 30 AIME 2026 problems are test items, and the mathematics training tasks are compiled from earlier AIME problems\. The six OOD benchmarks \(TriviaQA, MuSiQue, GPQA, MATH\-Hard, SWE\-Bench Verified, WebShop\) are evaluation\-only\. Every other test set has 128 items, so one run’s 0/1 metric lies on ak/128k/128grid \(k/30k/30for AIME 2026\) and a reported five\-run mean on ak/640k/640grid \(k/150k/150\); GPQA uses 128 random Diamond questions, and MATH\-Hard is the level\-5 subset of MATH\. The exact training and test splits and the full configuration are released in our code repository\.

#### Metrics\.

Answer exact match \(Ans EM\) and token\-level F1 \(Ans F1\) use SQuAD\-style answer normalization and the maximum over reference aliases\. Accuracy \(Acc\.\) is the share of correct final answers: the chosen option on MedQA and GPQA and the final value on AIME 2026 and MATH\-Hard\. Pass@1 on MBPP\+ and the success rate \(SR\) on WebShop follow the official evaluators; SR on ALFWorld is the share of episodes that complete the task within 50 steps\. The resolved rate is the share of SWE\-Bench Verified issues whose patch passes the official tests\.

#### Baselines and fairness\.

Unless a column is labeled otherwise, every method uses Qwen3\.5\-9B both as the frozen executor and as the model it trains, and every trained method uses the same training tasks; SFT fits the reference answers of those tasks\. Each baseline runs in the best configuration reported by its authors\. The DeepSeek\-V4\-Flash column prompts that model directly, as a larger reference point\. GRPO†in Table[1](https://arxiv.org/html/2609.38661#S5.T1)fine\-tunes the backbone itself, whereas GRPO in Figure[6](https://arxiv.org/html/2609.38661#S5.F6)\(a,b\) trainsπθ\\pi\_\{\\theta\}inside the EvoSteer architecture with the same harness, actions, and tools\.

#### Ablation variants and fixed paradigms\.

Each ablation in Table[2](https://arxiv.org/html/2609.38661#S5.T2)removes one mechanism and restores the usual alternative:−\-interleaved execution builds the whole graph before running it;−\-execution features keeps interleaved execution and the textual feedbackotexeco\_\{t\}^\{\\mathrm\{exec\}\}but removes the feature vectorfffrom the state ofπθ\\pi\_\{\\theta\}, so the orchestrator reads the textual history and the value estimate but notff, while the value head andbψb\_\{\\psi\}are unchanged;−\-learnable repair masksrerunanddrop;−\-reference value head withholdsv^k\\hat\{v\}\_\{k\}andΔ​v^\\Delta\\hat\{v\}from the state;−\-measured flows learnsuqu\_\{q\}as a free scalar instead of reading it fromp^q\\hat\{p\}\_\{q\};−\-flow corrections turns offc⁡\(g\)c\(g\)andbψb\_\{\\psi\};−\-skill evolution removes the whole skill mechanism \(no author window, no candidates, no paired rollouts, and an empty skill library\), so the team uses roles and tools only;−\-sequential validation admits a candidate after a fixed number of successes\. The four fixed paradigms decide the team before execution, on the same executor and budget as EvoSteer: a ReAct\-style agent with all tools\([Yao et al\., 2022b](https://arxiv.org/html/2609.38661#bib.bib57)\), one hand\-designed multi\-agent template for all tasks, a planner that writes the graph once and then runs it, and a workflow searched offline on the training tasks and then frozen\.

#### Test\-time protocol and transfer\.

At test time EvoSteer runs the trainedπθ\\pi\_\{\\theta\}with the value head and the skill library frozen; no reference, paired, or validation rollouts are drawn\. The transfer study of Figure[5](https://arxiv.org/html/2609.38661#S5.F5)reuses this orchestrator unchanged and swaps only the executor, called through its API identifier: gpt\-5\.6\-luna \(GPT\-5\.6 Luna\), grok\-4\.5 \(Grok 4\.5\), claude\-haiku\-4\-5 \(Claude Haiku 4\.5\), deepseek\-v4\-flash \(DeepSeek\-V4\-Flash\), gemini\-3\.5\-flash \(Gemini 3\.5 Flash\), and glm\-5\.3\-flash \(GLM\-5\.3\-Flash\)\. Frozen\-backbone scores prompt each model directly, as in the v4\-flash column of Table[1](https://arxiv.org/html/2609.38661#S5.T1)\. Each OOD benchmark uses its IID counterpart’s task type \(e\.g\., MuSiQue that of HotpotQA, SWE\-Bench that of MBPP\+\), so OOD results test generalization across matched task types\.

#### Skill author\.

The skill author only writes candidate skills in author windows; validation, training, and testing are unchanged\. Table[4](https://arxiv.org/html/2609.38661#A4.T4)replaces DeepSeek\-V4\-Flash with a stronger author, GPT\-5\.6\-Luna, and with the executor itself, Qwen3\.5\-9B, and reports the averages of Table[1](https://arxiv.org/html/2609.38661#S5.T1)\. The stronger author raises every average, self\-written skills lower them by at most 1\.56 points, and all three settings stay above the strongest baseline of Table[1](https://arxiv.org/html/2609.38661#S5.T1)in every column\.

Table 4:EvoSteer with different skill authors: five\-run means averaged as in Table[1](https://arxiv.org/html/2609.38661#S5.T1)\(Ans EM over the question\-answering benchmarks, Acc\. over the others\)\.
#### Objectives and cost\.

In Figure[6](https://arxiv.org/html/2609.38661#S5.F6)\(a,b\) every objective trains the same untrained architecture of Table[2](https://arxiv.org/html/2609.38661#S5.T2), with the harness, action space, tool set, and skill library unchanged\. Tempered TB follows the implementation of[Zhang et al\. \(2026c\)](https://arxiv.org/html/2609.38661#bib.bib64): the log reward is tempered asβ​log⁡\(R\+ϵ\)\\beta\\log\(R\+\\epsilon\),log⁡Z\\log Zand the backward policy are learned, and each edge’s log\-probability is normalized by its token count\. GPU time per episode is the compute of the parameter update \(PPO’s counts actor and critic, GRPO’s includes its KL term\)\. Tokens per problem is the rollout cost, prefill plus decode; training rollouts run the same orchestration as inference, so it is also the inference cost\. Training EvoSteer for 240 steps used 95\.98 H800 GPU hours and 1,580,691,323 tokens in total\. The fixed\-count admission of Table[2](https://arxiv.org/html/2609.38661#S5.T2)admits a candidate afterk=5k=5successes; everykkfrom 1 to 10 gives nearly the same results\.

#### Diagnostics\.

In Figure[5](https://arxiv.org/html/2609.38661#S5.F5)\(b\), the reference curve is the accuracy of the frozenρ\\rhoon the natural reference rollouts of the same batch, which changes with the batch althoughρ\\rhodoes not; the anchor\-only level evaluates the AnchorTB loss withbψ=0b\_\{\\psi\}=0andΔ​ℓt=0\\Delta\\ell\_\{t\}=0, so that each flow equals its measured partclip\[0,β\]⁡\(uq,x\+c⁡\(g\)\)\\operatorname\{clip\}\_\{\[0,\\beta\]\}\(u\_\{q,x\}\+c\(g\)\)\. In Figure[6](https://arxiv.org/html/2609.38661#S5.F6)\(c\), the predictor is the leave\-one\-out task anchorp^q\\hat\{p\}\_\{q\}\(equivalentlyuqu\_\{q\}, a monotone transform with the same AUC\), the target is the final correctness of every natural reference rollout after step 2, and the prediction is made before the episode starts; the task\-family mean uses the same reference statistics at the same step\. In Figure[6](https://arxiv.org/html/2609.38661#S5.F6)\(d\), the gain of a skill is its paired effect: the same task and initial state, the first action forced to bind the skill or not, and both arms continued by the frozenρ\\rho\. In Figure[6](https://arxiv.org/html/2609.38661#S5.F6)\(e\), an edit is better, the same, or worse when the task grader’s score of the designated answer rises, stays, or falls from just before to just after the edit on the same trajectory\. In Figure[6](https://arxiv.org/html/2609.38661#S5.F6)\(f\), the orchestrator is trained once with each of the Qwen3\.5\-9B and DeepSeek\-V4\-Flash executors, and the gain of value\-guided replanning is measured during training\.

#### Runs and aggregation\.

Every test score is the mean over five independent runs\. Table 1 also reports their standard deviation; for the Avg\. rows it is∑bσb2/n\\sqrt\{\\sum\_\{b\}\\sigma\_\{b\}^\{2\}\}/novernnbenchmarks\.

## Appendix ECase Study

### E\.1Code That Passes the Visible Test

We present a case from MBPP\+ in which the executor returns code that looks finished: it is short, it runs, and it passes the only test given in the task\. The code is wrong, andv^\\hat\{v\}falls by0\.2160\.216at that point; the orchestrator reruns the agent instead of submitting, and the rerun passes every hidden test\.

MBPP\+ Problem 801 \(training step 87\)Write a python function to count the number of equal numbers from three given integers\. Your code should pass this test: assert test\_three\_equal\(1,1,1\) == 3 Ground Truth \(hidden tests; only the first is shown in the task\):test\_three\_equal\(1,1,1\) == 3,test\_three\_equal\(\-1,\-2,\-3\) == 0,test\_three\_equal\(1,2,2\) == 2

Final Team:planner n0with the skill below \(output agent\) \(6 rounds, 4 executor calls\)

Skill on the Menu: Implement small Python functions from behavioral specificationsPurpose:Translate the task wording into the simplest function that preserves the specified semantics, especially indexing, ordering, and duplicate\-handling rules, then return executable code rather than planning or verification prose\. Use when:Use when asked to write a short Python function from a natural\-language requirement and one or more example assertions\. Steps:1\.Extract the exact function name, parameters, return value, and indexing convention from the prompt\.2\.Identify semantic keywords before coding: positional “kth element” means access the original sequence unless sorting is explicitly requested; \[…\]3\.Choose the direct operation or standard algorithm matching those semantics\. \[…\]4\.Mentally test more than the supplied example, including empty or boundary positions where applicable, duplicates for insertion behavior, unsorted input for positional retrieval, and negative values for extrema\.5\.Return a complete, syntactically valid function definition with the exact requested name\. Keep the final response focused on executable code\.Pitfalls to avoid:Do not infer behavior solely from a single assertion: sorting an array can accidentally pass a positional\-access example while changing the required semantics\. Do not confuse right insertion with lower\-bound insertion; \[…\] Do not output only a plan, review, or statement that code is correct\. \[…\] Avoid adding arbitrary validation or changed exception behavior unless required by the specification\. Constraint:Preserve the input’s order unless the task explicitly requests sorting or an order statistic\. Match the requested function name and argument order exactly\. Provide closed, valid Python code and ensure the function returns the result rather than merely printing it\. Do not include source\-specific answers or rely only on the visible test case\.

Round\-by\-Round Interaction Log

Round 1–2: Planner n0 Writes a PlanRound 1 \- Observation:Empty team;v^=0\.605\\hat\{v\}=0\.605; 15 legal actions; the menu holds the skill above and one other Python skill\.Round 1 \- Action:\{"kind":"ADD\_AGENT", "node\_id":"n0", "role\_id":"planner"\}Round 1 \- Input Prompt: Role: Write a short step\-by\-step plan for the public task\. Do not solve it\. Task:⟨\\langleproblem above⟩\\rangle Perform your role using the task\. Return your result\.Round 2 \- Observation:n0 holds a plan without code;v^=0\.623\\hat\{v\}=0\.623\(Δ​v^=\+0\.018\\Delta\\hat\{v\}=\{\\color\[rgb\]\{0,0\.45,0\}\+0\.018\}\); 20 legal actions\.Round 2 \- Action:\{"kind":"RERUN\_AGENT", "node\_id":"n0"\}Round 1 \- Executor Output:\(951 characters; excerpt\) \[…\] 3\.Compare Pairs: Check if the first number equals the second\. \[…\]Note: Depending on the specific interpretation of “equal numbers” in the context of the test casetest\_three\_equal\(1,1,1\) == 3, the logic likely needs to count every matching pair found\.\[…\]Round 2 \- Executor Output:\(981 characters; excerpt\) \[…\]Note: To satisfy the test casetest\_three\_equal\(1,1,1\) == 3, the logic must count every unique pair comparison that results in equality\.4\.Increment Logic: If any of the pair comparisons are true, increment thecountvariable by 1\. \[…\]Round 1–2 \- Feedback:\[Status\]: answered, no code∣\\mid\[Tokens\]: 86 in, 243 out; 338 in, 256 outTeam State:\(empty\)→\\toplanner n0: planGrader \(hidden\):✗ no code

Round 3: Bind the Skill to n0; the Code Looks FinishedRound 3 \- Observation:n0 holds a plan that counts equal pairs;v^=0\.569\\hat\{v\}=0\.569\(Δ​v^=−0\.054\\Delta\\hat\{v\}=\{\\color\[rgb\]\{0\.7,0,0\}\-0\.054\}\); 20 legal actions\.Round 3 \- Action:\{"kind":"BIND\_SKILL", "node\_id":"n0", "skill\_id":"evosteer\-2744c58b1d3dd27b8cb7"\}Round 3 \- Input Prompt:\(the role asks for a plan and the skill asks for executable code\) Role: Write a short step\-by\-step plan for the public task\. Do not solve it\. Task:⟨\\langleproblem above⟩\\rangle Procedure to follow \(skill evosteer\-2744c58b1d3dd27b8cb7\):⟨\\langleskill above⟩\\rangle Your previous output:⟨\\langlethe plan of Round 2⟩\\rangle Perform your role using the task, the procedure and your previous output\. Return your result\.Round 3 \- Executor Output:\(177 characters; full\)def test\_three\_equal\(a, b, c\):count = 0if a == b:count \+= 1if b == c:count \+= 1if a == c:count \+= 1return countRound 3 \- Feedback:\[Status\]: answered∣\\mid\[Answer\]: code∣\\mid\[Tokens\]: 872 in, 65 outSurface Check:passes the visible test\(1,1,1\)↦3\(1,1,1\)\\mapsto 3; also\(−1,−2,−3\)↦0\(\-1,\-2,\-3\)\\mapsto 0Grader \(hidden\):✗ wrong:\(1,2,2\)↦1\(1,2,2\)\\mapsto 1, expected22

Round 4:v^\\hat\{v\}Falls, and n0 Is Rerun Instead of SubmittedRound 4 \- Observation:n0 holds the code above in full, with an answer present and no failure;v^=0\.353\\hat\{v\}=0\.353\(Δ​v^=−0\.216\\Delta\\hat\{v\}=\{\\color\[rgb\]\{0\.7,0,0\}\-0\.216\}\), the largest drop of the episode; 19 legal actions, includingset\_output\.Round 4 \- Action:\{"kind":"RERUN\_AGENT", "node\_id":"n0"\}Round 4 \- Input Prompt: Role:⟨\\langleplanner instruction⟩\\rangle Task:⟨\\langleproblem above⟩\\rangle Procedure to follow \(skill evosteer\-2744c58b1d3dd27b8cb7\):⟨\\langleskill above⟩\\rangle Your previous output:⟨\\langlethe code of Round 3⟩\\rangle Perform your role using the task, the procedure and your previous output\. Return your result\.Round 4 \- Executor Output:\(162 characters; full\)def test\_three\_equal\(a, b, c\):if a == b == c:return 3elif a == b or b == c or a == c:return 2else:return 0Round 4 \- Feedback:\[Status\]: answered∣\\mid\[Answer\]: code∣\\mid\[Tokens\]: 681 in, 59 outTeam State:planner n0: corrected codeGrader \(hidden\):✓ correct\(3 of 3 tests\)

Round 5–6: Set Output and StopRound 5 \- Observation:v^=0\.396\\hat\{v\}=0\.396\(Δ​v^=\+0\.043\\Delta\\hat\{v\}=\{\\color\[rgb\]\{0,0\.45,0\}\+0\.043\}\); 19 legal actions\.Round 5 \- Action:\{"kind":"SET\_OUTPUT", "node\_id":"n0"\}Round 6 \- Observation:n0 is the output agent;v^=0\.521\\hat\{v\}=0\.521\(Δ​v^=\+0\.125\\Delta\\hat\{v\}=\{\\color\[rgb\]\{0,0\.45,0\}\+0\.125\}\); 18 legal actions\.Round 6 \- Action:\{"kind":"STOP"\}Final Status:\[Output agent\]:planner n0∣\\mid\[Tests\]: 3 of 3∣\\mid\[Reward\]:1\.0

Key Observations:At Round 4 every visible sign says the task is done: the call returned an answer without failure, the code is short and valid, and it passes the only test in the task\. The code is still wrong, because it counts equal pairs and returns 1 on\(1,2,2\)\(1,2,2\)\.v^\\hat\{v\}, which reads execution features and the task type, not the code, falls by0\.2160\.216to0\.3530\.353at this state, its largest drop in the episode, and the orchestrator reruns the agent instead of setting it as the output\. The rerun rewrites the function to count equal numbers and passes all three hidden tests, andv^\\hat\{v\}rises by0\.0430\.043and0\.1250\.125\. The rewrite follows the bound skill’s first pitfall, which warns against inferring behavior from a single assertion\. Four of the other five rollouts of this problem in the same batch fail: three, including both reference rollouts, submit code that passes the visible test and fails a hidden one, and one returns a Boolean\. The only other rollout that passes is the paired one that binds the same skill\.

相似文章

SkillFlow:流程驱动的递归技能演化用于智能体编排

arXiv cs.AI

SkillFlow 提出了一种基于流程驱动的递归技能演化框架,用于基于大语言模型的智能体编排,采用 Tempered Trajectory Balance 来防止策略崩溃并提供透明的信用分配。在 14 个数据集上的实验表明,在问答、数学、代码和决策制定任务中,该框架显著优于基线方法。