在学生状态处学习教师续写
摘要
论文提出了 OLIVE,一种在线知识蒸馏方法,由学生生成前缀、教师在此基础上继续生成,旨在解决现有蒸馏方法中协变量偏移与监督信号碎片化的问题。在相当的训练预算下,OLIVE 在推理任务和智能体任务上的表现优于 on-policy 蒸馏和离线 SFT。
arXiv:2609.36246v1 Announce Type: new
Abstract: We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.
查看缓存全文
缓存时间: 2026/09/30 09:51
# Learning from Teacher Continuations at Student States Ongoing work
Source: [https://arxiv.org/html/2609.36246](https://arxiv.org/html/2609.36246)
\\paperurl
https://dylanzsz\.github\.io/olive/
Dylan Zhang\(Project lead\)Affiliation:University of Illinois at Urbana\-ChampaignHuaibo ChenAffiliation:Massachusetts Institute of TechnologySuhao YuAffiliation:University of PennsylvaniaYihang SunAffiliation:University of Illinois at Urbana\-ChampaignZhanyang JinAffiliation:University of Illinois at Urbana\-ChampaignJiaying YeAffiliation:University of WashingtonDianqi LiAffiliation:University of Illinois at Urbana\-ChampaignPrasanna SattigeriAffiliation:International Business MachinesKamal Youcef\-ToumiAffiliation:Massachusetts Institute of TechnologyHao PengAffiliation:University of Illinois at Urbana\-Champaign
###### Abstract
We presentOLIVE\(OnLine InterVEntion\)\. At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross\-entropy computed on the teacher\-generated tokens\. Each design choice targets a corresponding limitation of existing distillation methods: \(1\) sequential covariate shift in offline supervised fine\-tuning \(SFT\) on fixed teacher trajectories, \(2\) fragmented supervision under prefix failure in token\-level on\-policy distillation \(OPD\), and \(3\) the need for access to teacher token probabilities in distribution\-matching distillation\.OLIVEachieves higher reasoning performance than OPD \(with a top\-16 KL approximation\) at comparable GPU\-hour cost\. Our asynchronous implementation further reducesOLIVE’s total training time by 23\.8%\. We evaluateOLIVEon both hard reasoning tasks and agentic tasks which reflects modern post\-training scenarios, and it consistently outperforms existing distillation methods under the same training budget\. By regenerating prefixes from the evolving student,OLIVEcontinues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student\. Using only text from GPT\-5\.4\-mini, continuously training withOLIVEoutperforms offline SFT from the same teacher by 13% on ScienceWorld\. These results supportOLIVEas an effective and efficient approach to online language\-model distillation\.
## 1Introduction
Knowledge distillation transfers capability from a stronger teacher to a weaker student, either by matching the teacher’s token\-level distributions or by training on text the teacher generates\([Hinton et al\., 2015](https://arxiv.org/html/2609.36246#bib.bib22);[Kim and Rush, 2016](https://arxiv.org/html/2609.36246#bib.bib26);[West et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib21)\)\. Offline supervised fine\-tuning \(SFT\) trains the student on fixed teacher trajectories, whereas at inference it conditions on its own outputs\. This sequential covariate shift can cause errors to compound over long horizons\([Ross and Bagnell, 2010](https://arxiv.org/html/2609.36246#bib.bib44);[Bengio et al\., 2015](https://arxiv.org/html/2609.36246#bib.bib20)\)\. As the student policy changes during training, a fixed dataset also fails to track the states it currently visits\. Offline SFT can also degrade the student’s prior capabilities\([Shenfeld et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib38);[Chen et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib32)\)\. Prior work suggests that supervision close to the student’s own distribution can improve adaptation and reduce forgetting\([Zhang et al\., 2026a](https://arxiv.org/html/2609.36246#bib.bib1);[Chen et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib32)\)\.
Recently, on\-policy distillation \(OPD\) has become a promising paradigm for large language model \(LLM\) post\-training\([Agarwal et al\., 2024](https://arxiv.org/html/2609.36246#bib.bib11);[Yang et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib5);[Lu and Lab, 2025](https://arxiv.org/html/2609.36246#bib.bib6);[Xiao et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib16)\)\. By sampling rollouts from the student policy itself, OPD uses the teacher policy to calculate the reverse\-KL loss for each token in the rollout\. It thus pairs dense supervision with on\-policy states, anchoring learning where the student actually is rather than pulling it toward teacher trajectories\([Lu and Lab, 2025](https://arxiv.org/html/2609.36246#bib.bib6)\)\. Yet token\-level OPD computes teacher targets along each sampled student rollout without revising it\. Even when the teacher recommends changing a token, subsequent targets remain conditioned on the student’s original continuation\. The supervision therefore does not directly demonstrate how to continue from that correction\([Jiang et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib45)\)\. Recent works have also shown that as OPD transfers supervision from teacher at distribution level, it assumes the teacher places meaningful probability mass on the states the student reaches\([Zhu et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib17);[Li et al\., 2026c](https://arxiv.org/html/2609.36246#bib.bib7)\)\. Such assumption fails once the capability gap between the two policies is too large, and it extends to multi\-turn agentic tasks where the irreversible actions made by the student lead the whole trajectory out of the support from the teacher\([Wang et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib19)\)\. Distribution\-matching OPD also requires access to teacher token probabilities, limiting its use with teachers that expose only generated text\.
Figure 1:Overview ofOLIVE\.\(a\) Long chain\-of\-thought reasoning tasks:the student generates a reasoning prefix, which the teacher continues\.\(b\) Long\-horizon agentic tasks:the student interacts with the environment for part of an episode, then the teacher takes over\. In both settings, prefixes are refreshed as the student policy changes\.We proposeOLIVE\(OnLine InterVEntion\)\. It trains the evolving student on teacher continuations from student\-generated states \([Fig\.1](https://arxiv.org/html/2609.36246#S1.F1)\)\. At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross\-entropy computed only on the teacher\-generated tokens\. For agentic tasks, the prefix consists of the student’s actions and the resulting observations; the teacher then takes over interaction with the environment\. Repeating this process refreshes the prefixes as the student improves\. To reduce the time spent waiting for teacher generation, we implementOLIVEasynchronously, drawing on asynchronous reinforcement learning\([Mnih et al\., 2016](https://arxiv.org/html/2609.36246#bib.bib46);[Espeholt et al\., 2018](https://arxiv.org/html/2609.36246#bib.bib47)\)\. Teacher continuation for one batch overlaps with student prefix generation for the next \([Fig\.2](https://arxiv.org/html/2609.36246#S3.F2)\), reducing total training time by 23\.8% relative to synchronousOLIVE\([§4\.3](https://arxiv.org/html/2609.36246#S4.SS3)\)\.
We evaluateOLIVEin the two settings at the center of frontier post\-training\([Team et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib42);[Xu et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib24);[Zeng et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib43)\): long chain\-of\-thought reasoning and long\-horizon agentic tasks\. For reasoning, we target problems beyond the student’s capability, using synthetic tasks from RLVE\([Zeng et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib4)\)whose difficulty we can control\. For agentic tasks, we use multiple environments from AgentGym\([Xi et al\., 2025b](https://arxiv.org/html/2609.36246#bib.bib15)\)\. Empirically, we find thatOLIVElifts the performance of the student policy on hard reasoning tasks by 6% to 8% pass@8 points and 7% to 22% avg@4 gains on agentic benchmarks, while only introducing 0\.9% average performance drop on general benchmarks \([§4](https://arxiv.org/html/2609.36246#S4),[§5\.1](https://arxiv.org/html/2609.36246#S5.SS1)\)\. The teacher continuation enablesOLIVEto learn where OPD would fail, raising ScienceWorld success to 7\.5% from a near\-zero student that OPD fails to improve \([§4\.2](https://arxiv.org/html/2609.36246#S4.SS2)\)\. With the nature of CE loss,OLIVEdoes not require the teacher logits, enabling black\-box distillation \([§5\.2](https://arxiv.org/html/2609.36246#S5.SS2)\)\. Refreshing the student policy online, we show thatOLIVEenables progressive performance improvement as the training proceeds, while offline distillation plateaus and fails to maintain plasticity \([§5\.2](https://arxiv.org/html/2609.36246#S5.SS2)\)\.
Taken together, we presentOLIVE, an online distillation method that trains the rolling student policy on teacher continuations elicited at the states its own prefixes reach\.OLIVEmoves the optimization objective in SFT from offline to online, so the supervision is refreshed as the student improves\. On top of OPD,OLIVEdemonstrates what to do next from student\-visited states, rather than grading past tokens, and needs only teacher text\. Empirically,OLIVElets training keep improving where offline distillation plateaus, offering better plastcity while causing less forgetting of the capability it already has\. Our results suggest that where supervision is placed, and whether it is refreshed as the student changes, is an important axis of distillation design alongside the form the supervision takes\.
## 2Motivation
Distillation provides supervision through teacher distributions or generated texts\([Hinton et al\., 2015](https://arxiv.org/html/2609.36246#bib.bib22);[Kim and Rush, 2016](https://arxiv.org/html/2609.36246#bib.bib26);[West et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib21)\), but its usefulness also depends on the contexts at which that supervision is provided\. For difficult reasoning and interaction tasks, we ask:how should supervision from the teacher connect the student’s own attempts to behavior it does not yet generate reliably?
#### Teacher continuations at student\-generated contexts\.
Training only on teacher trajectories can leave the student unprepared for situations created by its own decisions, a source of compounding errors in behavioral cloning\([Pomerleau, 1991](https://arxiv.org/html/2609.36246#bib.bib33);[Ross et al\., 2011](https://arxiv.org/html/2609.36246#bib.bib8)\)\. OPD addresses this by supervising the student along its own rollouts\([Agarwal et al\., 2024](https://arxiv.org/html/2609.36246#bib.bib11)\), but later targets remain conditioned on the student’s earlier decisions even where the teacher recommends a different one, so the rollout never demonstrates what would follow that recommendation\([Jiang et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib45)\)\. On\-policy supervision also struggles under large capability gaps and when student errors derail multi\-turn interaction\([Li et al\., 2026c](https://arxiv.org/html/2609.36246#bib.bib7);[Zhu et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib17);[Wang et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib19)\)\. Teacher takeover instead lets the teacher continue from a student\-generated prefix, so later decisions, and in interactive environments later observations, follow the teacher’s own choices\. The student thus learns how to proceed from contexts it actually reaches, without first having to produce the corrective path itself\. This mirrors learner roll\-in with expert rollout in imitation learning\([Ross and Bagnell, 2014](https://arxiv.org/html/2609.36246#bib.bib13)\)and recent expert\-intervention methods for language\-model agents\([Lauffer et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib27);[Li et al\., 2026a](https://arxiv.org/html/2609.36246#bib.bib28)\)\.
#### Online intervention with a rolling policy\.
Prefixes collected once reflect the behavior of an earlier student\. As training proceeds to update the student policy, the student may encounter different situations and benefit from different continuations\. We therefore regenerate prefixes and teacher interventions throughout training, following the learner\-state supervision principle of DAgger\([Ross et al\., 2011](https://arxiv.org/html/2609.36246#bib.bib8)\)\.[Zhang et al\. \(2026a\)](https://arxiv.org/html/2609.36246#bib.bib1)and[Zhang et al\. \(2026b\)](https://arxiv.org/html/2609.36246#bib.bib31)likewise motivate assessing supervision in relation to the target student and its subsequent learning\. We examine whether refreshing intervention contexts improves learning over repeatedly using interventions collected from the initial policy\.
#### Learning through teacher\-generated text\.
Teacher intervention comes in textual forms, learning it only requires cross\-entropy loss\. The teacher need not expose token probabilities, and its output can be tokenized using the student’s tokenizer\. Teacher continuation determines the trajectory to learn from; suffix CE makes learning from that trajectory possible through a text\-only interface\.
Together, these choices motivateOLIVEas a continuously refreshed procedure for learning from teacher continuations of student attempts\. We evaluate its effectiveness on complex reasoning and long\-horizon interaction tasks, together with the cost of generating these interventions\.
## 3OLIVE: OnLine InterVEntion
We describeOLIVEin[§3\.1](https://arxiv.org/html/2609.36246#S3.SS1)and extend it to multi\-turn agentic tasks in[§3\.2](https://arxiv.org/html/2609.36246#S3.SS2)\. We then present an asynchronous implementation that overlaps student sampling with teacher generation to reduce idle time \([§3\.3](https://arxiv.org/html/2609.36246#S3.SS3)\)\.
### 3\.1Main Algorithm
Algorithm 1OLIVE\(OnLine InterVEntion\)1:Student
πθ\\pi\_\{\\theta\}, teacher
πT\\pi\_\{T\}, prompts
𝒟\\mathcal\{D\}, student prefix length
kk, teacher continuation budget
MM\.
2:foreach training stepdo
3:Sample prompts
x∼𝒟x\\sim\\mathcal\{D\}
4:Online Roll\-in:
y1:k∼πθ\(⋅∣x\)y\_\{1:k\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\)
5:Continue:
y~∼πT\(⋅∣x,y1:k\)\\tilde\{y\}\\sim\\pi\_\{T\}\(\\cdot\\mid x,y\_\{1:k\}\),
\|y~\|≤M\|\\tilde\{y\}\|\\leq M
6:Update:Update
θ\\thetausing the masked CE loss in[Eq\.1](https://arxiv.org/html/2609.36246#S3.E1)\.
7:endfor
8:return
πθ\\pi\_\{\\theta\}
OLIVEtrains a studentπθ\\pi\_\{\\theta\}from a teacherπT\\pi\_\{T\}by collecting student prefixes online, letting the teacher demonstrate how to continue, and learning from teacher text alone\. These three design choices address the limitations of offline SFT and token\-level OPD, as we explain below using[Alg\.1](https://arxiv.org/html/2609.36246#alg1)as a walkthrough\.
#### Collect student prefixes online\.
To obtain supervision at states reached by the current student, we sample a promptxxfrom the prompt distribution𝒟\\mathcal\{D\}and akk\-token prefixy1:k∼πθ\(⋅∣x\)y\_\{1:k\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\)\([Alg\.1](https://arxiv.org/html/2609.36246#alg1), lines 2–3\)\. We repeat this step after each student update, so the supervision tracks the evolving policy instead of remaining tied to fixed offline trajectories\. We use a fixed prefix lengthkkand retain all sampled prefixes without filtering\. We report the student and teacher generation lengths for reasoning and agentic tasks in[§§4\.1](https://arxiv.org/html/2609.36246#S4.SS1)and[4\.2](https://arxiv.org/html/2609.36246#S4.SS2), respectively\.
#### Demonstrate how to continue\.
Token\-level OPD leaves later targets conditioned on the student’s original tokens even when an earlier target recommends a correction\. In line 4 of[Alg\.1](https://arxiv.org/html/2609.36246#alg1), the teacher receives the prompt and student prefix as the input and generates a continuationy~∼πT\(⋅∣x,y1:k\)\\tilde\{y\}\\sim\\pi\_\{T\}\(\\cdot\\mid x,y\_\{1:k\}\)of at mostMMtokens under its own tokenizer\. The teacher conditions each new token on its preceding choices, allowing the continuation to demonstrate how to follow a correction when the student prefix remains recoverable\. We use partial continuations to bound the cost of online teacher generation, without requiring a completed, verified solution\. As we show in[§4\.3](https://arxiv.org/html/2609.36246#S4.SS3),OLIVEachieves higher reasoning performance than OPD at comparable GPU\-hour cost with this limited continuation budget\.
#### Learn from teacher text alone\.
To avoid requiring teacher token probabilities, we train the student with CE on the generated continuation \([Alg\.1](https://arxiv.org/html/2609.36246#alg1), line 5\)\. For the student update, we tokenize the teacher continuation with the student’s tokenizer, obtainingy~1:ℓ\\tilde\{y\}\_\{1:\\ell\}, whereℓ\\ellis its length in student tokens\. This length may differ from the number of teacher tokens\. We feed the full sequence\(x,y1:k,y~1:ℓ\)\(x,y\_\{1:k\},\\tilde\{y\}\_\{1:\\ell\}\)to the student and compute
ℒ\(θ;x,y1:k,y~1:ℓ\)=−∑j=1ℓlogπθ\(y~j∣x,y1:k,y~<j\)\.\\mathcal\{L\}\(\\theta;x,y\_\{1:k\},\\tilde\{y\}\_\{1:\\ell\}\)=\-\\sum\_\{j=1\}^\{\\ell\}\\log\\pi\_\{\\theta\}\\\!\\left\(\\tilde\{y\}\_\{j\}\\mid x,y\_\{1:k\},\\tilde\{y\}\_\{<j\}\\right\)\.\(1\)The prompt and student prefix remain in the conditioning context but are masked out of the loss; the sampled sequences are held fixed during the update\. This objective requires only teacher\-generated text, so it supports black\-box teachers and different teacher and student tokenizers\.
Figure 2:Synchronous and asynchronous implementations ofOLIVE\. \(a\) The student waits for the teacher continuation before updating on the same batch\. \(b\) Prefix sampling for later batches overlaps with teacher generation, while student updates use completed continuations from earlier batches\. Arrows connect teacher continuations to the updates that use them\.
### 3\.2Multi\-turnOLIVE
OLIVEextends to multi\-turn interaction by applying the same prefix–continuation split at the level of turns \([Fig\.1](https://arxiv.org/html/2609.36246#S1.F1)b\)\. Here,kkandMMcount interaction turns rather than tokens\. The student interacts with the environment forkkturns, producing a history of actions and observations\. The teacher then takes over for up toMMturns, with each action conditioned on the full interaction history and executed in the environment to obtain the next observation\. During the student update, we retain the full interaction history as context, mask the student’s turns and all environment observations, and apply CE only to the teacher’s actions\. We find that student prefixes ofk=5k=5or1010turns, depending on the environment, followed by up toM=5M=5teacher turns work well \([§4\.2](https://arxiv.org/html/2609.36246#S4.SS2)\)\.
### 3\.3AsynchronousOLIVEfor Scalable Training
To reduce student GPU idle time, we overlap student prefix sampling with teacher generation \([Fig\.2](https://arxiv.org/html/2609.36246#S3.F2)\), drawing on asynchronous reinforcement learning\([Mnih et al\., 2016](https://arxiv.org/html/2609.36246#bib.bib46);[Espeholt et al\., 2018](https://arxiv.org/html/2609.36246#bib.bib47)\)\. While the teacher generates continuations for one batch, the student samples prefixes for the next; completed traces provide training data for the same masked CE objective in[Eq\.1](https://arxiv.org/html/2609.36246#S3.E1)\. This overlap means that a prefix may come from an earlier student policy than the one being updated\. We bound this lag by an asynchronous depthdd, the maximum number of student updates between prefix generation and use of the resulting trace for training\. A largerddallows more overlap but permits greater mismatch between the policy that generated the prefix and the policy being trained\. We used=3d=3in the reasoning experiments \([Tab\.4](https://arxiv.org/html/2609.36246#A1.T4)\) and evaluate the efficiency–performance tradeoff in[§4\.3](https://arxiv.org/html/2609.36246#S4.SS3)\.
## 4Experiments
We evaluateOLIVEon two main post\-training scenarios\. We present the experiment in reasoning tasks in[§4\.1](https://arxiv.org/html/2609.36246#S4.SS1)and multi\-turn agentic tasks in[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2)\. We further study the training efficiency ofOLIVEin[§4\.3](https://arxiv.org/html/2609.36246#S4.SS3)\.
### 4\.1Reasoning Task
#### Task\.
We use RLVE\([Zeng et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib4)\)as our primary testbed for single\-turn reasoning tasks\. RLVE is a synthetic reasoning environment that consists of different reasoning environments, each with verifiable rewards and different difficulty parameters to construct problems with tunable difficulty\. It provides a noise\-free data collection, training and evaluation pipeline as the instances are not provided during pre\-training or post\-training of the models themselves to introduce contamination\([Shao et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib34)\)\. Since knowledge distillation targets at introducing new capabilities from the teacher policy to the student policy, we set the difficulty parameters to find the problems that are challenging enough for the student policy to solve\. We thus obtain a 18 different reasoning environments subset with 500 training problems for each game, yielding a pool of99K hard problems\. For each task, we pair with 10 test problems with the same difficulty as the training problems\.
Qwen3\-1\.7BQwen3\-4BMethodPass@8Avg@8Pass@8Avg@8Original11\.13\.346\.118\.1Teacher\-Gen\.15\.0\(\+3\.9\)5\.6\(\+2\.3\)53\.3\(\+7\.2\)19\.9\(\+1\.8\)OPD14\.4\(\+3\.3\)4\.6\(\+1\.3\)45\.6\(−\-0\.5\)21\.3\(\+3\.2\)OLIVE19\.4\(\+8\.3\)7\.8\(\+4\.5\)52\.2\(\+6\.1\)23\.5\(\+5\.4\)Table 1:Results on RLVE\. We report pass@8 and avg@8 on the test set, with two different student models thinking enabled\. We use identical teacher model Qwen3\-4B\-Thinking\-2507 for all distillation methods\.OLIVEuses the asynchronous implementation withd=3d\{=\}3\. All numbers are percentages \(%\\%\)\.
Figure 3:Top\-KKoverlap ratio between student and teacher on validation set during OPD training \(Qwen3\-1\.7B student, Qwen3\-4B\-Thinking\-2507 teacher\)\. It stays nearly flat \(0\.707→0\.7130\.707\\to 0\.713\), resonating the finding in[Li et al\. \(2026c\)](https://arxiv.org/html/2609.36246#bib.bib7)\.
#### Models and Evaluation\.
We consider using models with thinking capabilities that are able to solve the reasoning problems with long\-term reasoning generation\. To achive this, we use Qwen3\-1\.7B and Qwen3\-4B\([Yang et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib5)\)as two student models with their thinking enabled\. We use Qwen3\-4B\-Thinking\-2507\([Yang et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib5)\)as the teacher policy since it is the continued scaled model from Qwen3\-4B and has stronger reasoning capabilities\. We report the pass@8 and avg@8 on the test set\.
#### Training Setup\.
We report the results of comparison betweenOLIVEand other distillation methods, including offline distillation and OPD in[Tab\.1](https://arxiv.org/html/2609.36246#S4.T1)\. To compare distillation methods under a matched budget, we cap the total number of distilled tokens per rollout at 7168 for both the offline and online baselines\. The offline baseline applies cross\-entropy on unfiltered teacher trajectories, and the online baseline is OPD\. SinceOLIVEdoes not require the student policy to generate the entire rollout, we set the prefix length to 4096 and the continuation length of 1024 by default\. We set the asynchronous depthd=3d=3forOLIVE\. Detailed experimental settings are provided in the[App\.A](https://arxiv.org/html/2609.36246#A1)\.
#### Results\.
We report the results ofOLIVEand comparison across two different student models in[Tab\.1](https://arxiv.org/html/2609.36246#S4.T1)\. Across both students,OLIVEachieves the largest avg@8 gains among all distillation methods \(\+4\.42 and \+5\.35 points\)\. Compare to offline distillation where offline data is collected from the teacher policy without filtering,OLIVEachieves better improvements with less samling from the teacher policy and remain non\-filtering\. We also include a training dynamic visualiztion of top\-KKoverlap ratio between student and teacher on validation set during OPD training in[Fig\.3](https://arxiv.org/html/2609.36246#S4.F3)\. Resonating the finding in[Li et al\. \(2026c\)](https://arxiv.org/html/2609.36246#bib.bib7), the top\-KKoverlap ratio remains nearly flat during OPD training, and little improvement is observed in OPD during training due to different thinking behaviors between the student and the teacher\. Overall, supervising the student with teacher continuations from its own states transfers more capability than either training on full teacher trajectories or scoring the student’s rollouts, while requiring only a fraction of the teacher’s generation\.
### 4\.2Multi\-turn Agentic Task
#### Task\.
We select 5 multi\-turn agentic tasks from AgentGym\([Xi et al\., 2025a](https://arxiv.org/html/2609.36246#bib.bib18)\)\. AgentGym is a multi\-turn agentic environment that provides a diverse environments with turn\-level feedback for each action the agent takes\. Following[Xi et al\. \(2025b\)](https://arxiv.org/html/2609.36246#bib.bib15), we select 5 environments, including ALFWorld\([Shridhar et al\., 2020](https://arxiv.org/html/2609.36246#bib.bib23)\), ScienceWorld\([Wang et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib14)\), SearchQA\([Dunn et al\., 2017](https://arxiv.org/html/2609.36246#bib.bib29)\), TextCraft\([Prasad et al\., 2024](https://arxiv.org/html/2609.36246#bib.bib35)\)and BabyAI\([Chevalier\-Boisvert et al\., 2018](https://arxiv.org/html/2609.36246#bib.bib36)\)\. We use the same training and evaluation set for different tasks, except for SearchQA we construct a 6K problems training set with 400 held out problems for evaluation\.
ALFWorldScienceWorldTextCraftBabyAISearchQAMethodSR \(%\\%\)TurnsSR \(%\\%\)TurnsSR \(%\\%\)TurnsSR \(%\\%\)TurnsSR \(%\\%\)TurnsStudent19\.3826\.970\.1228\.6323\.0023\.8938\.3314\.5230\.5011\.91Teacher52\.1221\.2415\.6223\.6185\.5010\.9683\.336\.2255\.069\.29OPD22\.2526\.280\.0028\.7729\.5022\.2843\.0614\.3229\.5612\.07TCoD\-B2F37\.1023\.300\.7529\.2839\.5019\.9866\.309\.0337\.6911\.12TCoD\-F2B28\.0024\.890\.5029\.5545\.5018\.4662\.5010\.5137\.7511\.19Guided OPD27\.1225\.260\.7529\.0345\.5018\.5967\.788\.3037\.5011\.10OLIVE40\.0023\.157\.5023\.6455\.2516\.4567\.509\.9839\.0611\.13Table 2:Results on multi\-turn agentic benchmarks \(ALFWorld, ScienceWorld, TextCraft, BabyAI and SearchQA\) with Qwen3\-1\.7B as the student and Qwen3\-32B as the teacher\. We report avg@4 success rate \(SR,%\\%\) and average trajectory score, along with the average turns across the five benchmarks\.
#### Training Setup\.
We use Qwen3\-1\.7B as the student model and Qwen3\-32B as the teacher model\. We compare different online distillation methods on multi\-turn agentic tasks, including OPD and subsequent variants that tries to adapt OPD to the multi\-turn agentic setting, including two variants of TCoD\([Wang et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib19)\)and Guided OPD\([Li et al\., 2026b](https://arxiv.org/html/2609.36246#bib.bib30)\)\. We set the maximum number of turns for ALFWorld, TextCraft and ScienceWorld to 30, 20 for BabyAI and 16 for SearchQA\. For each turn we follow the ReAct\([Yao et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib37)\)framework to generate the action as AgentGym originally implemented\. We set the training epoch for ALFWorld, ScienceWorld and SearchQA to 1, and 3 for BabyAI and TextCraft respectively, since BabyAI and TextCraft have less training data to train the model\. For evaluation, we report the average@4 success rate \(SR,%\\%\), along with the average turns across the five benchmarks\. Empirically, we set 10 student turns as prefix forOLIVEfor environments with longer turns like ALFWorld and ScienceWorld, and 5 student turns as prefix for environments with shorter turns\. The teacher turns continuation is set to 5 forOLIVEby default\.
#### Results\.
We report the results of comparison betweenOLIVEand other online distillation methods in[Tab\.2](https://arxiv.org/html/2609.36246#S4.T2)\. Across different environments,OLIVEconsistently outperforms the other two distillation methods, with each environment showing less turns to achieve higher success rate\. Empirically, we observe the biggest performance gain on ScienceWorld, which is the most complex environment among the five according to initial student policy performance\. This also explains why OPD does not work well on this environment, as the student action at early turns are likely flawed, making the subsequent turns within the same episode to be incorrect and fail the task\. Instead,OLIVEeffectively distill the teacher capabilities to the student policy by directly introducing teacher intervention at given student states, thus correcting the student episode on the right track\.
### 4\.3Training Efficiency
In this section, we study the training efficiency ofOLIVE\. We conduct a training comparison between OPD, synchronousOLIVEand asynchronousOLIVEon a single turn reasoning task\. To be specific, following the setting in[Tab\.1](https://arxiv.org/html/2609.36246#S4.T1), we use the same teacher model Qwen3\-4B\-Thinking\-2507 for all distillation methods, and set student model to be Qwen3\-4B\. We run training on the99K training set of RLVE on 8 H200 GPUs, and report the total GPU hours for each method\.
Figure 4:Training comparison between OPD, asynchronousOLIVEand synchronousOLIVE\.We present the results in[Fig\.4](https://arxiv.org/html/2609.36246#S4.F4), and observe thatOLIVEcan match the training efficiency of OPD while introducing performance improvement\. SinceOLIVEdoes not require the teacher model to complete the reasoning process or to compute reverse KL on all student generated tokens, it introduces less computation overhead for student rollouts and teacher continuation\. We further improves the training efficiency ofOLIVEby using asynchronous training\. AsynchronousOLIVElets the student policy keep sampling the next batch of prefix rollouts in parallel, and the collected traces are used for the update once the staleness limit is reached\. In practice, this significantly reduces the total training time by 28% comparing with OPD\. Compared with synchronousOLIVE, asynchronousOLIVEonly introduces minimal performance degradation while reducing the total training time by 23\.8%, which mitigates the effect of additional computation overhead introduced by hosting the teacher model online for sampling\.
## 5Analysis
### 5\.1Online Prefixes Learn More while Forgetting Less
OEC\([Lauffer et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib27)\)offers an offline variant ofOLIVE, where prefix reasoning content is sampled from the student policy, and then teacher continuation is sampled from the teacher policy\. Then they apply a filter to select only the correct reasoning content for finetuning with CE loss and student generation masked\. This also resonates SFT baseline, where the teacher model generates full reasoning content and then finetune the student policy on the generated continuations\. In this section, we present an analysis study to advocate thatonline prefixes enable more effective distillation while forgetting less on general capabilities, mitigating the exposure bias of offline distillation we mentioned in[§1](https://arxiv.org/html/2609.36246#S1)\.
We first run teacher 4 times on the same prompts in RLVE, and filter by verifier to get the correct reasoning SFT data for finetuning\. We then collect 1 prefix rollout from the student policy for each problem in the training set, and then ask the teacher policy to carry on reasoning as continuation for 4 times, and then filter by verifier, yileding a continuation dataset for finetuning\. We finally compare this two baselines withOLIVE, where instead of only using partial teacher continuation as in[Tab\.1](https://arxiv.org/html/2609.36246#S4.T1), we let the teacher to generate full reasoning content and then filter during online rollouts\. The online rollout number for each problem is set to be 4, and we only compute CE loss on the correct reasoning content after filtering\.
Figure 5:Comparison ofOLIVEwith OEC and SFT baselines on RLVE\. Obtaining prefix reasoning content from online rolling policy instead of collecting offline prefix or full reasoning content creates effective distillation\.Figure 6:Evaluation of forgetting and task performance gains\. Purple bars show pass@8 on the RLVE test set after finetuning\. Grey bars show the change in average accuracy on general benchmarks relative to the plain student\.
We report the results in[Fig\.6](https://arxiv.org/html/2609.36246#S5.F6)\. We observe thatOLIVEachieves a better performance than both baselines, indicating the effectiveness of applying online intervention during training\. Instead of collecting the prefix reasoning content directly from the student policy as in OEC,OLIVEcollects the prefix reasoning from the updating student policy during online stage\. This effectively boost the performance of distillation by providing more effective supervision during online training\. We also measure the forgetting effect ofOLIVEon RLVE\. To directly measure the effect of such forgetting between online and offline variants, we test the performance of the finetuned policies on 4 different general benchmarks, including AIME25\([Balunovic et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib2)\)for math tasks, LiveCoding Bench v6\([Jain et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib25)\)for code tasks, IF\-Eval\([Zhou et al\., 2023](https://arxiv.org/html/2609.36246#bib.bib39)\)for instruction following tasks and GPQA Diamond\([Rein et al\., 2023](https://arxiv.org/html/2609.36246#bib.bib40)\)for science tasks\. We use avg@16 for math tasks, avg@8 for code tasks, and report the average performance change across the four benchmarks\. The results are shown in[Fig\.6](https://arxiv.org/html/2609.36246#S5.F6)\. We observe thatOLIVEachieves less forgetting on general capabilities while maintaining higher task performance, indicating the advantage of online distillation over offline distillation\.
Takeaway\.Online prefixes enable more effective distillation than its offline variants, and mitigate the forgetting on general capabilities compared to offline distillation\.
\(a\)Over training epochs\.\(b\)Over rollouts per prompt\.
Figure 7:Success rate on ScienceWorld over training 5 epochs and comparison with rollout number, with Qwen3\-1\.7B as the student and GPT\-5\.4\-mini as the teacher\.
### 5\.2Online Interventions Keep the Student Learning
In this section we discuss the advantage of optimization under moving policy\. With only sampled text from the teacher, suffix CE turns an API model into an online teacher\. We use Qwen3\-1\.7B as the student policy, and GPT 5\.4\-mini\([OpenAI, 2026](https://arxiv.org/html/2609.36246#bib.bib41)\)as the teacher policy\. OPD and logit\-based distillation are inapplicable here, leaving offline distillation from the same teacher as the baseline\. We use ScienceWorld as the primary benchmark for study this problem\. We first collect one episode for each problem in the training set from the teacher policy to get a training set for offline distillation\. We then use the same configuration as in[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2)to train the student policy usingOLIVE\. Offline distillation is conducted on static offline data, whileOLIVEcollects data as the prefix turns sampling from dynamic student policy\. We train the student policy for 5 epochs for both methods, and report the performance as the training progress\.
The results are shown in[Fig\.7](https://arxiv.org/html/2609.36246#S5.F7)\. Training on static offline data leads to a plateau after 2 epochs, as the static data is not updated to follow the moving student policy\. Unlike offline distillation where the supervision elicitation context comes solely from static offline prompts,OLIVEcollects supervision from rolling student policy\. Such rolling policy naturally introduces diverse and suitable supervision for the student policy\([Zhang et al\., 2026a](https://arxiv.org/html/2609.36246#bib.bib1)\)\. Empirically, althoughOLIVEdoes not perform as well as offline distillation at first two epochs, it keeps improving through the training process, and successfully surpasses the performance of offline distillation by 13% after 5 epochs\.OLIVEis also more flexible in single epoch training\. Here we compare the performance ofOLIVEusing multi\-rollout and compare with corresponding offline epoch checkpoint\. We hypothesize that adding more rollouts per prompt creates same update steps as the offline distillation, yet since directly sampling from the student policy,OLIVEcan collect more diverse and suitable supervision for the student policy\. We observe thatOLIVEwith multi\-rollout performs better than the offline checkpoints, indicating the advantage of distillation during online stage over offline distillation\.
Figure 8:Success rate after 5 epochs of sequential training on each agentic environment\.We also showOLIVEpreserves plasticity of the student policy than offline distillation methods\. For the agentic environments we used in[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2), we run offline distillation andOLIVEfor 5 epochs with the same configuration in[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2)sequentially on Qwen3\-4B\. After this sequential training, we evaluate the performance of the finetuned policy on the test set of each environment, obtaining[Fig\.8](https://arxiv.org/html/2609.36246#S5.F8)\. Sequentially training on offline data leaves the student less able to learn later environments: the gap is largest on SciWorld, the last environment in the sequence, whereOLIVEimproves over offline distillation by 13\.0 points\. Because later environments are learned by a policy already shifted by earlier stages, fixed offline trajectories increasingly mismatch the states that policy visits, whereasOLIVEelicits supervision from the current policy’s own prefixes\. This suggests that maintainingOLIVEonline benefits the student policy to learn more effectively as different tasks proceed, demonstrating the advantage for plasticity preservation for student policy\.
Takeaway\.The online interventions inOLIVEkeep the student policy learning progressively through black\-box level access to the teacher policy, while preserving the plasticity of the student policy\.
## 6Related Work
Knowledge distillation\([Hinton et al\., 2015](https://arxiv.org/html/2609.36246#bib.bib22)\)transfers knowledge from a teacher model to a student model by minimizing the KL divergence between the two policy distributions\. Subsequently, SeqKD\([Kim and Rush, 2016](https://arxiv.org/html/2609.36246#bib.bib26)\)transfers sequence\-level information by optimizing on textual generations from the teacher model\. Symbolic KD\([West et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib21)\)further extends the idea to selectively distill by designing desired prompts to elicit teacher demonstrations\. Recently, many works distill long chains of thought from stronger reasoners\([Guo et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib3);[Ye et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib9);[Guha et al\., 2025](https://arxiv.org/html/2609.36246#bib.bib10)\)to enhance the reasoning capability of the student model\. OPD\([Agarwal et al\., 2024](https://arxiv.org/html/2609.36246#bib.bib11);[Gu et al\., 2024](https://arxiv.org/html/2609.36246#bib.bib12);[Lu and Lab, 2025](https://arxiv.org/html/2609.36246#bib.bib6)\)transfers knowledge from the teacher model to the student model by utilizing teacher distribution over student generations, which makes use of student context during distillation to mitigate the distribution mismatch which is common in offline distillation\([Ross et al\., 2011](https://arxiv.org/html/2609.36246#bib.bib8);[Gu et al\., 2024](https://arxiv.org/html/2609.36246#bib.bib12)\)\. Instead of scoring student generations which fails when the capability gap is large\([Zhu et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib17);[Li et al\., 2026c](https://arxiv.org/html/2609.36246#bib.bib7)\)or at long\-horizon tasks\([Wang et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib19)\),OLIVEconstructs supervision with student context by online teacher interventions, making it useful when student online rollouts are unreliable and when logit\-level supervision is unavailable\.
## 7Conclusion
We presentOLIVE, a symbolic online distillation method that distills the teacher policy to student via online teacher intervention\.OLIVEfirst collects online student prefix rollouts, introduces supervision by letting the teacher policy continue, and then calculate loss while masking out the prefix from students\.OLIVEuses refreshed student policy at each gradient step to collect online prefix, and only uses teacher generated text as supervision source\. To boost the training efficiency, we further depoly an asynchronous implementation ofOLIVEwhich lets the student prefix generation and teacher continuation run in parallel\. Empirically,OLIVEachieves 6% to 8% performance improvement on hard reasoning tasks and up to 22% performance improvement on agentic benchmarks, demonstrating effective distillation performance improvement\. Further analysis also shows thatOLIVEmitigates the distillation plateauing problem in offline distillation by refreshing student context and introduces less forgetting on general benchmarks\. More broadly, our findings suggest that beyond the form of supervision, where and how supervision is placed is also an important axis of distillation design\.
## AI Use Statement
We used generative AI tools to assist with polishing the writing; the authors verified all content and take full responsibility for it\.
## Acknowledgments
Dylan Zhang thanks Ilgee Hong, Vashisth Tiwari, Yapei Chang, Zhaocheng Zhu, and Yuxiao Qu for helpful discussions\.
This work was supported by NSF Grant No\. CHE2505932, a grant from Coefficient Giving, an Amazon AICE award, a Capital One ASKS award, and gift funding from AI2\. This research also used the Delta advanced computing and data resources, which are supported by the National Science Foundation \(award OAC 2005572\) and the State of Illinois\. Delta is a joint effort of the University of Illinois Urbana\-Champaign and its National Center for Supercomputing Applications\. This research used the DeltaAI advanced computing and data resource, which is supported by the National Science Foundation \(award OAC 2320345\) and the State of Illinois\. DeltaAI is a joint effort of the University of Illinois Urbana\-Champaign and its National Center for Supercomputing Applications\.
## References
- Agarwalet al\.\(2024\)R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. R\. Garea, M\. Geist, and O\. BachemOn\-policy distillation of language models: learning from self\-generated mistakes\.InThe twelfth international conference on learning representations,Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Balunovicet al\.\(2026\)M\. Balunovic, J\. Dekoninck, I\. Petrov, N\. Jovanović, and M\. VechevMatharena: evaluating llms on uncontaminated math competitions\.Advances in Neural Information Processing Systems38\.Cited by:[§5\.1](https://arxiv.org/html/2609.36246#S5.SS1.p3.1)\.
- Bengioet al\.\(2015\)S\. Bengio, O\. Vinyals, N\. Jaitly, and N\. ShazeerScheduled sampling for sequence prediction with recurrent neural networks\.Advances in neural information processing systems28\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1)\.
- Chenet al\.\(2025\)H\. Chen, N\. Razin, K\. Narasimhan, and D\. ChenRetaining by doing: the role of on\-policy data in mitigating forgetting\.arXiv preprint arXiv:2510\.18874\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1)\.
- Chevalier\-Boisvertet al\.\(2018\)M\. Chevalier\-Boisvert, D\. Bahdanau, S\. Lahlou, L\. Willems, C\. Saharia, T\. H\. Nguyen, and Y\. BengioBabyai: a platform to study the sample efficiency of grounded language learning\.arXiv preprint arXiv:1810\.08272\.Cited by:[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Dunnet al\.\(2017\)M\. Dunn, L\. Sagun, M\. Higgins, V\. U\. Guney, V\. Cirik, and K\. ChoSearchqa: a new q&a dataset augmented with context from a search engine\.arXiv preprint arXiv:1704\.05179\.Cited by:[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Espeholtet al\.\(2018\)L\. Espeholt, H\. Soyer, R\. Munos, K\. Simonyan, V\. Mnih, T\. Ward, Y\. Doron, V\. Firoiu, T\. Harley, I\. Dunning,et al\.Impala: scalable distributed deep\-rl with importance weighted actor\-learner architectures\.InInternational conference on machine learning,pp\. 1407–1416\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p3.1),[§3\.3](https://arxiv.org/html/2609.36246#S3.SS3.p1.1)\.
- Guet al\.\(2024\)Y\. Gu, L\. Dong, F\. Wei, and M\. HuangMinillm: knowledge distillation of large language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 32694–32717\.Cited by:[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Guhaet al\.\(2025\)E\. Guha, R\. Marten, S\. Keh, N\. Raoof, G\. Smyrnis, H\. Bansal, M\. Nezhurina, J\. Mercat, T\. Vu, Z\. Sprague,et al\.Openthoughts: data recipes for reasoning models\.arXiv preprint arXiv:2506\.04178\.Cited by:[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Hintonet al\.\(2015\)G\. Hinton, O\. Vinyals, and J\. DeanDistilling the knowledge in a neural network\.arXiv preprint arXiv:1503\.02531\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1),[§2](https://arxiv.org/html/2609.36246#S2.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Ivisonet al\.\(2026\)H\. Ivison, J\. O\. Yin, R\. Shao, T\. Xiao, N\. Lambert, and H\. HajishirziTmax: a simple recipe for terminal agents\.arXiv preprint arXiv:2606\.23321\.Cited by:[§C\.3](https://arxiv.org/html/2609.36246#A3.SS3.p1.1)\.
- Jainet al\.\(2025\)N\. Jain, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. StoicaLivecodebench: holistic and contamination free evaluation of large language models for code\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 58791–58831\.Cited by:[§5\.1](https://arxiv.org/html/2609.36246#S5.SS1.p3.1)\.
- Jianget al\.\(2026\)L\. Jiang, H\. Xu, Y\. Ding, and A\. ZhangTrajectory\-refined distillation\.arXiv preprint arXiv:2606\.08432\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1)\.
- Kim and Rush \(2016\)Y\. Kim and A\. M\. RushSequence\-level knowledge distillation\.InProceedings of the 2016 conference on empirical methods in natural language processing,pp\. 1317–1327\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1),[§2](https://arxiv.org/html/2609.36246#S2.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Laufferet al\.\(2025\)N\. Lauffer, X\. Deng, S\. Kundurthy, B\. Kenstler, and J\. DaImitation learning for multi\-turn lm agents via on\-policy expert corrections\.arXiv preprint arXiv:2512\.14895\.Cited by:[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1),[§5\.1](https://arxiv.org/html/2609.36246#S5.SS1.p1.1)\.
- Liet al\.\(2026a\)C\. Li, R\. Qiang, J\. Huang, C\. Gao, C\. Zhang, N\. He, and B\. DaiRevisiting dagger in the era of llm\-agents\.arXiv preprint arXiv:2605\.12913\.Cited by:[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2026b\)G\. Li, M\. Zheng, M\. Song, R\. Liu, T\. Yang, J\. Sun, Q\. Zhong, H\. Guo, J\. Fang, D\. Zhang,et al\.On\-policy distillation with curriculum turn\-level guidance for multi\-turn agents\.arXiv preprint arXiv:2606\.15912\.Cited by:[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px2.p1.1)\.
- Liet al\.\(2026c\)Y\. Li, Y\. Zuo, B\. He, J\. Zhang, C\. Xiao, C\. Qian, T\. Yu, H\. Gao, W\. Yang, Z\. Liu, and N\. DingRethinking on\-policy distillation of large language models: phenomenology, mechanism, and recipe\.arXiv preprint arXiv:2604\.13016\.Cited by:[Appendix A](https://arxiv.org/html/2609.36246#A1.p1.1),[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1),[Figure 3](https://arxiv.org/html/2609.36246#S4.F3),[Figure 3](https://arxiv.org/html/2609.36246#S4.F3.4),[§4\.1](https://arxiv.org/html/2609.36246#S4.SS1.SSS0.Px4.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Lu and Lab \(2025\)K\. Lu and T\. M\. LabOn\-policy distillation\.Thinking Machines Lab: Connectionism\.Note:https://thinkingmachines\.ai/blog/on\-policy\-distillationExternal Links:[Document](https://dx.doi.org/10.64434/tml.20251026)Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Mnihet al\.\(2016\)V\. Mnih, A\. P\. Badia, M\. Mirza, A\. Graves, T\. Lillicrap, T\. Harley, D\. Silver, and K\. KavukcuogluAsynchronous methods for deep reinforcement learning\.InInternational conference on machine learning,pp\. 1928–1937\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p3.1),[§3\.3](https://arxiv.org/html/2609.36246#S3.SS3.p1.1)\.
- OpenAI \(2026\)OpenAIIntroducing GPT\-5\.4 mini and nano\.Note:[https://openai\.com/index/introducing\-gpt\-5\-4\-mini\-and\-nano/](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/)Accessed: 2026\-09\-21Cited by:[§5\.2](https://arxiv.org/html/2609.36246#S5.SS2.p1.1)\.
- OpenThoughts\-Agent team \(2026\)OpenThoughts\-TBLite: A High\-Signal Benchmark for Iterating on Terminal AgentsNote:https://www\.openthoughts\.ai/blog/openthoughts\-tbliteCited by:[§C\.3](https://arxiv.org/html/2609.36246#A3.SS3.p1.1)\.
- Pomerleau \(1991\)D\. A\. PomerleauEfficient training of artificial neural networks for autonomous navigation\.Neural Computation3\(1\),pp\. 88–97\.External Links:[Document](https://dx.doi.org/10.1162/neco.1991.3.1.88)Cited by:[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1)\.
- Prasadet al\.\(2024\)A\. Prasad, A\. Koller, M\. Hartmann, P\. Clark, A\. Sabharwal, M\. Bansal, and T\. KhotADaPT: as\-needed decomposition and planning with language models\.InFindings of the Association for Computational Linguistics: NAACL 2024,K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 4226–4252\.External Links:[Link](https://aclanthology.org/2024.findings-naacl.264/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-naacl.264)Cited by:[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Reinet al\.\(2023\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGpqa: a graduate\-level google\-proof q&a benchmark\.arXiv preprint arXiv:2311\.12022\.Cited by:[§5\.1](https://arxiv.org/html/2609.36246#S5.SS1.p3.1)\.
- Ross and Bagnell \(2010\)S\. Ross and D\. BagnellEfficient reductions for imitation learning\.InProceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics,Y\. W\. Teh and M\. Titterington \(Eds\.\),Proceedings of Machine Learning Research, Vol\.9,Chia Laguna Resort, Sardinia, Italy,pp\. 661–668\.External Links:[Link](https://proceedings.mlr.press/v9/ross10a.html)Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1)\.
- Ross and Bagnell \(2014\)S\. Ross and J\. A\. BagnellReinforcement and imitation learning via interactive no\-regret learning\.arXiv preprint arXiv:1406\.5979\.Cited by:[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1)\.
- Rosset al\.\(2011\)S\. Ross, G\. Gordon, and D\. BagnellA reduction of imitation learning and structured prediction to no\-regret online learning\.InProceedings of the fourteenth international conference on artificial intelligence and statistics,pp\. 627–635\.Cited by:[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Shaoet al\.\(2025\)R\. Shao, S\. S\. Li, R\. Xin, S\. Geng, Y\. Wang, S\. Oh, S\. S\. Du, N\. Lambert, S\. Min, R\. Krishna,et al\.Spurious rewards: rethinking training signals in rlvr\.arXiv preprint arXiv:2506\.10947\.Cited by:[§4\.1](https://arxiv.org/html/2609.36246#S4.SS1.SSS0.Px1.p1.1)\.
- Shenfeldet al\.\(2026\)I\. Shenfeld, J\. Pari, and P\. AgrawalRl’s razor: why online reinforcement learning forgets less\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 59839–59864\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1)\.
- Shridharet al\.\(2020\)M\. Shridhar, X\. Yuan, M\. Côté, Y\. Bisk, A\. Trischler, and M\. HausknechtAlfworld: aligning text and embodied environments for interactive learning\.arXiv preprint arXiv:2010\.03768\.Cited by:[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Teamet al\.\(2026\)K\. Team, T\. Bai, Y\. Bai, Y\. Bao, J\. Cai, X\. Cai, P\. Cao, Y\. Cao, Z\. Chai, Y\. Charles,et al\.Kimi k3: open frontier intelligence\.arXiv preprint arXiv:2607\.24653\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p4.1)\.
- Team \(2026\)Q\. TeamQwen3\.5: towards native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§C\.3](https://arxiv.org/html/2609.36246#A3.SS3.p1.1)\.
- Wanget al\.\(2026\)J\. Wang, W\. Zhang, W\. Shi, Y\. Li, and J\. ChengTCOD: exploring temporal curriculum in on\-policy distillation for multi\-turn autonomous agents\.arXiv preprint arXiv:2604\.24005\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Wanget al\.\(2022\)R\. Wang, P\. Jansen, M\. Côté, and P\. AmmanabroluScienceworld: is your agent smarter than a 5th grader?\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 11279–11298\.Cited by:[§B\.1](https://arxiv.org/html/2609.36246#A2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Westet al\.\(2022\)P\. West, C\. Bhagavatula, J\. Hessel, J\. D\. Hwang, L\. Jiang, R\. Le Bras, X\. Lu, S\. Welleck, and Y\. ChoiSymbolic knowledge distillation: from general language models to commonsense models\.InProceedings of the 2022 conference of the North American chapter of the association for computational linguistics: Human language technologies,pp\. 4602–4625\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1),[§2](https://arxiv.org/html/2609.36246#S2.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Xiet al\.\(2025a\)Z\. Xi, Y\. Ding, W\. Chen, B\. Hong, H\. Guo, J\. Wang, X\. Guo, D\. Yang, C\. Liao, W\. He,et al\.Agentgym: evaluating and training large language model\-based agents across diverse environments\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 27914–27961\.Cited by:[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Xiet al\.\(2025b\)Z\. Xi, J\. Huang, C\. Liao, B\. Huang, H\. Guo, J\. Liu, R\. Zheng, J\. Ye, J\. Zhang, W\. Chen,et al\.Agentgym\-rl: training llm agents for long\-horizon decision making through multi\-turn reinforcement learning\.arXiv preprint arXiv:2509\.08755\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p4.1),[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px1.p1.1)\.
- Xiaoet al\.\(2026\)B\. Xiao, B\. Xia, B\. Yang, B\. Gao, B\. Shen, C\. Zhang, C\. He, C\. Lou, F\. Luo, G\. Wang,et al\.Mimo\-v2\-flash technical report\.arXiv preprint arXiv:2601\.02780\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1)\.
- Xuet al\.\(2026\)A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.Deepseek\-v4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p4.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.36246#S4.SS1.SSS0.Px2.p1.1)\.
- Yaoet al\.\(2022\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReact: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[Appendix D](https://arxiv.org/html/2609.36246#A4.p1.1),[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2.SSS0.Px2.p1.1)\.
- Yeet al\.\(2025\)Y\. Ye, Z\. Huang, Y\. Xiao, E\. Chern, S\. Xia, and P\. LiuLimo: less is more for reasoning\.arXiv preprint arXiv:2502\.03387\.Cited by:[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
- Zenget al\.\(2026\)A\. Zeng, X\. Lv, Z\. Hou, Z\. Du, Q\. Zheng, B\. Chen, D\. Yin, C\. Ge, C\. Huang, C\. Xie,et al\.Glm\-5: from vibe coding to agentic engineering\.arXiv preprint arXiv:2602\.15763\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p4.1)\.
- Zenget al\.\(2025\)Z\. Zeng, H\. Ivison, Y\. Wang, L\. Yuan, S\. S\. Li, Z\. Ye, S\. Li, J\. He, R\. Zhou, T\. Chen,et al\.Rlve: scaling up reinforcement learning for language models with adaptive verifiable environments\.arXiv preprint arXiv:2511\.07317\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p4.1),[§4\.1](https://arxiv.org/html/2609.36246#S4.SS1.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2026a\)D\. Zhang, Q\. Dai, and H\. PengThe best instruction\-tuning data are those that fit\.Advances in Neural Information Processing Systems38,pp\. 141172–141208\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p1.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2609.36246#S5.SS2.p2.1)\.
- Zhanget al\.\(2026b\)D\. Zhang, Y\. Xu, H\. Wang, Q\. Chen, and H\. PengGood sft optimizes for sft, better sft prepares for reinforcement learning\.arXiv preprint arXiv:2602\.01058\.Cited by:[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px2.p1.1)\.
- Zhouet al\.\(2023\)J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. HouInstruction\-following evaluation for large language models\.arXiv preprint arXiv:2311\.07911\.Cited by:[§5\.1](https://arxiv.org/html/2609.36246#S5.SS1.p3.1)\.
- Zhuet al\.\(2026\)S\. Zhu, X\. Ye, H\. Lu, W\. Shi, and G\. LiuThe many faces of on\-policy distillation: pitfalls, mechanisms, and fixes\.arXiv preprint arXiv:2605\.11182\.Cited by:[§1](https://arxiv.org/html/2609.36246#S1.p2.1),[§2](https://arxiv.org/html/2609.36246#S2.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2609.36246#S6.p1.1)\.
## Appendix AImplementation Details
Hyper\-parameterValueTraining temperature1\.01\.0Global batch size6464Mini batch size6464Student rollouts per prompt44LogProb top\-KK1616Max response length71687168Learning rate1×10−61\\times 10^\{\-6\}Train epochs11KL Coefficient0\.00\.0Table 3:Default configuration of OPD used in our reasoning tasks experiments\.Hyper\-parameterValueSupervision signalToken\-level CE on teacher continuationRollout modeOnline, asynchronousMaximum staleness33Student rollouts per prompt44Prefix truncation point40964096Teacher continuation length10241024Learning rate1×10−51\\times 10^\{\-5\}Train batch size6464Train epochs11Table 4:Default configuration ofOLIVEused in our reasoning tasks experiments\.We provide the default configuration of OPD andOLIVEin[Tab\.3](https://arxiv.org/html/2609.36246#A1.T3)and[Tab\.4](https://arxiv.org/html/2609.36246#A1.T4)on reasoning tasks, respectively\. We follow the default configuration of OPD as[Li et al\. \(2026c\)](https://arxiv.org/html/2609.36246#bib.bib7)did\.[Tab\.4](https://arxiv.org/html/2609.36246#A1.T4)lists the full configuration used forOLIVEin our main experiments\. The upper block covers the roadside supervision procedure itself \(student rollout, truncation, teacher continuation, and the asynchronous pipeline\), and the lower block covers the optimization setup used to train the student on the resulting stitched traces\.
## Appendix BCase Studies
### B\.1Multi\-turn Agentic Task
We show a ScienceWorld\([Wang et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib14)\)trajectory collected duringOLIVEtraining \(Qwen3\-1\.7B student, GPT\-5\.4\-mini teacher, 10 student turns followed by 5 teacher turns\) to demonstrate online intervention works\. The student prefix and the teacher continuation are stitched into one trace\. The task asks the agent to focus on the longest\-lived and then the shortest\-lived animal, and the animals are placedoutside\. For ten turns the student cycles throughgo to outside,open outside, andlook around: it treats the location as the object to open and never targets the door, so the score stays at00\. Starting from this prefix, the teacher corrects the action toopen door to the outsidein its first turn and finishes the task four turns later\. The student therefore receives supervision on how to recover from the exact failure state it reached on its own, which a teacher\-only trajectory starting from the initial observation would not contain\.
Case Study on ScienceWorldTaskFind the animal with the longest life span, then the shortest life span\. First, focus on the animal with the longest life span\. Then, focus on the animal with the shortest life span\. The animals are in theoutsidelocation\.SetupStudentQwen3\-1\.7B\(turns 1–10\)→\\rightarrowTeacherGPT\-5\.4\-mini\(turns 11–15\)OutcomeScore0\.000\.00at handoff→\\rightarrowfinal score1\.001\.00Initial observationThis room is called the greenhouse\. In it, you see: the agent; a substance called air; a bee hive \(door closed\); a jug \(containing nothing\); a shovel; a sink, which is turned off\. You also see: a door to the hallway \(closed\); a door to the outside \(closed\)\.Turn 1 · Studentscore 0\.00Thought:I need to find the animal with the longest life span and then the shortest life span\. The animals are in the ‘outside’ location\. Let me first check the ‘outside’ location\.Action:go to outsideObs:The door is not open\.Turn 2 · Studentscore 0\.00Thought:The door to the outside is closed\. I need to open it to access the animals\.Action:open outsideObs:The outside is not openable\.Turn 3 · Studentscore 0\.00Thought:The door to the outside is not openable\. I need to check if it can be opened using another method\.Action:look aroundObs:\(same greenhouse description as the initial observation\)Turns 4–10 · Student \(loop\)score 0\.00The student repeats turns 1–3 verbatim, with identical thoughts and observations:Turns 4–6:go to outside→\\rightarrowopen outside→\\rightarrowlook aroundTurns 7–9:go to outside→\\rightarrowopen outside→\\rightarrowlook aroundTurn 10:go to outside\(The door is not open\.\)Handoff: teacher continues from the student prefix \(score 0\.00\)Turn 11 · Teacherscore 0\.00Thought:I need to open the door to the outside, not the location itself\.Action:open door to the outsideObs:The door is now open\.Turn 12 · Teacherscore 0\.33Thought:The door to the outside is now open, so I can access the outside area and find the animals\.Action:go to outsideObs:You move to the outside\.Turn 13 · Teacherscore 0\.33Thought:I’m in the outside area now\. I need to see which animals are present and determine their life spans\.Action:look aroundObs:This outside location is called the outside\. Here you see: the agent; air; an axe;a baby dragonfly; a fire pit; a fountain \(containing water\);a giant tortoise egg; the ground;a baby wolf; wood\. You also see: doors to the foundry \(closed\), the greenhouse \(open\), and the kitchen \(closed\)\.Turn 14 · Teacherscore 0\.83Thought:The outside contains a baby dragonfly, a giant tortoise egg, and a baby wolf\. The giant tortoise egg is the best candidate for the longest lifespan, so I’ll focus on it first\.Action:focus on giant tortoise eggObs:You focus on the giant tortoise egg\.Turn 15 · Teacherscore 1\.00Thought:Now I need to focus on the animal with the shortest life span\. The baby dragonfly is the most likely shortest\-lived animal here\.Action:focus on baby dragonflyObs:You focus on the baby dragonfly\.
### B\.2Single\-turn Reasoning on RLVE
We also show a stitched trace on an RLVE reasoning task\. The student prefix \(Qwen3\-1\.7B\) is cut at 4096 tokens, and the teacher \(Qwen3\-4B\-Thinking\-2507\) continues from the same character\.\[…\]marks trimmed text and everything else is verbatim\. We highlight the wrong step in the student prefix inpinkand the correction in the teacher continuation ingreen\. This case illustrates why a short teacher continuation \(1024 tokens in our main experiments\) is sufficient as supervision: the teacher first completes the half\-written path of the student, then corrects the error within its first few sentences\.
Case Study on RLVE: Missed Triangle in a Cycle CountProblemgold answer: 3You are given anundirected graphwith 9 vertices, labeled from0to8\. The graph contains the following undirected edges: \(2, 3\), \(2, 5\), \(0, 8\), \(1, 3\), \(0, 4\), \(0, 1\), \(0, 7\), \(3, 8\), \(1, 8\)\. Please count the number of simple cycles in the graph\. A simple cycle is a cycle with at least 3 vertices, with no repeated vertices or edges\.\[…\]Student prefix \(Qwen3\-1\.7B\), last part before the cut\[…\]Wait, what about 0\-4\-…? No\. So maybe there are only two cycles? But that seems low\. Maybe I’m missing some\.\[…\]Vertex 1 is connected to 0,3,8\. Vertex 3 is connected to 1,2,8\.\[…\]Vertex 8 is connected to 0,3,1\.\[…\]So maybe there are no other cycles\. Then, the answer would be 2 cycles?But that seems low\. Let me check again\.\[…\]Another path: 0\-1\-3\-8\-0 \(cycle\), 0\-1\-8\-0 \(cycle\), 0\-1\-3\-2\-5\-Cut at student token 4096: teacher continues from the same characterTeacher continuation \(Qwen3\-4B\-Thinking\-2507\), first∼\\sim700 charactersanswer: 32\-… no\.Wait, here’s an idea:what about the cycle 1\-3\-8\-1? Let’s check\. 1 is connected to 3, 3 is connected to 8, 8 is connected to 1\. So yes\! That’s a triangle\.So 1\-3\-8\-1 is a cycle of length 3\. Oh\! I didn’t think of that\.\[…\]Cycle 1: 0\-1\-8\-0 \(triangle\); Cycle 2: 1\-3\-8\-1 \(triangle\); Cycle 3: 0\-1\-3\-8\-0 \(quadrilateral\)\.\[…\]Is there a triangle 0\-3\-8?\[…\]0 and 3 are not directly connected\.\[…\]\(Within these first∼\\sim700 characters, the teacher has already led the reasoning onto the correct path: the three cycles it lists are exactly the gold answer\.\)
## Appendix CAdditional Results
### C\.1Intervention Time
Figure 9:Avg@4 success rate ofOLIVEon ALFWorld with different numbers of student turns before the handoff\.We perform ablation studies onOLIVEduring agentic interaction tasks on the length of student turns\. In this study, we vary the number of student turns while keeping the number of teacher turns fixed\. We use the same setup as in[Tab\.2](https://arxiv.org/html/2609.36246#S4.T2)with Qwen3\-1\.7B as the student and Qwen3\-32B as the teacher on AlfWorld, and report the results in[Fig\.9](https://arxiv.org/html/2609.36246#A3.F9)\. We set the number of teacher turns to 5 to study the effect of student prefix on the performance ofOLIVE\. Adding any student prefix improves the success rate, to 35\.6–40\.0%, consistent with our main finding that supervision anchored at student\-reachable states is more effective than supervision on off\-policy teacher states\. Beyond this, performance does not increase monotonically with prefix length: 10 student turns performs best \(40\.0%\), while 15 and 20 turns give 35\.6% and 39\.0%\.
### C\.2Generalization to Other Tasks
Figure 10:Success\-rate gain ofOLIVEover offline distillation\.OLIVEgeneralizes better\.We perform another sequential training experiment to showOLIVEcan generalize to other tasks\. We start with Qwen3\-1\.7B as the student and GPT\-5\.4\-mini as the teacher, and sequentially train on BabyAI, TextCraft, SearchQA and ScienceWorld\. We did not train on ALFWorld during this experiment, and use it as the held out task to evaluate the generalization ability ofOLIVE\. We report the success\-rate difference betweenOLIVEand offline distillation in[Fig\.10](https://arxiv.org/html/2609.36246#A3.F10)\. We observe thatOLIVEdoes not only persist the plasticity on sequentially trained environments, but also generalizes to other tasks\. After same amount of sequential training,OLIVEcan generalize to other tasks, with average success\-rate gain over offline distillation on ALFWorld of\+2\.8\+2\.8\.
### C\.3Additional Results on Terminal Agents
Figure 11:pass@4 on TBLite subsets with Qwen3\.5\-2B as the student\.We further extend the experiments on terminal agent tasks\. We select 500 training examples from TMax\([Ivison et al\., 2026](https://arxiv.org/html/2609.36246#bib.bib48)\), a terminal agentic dataset that contains over 10,000 terminal agent tasks\. We use Qwen3\.5\-2B as the student and Qwen3\.5\-9B as the teacher\([Team, 2026](https://arxiv.org/html/2609.36246#bib.bib49)\)\. For OPD, the student rolls out the first 20 turns and the teacher supervises these turns\. ForOLIVE, the student rolls out 10 turns and the teacher continues for another 10 turns, matching the 20\-turn budget of OPD\. Both methods use 8 rollouts per prompt and a batch size of 16\. For evaluation, we randomly sample 50 tasks from TBLite\([OpenThoughts\-Agent team, 2026](https://arxiv.org/html/2609.36246#bib.bib50)\), run 4 attempts per task, and report pass@4 in[Fig\.11](https://arxiv.org/html/2609.36246#A3.F11)\.
As shown in[Fig\.11](https://arxiv.org/html/2609.36246#A3.F11), OPD does not improve over the base student \(10\.0% for both\), whileOLIVEraises pass@4 to 12\.5%\. Terminal tasks require long horizons, and a 2B student often drifts into states from which it cannot finish the task within the first few turns\. OPD only provides token\-level corrections on the student’s own 20 turns, so when the student is stuck, the teacher distribution conditioned on these failing states provides little signal toward completing the task\. In contrast, the teacher continuation inOLIVEstarts from the state the student actually reaches and carries the trajectory forward, which gives the student a demonstration of how to proceed from its own intermediate states\. This is consistent with our findings on the other agentic environments in[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2)\. Given the small size of the evaluation subset, we view this result as preliminary evidence thatOLIVEtransfers to terminal agents\.
## Appendix DPrompts for Agentic Environments
We describe the ReAct\([Yao et al\., 2022](https://arxiv.org/html/2609.36246#bib.bib37)\)interaction protocol shared by all five agentic environments in[§4\.2](https://arxiv.org/html/2609.36246#S4.SS2), and list the instruction each environment gives to the model\. The same prompts and action parsers are used for student rollouts, teacher continuations, and evaluation\.
Interactwithahouseholdtosolveatask\.Imagineyouareanintelligentagentinahouseholdenvironmentandyourtargetistoperformactionstocompletethetaskgoal\.Atthebeginningofyourinteractions,youwillbegiventhedetaileddescriptionofthecurrentenvironmentandyourgoaltoaccomplish\.Foreachofyourturn,youwillbegivenalistofactionswhichyoucanchooseonetoperforminthisturn\.Youshouldchoosefromtwoactions:"THOUGHT"or"ACTION"\.Ifyouchoose"THOUGHT",youshouldfirstthinkaboutthecurrentconditionandplanforyourfutureactions,andthenoutputyouractioninthisturn\.Youroutputmuststrictlyfollowthisformat:"Thought:
yourthoughts\.
Action:
yournextaction";Ifyouchoose"ACTION",youshoulddirectlyoutputtheactioninthisturn\.Youroutputmuststrictlyfollowthisformat:"Action:
yournextaction"\.Afteryoureachturn,theenvironmentwillgiveyouimmediatefeedbackbasedonwhichyouplanyournextfewsteps\.iftheenvrionmentoutput"Nothinghappened",thatmeansthepreviousactionisinvalidandyoushouldtrymoreoptions\.
Reminder:
1\.theactionmustbechosenfromthegivenavailableactions\.Anyactionsexceptprovidedavailableactionswillberegardedasillegal\.
2\.Thinkwhennecessary,trytoactdirectlymoreintheprocess\.
Youareanagentforscienceworld\.EveryroundIwillgiveyouanobservation,youhavetorespondanactionbasedontheobservationtofinishthegiventask\.Herearetheactionsyoumaytake:\[\{"action":"open/closeOBJ","description":"open/closeacontainer"\},\{"action":"de/activateOBJ","description":"activate/deactivateadevice"\},\{"action":"connectOBJtoOBJ","description":"connectelectricalcomponents"\},\{"action":"disconnectOBJ","description":"disconnectelectricalcomponents"\},\{"action":"useOBJ\[onOBJ\]","description":"useadevice/item"\},\{"action":"lookaround","description":"describethecurrentroom"\},\{"action":"lookatOBJ","description":"describeanobjectindetail"\},\{"action":"lookinOBJ","description":"describeacontainer’scontents"\},\{"action":"readOBJ","description":"readanoteorbook"\},\{"action":"moveOBJtoOBJ","description":"moveanobjecttoacontainer"\},\{"action":"pickupOBJ","description":"moveanobjecttotheinventory"\},\{"action":"putdownOBJ","description":"dropaninventoryitem"\},\{"action":"pourOBJintoOBJ","description":"pouraliquidintoacontainer"\},\{"action":"dunkOBJintoOBJ","description":"dunkacontainerintoaliquid"\},\{"action":"mixOBJ","description":"chemicallymixacontainer"\},\{"action":"gotoLOC","description":"movetoanewlocation"\},\{"action":"eatOBJ","description":"eatafood"\},\{"action":"flushOBJ","description":"flushatoilet"\},\{"action":"focusonOBJ","description":"signalintentonataskobject"\},\{"action":"wait","description":"takenoactionfor10iterations"\},\{"action":"wait1","description":"takenoactionfor1iteration"\},\{"action":"examineOBJ","description":"providesadescriptionoftheobjectspresentonorinareceptacle\."\},\{"action":"task","description":"describecurrenttask"\},\{"action":"inventory","description":"listyourinventory"\}\]
Yourresponseshouldusethefollowingformat:
Thought:
yourthoughts\.
Action:
yournextaction
YouaregivenfewusefulcraftingrecipestocraftitemsinMinecraft\.Craftingcommandsareoftheformat"craft\[targetobject\]using\[inputingredients\]"\.
EveryroundIwillgiveyouanobservation,youhavetorespondanactionbasedonthestateandinstruction\.Youcan"get"anobject\(ingredients\)fromtheinventoryortheenvironment,look\-upthegameinventoryby"inventory",or"craft"\(target\)usinganyofthecraftingcommands\.
Youroutputmuststrictlyfollowthisformat:"Thought:
yourthoughts\.
Action:
yournextaction"
Reminder:
1\.Alwaysspecifythequantitywhenusing"get"and"craft"commands\.\-Exampleofget:get1lapislazuli\-Example1ofcraft:craft1bluedyeusing1lapislazuli\-Example2ofcraft:craft1goldencarrotusing8goldnugget,1carrot
2\.Whenusing"get"command,donotspecifywhethertheitemcomesfromtheinventoryortheenvironment\.
3\.YoucanuseONLYcraftingcommandsprovided,donotuseyourowncraftingcommands\.However,ifthecraftingcommandusesagenericingredientlike"planks",youcanusespecialtypesofthesameingrediente\.g\."darkoakplanks"inthecommandinstead\.
Youareanexplorationmasterthatwantstofinisheverygoalyouaregiven\.EveryroundIwillgiveyouanobservation,andyouhavetorespondanactionandyourthoughtbasedontheobservationtofinishthegiventask\.Youareplacedinaroomandyouneedtoaccomplishthegivengoalwithactions\.
Youcanusethefollowingactions:
\-turnright
\-turnleft
\-moveforward
\-goto<obj\><id\>
\-pickup<obj\><id\>
\-gothrough<door\><id\>:<door\>mustbeanopendoor\.
\-toggleandgothrough<door\><id\>:<door\>canbeacloseddoororalockeddoor\.Ifyouwanttoopenalockeddoor,youneedtocarryakeythatisofthesamecolorasthelockeddoor\.
\-toggle:thereisaclosedorlockeddoorrightinfrontofyouandyoucantoggleit\.
Yourresponseshouldusethefollowingformat:
Thought:
<YourThought\>
Action:
<YourAction\>
Youareaquestion\-answeringagentwithaccesstoasearchengineoverWikipedia\.AnswerthequestionbyinterleavingThoughtandActionsteps\.
Everyresponsemustuseexactlythisformat:
Thought:<yourreasoning\>
Action:<oneaction\>
Therearetwoactions:
search\[query\]:searchWikipedia\.Thetop3passagescomebackasanObservation\.
answer\[answer\]:giveyourfinalanswer,asashortphrase\(anentity,name,dateornumber\)withnoexplanation,e\.g\.answer\[Beijing\]\.
Giveexactlyoneactionperresponseandthenstop\.DonotwritetheObservationyourself\.相似文章
OPRD:在策略表示蒸馏
OPRD提出了一种新的知识蒸馏方法,该方法在策略部署期间跨层对齐学生和教师的隐藏状态,消除了来自词空间KL估计的采样方差。实验表明,OPRD在数学推理基准(AIME 2024/2025、AIMO)上优于输出空间基线,同时速度快1.44倍,内存使用减少54%。
同策略蒸馏(5分钟阅读)
本文引入同策略蒸馏,通过在教师提供的token级KL正则化下,在学生自身轨迹上训练学生模型,解决训练-推理分布不匹配问题,统一了前向KL、反向KL和JSD损失,其中反向KL更适用于较小的学生模型。
潜在同策略自蒸馏 (LOPD)
本文介绍了潜在同策略自蒸馏(LOPD),该方法使教师的特权上下文能够从经验中端到端学习,提供密集的token级别监督,以提升智能体在智能工具使用和代码生成中的性能和效率。
策略内蒸馏真的在蒸馏吗?从嘈杂教师到自我改进
本文分析了策略内蒸馏,揭示其主要通过抑制低概率token而非依赖教师指导来实现改进,并提出了无需监督的OPSA方法,显著提升了推理性能。
当教师误导:虚假信号感知的在线策略蒸馏
本文介绍了 SA-OPD,一种虚假信号感知的在线策略蒸馏框架,它基于输入依据性和优化影响过滤误导性的词元级教师监督,从而提升 LLM 和 VLM 的蒸馏性能。