Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics
Summary
This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.
View Cached Full Text
Cached at: 09/30/26, 04:20 AM
Paper page - Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics
Source: https://huggingface.co/papers/2609.36608
Abstract
On-policydistillation(OPD)trainsmulti-turnlanguageagentswithdenseteachersupervisiononstudent-generatedresponses.However,standardthink-then-actrolloutsrequirelengthyreasoningbeforeeachshortaction,delayingenvironmenttransitionsandexperiencecollection.Generatingactionsdirectlyreducesthisdelaybutcandegraderolloutquality.Toaddressthis,weproposeActFirst-OPD,anact-first,reason-latertrainingframeworkthatdecouplesenvironmentinteractionfromfull-responsegeneration.Thestudentinfersandexecutesactionsthroughreference-conditionedinversedynamicsusingitscurrentinteractioncontextandareferencenextobservation,andswitchestoautonomousnext-actionpredictionwhentheresultingtransitiondeviatesfromthereferencetrajectory.Fromthecollectedinteractioncontexts,thestudentasynchronouslygeneratesfullthink-then-actresponsesfortoken-levelteachersupervision.Experimentsacross0.6B-,1.7B-,and4B-parameterQwen3studentsshowthatActFirst-OPDachievesaveragewall-clocktrainingspeedupsof2.3timesonALFWorld,1.8timesonWebShop,and4.9timesonScienceWorldoverVanillaOPD.ItmatchesorexceedsallcomparedOPDbaselinesinmeantasksuccessrateacrosseightofninebenchmark-modelsettings.Theseresultsdemonstratethatreasoningneednotblockactingduringmulti-turnagentdistillation.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.36608
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.36608 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.36608 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.36608 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
TurnOPD introduces turn-level budgeting for on-policy distillation of long-horizon agents, addressing inefficiencies in vanilla OPD by adaptive rollout-depth and progressive turn-normalized loss budgeting, achieving better accuracy under equal training budgets.
Multi-Turn On-Policy Distillation with Prefix Replay
This paper proposes ReOPD, a method for on-policy distillation of LLM agents that reuses pre-collected teacher trajectories as replayed prefixes, achieving improved efficiency and accuracy without new environment interactions.
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
The paper introduces ATOD, a hybrid online distillation algorithm combining on-policy distillation and reinforcement learning for training small language model agents in multi-turn tasks, featuring an annealed OPD-RL schedule and Turn-level Disagreement-Uncertainty Reweighting to improve dense supervision.
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
This paper investigates the parameter-level mechanisms behind the efficiency of On-Policy Distillation (OPD) for large language models, attributing it to early 'foresight' in module allocation and update direction. It proposes EffOPD, a plug-and-play method that accelerates OPD training by 3x without compromising final performance.
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Proposes Relay On-Policy Distillation (Relay-OPD) that addresses prefix failure in on-policy distillation by having the teacher briefly take over at detected handoff triggers to correct reasoning trajectories. Achieves consistent improvements on mathematical reasoning benchmarks and reduces training trajectory length by over 50%.