Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

Hugging Face Daily Papers Papers

Summary

This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.

On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of 2.3times on ALFWorld, 1.8times on WebShop, and 4.9times on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:20 AM

Paper page - Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics

Source: https://huggingface.co/papers/2609.36608

Abstract

On-policydistillation(OPD)trainsmulti-turnlanguageagentswithdenseteachersupervisiononstudent-generatedresponses.However,standardthink-then-actrolloutsrequirelengthyreasoningbeforeeachshortaction,delayingenvironmenttransitionsandexperiencecollection.Generatingactionsdirectlyreducesthisdelaybutcandegraderolloutquality.Toaddressthis,weproposeActFirst-OPD,anact-first,reason-latertrainingframeworkthatdecouplesenvironmentinteractionfromfull-responsegeneration.Thestudentinfersandexecutesactionsthroughreference-conditionedinversedynamicsusingitscurrentinteractioncontextandareferencenextobservation,andswitchestoautonomousnext-actionpredictionwhentheresultingtransitiondeviatesfromthereferencetrajectory.Fromthecollectedinteractioncontexts,thestudentasynchronouslygeneratesfullthink-then-actresponsesfortoken-levelteachersupervision.Experimentsacross0.6B-,1.7B-,and4B-parameterQwen3studentsshowthatActFirst-OPDachievesaveragewall-clocktrainingspeedupsof2.3timesonALFWorld,1.8timesonWebShop,and4.9timesonScienceWorldoverVanillaOPD.ItmatchesorexceedsallcomparedOPDbaselinesinmeantasksuccessrateacrosseightofninebenchmark-modelsettings.Theseresultsdemonstratethatreasoningneednotblockactingduringmulti-turnagentdistillation.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.36608

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.36608 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.36608 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.36608 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Multi-Turn On-Policy Distillation with Prefix Replay

Hugging Face Daily Papers

This paper proposes ReOPD, a method for on-policy distillation of LLM agents that reuses pre-collected teacher trajectories as replayed prefixes, achieving improved efficiency and accuracy without new environment interactions.

Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation

arXiv cs.CL

This paper investigates the parameter-level mechanisms behind the efficiency of On-Policy Distillation (OPD) for large language models, attributing it to early 'foresight' in module allocation and update direction. It proposes EffOPD, a plug-and-play method that accelerates OPD training by 3x without compromising final performance.

Pass the Baton: Trajectory-Relayed On-Policy Distillation

arXiv cs.CL

Proposes Relay On-Policy Distillation (Relay-OPD) that addresses prefix failure in on-policy distillation by having the teacher briefly take over at detected handoff triggers to correct reasoning trajectories. Achieves consistent improvements on mathematical reasoning benchmarks and reduces training trajectory length by over 50%.