Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
Summary
This paper introduces AlignOPSD to address decision-timestamp mismatch in on-policy self-distillation for long-horizon agents, improving performance on benchmarks like ALFWorld, WebShop, and Search-QA compared to baselines.
View Cached Full Text
Cached at: 09/29/26, 04:10 AM
Paper page - Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
Source: https://huggingface.co/papers/2609.33391
Abstract
Reinforcementlearningwithverifiablerewards(RLVR)oftenreliesonsparseoutcomerewards,providingcoarsesupervisionforlong-horizonagents.On-policyself-distillation(OPSD)complementsthissignalwithdenseprivilegedfeedback.However,weidentifyDecision--TimestampMismatch:privilegedguidancemaybemisalignedwiththestudent’sfunctionaldecisionbecausethecorrespondingdecisioncanoccuratadifferenttimestep,whilethestudent’sdecisionitselfmayspanmultipletimestepsratherthanbeingtiedtoasingletimestamp.Thus,timestamp-localsupervisioncanmisalignboththecontextandthetemporalscopeofcredit.Toaddressthismismatch,weintroduceAlignOPSD,followingtheprincipleofaligningsupervisionbeforeassigningcredit.Decision-AlignedSupervisionRectificationre-scoresthesamestudent-sampledresponseinfunctionallymatchedcontextsacrosssiblingrolloutstocalibratelocalteacherevidence.Semi-MarkovHierarchicalCreditAssignmentthenderivesvariable-durationdecisionspansfromcorrespondencechangesandusesrectifiedevidencetoallocateoutcome-groundedcreditacrossspansandtheirconstituentturns.WeevaluateAlignOPSDwithQwen2.5-3BandQwen2.5-7BonALFWorld,WebShop,andSearch-QAagainstrepresentativebaselines.AlignOPSDoutperformsbothGRPOandStepOPSDacrossalleightbackbone--aggregate-metriccomparisons,improvingonGRPOby5.5--8.7\%andrankingfirstinsix.Additionalanalyzesexaminethetwoalignmentstagesandhyperparametersensitivitybetweentasks.Ourcodeisavaliableathttps://github.com/mingju-c/Align-OPSD
View arXiv pageView PDFGitHub3Add to collection
Get this paper in your agent:
hf papers read 2609\.33391
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33391 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33391 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33391 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
TurnOPD introduces turn-level budgeting for on-policy distillation of long-horizon agents, addressing inefficiencies in vanilla OPD by adaptive rollout-depth and progressive turn-normalized loss budgeting, achieving better accuracy under equal training budgets.
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
The paper proposes RetireOPD, a method for training multi-turn agents using reinforcement learning with self-retiring on-policy distillation, improving performance on ALFWorld and WebShop benchmarks.
AsyncOPD: How Stale Can On-Policy Distillation Be?
This paper presents AsyncOPD, a fully asynchronous on-policy distillation pipeline for LLMs, systematically studying the effects of stale-policy data and proposing estimator designs that improve training throughput by 1.6-3.8x while maintaining comparable accuracy.
Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics
This paper introduces ActFirst-OPD, a framework that accelerates on-policy distillation for multi-turn language agents by decoupling action execution from full reasoning, achieving significant training speedups while maintaining performance across benchmarks.
β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
This paper introduces β-OPSD, a generalization of on-policy self-distillation that frames it as a policy optimization family with a controllable KL penalty. The method uses distillation to approximate expensive policy optimization, improving stability and reasoning performance on mathematical benchmarks.