Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

Hugging Face Daily Papers Papers

Summary

This paper introduces AlignOPSD to address decision-timestamp mismatch in on-policy self-distillation for long-horizon agents, improving performance on benchmarks like ALFWorld, WebShop, and Search-QA compared to baselines.

Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify Decision--Timestamp Mismatch: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce AlignOPSD, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate AlignOPSD with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. AlignOPSD outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:10 AM

Paper page - Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

Source: https://huggingface.co/papers/2609.33391

Abstract

Reinforcementlearningwithverifiablerewards(RLVR)oftenreliesonsparseoutcomerewards,providingcoarsesupervisionforlong-horizonagents.On-policyself-distillation(OPSD)complementsthissignalwithdenseprivilegedfeedback.However,weidentifyDecision--TimestampMismatch:privilegedguidancemaybemisalignedwiththestudent’sfunctionaldecisionbecausethecorrespondingdecisioncanoccuratadifferenttimestep,whilethestudent’sdecisionitselfmayspanmultipletimestepsratherthanbeingtiedtoasingletimestamp.Thus,timestamp-localsupervisioncanmisalignboththecontextandthetemporalscopeofcredit.Toaddressthismismatch,weintroduceAlignOPSD,followingtheprincipleofaligningsupervisionbeforeassigningcredit.Decision-AlignedSupervisionRectificationre-scoresthesamestudent-sampledresponseinfunctionallymatchedcontextsacrosssiblingrolloutstocalibratelocalteacherevidence.Semi-MarkovHierarchicalCreditAssignmentthenderivesvariable-durationdecisionspansfromcorrespondencechangesandusesrectifiedevidencetoallocateoutcome-groundedcreditacrossspansandtheirconstituentturns.WeevaluateAlignOPSDwithQwen2.5-3BandQwen2.5-7BonALFWorld,WebShop,andSearch-QAagainstrepresentativebaselines.AlignOPSDoutperformsbothGRPOandStepOPSDacrossalleightbackbone--aggregate-metriccomparisons,improvingonGRPOby5.5--8.7\%andrankingfirstinsix.Additionalanalyzesexaminethetwoalignmentstagesandhyperparametersensitivitybetweentasks.Ourcodeisavaliableathttps://github.com/mingju-c/Align-OPSD

View arXiv pageView PDFGitHub3Add to collection

Get this paper in your agent:

hf papers read 2609\.33391

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33391 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33391 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33391 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AsyncOPD: How Stale Can On-Policy Distillation Be?

arXiv cs.LG

This paper presents AsyncOPD, a fully asynchronous on-policy distillation pipeline for LLMs, systematically studying the effects of stale-policy data and proposing estimator designs that improve training throughput by 1.6-3.8x while maintaining comparable accuracy.

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Hugging Face Daily Papers

This paper introduces β-OPSD, a generalization of on-policy self-distillation that frames it as a policy optimization family with a controllable KL penalty. The method uses distillation to approximate expensive policy optimization, improving stability and reasoning performance on mathematical benchmarks.