Tag
This paper introduces AlignOPSD to address decision-timestamp mismatch in on-policy self-distillation for long-horizon agents, improving performance on benchmarks like ALFWorld, WebShop, and Search-QA compared to baselines.