Tag
TurnOPD introduces turn-level budgeting for on-policy distillation of long-horizon agents, addressing inefficiencies in vanilla OPD by adaptive rollout-depth and progressive turn-normalized loss budgeting, achieving better accuracy under equal training budgets.
This paper introduces 'memory in the loop', where language agents repeatedly access an in-process associative store on every reasoning step. By using a fast (∼100μs) in-process store, the per-step retrieval cost is reduced by three orders of magnitude compared to networked stores, eliminating redundant actions and improving recall across GPT-5-class models.
研究语言智能体在长期任务中世界模型塌缩的相变现象,发现状态负载和依赖密度等参数在临界点附近导致模型突然崩溃,而非逐渐退化。
This paper investigates whether natural-language feedback leads to improvement beyond repeated attempts alone in multi-turn language agent settings. Using a controlled student-teacher protocol across multiple benchmarks, the authors find that self-generated feedback adds little, while strong external teachers yield larger gains, and that the student's ability to act on feedback is a key bottleneck.
The paper introduces ATOD, a hybrid online distillation algorithm combining on-policy distillation and reinforcement learning for training small language model agents in multi-turn tasks, featuring an annealed OPD-RL schedule and Turn-level Disagreement-Uncertainty Reweighting to improve dense supervision.
This paper introduces Grounded Iterative Language Planning (GILP), a method that combines a small parameterized world model with LLM-based reasoning to reduce hallucination propagation in LLM agents. Experiments show GILP reduces hallucinated-state rate from 0.176 to 0.035 and raises task success from 0.668 to 0.838 on graph-structured planning benchmarks.
OPID is a framework that extracts dense token-level supervision from completed on-policy trajectories for reinforcement learning of language agents, using hierarchical skills (episode-level and step-level) to improve sample efficiency and robustness.
This paper introduces the concept of memory depth for long-running language agents, distinguishing it from retrieval-based memory access, and proposes EVAF, a selective parametric consolidation mechanism using surprise- and valence-gated LoRA updates. Experiments across multiple models show EVAF improves goal persistence after context unload with minimal parametric writes.
OPID proposes an on-policy skill distillation framework that extracts dense hindsight supervision from completed trajectories, combining outcome-based RL with token-level self-distillation to improve language agent training efficiency and performance on multi-turn tasks.
This paper formulates memory retention for long-horizon language agents as a constrained stochastic optimization problem, introducing OSL-MR, a framework that enforces observability-safe learning with a Mixed-Score heuristic. Experiments show consistent improvements over existing heuristic baselines under tight memory budgets.
This paper investigates whether spatial geometry improves language-agent memory recall, demonstrating that geometry must lead recall over recency and importance, and that a ray-tracing visibility predicate is crucial for occlusion handling in 3D voxel worlds.
Proposes CVT-RL, a constrained policy-gradient algorithm with policy-conditioned counterfactual contribution estimation and verifiable rewards, improving long-horizon language agent reliability and reducing reward hacking.
This paper introduces ArcANE, an automatically constructed benchmark for evaluating role-playing language agents' alignment with character psychological trajectories across narrative phases, showing that conditioning on character arc information improves performance, especially in scenarios beyond the source text.
A comprehensive evaluation framework for continual learning in language agents is introduced, emphasizing controlled task streams and memory design analysis to better assess reusable experience and learning stability.
This paper introduces PaW, a co-training framework that adds auxiliary world modeling supervision to policy learning during on-policy RL rollouts, improving language agent training without additional computational overhead.
LiteCoder-Terminal-Gen introduces a zero-dependency synthetic pipeline that generates executable terminal training environments, producing SFT and RL datasets that enable language agents to achieve significant performance gains on Terminal Bench benchmarks.
RICE-PO is a critic-free policy optimization framework that turns retrieval interactions into localized credit signals for training reasoning agents, outperforming prompt-based and group-based RL baselines on BRIGHT and BEIR benchmarks.
This paper systematically evaluates model-generated skills for language agents across the full lifecycle of experience generation, extraction, and consumption, finding that skills are beneficial on average but exhibit non-trivial negative transfer, leading to a meta-skill that improves skill quality.
Auto-Dreamer introduces a learned offline memory consolidation method for language agents, decoupling fast memory acquisition from slow cross-session consolidation, and achieving higher performance with smaller memory banks, generalizing to unseen environments.
OpenAgents is an open platform for using and hosting language agents in everyday life, featuring agents for data analysis, plugins, and web browsing, with open code and a demo.