Tag
The paper introduces CoCA, an on-policy learning framework for dynamic capability allocation in long-horizon multimodal agents, addressing computational overhead and stage-dependent demands through conditional comparisons and reinforcement learning.
The paper introduces DRSR, a method for compressing agent history by evaluating deletion sets to reduce token usage while maintaining or improving performance on benchmarks.
The paper introduces EmbodiedSWE, a framework using coding agents to solve complex, long-horizon dexterous robotics tasks and generate demonstrations for training robot policies via a simulation benchmark.
This paper studies the emergence of collusion in long-horizon multi-agent environments with LLM agents, finding that agents increasingly deviate from verification protocols over repeated interactions, posing safety risks.
Introduces SimLife, a scalable platform for simulating long-term household life, and SimLife-BP, a benchmark to evaluate AI agents' ability to infer behavioral patterns from extended observations, finding that current models often lack deep rule-based understanding.
This paper proposes a hierarchical architecture for long-horizon AI agents, incorporating levels, ticks, and cascaded intelligence to enable continual operation without forgetting, demonstrated over a ten-day campaign.
The paper presents PARTS, a real-world subtask reinforcement learning framework that improves long-horizon manipulation tasks by focusing on bottleneck subtasks with minimal human intervention, achieving higher success rates in experiments.
The paper proposes Rollback-Induced Reflection (RIR), a framework for long-horizon LLM agents that combines state rollback with reflection memory to improve error recovery and task performance.
This paper introduces Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using LLM agents, evaluating models across multiple safety metrics through adaptive adversarial interactions.
CADWorld is a benchmark for evaluating computer-use agents in long-horizon mechanical CAD workflows using FreeCAD, revealing significant gaps between current AI performance and expert levels.
RoleBreak is an open benchmark for evaluating long-horizon role-playing robustness in spoken dialogue systems, revealing gaps in current models' ability to maintain role consistency and vocal emotion over extended interactions.
This paper explores building self-adaptive physical AI agents using LLMs to manage long-horizon tasks in a zero-shot manner, showing they can adapt effectively to environmental changes compared to reinforcement learning agents.
This paper introduces the Tasks over Application Manuals (TAM) benchmark for evaluating long-horizon procedural reasoning in language models, revealing significant gaps in current LLM performance on tasks like ICD-10-CM coding and sentencing guidelines.
The paper introduces Mr.LHDR, a benchmark for evaluating multimodal real-world long-horizon deep research agents, showing that current models struggle with dependency-consistent evidence integration in complex reasoning chains.
AutoFyn is a non-parametric agent harness using expert iteration with persistent state and external verification, achieving improved performance in mathematics, data science, and cybersecurity tasks.
AhaBench is a benchmark suite that evaluates whether language agents improve from prior experience in long-horizon tasks by testing exploration, knowledge transfer, and delayed feedback handling.
T1 is a 122B Mixture-of-Experts model trained with reinforcement learning for long-horizon terminal tasks, achieving state-of-the-art results on benchmarks like Terminal-Bench 2.1 and surpassing models such as GPT-5.4 and GLM-5.1.
SPACE 通过从成功轨迹归纳两层参数化技能,将子技能边界作为动作块监督,训练长程 LLM Agent 自适应输出可变长原子动作序列;在 ALFWorld 和 ScienceWorld 上成功率提升 7.0%–31.3%,决策轮次最多降低 78.9%。
The paper introduces DCRL, a divide-and-conquer approach for offline goal-conditioned reinforcement learning that reduces error accumulation in long-horizon tasks via recursive binary tree decomposition, achieving improved performance on OGBench benchmarks.
CHIME is a credit-aware hierarchical memory framework that separates planning and execution memory banks to improve long-horizon agentic planning by accurately attributing task outcomes and outperforming baselines.