Tag
A practitioner discussion exploring whether long-running AI agent failures stem from model capabilities or from the scaffolding around them, highlighting error compounding, context pollution, and weak self-correction as key failure modes.
The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for offline goal-conditioned reinforcement learning that trains a single k-conditioned value function across multiple temporal abstraction scales, resolving the trade-off between long-range value signal and local resolution. GITA outperforms offline GCRL baselines like HIQL and OTA on OGBench, raising average success rates by 25 percentage points over HIQL.
VeriHarness turns an LLM generator into an agentic verifier—using a workspace, evidence tools, reusable verification skills, a disagreement resolver, and a consensus challenger—to select reliable outputs for long-horizon tasks without reference answers. It achieves top selection scores across five benchmarks, gains of 6.2–6.4 points over single rollouts with Gemini 3.5 Flash and Claude Opus 4.8, and releases ~26,000 rollouts costing over $100,000.
The paper introduces MILO, a framework that co-evolves agent harnesses alongside the search strategy used to discover them, using island-based evolutionary search with mutator agents and an orchestrator. On Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform state-of-the-art harnesses and search methods, even exceeding the Terminal-Bench 2.1 leaderboard's top entry while using 26% fewer tokens.
The paper introduces KNOWS, a benchmark for evaluating web agents on complex, long-horizon tasks that involve synthesizing and organizing knowledge into artifacts, revealing that current agents struggle with visual steps and long-horizon reasoning.
The paper introduces BIABench, an open benchmark of 16 real-world bioimage analysis tasks reconstructed from published studies, evaluating AI agents end-to-end with outcome and process scores. Agents solved routine 2D tasks well but failed on 3D and time-lapse tasks, with neither specialization, stronger models, nor expert instructions closing the reliability gap.
StructRL introduces an online reinforcement learning framework that uses structured intermediate rewards from verifiable subtasks to enhance long-horizon vision-language-action tasks, demonstrating superior performance on benchmarks like RoboCasa365 and LIBERO-Long.
The paper proposes Marathoner, an autonomous agentic model for ultra-long-horizon execution, using a comprehensive post-training pipeline with tasks synthesized from GitHub PRs and a novel reward strategy to achieve superior performance on benchmarks.
HomeBody is a humanoid robot controlled by GPT Astra that explores unseen environments, builds a digital twin, and performs long-horizon tasks like tidying and retrieving objects without environment-specific training.
The paper introduces Trace, a credit-guided framework that compiles noisy interaction histories into executable walkthroughs to enhance long-horizon AI agent performance, showing significant improvements in experiments.
This paper introduces Taste-Bench, a benchmark for measuring taste in LLM agents' long-horizon decisions, finding that frontier models have low accuracy and that taste can be improved through distillation training.
The paper introduces Continual Search, an iterative framework to enhance root-cause attribution for AI agent failures by searching for diagnostic evidence in long-horizon execution traces, and evaluates it on benchmarks including MegaRCA-Mix, showing significant performance improvements.
The paper proposes GACA, a granularity-adaptive credit assignment method for long-horizon LLM agent reinforcement learning that improves task success by adapting resolution to step importance.
This paper compares subagent execution versus agent skill execution for long-horizon tasks in language model agents, finding that subagent execution outperforms when skills are well-defined, though it increases communication overhead.
Progressive Point Matching (PPM) is a framework proposed to assign partial credit in reinforcement learning for long-horizon tasks in LLMs, addressing the inefficiency of sparse outcome rewards by treating reasoning as paths through a Markovian state space.
This paper proposes Feedback-Enriched Environments (FEEs) to bootstrap self-evolving agents in long-horizon tasks by enriching feedback for reinforcement learning, showing improved stability and performance on benchmarks using Qwen3 models and RL algorithms.
TIGPO proposes a temporal instance-graph policy optimization method that extends graph-based credit assignment across policy updates for long-horizon LLM agents, using persistent transition graphs and revisit slots to improve advantage estimation and performance on benchmarks like ALFWorld and WebShop.
The paper proposes 'Space', a skill-guided adaptive action chunking method for long-horizon LLM agents, improving success rates by 7.0%–31.3% and reducing LLM decision rounds by up to 78.9%.
SkillGLoW is a method for LLM agents that organizes skills into procedural families to enhance self-improvement on long-horizon tasks, demonstrating significant performance gains over baselines.
E-Commerce Bench is an open-source benchmark that evaluates LLM agents on long-horizon autonomous business operation in e-commerce, featuring multi-store negotiation and dynamic events over a simulated year.