long-horizon-agents

Tag

Cards List
#long-horizon-agents

@rohanpaul_ai: Long-horizon agent reliability has not arrived yet with better models. On WeaveBench's 114 hybrid GUI-CLI tasks, the be…

X AI KOLs Timeline · yesterday Cached

This paper argues that long-horizon AI agent reliability is lacking despite better models, as seen on WeaveBench with only 41.2% pass rate, and proposes the LongHorizon-Harness to manage task state for improved performance.

0 favorites 0 likes
#long-horizon-agents

@matanSF: ProgramBench is a benchmark where agents must reproduce the observable behavior of real software, completely from scrat…

X AI KOLs Following · 2d ago Cached

ProgramBench is a benchmark for agents to reproduce real software behavior from scratch. Factory is announced as the most advanced reverse-engineering system with breakthroughs in model-agnostic long-horizon agents.

0 favorites 0 likes
#long-horizon-agents

Thoughts on Long-Horizon Agents

Reddit r/AI_Agents · 4d ago

The author discusses insights from rebuilding self-service onboarding as agentic flows, emphasizing how long-horizon agents that combine deterministic task management with non-deterministic LLM actions improve reliability and reduce hallucinations in enterprise AI deployments.

0 favorites 0 likes
#long-horizon-agents

ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents

arXiv cs.AI · 5d ago Cached

ECHO is a cognitively inspired, auditable memory architecture for long-horizon agents, evaluated on benchmarks like LoCoMo and LongMemEval-S with high retrieval performance.

0 favorites 0 likes
#long-horizon-agents

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Hugging Face Daily Papers · 5d ago Cached

The paper introduces Recuris, a recursive memory architecture that improves long-horizon agent success by tracking progress and guiding skill selection through localized, validation-gated updates.

0 favorites 0 likes
#long-horizon-agents

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

arXiv cs.CL · 2026-08-18 Cached

This paper presents a holistic evaluation of memory substrates for memory-augmented LLM agents, finding that no single substrate dominates across all regimes and advocating for adaptive substrate routing to optimize performance in different operating conditions.

0 favorites 0 likes
#long-horizon-agents

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

Hugging Face Daily Papers · 2026-08-15 Cached

LongRCA Bench introduces a benchmark for diagnosing failures in long-horizon agent trajectories, and the RCTA method improves responsible role and root-cause step attribution.

0 favorites 0 likes
#long-horizon-agents

Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents

arXiv cs.AI · 2026-08-14 Cached

Introduces Governed Persistent Memory (GPM), a bitemporal state-transition model for auditable long-horizon agent memory with source-bound semantics and fail-closed release, validated on benchmarks and sealed evaluations.

0 favorites 0 likes
#long-horizon-agents

MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory

arXiv cs.AI · 2026-08-12 Cached

This paper introduces MESA, a framework that dynamically selects and fuses a query-adaptive subset of multiple structural memory views for long-horizon agents, achieving 8.5% accuracy improvement over the strongest baseline while using 41% fewer evidence tokens.

0 favorites 0 likes
#long-horizon-agents

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

arXiv cs.CL · 2026-08-10 Cached

This arXiv survey (1,547 papers, 2024-2026) systematically maps the field of long-horizon LLM agents, disambiguating long-horizon, long-context, and long-term memory, and organizing research into six lifecycle categories while identifying the core 'horizon gap' and open measurement problems.

0 favorites 0 likes
#long-horizon-agents

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability

arXiv cs.LG · 2026-08-10 Cached

This empirical study investigates how recurrent context compression affects long-horizon agent behavior, showing that compression can weaken recent interaction influence and cause instability. The authors introduce TRACE, a verifier-guided framework that improves compression reliability and performance on AppWorld.

0 favorites 0 likes
#long-horizon-agents

MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents

arXiv cs.AI · 2026-08-10 Cached

MemPrism proposes a task-conditioned relational memory framework that separates persistent experience storage from decision-time working memory, enabling long-horizon agents to dynamically construct relational views for improved performance and reduced token usage.

0 favorites 0 likes
#long-horizon-agents

StepReflect: Structured UI Transition Reflection for Mobile GUI Agents

arXiv cs.AI · 2026-08-07 Cached

StepReflect reformulates per-step GUI reflection for mobile agents as supervised structured prediction, achieving higher transition accuracy than GPT-5.2 on AndroidWorld while reducing API costs.

0 favorites 0 likes
#long-horizon-agents

The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

arXiv cs.AI · 2026-08-06 Cached

Presents a verification instrument for long-horizon agents that structurally separates commitment drift from binding drift, using a deterministic executive and pre-registered predictions. Reports ablation results showing commitment mechanism removal flips goal abandonment from 0 to 1 while binding error stays flat, though task efficacy is null on ARC-AGI-3.

0 favorites 0 likes
#long-horizon-agents

@ZhihuFrontier: Long-Horizon Agents Need More Than Bigger Context Windows AI Agents are moving from short conversations into software e…

X AI KOLs Timeline · 2026-08-03 Cached

A new survey from Renmin University reviews nearly 1,000 studies on long-horizon AI agents, arguing that reliable long-horizon intelligence depends on the whole model-harness system, not just larger context windows or stronger models.

0 favorites 0 likes
#long-horizon-agents

Just A Rather Very Intelligent Spoken Agent

arXiv cs.AI · 2026-07-21 Cached

This paper introduces JarvisBench, a benchmark for evaluating a continuous, real-time spoken mediation layer in long-horizon AI agent workflows, and presents a modular Jarvis prototype tested on WildClaw tasks with various LLM-based worker agents.

0 favorites 0 likes
#long-horizon-agents

@SharonYixuanLi: Scaling outcome-based RL won't solve long-horizon agentic tasks. Credit assignment is the bottleneck, and turn-level re…

X AI KOLs Timeline · 2026-07-19 Cached

TRACE introduces a turn-level reward assignment method using frozen reference model log-probabilities and temporal-difference learning to address credit assignment in long-horizon agentic tasks, achieving significant improvements in search benchmarks without critic or process labels.

0 favorites 0 likes
#long-horizon-agents

@sheriyuo: TRACE assigns dense credit at tool boundaries by asking a frozen reference model whether each new observation raises th…

X AI KOLs Timeline · 2026-07-16 Cached

TRACE is a dense credit assignment method for multi-turn agentic reinforcement learning that uses a frozen reference model to compute per-action rewards from log-probability changes at tool boundaries, eliminating the need for a critic or process reward model. It significantly improves long-horizon tool-use performance on benchmarks like BrowseComp-Plus.

0 favorites 0 likes
#long-horizon-agents

Proactive Memory for Long-Horizon Agents (16 minute read)

TLDR AI · 2026-07-13 Cached

This paper introduces a proactive memory agent that operates alongside a standard action agent to selectively inject memory-grounded reminders during long-horizon tasks, mitigating behavioral state decay. Experiments on Terminal-Bench and τ²-Bench show significant improvements in pass@1, and the approach is demonstrated with both weak and strong action agents.

0 favorites 0 likes
#long-horizon-agents

From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization

arXiv cs.CL · 2026-07-09 Cached

Introduces STRACE, a framework that performs structural trajectory analysis and causal extraction to construct high signal-to-noise optimization contexts for improving long-horizon agents, outperforming baselines on a formal verification task.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback