long-horizon-tasks

Tag

Cards List
#long-horizon-tasks

Long-running agents: is the bottleneck the model or the scaffolding around it?

Reddit r/artificial ↗ · 5h ago

A practitioner discussion exploring whether long-running AI agent failures stem from model capabilities or from the scaffolding around them, highlighting error compounding, context pollution, and weak self-correction as key failure modes.

0 favorites 0 likes
#long-horizon-tasks

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

arXiv cs.AI ↗ · 3d ago Cached

The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for offline goal-conditioned reinforcement learning that trains a single k-conditioned value function across multiple temporal abstraction scales, resolving the trade-off between long-range value signal and local resolution. GITA outperforms offline GCRL baselines like HIQL and OTA on OGBench, raising average success rates by 25 percentage points over HIQL.

0 favorites 0 likes
#long-horizon-tasks

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Hugging Face Daily Papers ↗ · 4d ago Cached

VeriHarness turns an LLM generator into an agentic verifier—using a workspace, evidence tools, reusable verification skills, a disagreement resolver, and a consensus challenger—to select reliable outputs for long-horizon tasks without reference answers. It achieves top selection scores across five benchmarks, gains of 6.2–6.4 points over single rollouts with Gemini 3.5 Flash and Claude Opus 4.8, and releases ~26,000 rollouts costing over $100,000.

0 favorites 0 likes
#long-horizon-tasks

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Hugging Face Daily Papers ↗ · 6d ago Cached

The paper introduces MILO, a framework that co-evolves agent harnesses alongside the search strategy used to discover them, using island-based evolutionary search with mutator agents and an orchestrator. On Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform state-of-the-art harnesses and search methods, even exceeding the Terminal-Bench 2.1 leaderboard's top entry while using 26% fewer tokens.

0 favorites 0 likes
#long-horizon-tasks

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

arXiv cs.CL ↗ · 2026-09-28 Cached

The paper introduces KNOWS, a benchmark for evaluating web agents on complex, long-horizon tasks that involve synthesizing and organizing knowledge into artifacts, revealing that current agents struggle with visual steps and long-horizon reasoning.

0 favorites 0 likes
#long-horizon-tasks

BIABench: Evaluating AI agents on real-world bioimage analysis tasks

Hugging Face Daily Papers ↗ · 2026-09-28 Cached

The paper introduces BIABench, an open benchmark of 16 real-world bioimage analysis tasks reconstructed from published studies, evaluating AI agents end-to-end with outcome and process scores. Agents solved routine 2D tasks well but failed on 3D and time-lapse tasks, with neither specialization, stronger models, nor expert instructions closing the reliability gap.

0 favorites 0 likes
#long-horizon-tasks

StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

Hugging Face Daily Papers ↗ · 2026-09-28 Cached

StructRL introduces an online reinforcement learning framework that uses structured intermediate rewards from verifiable subtasks to enhance long-horizon vision-language-action tasks, demonstrating superior performance on benchmarks like RoboCasa365 and LIBERO-Long.

0 favorites 0 likes
#long-horizon-tasks

Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Hugging Face Daily Papers ↗ · 2026-09-28 Cached

The paper proposes Marathoner, an autonomous agentic model for ultra-long-horizon execution, using a comprehensive post-training pipeline with tasks synthesized from GitHub PRs and a novel reward strategy to achieve superior performance on benchmarks.

0 favorites 0 likes
#long-horizon-tasks

@giohuh_: What can Astra do when given a humanoid embodiment? We built HomeBody to find out. Controlled by GPT Astra, it carries …

X AI KOLs Timeline ↗ · 2026-09-26 Cached

HomeBody is a humanoid robot controlled by GPT Astra that explores unseen environments, builds a digital twin, and performs long-horizon tasks like tidying and retrieving objects without environment-specific training.

0 favorites 0 likes
#long-horizon-tasks

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

arXiv cs.LG ↗ · 2026-09-22 Cached

The paper introduces Trace, a credit-guided framework that compiles noisy interaction histories into executable walkthroughs to enhance long-horizon AI agent performance, showing significant improvements in experiments.

0 favorites 0 likes
#long-horizon-tasks

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

This paper introduces Taste-Bench, a benchmark for measuring taste in LLM agents' long-horizon decisions, finding that frontier models have low accuracy and that taste can be improved through distillation training.

0 favorites 0 likes
#long-horizon-tasks

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

arXiv cs.AI ↗ · 2026-09-15 Cached

The paper introduces Continual Search, an iterative framework to enhance root-cause attribution for AI agent failures by searching for diagnostic evidence in long-horizon execution traces, and evaluates it on benchmarks including MegaRCA-Mix, showing significant performance improvements.

0 favorites 0 likes
#long-horizon-tasks

Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning

arXiv cs.LG ↗ · 2026-09-14 Cached

The paper proposes GACA, a granularity-adaptive credit assignment method for long-horizon LLM agent reinforcement learning that improves task success by adapting resolution to step importance.

0 favorites 0 likes
#long-horizon-tasks

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

arXiv cs.AI ↗ · 2026-09-11 Cached

This paper compares subagent execution versus agent skill execution for long-horizon tasks in language model agents, finding that subagent execution outperforms when skills are well-defined, though it increases communication overhead.

0 favorites 0 likes
#long-horizon-tasks

Progressive Point Matching (8 minute read)

TLDR AI ↗ · 2026-09-09 Cached

Progressive Point Matching (PPM) is a framework proposed to assign partial credit in reinforcement learning for long-horizon tasks in LLMs, addressing the inefficiency of sparse outcome rewards by treating reasoning as paths through a Markovian state space.

0 favorites 0 likes
#long-horizon-tasks

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper proposes Feedback-Enriched Environments (FEEs) to bootstrap self-evolving agents in long-horizon tasks by enriching feedback for reinforcement learning, showing improved stability and performance on benchmarks using Qwen3 models and RL algorithms.

0 favorites 0 likes
#long-horizon-tasks

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

arXiv cs.LG ↗ · 2026-09-04 Cached

TIGPO proposes a temporal instance-graph policy optimization method that extends graph-based credit assignment across policy updates for long-horizon LLM agents, using persistent transition graphs and revisit slots to improve advantage estimation and performance on benchmarks like ALFWorld and WebShop.

0 favorites 0 likes
#long-horizon-tasks

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

arXiv cs.LG ↗ · 2026-09-03 Cached

The paper proposes 'Space', a skill-guided adaptive action chunking method for long-horizon LLM agents, improving success rates by 7.0%–31.3% and reducing LLM decision rounds by up to 78.9%.

0 favorites 0 likes
#long-horizon-tasks

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

arXiv cs.AI ↗ · 2026-09-03 Cached

SkillGLoW is a method for LLM agents that organizes skills into procedural families to enhance self-improvement on long-horizon tasks, demonstrating significant performance gains over baselines.

0 favorites 0 likes
#long-horizon-tasks

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

E-Commerce Bench is an open-source benchmark that evaluates LLM agents on long-horizon autonomous business operation in e-commerce, featuring multi-store negotiation and dynamic events over a simulated year.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback