long-horizon-tasks

Tag

Cards List
#long-horizon-tasks

@rohanpaul_ai: "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid ev…

X AI KOLs Timeline ↗ · 2026-08-28 Cached

Chamath criticizes long-horizon tasks in AI as ineffective, predicts a hype cycle leading to disillusionment, and suggests using symbolic spaces to guide embedded spaces for better AI performance.

0 favorites 0 likes
#long-horizon-tasks

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Hugging Face Daily Papers ↗ · 2026-08-28 Cached

LoopArena introduces a benchmark to evaluate models acting as runtime controllers for loop engineering in coding agent tasks, revealing low strict success rates but significant cost reductions.

0 favorites 0 likes
#long-horizon-tasks

@FactoryAI: We built the world's most advanced program reverse-engineering system. It's model-independent and demonstrated improved…

X AI KOLs Following ↗ · 2026-08-27 Cached

FactoryAI has developed a model-independent program reverse-engineering system that improves performance on long-horizon software tasks for all tested frontier models by providing an independent definition of done.

0 favorites 0 likes
#long-horizon-tasks

@usenaive: Introducing Vetta, the most efficient harness for long-horizon agent tasks. Same model, same tasks. Only the harness ch…

X AI KOLs Following ↗ · 2026-08-24 Cached

Vetta is introduced as a cost-efficient harness for long-horizon agent tasks, reducing per-task cost to $0.298 compared to $0.872 for claude-code and $1.095 for hermes.

0 favorites 0 likes
#long-horizon-tasks

Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

Agent-G^2 introduces a Gaussian guidance framework for hint depth in reinforcement learning, enhancing performance on long-horizon agentic tasks without extra probing rollouts, with superior results on ALFWorld and WebShop benchmarks.

0 favorites 0 likes
#long-horizon-tasks

Nvidia just showed that the harness, not the AI model, is now the real hero

TechCrunch AI ↗ · 2026-08-21 Cached

Nvidia's research demonstrates that a well-designed harness around an AI model, rather than the model itself, significantly improves performance on long-horizon tasks, with Claude Opus 5 achieving a perfect score on the ARC-AGI-3 benchmark.

0 favorites 0 likes
#long-horizon-tasks

Casi 22 días, un solo objetivo y un repositorio de 1,8 millones de líneas: ¿estamos midiendo mal la autonomía de los agentes?

Reddit r/AI_Agents ↗ · 2026-08-20

AutoNodo ha logrado una ejecución autónoma continua de casi 22 días sobre un único objetivo técnico en un repositorio de 1.8 millones de líneas, cuestionando las métricas actuales de autonomía en agentes de programación.

0 favorites 0 likes
#long-horizon-tasks

@svpino: Progress in open models is keeping Big AI labs up at night, and I'm here for it! We have a brand new open-weight multim…

X AI KOLs Following ↗ · 2026-08-19 Cached

The article introduces the dots3-note Preview model, an open-weight multimodal AI model optimized for long-horizon tasks with TEMPO, a reinforcement learning technique that enables self-critique and adaptation.

0 favorites 0 likes
#long-horizon-tasks

TL;DR Of why dsh/cordis is a big deal: LH tasks and harness meta tuning

Reddit r/singularity ↗ · 2026-08-15

Cordis is a framework that supports long-horizon tasks by enabling crash recovery and dynamic plugin addition through a reversible, tracked context system, ensuring efficient state management with minimal loss.

0 favorites 0 likes
#long-horizon-tasks

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Hugging Face Daily Papers ↗ · 2026-08-13 Cached

This paper systematically evaluates seven frontier AI agents on long-horizon tasks, revealing they function more as engineering optimizers than autonomous researchers, with recommendations for improving training and experience management.

0 favorites 0 likes
#long-horizon-tasks

@Xudong07452910: An Agent completed a 20-step task and only received a "success/failure" at the end. During training, how do you know which step actually saved the task? This AgentOPSD paper by Tsinghua, Zhejiang University, and Meituan team studies the credit assignment problem for long-horizon agents. GRPO usually assigns the final reward…

X AI KOLs Timeline ↗ · 2026-08-10 Cached

Tsinghua, Zhejiang University, and Meituan team propose AgentOPSD, a recursive self-distillation credit assignment method that converts sparse final rewards into step-wise credit signals, improving reinforcement learning performance for long-horizon agents. It significantly outperforms the GRPO baseline on tasks like ALFWorld.

0 favorites 0 likes
#long-horizon-tasks

Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents

arXiv cs.AI ↗ · 2026-08-05 Cached

VerMem is a framework for unified memory management in LLM agents, using local and global verifiers with reinforcement learning to jointly control long-term and short-term memory. It outperforms strong baselines across five benchmarks with improved efficiency-performance trade-offs.

0 favorites 0 likes
#long-horizon-tasks

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Hugging Face Daily Papers ↗ · 2026-07-29 Cached

OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon office-suite tasks with economic grounding, comparing human costs and LLM inference costs. It includes 100 tasks requiring ~2.3 hours of human labor each, and finds that frontier LLMs are cheaper and faster but still below human quality.

0 favorites 0 likes
#long-horizon-tasks

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

arXiv cs.LG ↗ · 2026-07-28 Cached

Proposes Progress-conditioned Group Policy Optimization (ProGPO) to overcome credit assignment issues in long-horizon agentic tasks by using first-visit observation coverage when all group samples fail, improving performance on ALFWorld and WebShop.

0 favorites 0 likes
#long-horizon-tasks

RoboTTT: Context Scaling for Robot Policies

Hugging Face Daily Papers ↗ · 2026-07-16 Cached

RoboTTT scales visuomotor context to 8K timesteps for robot policies, enabling one-shot imitation from human video demonstrations, on-the-fly policy improvement, and robustness to perturbations. It achieves an 87% improvement over baselines and completes a five-minute, ten-stage assembly task that no baseline could.

0 favorites 0 likes
#long-horizon-tasks

SelfMem: Self-Optimizing Memory for AI Agents

arXiv cs.CL ↗ · 2026-07-07 Cached

SelfMem introduces a self-optimizing memory framework for AI agents that allows them to explore, evaluate, and refine their own memory strategies through memory tools and feedback signals, achieving significant improvements over baselines on the BEAM benchmark across large conversation scales.

0 favorites 0 likes
#long-horizon-tasks

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

RoboTALES introduces a two-stage framework combining LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training, significantly outperforming existing methods on long-horizon manipulation tasks.

0 favorites 0 likes
#long-horizon-tasks

@omarsar0: // AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory…

X AI KOLs Timeline ↗ · 2026-07-02 Cached

This Stanford research paper introduces AutoMem, a framework that treats agent memory management as a trainable skill. By optimizing memory structure and proficiency separately, AutoMem improves base agent performance 2x-4x on long-horizon tasks, enabling a 32B open-weight model to compete with frontier systems like Claude Opus 4.5 and Gemini 3.1 Pro Thinking.

0 favorites 0 likes
#long-horizon-tasks

AutoMem: Automated Learning of Memory as a Cognitive Skill

arXiv cs.AI ↗ · 2026-07-02 Cached

AutoMem introduces a framework that automates learning of memory management as a trainable skill for LLMs, improving performance on long-horizon tasks by 2x-4x through optimizing memory structure and proficiency.

0 favorites 0 likes
#long-horizon-tasks

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Hugging Face Daily Papers ↗ · 2026-06-28 Cached

OSWorld 2.0 is a new benchmark for evaluating computer-use agents on 108 long-horizon, real-world workflows. Current agents like Claude Opus 4.8 and GPT-5.5 achieve low completion rates, highlighting significant limitations in handling complex, multi-step tasks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback