long-horizon-reasoning

Tag

Cards List
#long-horizon-reasoning

@rohanpaul_ai: The paper shows that agents reason better over long periods when no past information is thrown away. Keeping every past…

X AI KOLs Timeline · 2026-07-24 Cached

The paper introduces PRO-LONG, which uses a programmatic, lossless memory system for AI agents, storing all past actions in a structured text log. This approach significantly improves long-horizon reasoning and performance on ARC-AGI-3 games while using fewer tokens than stronger specialized systems.

0 favorites 0 likes
#long-horizon-reasoning

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

arXiv cs.AI · 2026-07-13 Cached

LongMedBench is a new benchmark for evaluating LLM-based medical agents on long-horizon clinical decision-making using real EHR data from MIMIC-IV. It includes 335 patients with multiple visits and proposes evaluation suites for fact-based QA, temporal reasoning, and long-horizon decision-making.

0 favorites 0 likes
#long-horizon-reasoning

@chenxiao_yang_: For longer-horizon tasks, we often think about using a long-context model. But harnesses also matter! In fact, they are…

X AI KOLs Timeline · 2026-07-08 Cached

This ICML paper introduces recursive models that recursively invoke themselves to solve subtasks in isolated contexts, proving they can surpass context-bounded autoregressive models for long-horizon reasoning. Experiments on SAT solving and Go game-tree search show improved accuracy with small active contexts.

0 favorites 0 likes
#long-horizon-reasoning

Sakana Marlin (4 minute read)

TLDR AI · 2026-06-16 Cached

Sakana AI launches its first commercial product, Sakana Marlin, an autonomous research assistant that completes strategy work in hours by generating structured slides and detailed reports.

0 favorites 0 likes
#long-horizon-reasoning

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

arXiv cs.CL · 2026-06-02 Cached

This paper presents SPADER, a reinforcement learning framework for multi-answer QA that uses step-wise peer advantage for credit assignment and diversity-aware exploration rewards to improve recall of long-tail entities, achieving better performance on several benchmarks.

0 favorites 0 likes
#long-horizon-reasoning

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

Hugging Face Daily Papers · 2026-05-23 Cached

This paper proposes SAM, a state-adaptive memory framework that dynamically manages interaction histories for long-horizon agentic reasoning, enabling intent-driven recall without retraining the backbone model. It outperforms strong baselines across multiple benchmarks like BrowseComp and HLE.

0 favorites 0 likes
#long-horizon-reasoning

MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning

Hugging Face Daily Papers · 2026-05-13 Cached

The paper proposes the Map-then-Act Paradigm (MAP), a plug-and-play framework that shifts environmental understanding before execution in interactive LLM agents, achieving consistent gains across benchmarks and enabling frontier models to surpass near-zero baseline performance in 22 of 25 game environments.

0 favorites 0 likes
#long-horizon-reasoning

HAGE: Harnessing Agentic Memory via RL-Driven Weighted Graph Evolution

Hugging Face Daily Papers · 2026-05-11 Cached

HAGE introduces a weighted multi-relational memory framework that enables query-conditioned traversal over unified relational memory graphs, improving long-horizon reasoning accuracy through adaptive memory retrieval and reinforcement learning-based optimization.

0 favorites 0 likes
← Back to home

Submit Feedback