Marathoner: Ultra-Long-Horizon Autonomous Intelligence
Summary
The paper proposes Marathoner, an autonomous agentic model for ultra-long-horizon execution, using a comprehensive post-training pipeline with tasks synthesized from GitHub PRs and a novel reward strategy to achieve superior performance on benchmarks.
View Cached Full Text
Cached at: 09/30/26, 04:19 AM
Paper page - Marathoner: Ultra-Long-Horizon Autonomous Intelligence
Source: https://huggingface.co/papers/2609.34378
Abstract
Humansnaturallypossesstheabilitytoworkpersistentlytowardlong-termgoals.Givenachallengingtask,humanscancontinuouslyworkformonthsorevenyearstoaccomplishaspecificobjective.Inthispaper,weproposeMarathoner,anautonomousagenticmodelpossessingtheabilityofultra-long-horizonexecution.Specifically,weproposeacomprehensivepost-trainingpipelinetoinstillthiscriticalcapabilityintobasemodel.ForUltra-Long-HorizonTaskSynthesis,weleveragemajorreleasePRscontaining1000+linesofnewcodefromdiverseGitHubrepositoriesastheprimarysourceforsynthesizingchallengingtask-leveldata.Additionally,weintroduceMulti-TaskChaining,whichchainsmultiplegeneratedtasksintoasinglemorechallengingtask,enablingthesynthesisoftaskswithfrontier-leveldifficulty.Forrejectionsamplingfinetuning,wecombinestrongteachermodelwithdiverseharnessestogeneratetrajectoriesonoursynthesizedtasksandconductsupervisedfinetuningonbasemodelwithrejectionsampledtrajectories.Forreinforcementlearning,cold-startedmodelperformsreal-worldexecutionthroughharnessesinindependentsandboxesduringrolloutprocess,effectivelyfacilitatingtheacquisitionofgenuineultra-long-horizonexecutioncapability.Wefurtherproposeanovelrewardstrategy,LaterStageBonusReward,whichexplicitlyencouragesmodeltoperformmeaningfulmaneuversduringlaterstagesofexecution.Throughextensiveevaluationon5benchmarkscontainingultra-long-horizontasks,Marathonerachievesconsistentandsubstantialperformanceimprovementsoverbasemodelandevensurpassesperformanceofstrongproprietarymodel.FurtheranalysisshowsthatMarathonercanconsistentlyworkfor10+hoursandconduct1000+toolcallsonhighlychallengingtasks.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.34378
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34378 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34378 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34378 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@dair_ai: Super interesting paper from Meta Superintelligence Labs on controlling long agent runs. They use the same workers and …
Meta Superintelligence Labs introduces agentic meta-reasoning, an inference-time harness where a controller decides what work to run next in long agent runs, boosting ProgramBench from 63.7% to 71.5% with GPT-5.5 under the same workers and budget. Gains of 3.6-4.2 points hold across ProofBench, ARC-AGI-2 and LongCoT-mini when averaged over frontier models.
Meta-Harness R&D: Enterprise-Grade Self-Improvement for Long-Horizon AI Workflows
This article explores research and development into making autonomous code improvement disciplined and enterprise-ready for long-horizon AI workflows.
@dair_ai: Outstanding paper on long-horizon agents. (bookmark it) Similar to humans, how do you make agents persist on a difficul…
AutoLab is a new benchmark evaluating 17 frontier models on 36 expert-curated long-horizon tasks (system optimization, model development, CUDA kernels, puzzles), finding that persistence—not initial attempt quality—is the dominant predictor of success. Claude-opus-4.6 led all categories, while most other models terminated prematurely or exhausted budgets with minimal progress.
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Introduces LongHorizon-Harness, a task-state management approach for long-horizon LLM agents using a Manage-Execute-Audit loop, showing consistent improvements across models and benchmarks like WeaveBench and OSWorld.
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Introduces Long-Horizon-Terminal-Bench, a benchmark of 46 long-horizon terminal tasks with dense reward-based grading, evaluating AI agents on planning, long-context, and debugging. Even the strongest model achieves only 15.2% pass@1, showing significant room for improvement.