Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Hugging Face Daily Papers Papers

Summary

The paper proposes Marathoner, an autonomous agentic model for ultra-long-horizon execution, using a comprehensive post-training pipeline with tasks synthesized from GitHub PRs and a novel reward strategy to achieve superior performance on benchmarks.

Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:19 AM

Paper page - Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Source: https://huggingface.co/papers/2609.34378

Abstract

Humansnaturallypossesstheabilitytoworkpersistentlytowardlong-termgoals.Givenachallengingtask,humanscancontinuouslyworkformonthsorevenyearstoaccomplishaspecificobjective.Inthispaper,weproposeMarathoner,anautonomousagenticmodelpossessingtheabilityofultra-long-horizonexecution.Specifically,weproposeacomprehensivepost-trainingpipelinetoinstillthiscriticalcapabilityintobasemodel.ForUltra-Long-HorizonTaskSynthesis,weleveragemajorreleasePRscontaining1000+linesofnewcodefromdiverseGitHubrepositoriesastheprimarysourceforsynthesizingchallengingtask-leveldata.Additionally,weintroduceMulti-TaskChaining,whichchainsmultiplegeneratedtasksintoasinglemorechallengingtask,enablingthesynthesisoftaskswithfrontier-leveldifficulty.Forrejectionsamplingfinetuning,wecombinestrongteachermodelwithdiverseharnessestogeneratetrajectoriesonoursynthesizedtasksandconductsupervisedfinetuningonbasemodelwithrejectionsampledtrajectories.Forreinforcementlearning,cold-startedmodelperformsreal-worldexecutionthroughharnessesinindependentsandboxesduringrolloutprocess,effectivelyfacilitatingtheacquisitionofgenuineultra-long-horizonexecutioncapability.Wefurtherproposeanovelrewardstrategy,LaterStageBonusReward,whichexplicitlyencouragesmodeltoperformmeaningfulmaneuversduringlaterstagesofexecution.Throughextensiveevaluationon5benchmarkscontainingultra-long-horizontasks,Marathonerachievesconsistentandsubstantialperformanceimprovementsoverbasemodelandevensurpassesperformanceofstrongproprietarymodel.FurtheranalysisshowsthatMarathonercanconsistentlyworkfor10+hoursandconduct1000+toolcallsonhighlychallengingtasks.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.34378

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.34378 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.34378 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.34378 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

@dair_ai: Super interesting paper from Meta Superintelligence Labs on controlling long agent runs. They use the same workers and …

X AI KOLs Timeline

Meta Superintelligence Labs introduces agentic meta-reasoning, an inference-time harness where a controller decides what work to run next in long agent runs, boosting ProgramBench from 63.7% to 71.5% with GPT-5.5 under the same workers and budget. Gains of 3.6-4.2 points hold across ProofBench, ARC-AGI-2 and LongCoT-mini when averaged over frontier models.

@dair_ai: Outstanding paper on long-horizon agents. (bookmark it) Similar to humans, how do you make agents persist on a difficul…

X AI KOLs Following

AutoLab is a new benchmark evaluating 17 frontier models on 36 expert-curated long-horizon tasks (system optimization, model development, CUDA kernels, puzzles), finding that persistence—not initial attempt quality—is the dominant predictor of success. Claude-opus-4.6 led all categories, while most other models terminated prematurely or exhausted budgets with minimal progress.