MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Summary
The paper introduces MILO, a framework that co-evolves agent harnesses alongside the search strategy used to discover them, using island-based evolutionary search with mutator agents and an orchestrator. On Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform state-of-the-art harnesses and search methods, even exceeding the Terminal-Bench 2.1 leaderboard's top entry while using 26% fewer tokens.
View Cached Full Text
Cached at: 10/01/26, 08:24 PM
Paper page - MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Source: https://huggingface.co/papers/2609.38349 Authors:
,
,
,
,
,
,
,
,
,
Abstract
ModernagenticsystemscombineanAImodelwithaharnessthatcontrolsexecutionandenvironmentalinteractions.Harnessdesignstronglyaffectslong-horizonperformance,yetitscombinatorialsearchspacedemandssubstantialhumaneffortthatmustberepeatedasmodelschange.Existingautomatedmethodsexplorethisspacenarrowly,optimizingonlycomponentssuchaspromptsorskillsorbecomingtrappedbyfixed,exploitativesearchstrategies.WeintroduceMILO(Meta-evolutionaryIslandOrchestration),aframeworkthatco-evolvesagentharnessesandthestrategyusedtodiscoverthem.MILOcombines:(i)hierarchicallineagememoryoverisland-basedtrees,usingrejectedmutationsasnegativeevidence;(ii)per-islandmutatoragentsthatrewritecompleteharnessesusingglobalsearchhistoryandparent-specificfeedback;and(iii)anorchestratorthatadaptssearchthroughlineagegraftingandspeciation,mutatorreassignmentandcurriculumrevision.AcrossTerminal-Bench2.1,PaperBench,andDeepSWE,MILO-discoveredharnessesoutperformeightstate-of-the-artharnessesandsixsearchmethodsusingfrontier(Opus4.8)andopen-weight(gpt-oss-120b)models.WithOpus4.8,MILOimprovesresolutionoveritsinitialharnessby+12.0%,+28.3%,and+10.3%,respectively,comparedwithbestprior-searchgainsof+4.5%,+18.3%,and0%.OnTerminal-Bench2.1,itachieves86.1pm2.0%,exceedingtheofficialleaderboard’stopentry(83.8pm2.3%)whileusing26\%fewertokensthanitsinitialharness.OnEinsteinArenaopenproblems,MILOimprovesbest-knownupperboundsforErdősminimum-overlap(0.3808586to0.3808568)andthefirstandthirdautocorrelationinequalities(1.50274365to1.50274360;1.45081to1.44889).
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.38349
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38349 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.38349 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38349 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Self-Harness: Harnesses That Improve Themselves
Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.
@dair_ai: Great paper from Meta on agent harness optimization. Meta-Harness-style search uses one development set and one proposa…
Meta (with Duke and UC Davis) proposes a Mixture of Self-Improving Branches framework for agent harness optimization, splitting the single-trajectory Meta-Harness search into adaptive branches with evolving development subsets and proposal policies, plus a router that selects the best branch per input, achieving up to +34.8% relative gains on Olympiad-level math, +11.6% on Terminal-Bench 2.0, and +3.8% on SWE-bench Lite.
Rethinking the Evaluation of Harness Evolution for Agents
This paper re-evaluates the methodology of automatic harness evolution for LLM agents, highlighting that its gains may stem from additional test-time search rather than improved harness design, and that evaluation on the same benchmark risks overfitting. Experiments show that harness evolution does not consistently outperform simpler test-time scaling methods.
@omarsar0: // Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today …
This paper introduces Self-Harness, a new paradigm where LLM-based agents iteratively improve their own operating harness—prompts, tools, and control flow—without human engineers or stronger external agents, achieving significant performance gains across multiple models.
EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
The paper introduces EVOHARNESSBENCH, a benchmark for evaluating LLM agents under evolving tool, skill, and agent harnesses, revealing gaps in retention and adaptation.