MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Hugging Face Daily Papers Papers

Summary

The paper introduces MILO, a framework that co-evolves agent harnesses alongside the search strategy used to discover them, using island-based evolutionary search with mutator agents and an orchestrator. On Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform state-of-the-art harnesses and search methods, even exceeding the Terminal-Bench 2.1 leaderboard's top entry while using 26% fewer tokens.

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by +12.0%, +28.3%, and +10.3%, respectively, compared with best prior-search gains of +4.5%, +18.3%, and 0%. On Terminal-Bench 2.1, it achieves 86.1 pm 2.0%, exceeding the official leaderboard's top entry (83.8 pm 2.3%) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap (0.3808586 to 0.3808568) and the first and third autocorrelation inequalities (1.50274365 to 1.50274360; 1.45081 to 1.44889).
Original Article
View Cached Full Text

Cached at: 10/01/26, 08:24 PM

Paper page - MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Source: https://huggingface.co/papers/2609.38349 Authors:

,

,

,

,

,

,

,

,

,

Abstract

ModernagenticsystemscombineanAImodelwithaharnessthatcontrolsexecutionandenvironmentalinteractions.Harnessdesignstronglyaffectslong-horizonperformance,yetitscombinatorialsearchspacedemandssubstantialhumaneffortthatmustberepeatedasmodelschange.Existingautomatedmethodsexplorethisspacenarrowly,optimizingonlycomponentssuchaspromptsorskillsorbecomingtrappedbyfixed,exploitativesearchstrategies.WeintroduceMILO(Meta-evolutionaryIslandOrchestration),aframeworkthatco-evolvesagentharnessesandthestrategyusedtodiscoverthem.MILOcombines:(i)hierarchicallineagememoryoverisland-basedtrees,usingrejectedmutationsasnegativeevidence;(ii)per-islandmutatoragentsthatrewritecompleteharnessesusingglobalsearchhistoryandparent-specificfeedback;and(iii)anorchestratorthatadaptssearchthroughlineagegraftingandspeciation,mutatorreassignmentandcurriculumrevision.AcrossTerminal-Bench2.1,PaperBench,andDeepSWE,MILO-discoveredharnessesoutperformeightstate-of-the-artharnessesandsixsearchmethodsusingfrontier(Opus4.8)andopen-weight(gpt-oss-120b)models.WithOpus4.8,MILOimprovesresolutionoveritsinitialharnessby+12.0%,+28.3%,and+10.3%,respectively,comparedwithbestprior-searchgainsof+4.5%,+18.3%,and0%.OnTerminal-Bench2.1,itachieves86.1pm2.0%,exceedingtheofficialleaderboard’stopentry(83.8pm2.3%)whileusing26\%fewertokensthanitsinitialharness.OnEinsteinArenaopenproblems,MILOimprovesbest-knownupperboundsforErdősminimum-overlap(0.3808586to0.3808568)andthefirstandthirdautocorrelationinequalities(1.50274365to1.50274360;1.45081to1.44889).

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.38349

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38349 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38349 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38349 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Self-Harness: Harnesses That Improve Themselves

Hacker News Top

Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.

@dair_ai: Great paper from Meta on agent harness optimization. Meta-Harness-style search uses one development set and one proposa…

X AI KOLs Timeline

Meta (with Duke and UC Davis) proposes a Mixture of Self-Improving Branches framework for agent harness optimization, splitting the single-trajectory Meta-Harness search into adaptive branches with evolving development subsets and proposal policies, plus a router that selects the best branch per input, achieving up to +34.8% relative gains on Olympiad-level math, +11.6% on Terminal-Bench 2.0, and +3.8% on SWE-bench Lite.

Rethinking the Evaluation of Harness Evolution for Agents

arXiv cs.AI

This paper re-evaluates the methodology of automatic harness evolution for LLM agents, highlighting that its gains may stem from additional test-time search rather than improved harness design, and that evaluation on the same benchmark risks overfitting. Experiments show that harness evolution does not consistently outperform simpler test-time scaling methods.