@dair_ai: Great paper from Meta on agent harness optimization. Meta-Harness-style search uses one development set and one proposa…
Summary
Meta (with Duke and UC Davis) proposes a Mixture of Self-Improving Branches framework for agent harness optimization, splitting the single-trajectory Meta-Harness search into adaptive branches with evolving development subsets and proposal policies, plus a router that selects the best branch per input, achieving up to +34.8% relative gains on Olympiad-level math, +11.6% on Terminal-Bench 2.0, and +3.8% on SWE-bench Lite.
View Cached Full Text
Cached at: 10/02/26, 02:32 AM
Great paper from Meta on agent harness optimization.
Meta-Harness-style search uses one development set and one proposal policy, so every edit follows a single path and can get stuck in a local optimum.
This work splits the search into branches.
Each branch keeps the development cases its harnesses solve better than other branches, drops cases every branch already solves, and rewrites its own proposal policy from its history. A router then picks one branch’s harness for each new input before it runs.
Relative to Meta-Harness, that gives +34.8% on Olympiad-level math, +11.6% on Terminal-Bench 2.0 and +3.8% on SWE-bench Lite. Harness selection and the router use development data only.
Paper: https://academy.dair.ai/papers/mixture-of-self-improving-branches-for-agent-harness-optimization-2609.37834…
Mixture of Self-Improving Branches For Agent Harness Optimization
Source: https://academy.dair.ai/papers/mixture-of-self-improving-branches-for-agent-harness-optimization-2609.37834 AgentsChat with Paper
First page

The curator’s take
Haoyu Dong, Zihao Lin, Lizhu Zhang, Zhuokai Zhao and colleagues at Meta (with Duke and UC Davis) extend Meta-Harness-style harness optimization by splitting the search into branches, each with its own evolving development subset and proposal policy, and then routing each new input to one branch’s best harness.
Ask this paper
Question about this paper Key points01
Problem with one trajectory. Meta-Harness keeps a fixed development set and proposal policy, so all edits follow one search path and can settle in a local optimum.
02
Branch objectives. Each branch keeps the development cases its leading harnesses solve more often than other branches do, drops cases that every branch already solves, and rewrites its proposal policy from its own search history.
03
Router. Before execution a router picks one development-selected branch head per input, combining complementary harnesses without looking at test outcomes.
04
Results. Relative gains over Meta-Harness of 34.8% on Olympiad-level math, 11.6% on Terminal-Bench 2.0 and 3.8% on SWE-bench Lite, with harness selection and router configuration using development data only.
AbstractHarness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.
Similar Articles
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
This paper introduces Meta-Skill, a method letting a frozen Builder model learn reusable principles from Target execution feedback to construct better agent harnesses for unseen tasks, improving performance by 8.95 points on Harness-Bench and NewtonBench. It suggests a path toward system-level self-improvement by learning to build better environments rather than changing model weights.
@omarsar0: // Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today …
This paper introduces Self-Harness, a new paradigm where LLM-based agents iteratively improve their own operating harness—prompts, tools, and control flow—without human engineers or stronger external agents, achieving significant performance gains across multiple models.
@dair_ai: Great paper on self-evolving agent harnesses. Self-evolving agent harnesses have two practical problems: 1. Search is s…
The paper proposes Ecdysis, a framework for training runtime harnesses for LLM agents that identifies recurring failure patterns to improve efficiency and accuracy, achieving 1.84x faster training and 18.56% higher reasoning accuracy.
@dair_ai: // State-Externalizing Harnesses // A new paradigm is emerging on how to effectively build agents and harnesses. If the…
Harness-1 introduces a state-externalizing harness that separates routine bookkeeping from policy decisions in search agents, enabling a 20B model to outperform larger frontier searchers across multiple benchmarks.
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
AutoDesign introduces a meta-harness optimizer that recursively improves a code agent for long-horizon structured media generation, achieving state-of-the-art results on paper-to-poster synthesis and outperforming commercial systems like Claude Design on the new PosterBench benchmark.