Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Summary
This paper introduces Meta-Skill, a method letting a frozen Builder model learn reusable principles from Target execution feedback to construct better agent harnesses for unseen tasks, improving performance by 8.95 points on Harness-Bench and NewtonBench. It suggests a path toward system-level self-improvement by learning to build better environments rather than changing model weights.
View Cached Full Text
Cached at: 10/01/26, 04:20 AM
Paper page - Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Source: https://huggingface.co/papers/2609.38143
Abstract
Agentperformancedependsonbothreasoningabilityandtheenvironmentinwhichitacts.Westudytest-timeAI-for-AI,askinghowaBuildercanlearntoconstructbetterexecutionenvironmentsforaTargetwhilebothmodels’weightsremainfixed.TomaketheBuilder’sexperiencereusable,weintroduceMeta-Skill:principlesspecifyingwhensupportisneededandwhatresourcestoprovide.TheBuilderlearnstheseprinciplesfromTarget’sexecutionfeedbackonthedevelopmentset,thenusesthefrozenskillbanktoconstructharnessesforunseentasks.AcrossHarness-BenchandNewtonBench,full-bankmeta-skillsimprovemacro-averageperformanceby8.95percentagepointsoverno-skillconstruction,and12.02pointsoverdirectdeliveryofthesamebanktotheTarget.Theseresultshighlightthevalueoftranslatingexperienceintoexecutablesupport.Gainswhenthesamemodelservesbothrolesfurthersuggestapathtosystemlevelself-improvementthroughlearningtobuildbetterenvironments.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.38143
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38143 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.38143 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38143 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@dair_ai: Great paper from Meta on agent harness optimization. Meta-Harness-style search uses one development set and one proposa…
Meta (with Duke and UC Davis) proposes a Mixture of Self-Improving Branches framework for agent harness optimization, splitting the single-trajectory Meta-Harness search into adaptive branches with evolving development subsets and proposal policies, plus a router that selects the best branch per input, achieving up to +34.8% relative gains on Olympiad-level math, +11.6% on Terminal-Bench 2.0, and +3.8% on SWE-bench Lite.
@dair_ai: // Evolving Meta-Skill for Multi-Agent Systems // Can a multi-agent system get better at orchestration without touching…
Skill-MAS introduces a method for evolving meta-skills in multi-agent systems to improve orchestration without modifying model weights, achieving transferable performance gains across tasks and LLMs.
Harness Engineering for Self-Improvement (28 minute read)
This blog post by Lilian Weng explores the concept of recursive self-improvement in AI, focusing on how harness engineering—the system surrounding base models—enables automation and improvement of AI agents through workflow design and evaluation.
@omarsar0: // Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today …
This paper introduces Self-Harness, a new paradigm where LLM-based agents iteratively improve their own operating harness—prompts, tools, and control flow—without human engineers or stronger external agents, achieving significant performance gains across multiple models.
Self-Harness: Harnesses That Improve Themselves
Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.