Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Hugging Face Daily Papers Papers

Summary

This paper introduces Meta-Skill, a method letting a frozen Builder model learn reusable principles from Target execution feedback to construct better agent harnesses for unseen tasks, improving performance by 8.95 points on Harness-Bench and NewtonBench. It suggests a path toward system-level self-improvement by learning to build better environments rather than changing model weights.

Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:20 AM

Paper page - Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Source: https://huggingface.co/papers/2609.38143

Abstract

Agentperformancedependsonbothreasoningabilityandtheenvironmentinwhichitacts.Westudytest-timeAI-for-AI,askinghowaBuildercanlearntoconstructbetterexecutionenvironmentsforaTargetwhilebothmodels’weightsremainfixed.TomaketheBuilder’sexperiencereusable,weintroduceMeta-Skill:principlesspecifyingwhensupportisneededandwhatresourcestoprovide.TheBuilderlearnstheseprinciplesfromTarget’sexecutionfeedbackonthedevelopmentset,thenusesthefrozenskillbanktoconstructharnessesforunseentasks.AcrossHarness-BenchandNewtonBench,full-bankmeta-skillsimprovemacro-averageperformanceby8.95percentagepointsoverno-skillconstruction,and12.02pointsoverdirectdeliveryofthesamebanktotheTarget.Theseresultshighlightthevalueoftranslatingexperienceintoexecutablesupport.Gainswhenthesamemodelservesbothrolesfurthersuggestapathtosystemlevelself-improvementthroughlearningtobuildbetterenvironments.

View arXiv pageView PDFGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.38143

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38143 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38143 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38143 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

@dair_ai: Great paper from Meta on agent harness optimization. Meta-Harness-style search uses one development set and one proposa…

X AI KOLs Timeline

Meta (with Duke and UC Davis) proposes a Mixture of Self-Improving Branches framework for agent harness optimization, splitting the single-trajectory Meta-Harness search into adaptive branches with evolving development subsets and proposal policies, plus a router that selects the best branch per input, achieving up to +34.8% relative gains on Olympiad-level math, +11.6% on Terminal-Bench 2.0, and +3.8% on SWE-bench Lite.

Harness Engineering for Self-Improvement (28 minute read)

TLDR AI

This blog post by Lilian Weng explores the concept of recursive self-improvement in AI, focusing on how harness engineering—the system surrounding base models—enables automation and improvement of AI agents through workflow design and evaluation.

Self-Harness: Harnesses That Improve Themselves

Hacker News Top

Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.