RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Hugging Face Daily Papers Papers

Summary

This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.

An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
Original Article
View Cached Full Text

Cached at: 09/22/26, 03:24 AM

Paper page - RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Source: https://huggingface.co/papers/2609.24972 Published on Sep 21

#2 Paper of the day Authors:

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

AnLLMagent’scapabilityislargelymagnifiedbyitsharness,namelytheprompts,controlflow,tooling,memory,andcontextmanagementsurroundingthefrozenbackbonemodel.Recentmethodsincreasinglyautomatethisprocessbyiterativelyproposingandselectingcomponent-wiseeditsofanagentharness,practicallyestablishingaformofrecursiveself-improvement(RSI)attheagent-systemlevel.However,suchrecursiveevolutionmayoverfitbymemorizingthetrainingtasks,showinglargein-distributiongainsthatshrinkorevenvanishonout-of-distributionbenchmarks.WeintroduceRegularizedRecursiveSelf-ImprovementofAgentHarnesses(RRSI),whichincorporatestheprinciplesofregularizationsintoharnessself-improvementbyconstrainingtheevolutioncandidateproposalandselection.Theproposeroperateswithatemporallyannealedbudget,limitinghowmanyeditsacandidatecanbundle,anditencouragesunexploredtrajectoriesbasedonevolutionhistory.Theselectorisequippedwithacriticandapruner:thecriticscreensbenchmark-specificproposals,whilethepruner,removeschangesthataretoosmall,tooexpensive,ornolongeruseful.Togethertheseconstraintsfavorreusableagentmechanismsoverbenchmark-specificonesorevennoises.Acrosseightbenchmarksspanningcoding,agenticworkspaceandengineeringdesigntasks,RRSIgainsupto14.1pointsonthesplititevolvesagainstandupto4.7pointsonthefiveout-of-distributionbenchmarks,whileproducingaharnessthatrunson30%fewerpolicytokensthantheunregularizedevolution.Codeisavailableathttps://github.com/google-research/rrsiandprojectpageishttps://regularized-rsi.com/.

View arXiv pageView PDFProject pageGitHub2Add to collection

Get this paper in your agent:

hf papers read 2609\.24972

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.24972 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.24972 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.24972 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Recursive Harness Self-Improvement

arXiv cs.AI

Introduces Recursive Harness Self-Improvement (RHI), a method that iteratively refines prompt-level harness specifications for AI agents using pairwise feedback, improving performance and reducing inference cost by up to 60% on diverse machine learning research tasks.

Can AI Improve Itself? RSI Might Be the Answer [R]

Reddit r/MachineLearning

Introduces HarnessOpt-Bench to measure recursive self-improvement in AI, evaluating 5 frontier models on 4 tasks and finding that model choice has a greater impact than coding harness choice.