SkillEvoReg: 正则化智能体技能演化以防止过拟合
摘要
SkillEvoReg引入了一个正则化框架,以防止语言模型代理的技能演化中的过拟合,结合了dropout、局部正则化和因果验证,以在控制技能增长的同时保持性能。
查看缓存全文
缓存时间: 2026/09/28 09:51
# SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting
Source: [https://arxiv.org/html/2609.30861](https://arxiv.org/html/2609.30861)
Guanyu Nie1Fangzhou Zhu1Shixiong Kai1Xiongwei Han1Tao Zhong1Mingxuan Yuan11Huawei Noah’s Ark Labnieguanyu@huawei\.com
###### Abstract
Language\-model agents increasingly improve by converting execution experience into reusable external skills\. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task\-specific instructions, while new updates can disrupt behavior that previously worked\. We study this problem as*skill\-evolution overfitting*and introduceSkillEvoReg, a general regularization framework for skill evolution inspired by anti\-overfitting techniques in neural\-network training\.SkillEvoRegcombines training\-time skill dropout, which perturbs update generation, and complexity\-aware local regularization, which controls unnecessary structural growth, with causal counterexample validation \(CCV\), which provides targeted behavioral validation of candidate\-specific regressions\. We instantiate the framework across heterogeneous skill\-evolution systems while retaining each system’s native skill evolver and task evaluator\. Across SkillOpt, SkillEvolBench, and ContinualSkillBench,SkillEvoRegconsistently controls skill\-state growth while preserving competitive downstream capability, improves several transfer and later\-stage evolution outcomes, and identifies update\-level regressions that structural metrics alone cannot reveal\. These results suggest that explicit regularization is a useful complement to increasingly capable skill updaters\.
## 1Introduction
Large language models \(LLMs\) are increasingly deployed as agents that interact with environments, invoke tools, and improve from experience through reusable external state\([Shinn et al\., 2023](https://arxiv.org/html/2609.30861#bib.bib2);[Zhao et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib3);[Wang et al\., 2024a](https://arxiv.org/html/2609.30861#bib.bib4);[Zhang et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib5);[Wang et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib6);[Xu et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib7)\)\. In particular,*agent skills*—reusable instructions, procedures, code, or structured resources that guide future executions—have emerged as a practical substrate for continual self\-improvement\. Recent systems therefore move beyond one\-shot skill generation toward repeated*skill evolution*, treating external skill state as an object that can be revised from execution experience\([Yang et al\., 2026b](https://arxiv.org/html/2609.30861#bib.bib8);[Ni et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib9);[Zhang et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib10);[Lei et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib11);[Guan et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib12)\)\.
Repeated skill optimization, however, creates a generalization problem of its own\. Each update is inferred from finite and often narrow recent experience, yet the resulting content persists and influences future behavior\. Over time, this can produce*structural accumulation*of redundant or narrow rules,*semantic specialization*to incidental properties of recent tasks, or*update\-induced behavioral regression*that disrupts behavior already supported by the previous skill\. We refer to these phenomena collectively as*skill\-evolution overfitting*\. Importantly, they need not coincide: substantial structural growth can precede measurable performance loss, while a concise update can introduce a narrow regression without materially increasing skill size\.
Recent work has begun to expose or mitigate individual aspects of this problem through consolidation, replay\-based validation, and adaptive update control\([Ni et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib9);[Guan et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib12);[Yang et al\., 2026a](https://arxiv.org/html/2609.30861#bib.bib13);[He and Yang, 2026](https://arxiv.org/html/2609.30861#bib.bib14);[Li et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib15)\)\. We ask a broader question:*can general anti\-overfitting principles be instantiated across otherwise heterogeneous skill\-evolution procedures?*
This question has a close analogue in neural\-network training, where optimization is complemented by mechanisms that perturb learning, control capacity, and test generalization beyond the examples directly optimized\. Dropout reduces co\-adaptation\([Srivastava et al\., 2014](https://arxiv.org/html/2609.30861#bib.bib23)\), complexity penalties discourage unnecessary capacity\([Krogh and Hertz, 1991](https://arxiv.org/html/2609.30861#bib.bib24)\), and augmentation or adversarial training exposes models to informative input variations\([Goodfellow et al\., 2015](https://arxiv.org/html/2609.30861#bib.bib25);[Zhang et al\., 2018](https://arxiv.org/html/2609.30861#bib.bib26)\)\. Agent skills are discrete, language\-generated artifacts rather than differentiable parameter vectors, so these techniques cannot be transferred mechanically\. Their underlying regularization principles, however, provide a useful lens for controlling repeated skill updates\.
We introduceSkillEvoReg, a general regularization framework for skill evolution\. Rather than prescribing a new updater,SkillEvoRegregularizes both how candidate updates are generated and what structure becomes persistent\.*Training\-time skill dropout*perturbs the skill context used to generate updates, while*complexity\-aware local regularization*controls unnecessary structural growth around candidate changes\. Complementing these trajectory\-shaping mechanisms,*causal counterexample validation*\(CCV\) uses candidate\-conditioned paired validation to identify update\-induced behavioral regressions that structural signals alone cannot reveal\. The framework retains each system’s native skill evolver and evaluator while adapting these regularization principles to its update semantics\.
We evaluateSkillEvoRegacross SkillOpt, SkillEvolBench, and ContinualSkillBench, covering single\-document optimization, later\-stage multi\-skill evolution, and continual library growth\. Across these settings,SkillEvoRegproduces more compact skill states while preserving competitive downstream capability, improves several transfer and later\-stage outcomes, and identifies behavioral regressions that structural metrics alone cannot reveal\. Component analyses further support the complementary roles of generation\-time perturbation, structural regularization, and behavioral validation\.
Our contributions are:
- •We introduceSkillEvoReg, a general regularization framework for repeated skill evolution that targets overfitting in the evolution process while retaining the native updater\.SkillEvoRegcombines training\-time skill dropout and complexity\-aware local regularization with CCV, which provides targeted behavioral validation for candidate\-specific regressions\.
- •We instantiateSkillEvoRegacross SkillOpt, SkillEvolBench, and ContinualSkillBench, showing that the same regularization principles can be adapted to heterogeneous update semantics while controlling skill\-state growth, preserving competitive downstream capability, and improving several transfer and later\-stage evolution outcomes\.
## 2Related Work
##### External agent state and skill evolution\.
Language\-model agents increasingly improve through reusable external state, including reflections, memories, workflows, and executable skills\([Shinn et al\., 2023](https://arxiv.org/html/2609.30861#bib.bib2);[Zhao et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib3);[Wang et al\., 2024a](https://arxiv.org/html/2609.30861#bib.bib4);[Wang et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib6);[Xu et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib7)\)\. Recent systems go further by treating procedural skill state as an object of iterative optimization or continual maintenance\([Yang et al\., 2026b](https://arxiv.org/html/2609.30861#bib.bib8);[Ni et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib9);[Zhang et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib10);[Lei et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib11);[Guan et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib12)\)\. Related prompt\- and program\-optimization methods similarly search over external language\-model state\([Pryzant et al\., 2023](https://arxiv.org/html/2609.30861#bib.bib16);[Yang et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib17);[Guo et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib18);[Fernando et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib20);[Wang et al\., 2024b](https://arxiv.org/html/2609.30861#bib.bib19);[Khattab et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib21);[Yuksekgonul et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib22)\)\. These lines of work primarily study how to acquire or improve external state; we study how repeated revision of that state should be regularized\.
##### Reliable skill evolution\.
Several concurrent methods introduce safeguards against brittle skill updates\. GSE combines cross\-task consolidation with replay verification\([Yang et al\., 2026a](https://arxiv.org/html/2609.30861#bib.bib13)\); SkillCommit validates broader abstractions before commitment\([He and Yang, 2026](https://arxiv.org/html/2609.30861#bib.bib14)\); SkillAdam uses optimization history and adaptive edit budgets\([Li et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib15)\); and Trace2Skill consolidates trajectory\-local lessons before skill formation\([Ni et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib9)\)\. These approaches improve reliability within particular evolution procedures\.SkillEvoReginstead factors anti\-overfitting into a complementary layer that can be instantiated around heterogeneous native updaters, targeting structural accumulation, specialization, and update\-induced regression as distinct failure modes\.
##### Regularization and behavioral validation\.
Our design draws on the complementary roles of dropout, capacity regularization, and robust validation in conventional learning\([Srivastava et al\., 2014](https://arxiv.org/html/2609.30861#bib.bib23);[Krogh and Hertz, 1991](https://arxiv.org/html/2609.30861#bib.bib24);[Goodfellow et al\., 2015](https://arxiv.org/html/2609.30861#bib.bib25);[Zhang et al\., 2018](https://arxiv.org/html/2609.30861#bib.bib26)\)\. CCV is also related to counterexample\-guided repair, where targeted failures expose weaknesses in candidate behavior\([Orvalho et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib28)\)\. Rather than reproducing these techniques mechanically,SkillEvoRegadapts their underlying roles to discrete, language\-generated skill transitions\. Appendix[G](https://arxiv.org/html/2609.30861#A7)discusses these connections in greater detail\.
## 3Skill Evolution as a Regularized Learning Problem
We consider a generic skill\-evolution system with skill stateStS\_\{t\}, ranging from a single instruction document to a skill library or structured directory\. Given an execution trajectoryτt\\tau\_\{t\}, the native updaterUUproposes
S~t\+1=U\(St,τt\)\.\\widetilde\{S\}\_\{t\+1\}=U\(S\_\{t\},\\tau\_\{t\}\)\.\(1\)The host framework then applies its own validation and commit semantics; we assume no particular skill representation, updater, or commit rule\.
Repeated application of this update creates a learning problem beyond immediate task performance: each update is inferred from finite recent trajectories yet persists to influence later tasks, allowing locally useful changes to accumulate, specialize, or interfere with previously supported behavior over time\.
We use*skill\-evolution overfitting*to describe this mismatch between narrow update evidence and persistent downstream influence\. Three manifestations are particularly relevant\.*Structural accumulation*adds rules, examples, or narrow exceptions without commensurate transferable benefit;*semantic specialization*promotes properties useful for recent trajectories into broadly applied instructions; and*update\-induced regression*occurs when a new edit disrupts behavior already supported by the previous skill\. These manifestations need not coincide: substantial structure may accumulate before aggregate performance deteriorates, while a concise edit may create a narrow regression without materially increasing skill complexity\. Consequently, neither performance degradation nor skill complexity alone is sufficient to characterize skill\-evolution overfitting\.
Figure 1:A motivating longitudinal study of native skill evolution on SpreadsheetBench\.\(a\)Performance across evolution epochs\.\(b\)The IID–transfer gap narrows and then re\-expands after the transfer peak\.\(c\)Skill structure continues to accumulate\.### 3\.1A Motivating Longitudinal Study
We trace a SkillOpt trajectory on SpreadsheetBench for eight epochs, evaluating each checkpoint on an IID test set and a structural transfer\-stress set with matched task semantics but greater workbook\-level variation\. Full split statistics and protocol details are provided in Appendix[A\.1](https://arxiv.org/html/2609.30861#A1.SS1)\.
Figure[1](https://arxiv.org/html/2609.30861#S3.F1)shows a clear post\-peak transfer pattern\. Transfer performance initially improves and peaks at epoch 4 \(46\.67%46\.67\\%\), with the IID–transfer gap shrinking to7\.507\.50points\. Continued evolution then favors familiar structure: by epoch 8, IID accuracy reaches61\.67%61\.67\\%while transfer falls to37\.50%37\.50\\%, widening the gap to24\.1724\.17points\.
Meanwhile, the skill continues to accumulate more specific structure after transfer has peaked, with lexical length reaching23\.6×23\.6\\timesits initial value by epoch 8\. Later updates thus keep reshaping the skill beyond its best transfer point, motivating explicit regularization of the evolution process\.
## 4SkillEvoReg
Figure 2:Overview ofSkillEvoReg\.The three regularizers act on update generation, skill complexity, and behavioral validation, respectively\. Colors distinguish thenative skill evolver,SkillEvoRegregularization modulesandskill\-states\.### 4\.1Regularizing Skill Evolution
The pilot suggests that the challenge is not simply producing useful skill updates, but controlling how repeated updates shape the skill over an evolution trajectory\. We introduceSkillEvoReg, a general framework that regularizes skill evolution while retaining the native updater\. It acts at three complementary points: update generation, skill complexity, and behavioral validation\.
The design draws inspiration from anti\-overfitting mechanisms in neural\-network training: dropout perturbs the information during learning, capacity regularization discourages unnecessary complexity, and validation under informative variations exposes brittle behavior\. Because agent skills are discrete language\-generated artifacts whose updates persist into future executions, these principles require operational adaptations\.SkillEvoReginstantiates the first two roles as training\-time skill dropout and complexity\-aware local regularization, which directly shape the evolution trajectory, and complements them with CCV, which validates the candidate behavior before persistence\.
LetStS\_\{t\}denote the incumbent skill state at update steptt\. We use distinct symbols at regularization\-stage boundaries:St\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}is the candidate after dropout reconciliation,St\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}is the candidate after complexity regularization, andSt\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}is the candidate after CCV\. The native framework then applies its own acceptance or commit semantics to obtain the next persistent stateSt\+1S\_\{t\+1\}\. Figure[2](https://arxiv.org/html/2609.30861#S4.F2)shows where the three regularizers intervene\. Across evolution systems, the main adaptation lies at the native update boundary, where candidate changes may appear as document edits, library patches, or explicit create/modify operations\. Full pseudocode is provided in Appendix[D](https://arxiv.org/html/2609.30861#A4)\.
### 4\.2Training\-Time Skill Dropout
Skill evolution repeatedly generates new updates while conditioning on the complete current skill state\. This can encourage successive updates to depend too strongly on the precise presence or wording of existing instructions\. We therefore perturb the skill context used during update generation\.
We represent the current skill state as locally meaningful*skill atoms*,
𝒜\(St\)=\{a1,…,am\},\\mathcal\{A\}\(S\_\{t\}\)=\\\{a\_\{1\},\\ldots,a\_\{m\}\\\},\(2\)where an atom is a self\-contained instruction, list item, procedural block, or another semantically coherent unit\. At update steptt, we sampleMt⊆𝒜\(St\)M\_\{t\}\\subseteq\\mathcal\{A\}\(S\_\{t\}\)and construct a temporary dropout view
Stdrop=Mask\(St,Mt\)\.S\_\{t\}^\{\\mathrm\{drop\}\}=\\operatorname\{Mask\}\(S\_\{t\},M\_\{t\}\)\.\(3\)The training rollout and native update are generated underStdropS\_\{t\}^\{\\mathrm\{drop\}\}\. If the native updater proposesS~t\+1drop\\widetilde\{S\}\_\{t\+1\}^\{\\mathrm\{drop\}\}, we extract only the change learned under that perturbed context and reconcile it with the complete incumbent:
Δtdrop=Diff\(Stdrop,S~t\+1drop\),St\+1drop=Reconcile\(St,Δtdrop\)\.\\Delta\_\{t\}^\{\\mathrm\{drop\}\}=\\operatorname\{Diff\}\\left\(S\_\{t\}^\{\\mathrm\{drop\}\},\\widetilde\{S\}\_\{t\+1\}^\{\\mathrm\{drop\}\}\\right\),\\qquad S\_\{t\+1\}^\{\\mathrm\{drop\}\}=\\operatorname\{Reconcile\}\\left\(S\_\{t\},\\Delta\_\{t\}^\{\\mathrm\{drop\}\}\\right\)\.\(4\)Thus, masking changes only the context in which an update is learned; masked atoms are restored before later regularization and are never treated as candidate deletions\. Frozen evaluation always uses the complete skill state\. Algorithm[1](https://arxiv.org/html/2609.30861#alg1)in Appendix[D](https://arxiv.org/html/2609.30861#A4)gives the complete procedure\.
##### Remark \(Relation to neural\-network dropout\)\.
The analogy lies in when the perturbation acts\. Neural dropout perturbs the representation through which a parameter update is learned; skill dropout perturbs the external skill context through which a skill update is generated\. In both cases, the perturbation affects learning but is absent from the persistent state used at inference\. Unlike inverted neural dropout, no activation\-rescaling analogue is needed because skill atoms are discrete contextual units rather than additive activations\.
### 4\.3Complexity\-Aware Local Regularization
Skill dropout changes how an update is generated but does not directly control how much structure becomes persistent\. Skill state has no differentiable parameter norm, and additional content is not inherently harmful because genuinely new capability may require additional structure\. We therefore use observable complexity to determine*when structural reconsideration is needed*, while separately constraining*what content may be changed*, rather than minimizing skill size itself\.
Given the incumbentStS\_\{t\}and reconciled candidateSt\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}, we define the change examined by this stage as
Δtcomp=Diff\(St,St\+1drop\)\.\\Delta\_\{t\}^\{\\mathrm\{comp\}\}=\\operatorname\{Diff\}\\left\(S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{drop\}\}\\right\)\.\(5\)Letϕ\(Z\)\\phi\(Z\)summarize observable structural properties of a skill state or update, such as lexical size, atom organization, structural fragmentation, and executable\-code structure\. The regularizer computes
ct=fcomp\(ϕ\(St\),ϕ\(St\+1drop\),ϕ\(Δtcomp\)\),c\_\{t\}=f\_\{\\mathrm\{comp\}\}\\left\(\\phi\(S\_\{t\}\),\\phi\(S\_\{t\+1\}^\{\\mathrm\{drop\}\}\),\\phi\(\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\)\\right\),\(6\)wherefcompf\_\{\\mathrm\{comp\}\}maps these measurements to the framework\-specific signal used to trigger or score regularization\. Exact feature definitions and realizations offcompf\_\{\\mathrm\{comp\}\}are given in Appendix[D\.3\.1](https://arxiv.org/html/2609.30861#A4.SS3.SSS1)\.
When structural reconsideration is triggered, we restrict editing to atoms introduced or modified by the candidate and a bounded set of related incumbent atoms:
𝒟t=SegmentAtoms\(Δtcomp\),Ωt=𝒟t∪Related\(St,𝒟t\)\.\\mathcal\{D\}\_\{t\}=\\operatorname\{SegmentAtoms\}\(\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\),\\qquad\\Omega\_\{t\}=\\mathcal\{D\}\_\{t\}\\cup\\operatorname\{Related\}\(S\_\{t\},\\mathcal\{D\}\_\{t\}\)\.\(7\)This locality constraint prevents a structurally large candidate from triggering an unrestricted rewrite of the complete skill state\. A diagnostic SpreadsheetBench case shows that prompt\-level locality alone can still yield broad, destructive rewrites, motivating the explicit candidate\-local edit scope used here; Appendix[E\.1](https://arxiv.org/html/2609.30861#A5.SS1)provides the detailed analysis\.
Inspired by operation\-based memory reconciliation in Mem0\([Chhikara et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib1)\), we adapt discrete edit decisions to post\-update skill reconciliation\. For each candidate atoma∈𝒟ta\\in\\mathcal\{D\}\_\{t\}, the local editor selects
oa∈\{NoOp,Merge,Rewrite,Delete\}\.o\_\{a\}\\in\\\{\\textsc\{NoOp\},\\textsc\{Merge\},\\textsc\{Rewrite\},\\textsc\{Delete\}\\\}\.\(8\)NoOppreserves useful new structure, while the other operations consolidate, revise, or remove candidate content withinΩt\\Omega\_\{t\}\. Similarity is used only to define what should be considered jointly, not as evidence that two atoms should be merged\. We denote the resulting candidate bySt\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}; if no intervention is triggered,St\+1comp=St\+1dropS\_\{t\+1\}^\{\\mathrm\{comp\}\}=S\_\{t\+1\}^\{\\mathrm\{drop\}\}\. Algorithm[2](https://arxiv.org/html/2609.30861#alg2)gives the complete procedure\.
##### Remark \(Relation to capacity regularization\)\.
The analogy to neural\-network capacity regularization is functional rather than literal\. Parametric penalties discourage unnecessary capacity through continuous norms;SkillEvoReginstead uses observable structural complexity to trigger local reconsideration of a discrete candidate update\. In both cases the goal is to limit unnecessary capacity accumulated from finite experience, whileNoOpexplicitly allows additional skill structure when it carries distinct useful behavior\. Complexity reduction is therefore a means of regularization, not the objective itself\.
### 4\.4Causal Counterexample Validation
Structural regularization cannot detect every harmful update\. A concise candidate may add little complexity while introducing an assumption or decision boundary that breaks behavior already supported by the incumbent\. CCV therefore asks whether the structurally regularized candidateSt\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}deteriorates on a targeted variation that the incumbentStS\_\{t\}can still handle\.
The key comparison is
original sourceetattacked casextccvincumbentSt✓✓candidateSt\+1comp✓×⟹CCV\-confirmed regression\.\\begin\{array\}\[\]\{c\|cc\}&\\text\{original source \}e\_\{t\}&\\text\{attacked case \}x\_\{t\}^\{\\mathrm\{ccv\}\}\\\\ \\hline\\cr\\text\{incumbent \}S\_\{t\}&\\checkmark&\\checkmark\\\\ \\text\{candidate \}S\_\{t\+1\}^\{\\mathrm\{comp\}\}&\\checkmark&\\times\\end\{array\}\\quad\\Longrightarrow\\quad\\text\{CCV\-confirmed regression\}\.\(9\)Here✓\\checkmarkand×\\timesdenote evaluator\-dependent qualification and degradation rather than necessarily binary success and failure\.
CCV constructs this comparison in three steps\. First, it selects an eligible source from the available training historyℋt\\mathcal\{H\}\_\{t\},
et=SelectSource\(ℋt,St,St\+1comp\)\.e\_\{t\}=\\operatorname\{SelectSource\}\\left\(\\mathcal\{H\}\_\{t\},S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{comp\}\}\\right\)\.\(10\)The incumbent must already supportete\_\{t\}, and the candidate must remain qualified on the same unmodified source\. This clean\-pair requirement rules out candidates that have already broken the original task\.
Second, because CCV acts after complexity regularization, it defines its own stage\-local candidate change and uses it to construct the attack:
Δtccv=Diff\(St,St\+1comp\),at=GenerateAttack\(et,Δtccv\),xtccv=ApplyAttack\(et,at\)\.\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}=\\operatorname\{Diff\}\(S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{comp\}\}\),\\quad a\_\{t\}=\\operatorname\{GenerateAttack\}\(e\_\{t\},\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}\),\\quad x\_\{t\}^\{\\mathrm\{ccv\}\}=\\operatorname\{ApplyAttack\}\(e\_\{t\},a\_\{t\}\)\.\(11\)Hereata\_\{t\}specifies a targeted perturbation of an assumption, dependency, or decision boundary exposed byΔtccv\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}, andxtccvx\_\{t\}^\{\\mathrm\{ccv\}\}is the resulting attacked case\.
Third, the incumbent and candidate are evaluated independently on that same case:
rold=\\displaystyle r\_\{\\mathrm\{old\}\}=Evaluate\(St,xtccv\),rnew=Evaluate\(St\+1comp,xtccv\),\\displaystyle\\operatorname\{Evaluate\}\(S\_\{t\},x\_\{t\}^\{\\mathrm\{ccv\}\}\),\\qquad r\_\{\\mathrm\{new\}\}=\\operatorname\{Evaluate\}\(S\_\{t\+1\}^\{\\mathrm\{comp\}\},x\_\{t\}^\{\\mathrm\{ccv\}\}\),\(12\)gt=RegressionCriterion\(rold,rnew,θccv\)\.\\displaystyle g\_\{t\}=\\operatorname\{RegressionCriterion\}\(r\_\{\\mathrm\{old\}\},r\_\{\\mathrm\{new\}\};\\theta\_\{\\mathrm\{ccv\}\}\)\.whereθccv\\theta\_\{\\mathrm\{ccv\}\}denotes evaluator\-specific qualification and degradation criteria\. A regression is confirmed only when the incumbent remains sufficiently successful on the attacked case while the candidate deteriorates\. Whengt=1g\_\{t\}=1, CCV performs at most one scope\-constrained repair over the implicated candidate region; otherwise the candidate is unchanged\. We denote the resulting CCV\-processed candidate bySt\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}\. Exact qualification rules, attack construction, and repair procedures are given in Appendix[D\.4](https://arxiv.org/html/2609.30861#A4.SS4)\.
##### Remark \(Relation to data augmentation\)\.
CCV adapts the augmentation principle to a skill transition: the candidate update itself reveals where generalization may have narrowed, and the attack perturbs the source along that direction\. The resulting case is therefore a targeted stress test of the behavioral boundary moved by the update, used primarily for paired validation and, when needed, scope\-constrained repair rather than simply being added back to training\.
##### Remark \(Why causal?\)\.
The four\-way comparison in Equation[9](https://arxiv.org/html/2609.30861#S4.E9)provides the basis for the causal terminology\. Qualification of both skills on the original source controls for ordinary source\-task failure, while incumbent success on the attacked case controls for an attack that is simply too difficult\. Candidate\-specific degradation on that same case therefore associates the regression with the candidate transition\. We use*CCV\-confirmed regression*in this operational sense, not as a claim of formal causal identification under stochastic rollouts or multi\-edit candidates\.
### 4\.5Adapting to Different Update Semantics
The regularizers retain the same roles across evolution systems; the main adaptation lies at the native update boundary\. SkillOpt exposes ranked edits to a single skill document, SkillEvolBench localized patches over a multi\-skill library, and ContinualSkillBench explicitCREATE/MODIFYoperations\. These semantics determine how stage\-specific candidate changes are exposed, how dropout updates are reconciled with the incumbent, how editable scopes are defined, and howSt\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}is returned to the native framework\. Appendix[D](https://arxiv.org/html/2609.30861#A4)details the framework\-specific integrations\.
## 5Experiments
We evaluate whether the same regularization principles improve three distinct forms of skill evolution: repeated optimization of a single skill, later\-stage revision of an established multi\-skill library, and native continual skill\-library growth\. We then analyze the roles of the individual regularization components\.
### 5\.1Experimental Setup
All compared methods use the same proprietary instruction\-following LLM under a fixed inference configuration and each benchmark’s native execution harness, task order, and evaluator\.SkillEvoRegis inserted at the candidate\-update boundary while retaining the native skill evolver and evaluation protocol\. We report downstream capability together with skill\-state size and structural complexityCC; where checkpoint selection is available, we distinguish the validation\-selected*delivery*checkpoint from the terminal*final*checkpoint\. Full benchmark protocols and measurement conventions are given in Appendix[A\.2](https://arxiv.org/html/2609.30861#A1.SS2)–[A\.3](https://arxiv.org/html/2609.30861#A1.SS3)\. Frozen\-evaluation inference usage is analyzed separately in Appendix[C\.4](https://arxiv.org/html/2609.30861#A3.SS4)as a deployment\-side by\-product of compact persistent skill states rather than an optimization objective ofSkillEvoReg\.
### 5\.2SkillOpt: Episodic Single\-Skill Optimization
SkillOpt\([Yang et al\., 2026b](https://arxiv.org/html/2609.30861#bib.bib8)\)repeatedly optimizes a single instruction document\. We evaluate SpreadsheetBench, SearchQA, and LiveMathBench over eight evolution epochs, retaining SkillOpt’s native updater and checkpointing procedure while regularizing candidate transitions\. Full protocol and checkpoint details are given in Appendix[A\.2](https://arxiv.org/html/2609.30861#A1.SS2)\.
Table 1:SkillOpt results\. Delivery denotes the validation\-selected checkpoint and Final the epoch\-8 checkpoint;Δ\\DeltaTest is Final minus Delivery\. Skill tokens and complexity describe the final skill state\.##### Capability and evolution quality\.
At the delivery checkpoint,SkillEvoRegmatches or improves all three tasks while producing substantially smaller final skill states\. SpreadsheetBench continues to improve through the terminal checkpoint, while SearchQA remains essentially stable\. LiveMath instead exposes a distribution\-sensitive trade\-off\. SkillOpt repeatedly reinforces a meta\-option prior favoring the answer that “a stronger result can be proved,” whereasSkillEvoRegconsolidates repeated reinforcement of this pattern\. Of the 217 test questions, 92 use this meta\-option as the correct answer, and this subset accounts for85\.7%85\.7\\%ofSkillEvoReg’s late\-stage decline\. The performance gap is therefore strongly concentrated in a specific test\-distribution pattern rather than distributed uniformly across questions\. Appendix[C\.1](https://arxiv.org/html/2609.30861#A3.SS1)provides the item\-level analysis\.
### 5\.3SkillEvolBench: Regularizing Later\-Stage Skill Evolution
SkillEvolBench\([Lei et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib11)\)evaluates transfer under context shift, adversarial variation, and skill composition\. Its native protocol uses a single acquisition pass\. We add a controlled second pass to emulate later\-stage skill evolution, where an already formed library continues to absorb repeated experience and may begin to over\-specialize\. Starting from the same native one\-pass library, we compare freezing it \(1\-pass\), continuing with a second unregularized pass \(2\-pass\), and regularizing that same second pass withSkillEvoReg\. The two second\-pass conditions therefore receive identical additional experience, while deployment tasks remain unseen throughout acquisition\. Details are given in Appendix[B](https://arxiv.org/html/2609.30861#A2)and Appendix[C\.2](https://arxiv.org/html/2609.30861#A3.SS2)\.
##### Later\-stage transfer and controlled growth\.
A second unregularized pass illustrates the risk of repeated skill evolution: the library continues to grow, while deployment success decreases from31\.67%31\.67\\%to29\.44%29\.44\\%\. ApplyingSkillEvoRegto the same additional experience instead reaches33\.33%33\.33\\%deployment success while substantially controlling persistent\-state growth\. Relative to the one\-pass anchor,SkillEvoRegreduces second\-pass token growth by34\.0%34\.0\\%and complexity growth by32\.7%32\.7\\%compared with unregularized two\-pass evolution\.
Table 2:SkillEvolBench results\. T1–T3 are replay tasks and T4–T6 are unseen deployment tasks\.##### Transfer behavior\.
The improvement is not uniform across every transfer mode\.SkillEvoReggives the strongest gain on composition \(26\.67%26\.67\\%versus18\.33%18\.33\\%after one pass and13\.33%13\.33\\%after two unregularized passes\) and improves adversarial transfer, while context\-shift success decreases\. Together with the controlled skill\-state growth, this suggests that regularization does not simply suppress second\-pass learning, but changes which parts of repeated experience are allowed to persist\. Continuous verifier scores and additional structural diagnostics are reported in Appendix[C\.2](https://arxiv.org/html/2609.30861#A3.SS2)\.
### 5\.4ContinualSkillBench: Continual Skill\-Library Evolution
ContinualSkillBench\([Guan et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib12)\)extends the setting further by allowing the evolving library to create new skills as well as modify existing ones through nativeCreate\-SkillandModify\-Skillactions\. This makes library proliferation itself part of the regularization problem\. Each condition learns sequentially on tasks 1–50, freezes the resulting library, and evaluates tasks 51–100 without further learning\.SkillEvoRegregularizes both CREATE and MODIFY outcomes before write\-back while retaining the benchmark’s native update semantics\. Complete protocol and online\-learning results are given in Appendix[C\.3](https://arxiv.org/html/2609.30861#A3.SS3)and Appendix[B](https://arxiv.org/html/2609.30861#A2)\.
Table 3:ContinualSkillBench frozen\-library results after learning tasks 1–50\. Scores average two evaluations of held\-out tasks 51–100; skill\-state statistics use the complete step\-50 library\.##### Held\-out capability and consolidation\.
SkillEvoRegmaintains comparable or better Overall held\-out performance across the five domains while consistently producing more compact persistent skill states\. Skill tokens, complexity, and skill count decrease in every domain\. The reduction is not simply uniform pruning: Mathematics primarily shortens existing skills, whereas the other domains show stronger suppression of library proliferation\. Together, these results show that regularization controls persistent\-state accumulation without broadly sacrificing held\-out capability\. Evaluator\-level breakdowns and sensitivity analyses are provided in Appendix[C\.3](https://arxiv.org/html/2609.30861#A3.SS3)\.
### 5\.5Component Analysis
We next examine how dropout and complexity regularization shape the evolution trajectory, and how CCV complements them with targeted behavioral validation\.
#### 5\.5\.1Dropout and Complexity Regularization Are Complementary
To isolate the roles of skill dropout and complexity\-aware regularization, we conduct a controlled SpreadsheetBench ablation with four variants: the native updater, complexity regularization only, skill dropout only, and the fullSkillEvoReg\. All variants use the same eight\-epoch evolution budget, training data, native updater, validation procedure, and frozen evaluation; only the active regularizers differ, and we compare final checkpoints to ensure equal evolution budgets\.
Table 4:SpreadsheetBench ablation at the final checkpoint\.The two regularizers exhibit complementary effects\. Complexity regularization substantially limits persistent\-state growth, while dropout alone attains higher test accuracy but allows a larger skill state\. Their combination yields both the smallest regularized skill and the strongest frozen score, consistent with dropout shaping how updates are generated and complexity regularization controlling what structure is allowed to persist\.CCVdoes not modify the persistent state on this particular trajectory, so we examine its behavioral role separately\.
#### 5\.5\.2CCVDetects Regressions Beyond Structural Signals
CCVprovides a complementary behavioral signal: even a concise update can disrupt behavior already supported by the incumbent without producing a distinctive structural signature\. Across the evaluated trajectories, we identify 12 distinct CCV regression signals and 9 state\-changing repairs\. A representative SkillEvolBench dependency\-resolution trace illustrates the mechanism: after the incumbent and candidate both qualify on the original source, a candidate\-conditioned package\-boundary attack yields reward1\.01\.0for the incumbent and0\.80\.8for the candidate\. CCV identifies candidate\-specific sensitivity to the package\-boundary condition and produces a scope\-constrained rewrite; a later replay recovers the affected process checks\. Appendix[E\.2](https://arxiv.org/html/2609.30861#A5.SS2)provides additional regression evidence and a detailed representative trace\.
## 6Conclusion
Repeated skill evolution creates a learning problem beyond the native updater\. We introducedSkillEvoReg, a general framework that translates complementary anti\-overfitting principles from neural\-network training to discrete skill evolution: training\-time skill dropout perturbs update generation, complexity\-aware local regularization controls persistent structural growth, and CCV provides targeted behavioral validation of candidate\-specific regressions\. Across SkillOpt, SkillEvolBench, and ContinualSkillBench,SkillEvoReglimits skill\-state growth while preserving downstream capability and improving several transfer and later\-stage outcomes\. More broadly, these results suggest that regularization principles developed for parametric learning provide a useful lens for controlling persistent, language\-generated skill updates, and support regularization as a first\-class component of skill evolution\.
## References
- P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready AI agents with scalable long\-term memory\.InProceedings of the 28th European Conference on Artificial Intelligence,Cited by:[§D\.3\.1](https://arxiv.org/html/2609.30861#A4.SS3.SSS1.Px4.p1.1),[§4\.3](https://arxiv.org/html/2609.30861#S4.SS3.p4.1)\.
- Fernandoet al\.\(2024\)C\. Fernando, D\. S\. Banarse, H\. Michalewski, S\. Osindero, and T\. RocktäschelPromptbreeder: self\-referential self\-improvement via prompt evolution\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 13481–13544\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Goodfellowet al\.\(2015\)I\. J\. Goodfellow, J\. Shlens, and C\. SzegedyExplaining and harnessing adversarial examples\.InInternational Conference on Learning Representations,Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p4.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px3.p1.1)\.
- Guanet al\.\(2026\)T\. Guan, Y\. Wang, H\. Yang, S\. Cao, S\. Liu, Y\. Hu, J\. Li, and M\. ZhangContinualSkillBench: can llm agents truly evolve their capabilities?\.External Links:2608\.03874,[Link](https://arxiv.org/abs/2608.03874)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p3.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1),[§5\.4](https://arxiv.org/html/2609.30861#S5.SS4.p1.1)\.
- Guoet al\.\(2024\)Q\. Guo, R\. Wang, J\. Guo, B\. Li, K\. Song, X\. Tan, G\. Liu, J\. Bian, and Y\. YangConnecting large language models with evolutionary algorithms yields powerful prompt optimizers\.InInternational Conference on Learning Representations,Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- He and Yang \(2026\)Y\. He and W\. YangSkillCommit: evolving agent skills through behaviorally validated scope expansion\.External Links:2608\.15165,[Link](https://arxiv.org/abs/2608.15165)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p3.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px2.p1.1)\.
- Khattabet al\.\(2024\)O\. Khattab, A\. Singhvi, P\. Maheshwari, Z\. Zhang, K\. Santhanam, S\. V\. A, S\. Haq, A\. Sharma, T\. T\. Joshi, H\. Moazam, H\. Miller, M\. Zaharia, and C\. PottsDSPy: compiling declarative language model calls into state\-of\-the\-art pipelines\.InInternational Conference on Learning Representations,Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Krogh and Hertz \(1991\)A\. Krogh and J\. A\. HertzA simple weight decay can improve generalization\.InAdvances in Neural Information Processing Systems,Vol\.4,pp\. 950–957\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p4.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px3.p1.1)\.
- Leiet al\.\(2026\)Y\. Lei, Z\. Wan, J\. Zhang, S\. Alam, Z\. Zhong, P\. Huang, X\. Wang, J\. Zhang, D\. Zhou, Y\. Hsieh, Z\. Dou, H\. Shen, Y\. Xu, D\. Dimitriadis, T\. Zhang, and M\. ZhangSkillEvolBench: benchmarking the evolution from episodic experience to procedural skills\.External Links:2605\.24117,[Link](https://arxiv.org/abs/2605.24117)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1),[§5\.3](https://arxiv.org/html/2609.30861#S5.SS3.p1.1)\.
- Liet al\.\(2026\)G\. Li, M\. Fan, Y\. Liu, S\. Zhang, J\. Fan, S\. Wang, J\. Hou, X\. Weng, H\. Tian, and Z\. LiSkillAdam: stable and efficient skill evolution for agents\.External Links:2609\.08944,[Link](https://arxiv.org/abs/2609.08944)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p3.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px2.p1.1)\.
- Niet al\.\(2026\)J\. Ni, Y\. Liu, X\. Liu, Y\. Sun, M\. Zhou, P\. Cheng, D\. Wang, E\. Zhao, X\. Jiang, and G\. JiangTrace2Skill: distill trajectory\-local lessons into transferable agent skills\.External Links:2603\.25158,[Link](https://arxiv.org/abs/2603.25158)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p3.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px2.p1.1)\.
- Orvalhoet al\.\(2025\)P\. Orvalho, M\. Janota, and V\. M\. ManquinhoCounterexample guided program repair using zero\-shot learning and MaxSAT\-based fault localization\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 649–657\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px3.p1.1)\.
- Prechelt \(1998\)L\. PrecheltEarly stopping—but when?\.InNeural Networks: Tricks of the Trade,G\. B\. Orr and K\. Müller \(Eds\.\),Lecture Notes in Computer Science, Vol\.1524,pp\. 55–69\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px4.p1.1)\.
- Pryzantet al\.\(2023\)R\. Pryzant, D\. Iter, J\. Li, Y\. Lee, C\. Zhu, and M\. ZengAutomatic prompt optimization with “gradient descent” and beam search\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 7957–7968\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Srivastavaet al\.\(2014\)N\. Srivastava, G\. Hinton, A\. Krizhevsky, I\. Sutskever, and R\. SalakhutdinovDropout: a simple way to prevent neural networks from overfitting\.Journal of Machine Learning Research15\(56\),pp\. 1929–1958\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p4.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2024a\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.Transactions on Machine Learning Research\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024b\)X\. Wang, C\. Li, Z\. Wang, F\. Bai, H\. Luo, J\. Zhang, N\. Jojic, E\. P\. Xing, and Z\. HuPromptAgent: strategic planning with language models enables expert\-level prompt optimization\.InInternational Conference on Learning Representations,Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2025\)Z\. Z\. Wang, J\. Mao, D\. Fried, and G\. NeubigAgent workflow memory\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 63897–63911\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Xuet al\.\(2025\)W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. ZhangA\-Mem: agentic memory for LLM agents\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2026a\)C\. Yang, J\. Tian, Z\. Wang, X\. Liu, M\. Ye, and J\. ChenLearning globally reusable skills for coding agents\.External Links:2608\.06153,[Link](https://arxiv.org/abs/2608.06153)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p3.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2024\)C\. Yang, X\. Wang, Y\. Lu, H\. Liu, Q\. V\. Le, D\. Zhou, and X\. ChenLarge language models as optimizers\.InInternational Conference on Learning Representations,Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Yanget al\.\(2026b\)Y\. Yang, Z\. Gong, W\. Huang, Q\. Yang, Z\. Zhou, Z\. Huang, Y\. Li, X\. Gao, Q\. Dai, B\. Liu, K\. Qiu, Y\. Yang, D\. Chen, X\. Yang, and C\. LuoSkillOpt: executive strategy for self\-evolving agent skills\.External Links:2605\.23904,[Link](https://arxiv.org/abs/2605.23904)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1),[§5\.2](https://arxiv.org/html/2609.30861#S5.SS2.p1.1)\.
- Yuksekgonulet al\.\(2025\)M\. Yuksekgonul, F\. Bianchi, J\. Boen, S\. Liu, P\. Lu, Z\. Huang, C\. Guestrin, and J\. ZouOptimizing generative AI by backpropagating language model feedback\.Nature639,pp\. 609–616\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2026\)H\. Zhang, S\. Fan, H\. P\. Zou, Y\. Chen, Z\. Wang, J\. Zhou, C\. Li, W\. Huang, Y\. Yao, K\. Zheng, Xue, Liu, X\. Li, and P\. S\. YuCoEvoSkills: self\-evolving agent skills via co\-evolutionary verification\.External Links:2604\.01687,[Link](https://arxiv.org/abs/2604.01687)Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2018\)H\. Zhang, M\. Cissé, Y\. N\. Dauphin, and D\. Lopez\-PazMixup: beyond empirical risk minimization\.InInternational Conference on Learning Representations,Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p4.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px3.p1.1)\.
- Zhanget al\.\(2024\)W\. Zhang, K\. Tang, H\. Wu, M\. Wang, Y\. Shen, G\. Hou, Z\. Tan, P\. Li, Y\. Zhuang, and W\. LuAgent\-pro: learning to evolve via policy\-level reflection and optimization\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 5348–5375\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1)\.
- Zhaoet al\.\(2024\)A\. Zhao, D\. Huang, Q\. Xu, M\. Lin, Y\. Liu, and G\. HuangExpeL: LLM agents are experiential learners\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 19632–19642\.Cited by:[Appendix G](https://arxiv.org/html/2609.30861#A7.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.30861#S1.p1.1),[§2](https://arxiv.org/html/2609.30861#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix AAdditional Experimental Details
### A\.1Motivating SpreadsheetBench Pilot
##### Purpose\.
The study in Section[3\.1](https://arxiv.org/html/2609.30861#S3.SS1)is a longitudinal stress test of native skill evolution rather than a confirmatory evaluation ofSkillEvoReg\. The pilot is designed to examine how continued skill updates affect performance on familiar workbook structures and on structurally shifted workbooks while the skill state itself evolves\. We therefore use this study as motivating evidence for the regularization problem; all comparative claims aboutSkillEvoRegare based on the independently specified experiments in Section[5](https://arxiv.org/html/2609.30861#S5)\.
##### IID and structural\-transfer evaluation\.
We organize the held\-out evaluation into an in\-distribution \(IID\) set and a structural transfer\-stress set\. The two sets keep the primary task semantics closely matched while differing in workbook\-level structure\. They contain identical proportions of the benchmark’s primary operation families and the same proportions of cell\-level and sheet\-level tasks\. Their average operation counts, compositional\-depth proxies, and nearest\-training instruction similarities are also similar, indicating that the transfer set does not primarily introduce new task families, rarer operations, or substantially different instruction semantics\. Instead, the main shift lies in the structure of the input workbooks\.
Compared with the IID set, the structural\-transfer set contains more multi\-sheet workbooks \(12/4012/40versus5/405/40\), more workbooks with three or more sheets \(3/403/40versus0/400/40\), more workbooks containing existing cross\-sheet formulas \(8/408/40versus1/401/40\), more tasks requiring cross\-sheet reasoning \(13/4013/40versus6/406/40\), and more merged\-cell workbooks \(8/408/40versus2/402/40\)\. The mean number of sheets increases from1\.1251\.125to1\.4751\.475, and the mean nearest\-training workbook\-structure distance increases from0\.8820\.882to1\.6761\.676\. We therefore use this set to evaluate transfer across workbook structure rather than transfer to unseen semantic task families\.
Table 5:Characteristics of the two held\-out SpreadsheetBench pilot sets\. The two sets have closely matched task semantics, while the transfer set contains greater workbook\-level structural variation\.
##### Protocol\.
The frozen pilot contains 80 training tasks, 40 validation tasks, 40 IID\-test tasks, and 40 structural\-transfer tasks, with a further 200 SpreadsheetBench tasks excluded from this study\. SkillOpt is run for eight epochs with seed 41 and a batch size of 20, yielding 32 skill updates\. Intermediate validation\-based stopping and other dynamic regularization mechanisms are disabled so that the complete native evolution trajectory can be observed\.
We retain the initial skill and the final checkpoint of every epoch\. Each checkpoint is evaluated on the same frozen training, validation, IID\-test, and structural\-transfer sets\. The complete checkpoint trajectory is evaluated in three independent evaluation replays, and we report the mean across replays\. These replays evaluate the same learned trajectory and therefore should not be interpreted as independent training runs\.
Table 6:Checkpoint performance in the motivating SpreadsheetBench pilot\. Values are mean task accuracy across three independent evaluation replays of the same learned trajectory\.
##### Post\-peak transfer behavior\.
Structural\-transfer performance reaches its maximum of46\.67%46\.67\\%at E4\. From E4 to E8, train, validation, and IID accuracy change by\+1\.25\+1\.25,\+5\.00\+5\.00, and\+7\.50\+7\.50percentage points, respectively, whereas transfer accuracy decreases by9\.179\.17points\. The IID–transfer gap consequently widens from7\.507\.50to24\.1724\.17percentage points\. The decline is also distributed across tasks: among the 40 transfer tasks, 13 regress from E4 to E8, 22 remain unchanged, and 5 improve\.
##### Deterministic lexical diagnostics\.
To characterize how the skill changes along the same trajectory without relying on an LLM judge, we apply four deterministic lexical diagnostics to every checkpoint\. These measures are descriptive proxies used only in this motivating study and are distinct from the operational complexity measureC\(S\)C\(S\)used bySkillEvoRegin Appendix[A\.2](https://arxiv.org/html/2609.30861#A1.SS2)\.
*Lexical length*counts word\-like sequences, numeric tokens, and non\-whitespace punctuation over the full skill document\.*Strict\-directive lines*count substantive lines containing directive expressions such asmust,always,never, ordo not\.*Conditional\-rule lines*count substantive lines containing conditional expressions such asif,when,unless, orfallback\.*Concrete\-artifact matches*count spreadsheet\-specific cell or range references, error identifiers, selected code identifiers, spreadsheet\-specific phrases, and selected functions\. These diagnostics characterize observable textual structure rather than semantic quality\.
Table 7:Deterministic lexical diagnostics along the motivating pilot trajectory\. These are descriptive proxies and are distinct from the complexity measureC\(S\)C\(S\)used bySkillEvoReg\.Between the transfer\-performance peak at E4 and the final checkpoint E8, lexical length increases by69\.9%69\.9\\%and conditional\-rule lines increase by70\.0%70\.0\\%\. Over the full trajectory, lexical length, strict\-directive lines, conditional\-rule lines, and concrete\-artifact matches reach23\.6×23\.6\\times,19\.5×19\.5\\times,85×85\\times, and10\.7×10\.7\\timestheir initial values, respectively\.
##### Interpretation\.
Taken together, the pilot illustrates a post\-peak transfer pattern: after structural\-transfer performance peaks, continued skill evolution is accompanied by a widening IID–transfer gap and continued accumulation of increasingly specific skill structure\. This pattern motivates treating repeated skill updates as a regularized learning process rather than relying on the updater alone\.
### A\.2SkillOpt Measurement Conventions and Complexity
We evaluate SpreadsheetBench \(100/100/200100/100/200train/validation/test examples\), SearchQA \(400/200/1400400/200/1400\), and LiveMathBench \(100/80/217100/80/217\)\. All SkillOpt runs use eight evolution epochs, eight\-example reflection minibatches, and an initial edit budget of four\. The native baseline retains SkillOpt’s slow\-update and meta\-skill mechanisms; theSkillEvoRegintegration retains the same ranked\-edit proposal mechanism and applies regularization before the native acceptance decision\.
Final checkpoints are evaluated after the fixed eight\-epoch training budget\. Delivery selects among the eight epoch\-end snapshots using the incumbent validation score recorded at the last training update of each epoch, breaking ties in favor of the earlier epoch\. The selected epochs for SkillOpt/SkillEvoRegare 7/7 on SpreadsheetBench, 4/2 on SearchQA, and 5/1 on LiveMath\. Their selection scores are59\.5/63\.5%59\.5/63\.5\\%,80\.25/80\.25%80\.25/80\.25\\%, and43\.125/38\.125%43\.125/38\.125\\%, respectively\. Test scores average two frozen evaluations of the exact selected snapshot\.
For completeness, SearchQA token\-overlap F1 is87\.98/88\.68%87\.98/88\.68\\%at delivery and88\.26/88\.41%88\.26/88\.41\\%at the final checkpoint for SkillOpt/SkillEvoReg, respectively, averaged over the same two frozen evaluations as EM\.
Skill length uses theo200k\_basetokenizer on the full document\. We segment skill documents into deterministic exact\-span atoms and define
C\(S\)=∑ai∈𝒜\(S\)\(τi100\+0\.2\+Li20\+2Bi5\+2Di3\),C\(S\)=\\sum\_\{a\_\{i\}\\in\\mathcal\{A\}\(S\)\}\\left\(\\frac\{\\tau\_\{i\}\}\{100\}\+0\.2\+\\frac\{L\_\{i\}\}\{20\}\+2\\frac\{B\_\{i\}\}\{5\}\+2\\frac\{D\_\{i\}\}\{3\}\\right\),\(13\)whereτi\\tau\_\{i\},LiL\_\{i\},BiB\_\{i\}, andDiD\_\{i\}denote the lexical\-unit count, nonempty code\-line count, control\-flow count, and external\-dependency count of atomaia\_\{i\}, respectively; code terms are zero for non\-code atoms\. The exact\-span segmenter retains document scaffolding, and each returned atom receives the base cost\. For code,LiL\_\{i\}excludes empty lines and Markdown fence delimiters,BiB\_\{i\}counts common control\-flow constructs, andDiD\_\{i\}counts recognized imports and selected HTTP dependency calls\. The measure is an operational regularization signal rather than a semantic\-quality score\.
Table 8:Final composite complexityCCon the SkillOpt benchmarks\.
### A\.3Evaluation Protocol Conventions
All reported performance values use the benchmark\-native verifier and the same task order for compared methods\. Frozen evaluations do not update the skill state\. When a benchmark provides a validation stream, checkpoint selection uses only the recorded validation scores and the selected checkpoint is then evaluated on the frozen test stream\. Main comparisons use one evolved state per condition; repeated frozen evaluations measure execution\-time variation of that state rather than variation across independently trained trajectories\.
## Appendix BFramework\-Specific Instantiations ofSkillEvoReg
SkillEvoRegis a general regularization framework rather than a single shared software core or a fixed pipeline implementation\. The three benchmark integrations retain their native skill evolvers and task evaluators but differ in skill representation, complexity semantics, CCV case construction, repair scope, and final candidate handling\.
Table 9:Framework\-specific integration ofSkillEvoReg\.Table 10:Framework\-specific CCV instantiations used in our experiments\.Across all three integrations, the incumbent and candidate must first satisfy the framework\-specific clean qualification, after which both are evaluated on the same attacked case under the same model configuration and benchmark\-native evaluator\. SkillOpt uses a strict hard\-pass/hard\-fail attacked\-case criterion with additional attribution controls, whereas SkillEvolBench and ContinualSkillBench also allow a framework\-defined continuous\-score drop to trigger a regression signal\.
## Appendix CAdditional Benchmark Results
### C\.1LiveMath: A Distribution\-Sensitive Answer Heuristic
LiveMath provides a case in which suppressing repetitive specialization can trade off against accuracy on the current test distribution\. SkillOpt’s final skill repeatedly emphasizes the meta\-option stating that a stronger result can be proved\. The regularized final skill retains one explicit meta\-option rule but consolidates overlapping strongest\-statement and strongest\-existence guidance\.
Of the 217 test questions, 92 have the correct answer text “One of the remaining options is correct, but a stronger result can be proven\.” Across the two frozen evaluation repeats,SkillEvoReg’s delivery checkpoint answers 33 of 184 corresponding question–repeat instances correctly, whereas the final checkpoint answers 15\. The remaining 125 questions change only from 119/250 to 116/250 correct\. Thus, meta\-option questions account for 18 of the 21 net lost correct answers:4\.154\.15of the overall4\.844\.84percentage\-point decline, or85\.7%85\.7\\%\. Pairing delivery and final outcomes within each repeat yields 23 correct\-to\-incorrect transitions and five reverse transitions on this subset\.
The concentration is consistent with reduced reinforcement of the meta\-option heuristic, but the comparison does not isolate any individual rule edit as causal\. It illustrates that regularization can suppress repetitive specialization even when the specialized pattern remains useful under the current evaluation distribution\. The case therefore illustrates a genuine regularization trade\-off: suppressing repeated specialization can reduce benchmark accuracy when the evaluation distribution itself rewards the specialized answer prior\.
### C\.2SkillEvolBench Measurements and Diagnostics
SkillEvolBench contains six environments with five task families each; T1–T3 are acquisition tasks and T4–T6 are deployment tasks\. In the original benchmark protocol, the acquisition sequence is traversed once before the resulting skill library is evaluated on unseen deployment tasks\. Our 1\-pass condition corresponds to this native acquisition endpoint\.
To study later\-stage skill evolution after a library has already formed, we introduce a controlled second acquisition pass; this 2\-pass setting is not part of the original SkillEvolBench protocol\. Both the unregularized 2\-pass condition andSkillEvoRegstart from the exact same one\-pass library and process the same T1–T3 acquisition sequence again in the same order\. Thus, they receive identical additional experience, while T4–T6 remain unseen throughout acquisition; the comparison isolates how that repeated experience is incorporated into the existing skill library\.SkillEvoRegretains the native execution environment, verifier, SkillAuthor proposal, and library commit interface, inserting regularization only around the candidate update before write\-back\. Deployment and replay evaluations use frozen libraries and two independent sessions\.
Library tokens useo200k\_base, summed over the 32 finalSKILL\.mdfiles across six isolated environments\. Structural atoms use the same deterministic exact\-span segmentation as Equation[13](https://arxiv.org/html/2609.30861#A1.E13)\. We sumCCover individual files and report growth relative to the one\-pass library\. The original benchmark\-specific atom counts are 991, 1,113, and 1,070 for the one\-pass, two\-pass, and regularized libraries; the common structural counts are 1,422, 1,551, and 1,515, respectively\. Relative complexity growth is14\.32%14\.32\\%for two passes and9\.65%9\.65\\%withSkillEvoReg\.
#### C\.2\.1Continuous verifier scores
In addition to strict success, we report the benchmark’s normalized overall verifier score, which gives partial credit according to each task’s official rubric\. We average the two frozen evaluations of each task and then average across tasks in each block\.
Table 11:Mean normalized SkillEvolBench verifier scores, averaged over two frozen evaluations\. Scores range from 0 to 1\.The deployment means differ by less than 0\.004 despite larger differences in strict success\. The two composition evaluations yield 6/30 and 5/30 successes for one pass, 4/30 and 4/30 for two passes, and 8/30 and 8/30 withSkillEvoReg\. All conditions retrieve the required skills on all original composition tasks, making missing\-skill retrieval an unlikely explanation for the composition difference\.
### C\.3ContinualSkillBench Measurements and Complete Results
ContinualSkillBench differs from the other settings in that its native updater can expand the skill library through explicitCreate\-Skillactions in addition to revising existing skills throughModify\-Skill\. The evolving state therefore includes both the contents of individual skills and the library structure itself\. We consequently treat library proliferation as part of the regularization problem: CREATE introduces a new persistent skill, whereas MODIFY changes an existing one\.
Training processes tasks 1–50 sequentially while retaining the evolving library\. Frozen evaluation then uses the step\-50 library on tasks 51–100, with a fresh session for each task and no further skill updates\. Baseline andSkillEvoRegshare the same task order, model, and evaluation procedure\. The two evaluation repeats use the same frozen library and therefore measure evaluation variation rather than variation across training runs\.
##### Online learning performance\.
Table[12](https://arxiv.org/html/2609.30861#A3.T12)follows the source benchmark’s reporting format and evaluates tasks 1–50 as they are encountered during skill evolution\. Raw is the mean over all 50 tasks; Norm\. is the mean over tasks with scoreable outputs in both conditions\. All 50 tasks are common\-valid in every domain here, so Raw and Norm\. coincide\.
Table 12:ContinualSkillBench online learning performance on tasks 1–50, reported in the source benchmark’s format\. Entries are rewards in\[0,1\]\[0,1\]from a single sequential learning trajectory\.
##### Held\-out evaluator breakdown\.
Structured in Table[3](https://arxiv.org/html/2609.30861#S5.T3)averages individual non\-rubric task scores across both repeats rather than averaging evaluator\-category means equally\. The source\-category breakdown is shown in Table[13](https://arxiv.org/html/2609.30861#A3.T13)\.
Table 13:Mean frozen reward by the source benchmark’s evaluator categories\. Every score averages two repeats\.
##### CREATE/MODIFY regularization\.
The two native update modes induce different growth semantics\. ForCreate\-Skill,SkillEvoRegcontrols the absolute complexity of the newly introduced skill and considers overlap with the existing library; forModify\-Skill, it regularizes the candidate relative to the incumbent skill and its update\-induced growth\. In both cases, the final regularized state is written back through the benchmark’s native continual\-learning workflow\.
##### Library measurements\.
Skill count is the number of skill directories containing aSKILL\.mdin the delivered domain library, including retained initial domain skills\. Tokens useo200k\_baseand are summed over skill documents and bundled auxiliary text/code files\. Complexity uses Equation[13](https://arxiv.org/html/2609.30861#A1.E13); auxiliary files contribute lexical length and code cost\. All measurements use the step\-50 frozen snapshot\.
##### Office sensitivity analysis\.
Office’s1\.441\.44\-point Overall difference is concentrated rather than broad\. All but three of the 21 structured tasks are identical between conditions\. One rubric task additionally generates a complete briefing but writes it to a path different from the evaluated path\. We retain the strict end\-to\-end score in the main table\. Excluding this single task reduces the Overall gap to0\.630\.63points; excluding it together with the three table\-selection outliers leavesSkillEvoReg0\.420\.42points higher on the other 46 tasks\. We use this analysis only to localize the observed difference, not to replace the official score\.
##### Healthcare evaluator analysis\.
Healthcare gains2\.202\.20Overall points while reducing the library from 42 to 34 skills\. Its Rubric score improves by4\.214\.21points, whereas the Structured difference corresponds to one net exact\-match success across the 24 question–repeat instances\. The aggregate gain therefore comes primarily from open\-ended transfer rather than from easier exact\-match behavior\.
### C\.4Deployment Inference Usage and Skill Compactness
Inference usage is not an optimization target ofSkillEvoReg; we examine it as a deployment\-side by\-product of compacting the persistent skill state\. A shorter persistent skill can reduce the context consumed during downstream execution, but end\-to\-end model usage also depends on the execution trajectory, particularly when external search, document retrieval, or other tool interactions introduce substantial and variable context\. For each evaluation run, we count input plus output tokens from the completed model session, including cached input\. Judge calls, training, regularizer calls, and superseded failed attempts are excluded\. Every row uses the frozen evaluation segment corresponding to the reported checkpoint and pools the two independent repeats\.
Table 14:Frozen\-evaluation model\-token usage\. Values are total input\-plus\-output tokens per task, pooled over two repeats; trimmed means remove the top and bottom10%10\\%of tasks\.The main pattern is that skill compactness reduces a stable component of inference context, but whether this translates into lower end\-to\-end token usage depends on how much execution is dominated by trajectory\-dependent external context\. The effect is clearest on tasks with relatively controlled execution paths\. On SpreadsheetBench and LiveMath, where task context is fixed and execution does not depend on open\-ended external information gathering,SkillEvoRegreduces mean frozen\-evaluation token usage by29\.5%29\.5\\%and38\.6%38\.6\\%, respectively, with corresponding reductions in both the median and trimmed mean\. SearchQA also shows a substantial56\.1%56\.1\\%reduction despite using web search, indicating that external search does not by itself eliminate the benefit of a compact skill state\. SkillEvolBench likewise reduces mean usage by9\.4%9\.4\\%relative to the matched two\-pass condition\. Together, these results show that compacting the persistent skill can produce substantial deployment\-side savings when the skill state remains a meaningful and comparatively stable component of the inference context\.
ContinualSkillBench exposes a different regime, in which high\-variance information\-acquisition and tool\-use trajectories can dominate total inference usage\. Finance and Office decrease across the principal summaries, whereas Healthcare increases; Mathematics and Law have higher arithmetic means underSkillEvoRegdespite lower medians, indicating that a small number of long trajectories can dominate average usage\. To examine this pattern more closely, we pair the same frozen\-evaluation task and repeat under the baseline andSkillEvoRegfor Mathematics, Law, and Healthcare\. As one measurable indicator of externally mediated trajectories, we separate pairs according to whether either execution invokesweb\_search\. Across the resulting 300 pairs, the three domains increase by 6\.45 million tokens in aggregate; pairs involvingweb\_searchaccount for a 6\.66 million\-token increase, whereas pairs with noweb\_searchdecrease by 0\.21 million tokens overall\. The contrast is particularly pronounced in Mathematics and Healthcare, where no\-search pairs decrease by 6,754 and 1,005 tokens per task on average, respectively, while search\-associated pairs increase by 53,247 and 73,161 tokens\. Law is more mixed: search\-associated pairs contribute most of the net increase, but several high\-cost trajectories involve extensive local\-document processing without web search\.
Theweb\_searchsplit should therefore be interpreted as an observable indicator of externally mediated trajectories rather than as a causal attribution to search itself\. High\-usage trajectories can also involve substantially different amounts of PDF or document extraction, shell\-returned content, and prolonged tool interaction, while the available logs do not expose the exact token contribution of each source after context injection and truncation\. The resulting picture is therefore not that compact skills universally reduce total inference tokens, but that they reduce one persistent and controllable source of context consumption\. When execution context is relatively stable, this reduction can translate into substantial token savings; when external context dominates the trajectory, the same benefit can be obscured by much larger variation in tool\-mediated context\.
##### Takeaway\.
Reduced inference usage is a deployment\-side by\-product rather than an optimization objective ofSkillEvoReg\. Compact persistent skills reduce the controllable skill\-context component of inference usage, while end\-to\-end savings depend on how strongly the downstream trajectory is dominated by external information and tool interactions\.
## Appendix DAlgorithmic Details and Updater Integration
This appendix formalizes the three regularizers introduced in Section[4](https://arxiv.org/html/2609.30861#S4)and describes how they connect to different native skill updaters\. The regularization roles are shared across SkillOpt, SkillEvolBench, and ContinualSkillBench, while the update boundary is framework\-dependent: candidate changes may be represented as document edits, localized library patches, or changes between skill\-library snapshots\.
### D\.1Native Update Interface
We represent one native evolution step abstractly as
\(τt,S~t\+1\)=NativeUpdate\(St,xt\),\(\\tau\_\{t\},\\widetilde\{S\}\_\{t\+1\}\)=\\operatorname\{NativeUpdate\}\(S\_\{t\},x\_\{t\}\),\(14\)whereStS\_\{t\}is the incumbent skill state,xtx\_\{t\}is the current training task or batch,τt\\tau\_\{t\}is the resulting execution trajectory, andS~t\+1\\widetilde\{S\}\_\{t\+1\}is the native candidate\. When dropout is active, the native update is first generated from a temporary masked view and then reconciled with the complete incumbent, as described below\.
Throughout the appendix,Diff\(S,S′\)\\operatorname\{Diff\}\(S,S^\{\\prime\}\)denotes the change from a source stateSSto a candidate stateS′S^\{\\prime\}\. We use stage\-specific superscripts for the resulting deltas:Δtdrop\\Delta\_\{t\}^\{\\mathrm\{drop\}\}for the change learned under the dropout view,Δtcomp\\Delta\_\{t\}^\{\\mathrm\{comp\}\}for the reconciled candidate change examined by complexity regularization, andΔtccv\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}for the post\-complexity candidate change examined by CCV\. This keeps the underlying notion of a candidate delta shared while distinguishing the state boundary at which it is computed\.
The main framework\-dependent integration operations areReconcile, which transfers an update learned from a temporarily perturbed skill view onto the complete incumbent, andWriteBack, which returns the final regularized candidate through the native framework’s acceptance or commit interface\.
### D\.2Training\-Time Skill Dropout
A*skill atom*is a locally meaningful unit of reusable skill content, such as an individual directive, list item, compact procedural block, or another semantically coherent segment\. Atom segmentation provides stable units for temporary masking and later local structural comparison\.
Algorithm[1](https://arxiv.org/html/2609.30861#alg1)formalizes training\-time skill dropout\. Its central invariant is that masking affects update generation only\. Content removed from the temporary view is restored before the dropout\-stage candidate proceeds to subsequent regularization\.
Algorithm 1Training\-Time Skill Dropout1:Incumbent skill state
StS\_\{t\}, training task
xtx\_\{t\}
2:Dropout\-reconciled candidate
St\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}
3:
𝒜t←SegmentAtoms\(St\)\\mathcal\{A\}\_\{t\}\\leftarrow\\operatorname\{SegmentAtoms\}\(S\_\{t\}\)
4:
Mt←SampleMask\(𝒜t\)M\_\{t\}\\leftarrow\\operatorname\{SampleMask\}\(\\mathcal\{A\}\_\{t\}\)
5:
Stdrop←Mask\(St,Mt\)S\_\{t\}^\{\\mathrm\{drop\}\}\\leftarrow\\operatorname\{Mask\}\(S\_\{t\},M\_\{t\}\)
6:
\(τt,S~t\+1drop\)←NativeUpdate\(Stdrop,xt\)\(\\tau\_\{t\},\\widetilde\{S\}\_\{t\+1\}^\{\\mathrm\{drop\}\}\)\\leftarrow\\operatorname\{NativeUpdate\}\(S\_\{t\}^\{\\mathrm\{drop\}\},x\_\{t\}\)
7:
Δtdrop←Diff\(Stdrop,S~t\+1drop\)\\Delta\_\{t\}^\{\\mathrm\{drop\}\}\\leftarrow\\operatorname\{Diff\}\(S\_\{t\}^\{\\mathrm\{drop\}\},\\widetilde\{S\}\_\{t\+1\}^\{\\mathrm\{drop\}\}\)
8:
St\+1drop←Reconcile\(St,Δtdrop\)S\_\{t\+1\}^\{\\mathrm\{drop\}\}\\leftarrow\\operatorname\{Reconcile\}\(S\_\{t\},\\Delta\_\{t\}^\{\\mathrm\{drop\}\}\)
9:return
St\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}
##### Atom segmentation\.
For textual skill states, atoms are obtained from deterministic structural segmentation of the skill document, so that repeated references to the same instruction or procedural block remain stable across an update\. Multi\-skill libraries additionally retain skill\-file or skill\-path boundaries\. These higher\-level boundaries determine the region within which atom\-level operations are permitted\.
##### Restoration\.
Because the native updater observesStdropS\_\{t\}^\{\\mathrm\{drop\}\}rather than the complete incumbentStS\_\{t\},Δtdrop\\Delta\_\{t\}^\{\\mathrm\{drop\}\}captures only the change proposed under the perturbed context\.Reconcilerestores the masked atoms fromStS\_\{t\}and transfers that change onto the complete incumbent, producingSt\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}\. An atom is therefore never removed merely because it was absent from the temporary training\-time view\.
### D\.3Complexity\-Aware Local Regularization
Complexity\-aware local regularization operates on the dropout\-reconciled candidateSt\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}\. The stage\-local change is
Δtcomp=Diff\(St,St\+1drop\),\\Delta\_\{t\}^\{\\mathrm\{comp\}\}=\\operatorname\{Diff\}\\left\(S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{drop\}\}\\right\),\(15\)which is measured relative to the complete incumbent, unlikeΔtdrop\\Delta\_\{t\}^\{\\mathrm\{drop\}\}in the preceding stage\.
Algorithm[2](https://arxiv.org/html/2609.30861#alg2)separates complexity measurement, local\-scope construction, and discrete edit selection\.
Algorithm 2Complexity\-Aware Local Regularization1:Incumbent skill state
StS\_\{t\}, dropout\-reconciled candidate
St\+1dropS\_\{t\+1\}^\{\\mathrm\{drop\}\}
2:Complexity\-regularized candidate
St\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}
3:
Δtcomp←Diff\(St,St\+1drop\)\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\\leftarrow\\operatorname\{Diff\}\(S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{drop\}\}\)
4:
ct←fcomp\(ϕ\(St\),ϕ\(St\+1drop\),ϕ\(Δtcomp\)\)c\_\{t\}\\leftarrow f\_\{\\mathrm\{comp\}\}\\\!\\left\(\\phi\(S\_\{t\}\),\\phi\(S\_\{t\+1\}^\{\\mathrm\{drop\}\}\),\\phi\(\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\)\\right\)
5:
St\+1comp←St\+1dropS\_\{t\+1\}^\{\\mathrm\{comp\}\}\\leftarrow S\_\{t\+1\}^\{\\mathrm\{drop\}\}
6:if
NeedsRegularization\(ct,Δtcomp\)\\operatorname\{NeedsRegularization\}\(c\_\{t\},\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\)then
7:
𝒟t←SegmentAtoms\(Δtcomp\)\\mathcal\{D\}\_\{t\}\\leftarrow\\operatorname\{SegmentAtoms\}\(\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\)
8:
ℛt←RetrieveRelated\(St,𝒟t\)\\mathcal\{R\}\_\{t\}\\leftarrow\\operatorname\{RetrieveRelated\}\(S\_\{t\},\\mathcal\{D\}\_\{t\}\)
9:
Ωt←BoundScope\(𝒟t∪ℛt\)\\Omega\_\{t\}\\leftarrow\\operatorname\{BoundScope\}\(\\mathcal\{D\}\_\{t\}\\cup\\mathcal\{R\}\_\{t\}\)
10:for all
a∈𝒟ta\\in\\mathcal\{D\}\_\{t\}do
11:
oa←SelectOperation\(a,Ωt,ct\)o\_\{a\}\\leftarrow\\operatorname\{SelectOperation\}\(a,\\Omega\_\{t\},c\_\{t\}\)
12:
St\+1comp←ApplyLocalOperation\(St\+1comp,a,oa,Ωt\)S\_\{t\+1\}^\{\\mathrm\{comp\}\}\\leftarrow\\operatorname\{ApplyLocalOperation\}\(S\_\{t\+1\}^\{\\mathrm\{comp\}\},a,o\_\{a\},\\Omega\_\{t\}\)
13:endfor
14:endif
15:return
St\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}
#### D\.3\.1Complexity Signal Definitions
Equation[6](https://arxiv.org/html/2609.30861#S4.E6)leaves the structural feature extractor abstract in the main text\. We instantiate it as
ϕ\(Z\)=\[ϕsize\(Z\),ϕatom\(Z\),ϕstructure\(Z\),ϕcode\(Z\),…\]\\phi\(Z\)=\\big\[\\phi\_\{\\mathrm\{size\}\}\(Z\),\\phi\_\{\\mathrm\{atom\}\}\(Z\),\\phi\_\{\\mathrm\{structure\}\}\(Z\),\\phi\_\{\\mathrm\{code\}\}\(Z\),\\ldots\\big\]\(16\)for either a skill state or an update delta\. For a skill state, the components summarize lexical, atom\-level, structural, and executable\-code properties; for an update delta, they summarize the corresponding structural change introduced or removed by the candidate\.
The regularization signal is
ct=fcomp\(ϕ\(St\),ϕ\(St\+1drop\),ϕ\(Δtcomp\)\)\.c\_\{t\}=f\_\{\\mathrm\{comp\}\}\\left\(\\phi\(S\_\{t\}\),\\phi\(S\_\{t\+1\}^\{\\mathrm\{drop\}\}\),\\phi\(\\Delta\_\{t\}^\{\\mathrm\{comp\}\}\)\\right\)\.\(17\)The functionfcompf\_\{\\mathrm\{comp\}\}need not be identical across native update protocols; it determines how shared structural measurements are converted into the signal used to activate or score regularization\.
##### SkillOpt\.
For SkillOpt, the structural features are summarized by the composite complexity measureC\(S\)C\(S\)defined in Equation[13](https://arxiv.org/html/2609.30861#A1.E13)\. The correspondingfcompf\_\{\\mathrm\{comp\}\}compares incumbent and candidate complexity and supplies the complexity term used by regularized candidate evaluation\. Complexity can therefore influence whether a candidate transition is retained\.
##### SkillEvolBench\.
For SkillEvolBench,fcompf\_\{\\mathrm\{comp\}\}is used primarily as a growth trigger\. When the candidate exceeds the configured structural\-growth budget, it enters bounded local regularization\. The signal determines whether additional consolidation is invoked rather than acting as an independent hard rejection criterion\.
##### ContinualSkillBench\.
ContinualSkillBench conditionsfcompf\_\{\\mathrm\{comp\}\}on the native update type\. ForCreate, the signal uses an absolute complexity budget for the newly created skill together with its relation to existing library content\. ForModify, it emphasizes relative growth of the modified skill with respect to its pre\-update state\. These are soft regularization triggers: exceeding them can invoke structural reconsideration without automatically vetoing the final candidate\.
##### Discrete local operations\.
The operation\-based formulation is inspired by Mem0’s memory reconciliation mechanism\([Chhikara et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib1)\), in which newly extracted facts are compared with related memories and assigned one of four actions:Add,Update,Delete, orNoOp\. Our problem begins after the native skill updater has already proposed new persistent content, so we adapt the discrete\-action principle to post\-update skill reconciliation:
oa∈\{NoOp,Merge,Rewrite,Delete\}\.o\_\{a\}\\in\\\{\\textsc\{NoOp\},\\textsc\{Merge\},\\textsc\{Rewrite\},\\textsc\{Delete\}\\\}\.\(18\)NoOpretains the candidate atom unchanged\.Mergeconsolidates overlapping candidate and incumbent atoms\.Rewritemodifies useful but unnecessarily narrow, redundant, or overspecified content\.Deleteremoves candidate content that contributes no distinct capability\.
##### Related\-content retrieval\.
For each candidate atom, the regularizer retrieves a bounded number of semantically or structurally related incumbent atoms\. Retrieval defines the comparison neighborhood rather than the edit decision itself\. High similarity does not imply that two atoms are interchangeable, since apparently redundant instructions can play distinct operational roles in language\-model execution\.
##### Granularity across updaters\.
SkillOpt and SkillEvolBench expose the four operations directly over candidate\-relative local content\. ContinualSkillBench enforces the same locality principle through affected skill paths and can rewrite the full text of an implicated skill\. This changes the editing granularity but not the role of the regularizer\.
##### Complexity semantics\.
The complexity signal can combine textual size, structural fragmentation, executable\-code structure, and relative growth\. SkillOpt can consume it as part of candidate acceptance; SkillEvolBench uses excessive growth to trigger bounded local regularization; and ContinualSkillBench uses soft absolute and relative budgets forCreateandModify, respectively\. These differences determine when local reconsideration is activated rather than defining different regularization mechanisms\.
### D\.4CCV Implementation Details
Section[4\.4](https://arxiv.org/html/2609.30861#S4.SS4)presents the four\-way behavioral comparison\. Here we make the underlying checks explicit\. CCV validates the post\-complexity transition fromStS\_\{t\}toSt\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}, so its stage\-local change isΔtccv=Diff\(St,St\+1comp\)\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}=\\operatorname\{Diff\}\(S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{comp\}\}\)\. This is distinct fromΔtcomp\\Delta\_\{t\}^\{\\mathrm\{comp\}\}, which is computed before structural regularization\.
Algorithm 3Causal Counterexample Validation1:Incumbent skill state
StS\_\{t\}, complexity\-regularized candidate
St\+1compS\_\{t\+1\}^\{\\mathrm\{comp\}\}, training history
ℋt\\mathcal\{H\}\_\{t\}, CCV configuration
θccv\\theta\_\{\\mathrm\{ccv\}\}
2:CCV\-processed candidate
St\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}
3:
St\+1ccv←St\+1compS\_\{t\+1\}^\{\\mathrm\{ccv\}\}\\leftarrow S\_\{t\+1\}^\{\\mathrm\{comp\}\}
4:
Δtccv←Diff\(St,St\+1comp\)\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}\\leftarrow\\operatorname\{Diff\}\(S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{comp\}\}\)
5:
et←SelectSource\(ℋt,St,St\+1comp\)e\_\{t\}\\leftarrow\\operatorname\{SelectSource\}\(\\mathcal\{H\}\_\{t\},S\_\{t\},S\_\{t\+1\}^\{\\mathrm\{comp\}\}\)
6:if
et=∅e\_\{t\}=\\varnothingthen
7:return
St\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}
8:endif
9:
roldclean←Evaluate\(St,et\)r\_\{\\mathrm\{old\}\}^\{\\mathrm\{clean\}\}\\leftarrow\\operatorname\{Evaluate\}\(S\_\{t\},e\_\{t\}\)
10:
rnewclean←Evaluate\(St\+1comp,et\)r\_\{\\mathrm\{new\}\}^\{\\mathrm\{clean\}\}\\leftarrow\\operatorname\{Evaluate\}\(S\_\{t\+1\}^\{\\mathrm\{comp\}\},e\_\{t\}\)
11:
qt←CleanQualify\(roldclean,rnewclean\)q\_\{t\}\\leftarrow\\operatorname\{CleanQualify\}\(r\_\{\\mathrm\{old\}\}^\{\\mathrm\{clean\}\},r\_\{\\mathrm\{new\}\}^\{\\mathrm\{clean\}\}\)
12:if
qt=0q\_\{t\}=0then
13:return
St\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}
14:endif
15:
at←GenerateAttack\(et,Δtccv\)a\_\{t\}\\leftarrow\\operatorname\{GenerateAttack\}\(e\_\{t\},\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}\)
16:if
at=∅a\_\{t\}=\\varnothingthen
17:return
St\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}
18:endif
19:
xtccv←ApplyAttack\(et,at\)x\_\{t\}^\{\\mathrm\{ccv\}\}\\leftarrow\\operatorname\{ApplyAttack\}\(e\_\{t\},a\_\{t\}\)
20:if
xtccv=∅x\_\{t\}^\{\\mathrm\{ccv\}\}=\\varnothingthen
21:return
St\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}
22:endif
23:
roldattack←Evaluate\(St,xtccv\)r\_\{\\mathrm\{old\}\}^\{\\mathrm\{attack\}\}\\leftarrow\\operatorname\{Evaluate\}\(S\_\{t\},x\_\{t\}^\{\\mathrm\{ccv\}\}\)
24:
rnewattack←Evaluate\(St\+1comp,xtccv\)r\_\{\\mathrm\{new\}\}^\{\\mathrm\{attack\}\}\\leftarrow\\operatorname\{Evaluate\}\(S\_\{t\+1\}^\{\\mathrm\{comp\}\},x\_\{t\}^\{\\mathrm\{ccv\}\}\)
25:
gt←RegressionCriterion\(roldattack,rnewattack,θccv\)g\_\{t\}\\leftarrow\\operatorname\{RegressionCriterion\}\(r\_\{\\mathrm\{old\}\}^\{\\mathrm\{attack\}\},r\_\{\\mathrm\{new\}\}^\{\\mathrm\{attack\}\};\\theta\_\{\\mathrm\{ccv\}\}\)
26:if
gt=1g\_\{t\}=1then
27:
Ωtccv←ImplicatedScope\(Δtccv,at,roldattack,rnewattack\)\\Omega\_\{t\}^\{\\mathrm\{ccv\}\}\\leftarrow\\operatorname\{ImplicatedScope\}\(\\Delta\_\{t\}^\{\\mathrm\{ccv\}\},a\_\{t\},r\_\{\\mathrm\{old\}\}^\{\\mathrm\{attack\}\},r\_\{\\mathrm\{new\}\}^\{\\mathrm\{attack\}\}\)
28:
St\+1ccv←Repair\(St\+1comp,Ωtccv\)S\_\{t\+1\}^\{\\mathrm\{ccv\}\}\\leftarrow\\operatorname\{Repair\}\(S\_\{t\+1\}^\{\\mathrm\{comp\}\},\\Omega\_\{t\}^\{\\mathrm\{ccv\}\}\)
29:endif
30:return
St\+1ccvS\_\{t\+1\}^\{\\mathrm\{ccv\}\}
##### Source selection and clean qualification\.
CCV drawsete\_\{t\}from previously observed training trajectories rather than held\-out evaluation data\. The first two checks correspond to the first column of Equation[9](https://arxiv.org/html/2609.30861#S4.E9)\. The incumbent must succeed on the original source, establishing that the behavior is already supported byStS\_\{t\}, and the candidate must remain qualified on that source, ruling out the simpler case in which the update has already broken the original task\. Previously computed clean results are reused when available; the twoEvaluateoperations in Algorithm[3](https://arxiv.org/html/2609.30861#alg3)denote logical checks and do not require redundant re\-execution\.
##### Candidate\-conditioned attack generation\.
After the clean pair qualifies,GenerateAttackreceivesete\_\{t\}and the post\-complexity candidate changeΔtccv\\Delta\_\{t\}^\{\\mathrm\{ccv\}\}\. The generator identifies an assumption, dependency, or decision boundary that the candidate may have introduced, strengthened, or narrowed, and proposes a bounded perturbation designed to stress that property\. Generation may return no attack when the candidate change exposes no suitable target or when no valid perturbation can be constructed\.
##### Benchmark\-specific attack application\.
The attack specificationata\_\{t\}is converted into an executable instance throughApplyAttack\. SpreadsheetBench can perturb workbook cells, formulas, or task instructions subject to validity checks\. SearchQA introduces bounded distractor context while retaining the original question and gold answer\. LiveMath adds controlled prior scratch\-work context while retaining the original problem, answer options, and label\. SkillEvolBench preserves the original software\-task instruction and appends bounded robustness context targeted at the candidate change\. ContinualSkillBench uses both answer\-preserving transformations and nearby boundary cases with independently checkable oracles\. Although the surface transformations differ, every integration constructs one attacked instancextccvx\_\{t\}^\{\\mathrm\{ccv\}\}and holds it fixed for the incumbent–candidate comparison\.
##### Regression criteria\.
The attacked evaluations correspond to the second column of Equation[9](https://arxiv.org/html/2609.30861#S4.E9)\. A CCV\-confirmed regression requires the incumbent to remain sufficiently successful onxtccvx\_\{t\}^\{\\mathrm\{ccv\}\}while the candidate deteriorates according to the native evaluator\. SkillOpt uses a strict incumbent\-pass/candidate\-fail condition after clean qualification\. SkillEvolBench requires the incumbent attacked\-case score to be at least0\.80\.8and the candidate score to decrease by at least0\.050\.05relative to the incumbent; the candidate need not cross a binary failure threshold\. ContinualSkillBench requires incumbent success together with either candidate failure or a score decrease greater than0\.10\.1\. These differences affect the numerical realization ofRegressionCriterion, while the four\-way comparison remains unchanged\.
##### Scope\-constrained repair\.
When the paired criterion identifies a regression, CCV usesΔtccv\\Delta\_\{t\}^\{\\mathrm\{ccv\}\},ata\_\{t\}, and the attacked\-case outcomes to localize the candidate content most directly implicated by the failure\. SkillOpt and SkillEvolBench restrict repair to the relevant candidate change and related local content, using the local editing mechanism described in Section[4\.3](https://arxiv.org/html/2609.30861#S4.SS3)\. ContinualSkillBench follows the same locality principle at the level of implicated skill paths and may rewrite the full content of an affected skill\. Repair granularity therefore follows the native representation while remaining constrained to the region implicated by the candidate\-conditioned attack\. Each triggering CCV event performs at most one scope\-constrained repair\.
Table 15:Integration ofSkillEvoRegwith the three native skill\-update interfaces\. The regularization mechanisms remain largely unchanged; the main differences concern how the native framework represents, exposes, and writes back a candidate update\.
### D\.5Integration with Native Skill Updaters
The principal framework\-dependent implementation choices occur at the native update boundary\. Table[15](https://arxiv.org/html/2609.30861#A4.T15)summarizes these differences\.
The table highlights why the three benchmarks do not require three different formulations ofSkillEvoReg\. Their native updaters determine how each candidate transition becomes visible and how the final candidate is returned\. Once the stage\-specific changes are exposed, dropout reconciliation, local structural regularization, and paired counterexample validation retain the same conceptual roles\.
## Appendix EDetailed Component and Trajectory Analyses
### E\.1Failure Case: Non\-Local Whole\-Skill Rewriting
We analyze a SpreadsheetBench failure case to illustrate whySkillEvoRegrestricts complexity regularization to candidate\-local regions rather than allowing unrestricted whole\-skill rewriting\. In this case, a validation regression triggered a repair procedure that operated over the complete persistent skill, even though the complexity trigger itself was inactive\. The repaired skill was then written back without a post\-repair evaluation\. We therefore treat this example as a diagnostic case study of non\-local rewriting rather than as a controlled comparison between compression strategies\.
The original candidate update was comparatively local\. Relative to the preceding incumbent, it added four bullets, removed one inherited bullet and one third\-level heading, and altered content in only44of the1414major\#\#sections\. The subsequent repair, however, was permitted to rewrite the complete skill rather than being restricted to this candidate\-local change\. Although the repair prompt requested localized revision, it exposed the complete skill and imposed no programmatic boundary on the editable region\.
Table 16:Diagnostic failure case illustrating the scope and behavioral consequences of unrestricted whole\-skill rewriting\. Textual rewrite statistics characterize rewrite scope rather than semantic rule deletion\.The rewrite footprint was substantially broader than the source candidate delta\. Although the candidate affected only4/144/14major sections, the whole\-skill repair changed content in all1414sections, and seven original section titles no longer appeared verbatim\. Among the 47 bullets inherited by the candidate from the preceding skill, only six remained textually identical after the repair, while41/4741/47were no longer preserved verbatim\. These statistics quantify textual non\-locality rather than semantic deletion: they do not imply that 41 distinct rules were removed, since some content may have been merged, relocated, or paraphrased\.
The resulting checkpoint also exhibited strongly asymmetric behavioral change\. Test accuracy decreased from60\.5%60\.5\\%to48\.5%48\.5\\%: 32 previously solved tasks became failures, while only 8 previously failed tasks became correct\. Inspection of the trace indicated that the global rewrite altered established operational guidance rather than merely consolidating candidate\-local redundancy\. Because the transition also involved an erroneous trigger route and lacked post\-repair evaluation, these observations should not be interpreted as a controlled estimate of the causal effect of whole\-skill rewriting\.
This failure case motivates the separation between the*trigger*for structural reconsideration and the*scope*permitted to change\. InSkillEvoReg, complexity measurements can use the broader incumbent and candidate states to determine whether regularization is warranted, but editing is anchored to the candidate\-local delta𝒟t\\mathcal\{D\}\_\{t\}and a bounded related contextΩt\\Omega\_\{t\}\. Unrelated incumbent content is therefore preserved rather than being exposed to an unconstrained global rewrite\. Locality does not require every repair to be small; it requires the permissible edit region to remain tied to the candidate transition being regularized\.
### E\.2CCV Regression Evidence and Representative Trace
##### Regression evidence\.
Across the evaluated trajectories, CCV identifies 12 distinct regression signals, 9 of which lead to state\-changing repairs\. Each signal arises from a candidate\-conditioned comparison in which the incumbent and candidate remain qualified on the original source, while the candidate deteriorates relative to the incumbent on the same attacked case\. These events provide update\-level evidence for the type of behavioral regression targeted by CCV\. We next examine one representative trace in detail\.
##### Representative CCV trace\.
We use a representative SkillEvolBench dependency\-scope trace to illustrate the mechanism\. The source task asks the agent to repair a React 18 peer\-dependency conflict\. The incumbent contains a general instruction to regenerate and validate the repository lockfile\. The candidate adds a more specific rule that makes manifest ownership and package/workspace boundaries an explicit decision criterion\.
CCV applies the same package\-boundary attack to both incumbent and candidate runs\. The incumbent receives reward1\.01\.0, satisfying 7/7 outcome checks and 5/5 process checks\. The candidate receives0\.80\.8: all seven outcome checks remain satisfied, but two process checks fail after it replaces versioned React dependencies with localfile:shims\. The attacked comparison therefore reveals candidate\-specific sensitivity to the package\-boundary condition targeted by the candidate\-added rule\.
CCV rewrites the exact candidate\-added atom from a rule that emphasizes the directory owning the manifest to one that ties lockfile regeneration and validation to the dependency scope actually used by the resolver\. CCV therefore identifies candidate\-specific sensitivity to package\-boundary wording and produces a targeted scope\-constrained rewrite\. A later replay recovers the two affected process checks\.
### E\.3Residual Stochasticity and Validation\-Based Checkpointing
##### Shared\-prefix continuation\.
TwoSkillEvoRegSpreadsheetBench continuations share a byte\-identical skill state through step 30, the end of epoch six, and then process the same remaining batch sequence under the same configuration\. Steps 29 and 30 are rejected, so the shared step\-30 snapshot is also textually identical to the state produced at step 28\. The first different skill states appear at step 31, where independently sampled candidate updates are accepted\. These are controlled continuations of one learned prefix, not independent end\-to\-end training seeds\.
Table 17:SpreadsheetBench shared\-prefix continuation\. E1–E6 are common to both branches; E7–E8 differ after independent update sampling\. Test scores average two frozen evaluations\.The validation\-selected delivery checkpoint is E7 on the main continuation and E6 on the degrading continuation\. Although the two terminal states diverge substantially \(59\.25%59\.25\\%versus50\.00%50\.00\\%test accuracy\), checkpoint selection yields much more comparable delivered performance \(58\.50%58\.50\\%versus57\.25%57\.25\\%\)\. This illustrates that regularization and checkpoint selection address distinct sources of risk:SkillEvoRegshapes the repeated update process, whereas validation\-based checkpointing protects delivery from residual stochastic trajectory degradation\. We retain the latter as an established early\-stopping\-style safeguard rather than a contribution ofSkillEvoReg\.
##### Accepted update patterns\.
After the fork, each branch accepts five of the ten remaining proposals\. The main branch accepts steps 31, 32, 33, 35, and 37; the degrading branch accepts 31, 34, 36, 37, and 39\. The degrading branch accumulates broader instructions for preserving blank/nonmatching slots, clearing a prefilled destination block, blanking non\-terminal rows, and clearing unused spill cells\. The main branch is not free of blanking language, but combines it with explicit full\-range reconnaissance, blank\-versus\-zero disambiguation, and conjunctive deletion conditions\. The comparison therefore concerns the combination and scope of accepted rules, not a binary distinction between “blanking” and “no blanking\.”
##### Task\-level audit\.
The same 200 test tasks are evaluated twice at the shared E6 checkpoint and at the degrading E8 checkpoint\. Thirty\-six tasks regress, 12 improve, and 152 are unchanged, yielding a net decline of7\.257\.25percentage points\. Thirty\-one of the 36 regressions are cell\-level manipulation tasks\. Twenty\-one of the 36 produce a missing prediction in both E8 repeats, and 30 do so in at least one repeat\. The dominant observed failure pattern is therefore missing required cell values rather than uniform degradation across task families\.
The accepted clearing and blanking instructions are behaviorally compatible with this pattern, but they are not isolated as the cause\. Multiple skill edits and evaluation\-time randomness remain confounders\. Train and validation performance also decline along the degrading branch\. We therefore interpret this controlled fork as stochastic trajectory\-level degradation rather than conventional train–test overfitting\. The example illustrates a residual source of variation that update\-level regularization does not eliminate, and shows how validation\-based checkpoint selection can prevent such late\-stage degradation from determining the delivered state\.
## Appendix FPrompts and Scope Enforcement for LLM\-Based Regularization Operations
The algorithms in Appendix[D](https://arxiv.org/html/2609.30861#A4)define semantic operations such asSelectOperation\\operatorname\{SelectOperation\},GenerateAttack\\operatorname\{GenerateAttack\}, andRepair\\operatorname\{Repair\}rather than prescribing a single prompt shared across all host frameworks\. Their concrete realizations depend on the native skill representation and task\-valid perturbation space\. This section provides representative prompt excerpts and summarizes the corresponding writable\-scope constraints\. Native skill\-updater prompts are inherited from the host frameworks and are not reproduced here\.
For presentation consistency, implementation identifiers in the excerpts are normalized to the terminology used in this paper; the operational instructions and constraints are otherwise preserved\. Dynamic runtime fields are shown in braces\.
### F\.1Complexity\-Aware Local Editing
For SkillOpt and SkillEvolBench, the local editor receives broader skill context for interpretation, but only explicitly identified source\-delta regions are writable\. The core editing instruction defines four operations:
> Treat the listed source deltas as independent editable regions\. Similarity only nominates historical content for semantic comparison; it does not imply that two rules are interchangeable\. For every source delta, choose exactly one operation: NOOP: preserve the source unchanged when it adds a distinct or complementary capability, or whenever semantic preservation is uncertain\. REWRITE: replace only the source delta when it is overspecific, unclear, or internally redundant, while preserving its supported capability\. MERGE: use only when one coherent rule can preserve the full union of the source delta and linked historical content\. DELETE: remove only a source delta that is unsupported, harmful, or fully subsumed by retained content\. The complete skill is read\-only context\. Return modifications only for explicitly supplied editable regions; never return a rewritten complete skill\.
When complexity regularization is triggered, the editor additionally receives:
> Merge duplicates and semantic overlaps, remove narrow or unsupported rules, replace similar rule clusters with transferable general rules, and compress repeated explanations or examples\. Do not add new rules merely in response to excessive complexity, and preserve distinct useful behavior\.
The corresponding user request separates read\-only context from writable regions:
> \#\# Full Current Skill \(read\-only context\) \{current\_skill\} \#\# Editable Source\-Delta Regions \{editable\_regions\} \#\# Regularization Feedback \{regularization\_feedback\} Return one decision for every listed source delta and only the replacements required by those decisions\.
SkillEvolBench uses the same four\-operation semantics with an additionalskill\_id, since edits may belong to different localized skill files\. ContinualSkillBench instead follows its native CREATE/MODIFY representation: CREATE operates over explicitly supplied skill paths, whereas MODIFY restricts changes to a set of allowed window identifiers\. In all cases, the writable scope is constructed and validated by the framework adapter rather than inferred by the language model\.
### F\.2CCV Attack Generation
CCV constructs a candidate\-conditioned behavioral probe rather than an arbitrary harder example\. Because valid perturbations differ substantially across benchmarks,GenerateAttack\\operatorname\{GenerateAttack\}has framework\-specific realizations\.
SkillOpt uses a two\-stage Scout–Worker procedure\. The Scout observes the incumbent and candidate skills, their semantic delta, recent trajectories, eligible source tasks, and the verifier contract, and returns a candidate\-specific weakness together with a suitable source\. A Worker then materializes that weakness using the benchmark’s permitted perturbation space\. SpreadsheetBench applies bounded operations to a private workbook copy; SearchQA inserts bounded distractor context while preserving the question and gold answer; and LiveMath introduces bounded prior scratch\-work context while preserving the original problem, options, and gold label\.
SkillEvolBench uses a single bounded probe generator\. Its central instruction is:
> Generate one bounded CCV probe for a coding\-agent skill update\. The original task instruction and verifier are immutable\. Target a specific semantic addition, deletion, or strengthening introduced by the candidate skill\. The probe may add only neutral operational context, an answer\-free counterfactual condition, or a presentation variation targeted at an incidental assumption introduced by the candidate\. Do not change the requested deliverable, acceptance criteria, repository contents, tests, or environment\. Do not reveal a solution, expected patch, hidden test, or answer\. Select responsible skill identifiers only from skills changed by the candidate\.
The generator returns the attack type, targeted assumption, predicted failure region, responsible skill identifiers, a goal\-preservation justification, and the bounded robustness context appended to the original instruction\.
ContinualSkillBench follows the same candidate\-conditioned principle but uses domain\-specific probe generators for Mathematics, Finance, Law, Healthcare, and Office\. These generators target domain\-appropriate boundaries or invariance\-preserving transformations, with additional oracle or equivalence checks before the probe is used for paired incumbent–candidate evaluation\.
### F\.3CCV Repair
Once paired evaluation identifies a candidate\-specific regression, the repair stage is instructed to recover a transferable capability rather than encode the attacked example itself\. SkillOpt and SkillEvolBench reuse the local editor from Appendix[F\.1](https://arxiv.org/html/2609.30861#A6.SS1)with CCV\-specific feedback:
> A CCV regression has been detected\. Revise, generalize, or remove the incidental assumption responsible for the regression and restore the exposed transferable capability\. Never add a one\-off exception for a single counterexample\. Derive exactly one transferable procedural invariant from the CCV counterexample and use it to replace, merge, or generalize the responsible existing rule\. Do not copy counterexample\-specific field names, file names, schema keys, literal values, concrete formats, or wording; express the invariant at the capability level instead\.
SkillEvolBench additionally provides the measured score decrease, attack type, targeted assumption, and predicted failure region as diagnostic context, while retaining the same region\-level editing restriction\.
ContinualSkillBench uses a separate repair call at responsibility\-path granularity:
> Repair the cause of the detected CCV regression as a transferable invariant\. Prefer editing or merging the responsible rule\. Modify only the supplied responsibility skills, preserve unrelated semantics, and do not add source answers, task identifiers, or irrelevant probe\-specific details\.
Unlike the region\-local realizations above, this operation may return the complete text of an authorized responsibility skill\. The implementation therefore validates the returned paths and verifies the actual modified\-path set before write\-back\.
Table 18:Writable\-scope enforcement across host frameworks\. The model may inspect broader context, but the adapter determines which regions, windows, or skill paths are eligible for write\-back\.
### F\.4Programmatic Enforcement of Editable Scope
Prompt\-level locality is not itself an enforceable editing boundary\. As illustrated by the failure case in Appendix[E\.1](https://arxiv.org/html/2609.30861#A5.SS1), a model can perform a broad rewrite even when instructed to revise locally\. In the final implementation, the framework adapter determines the writable scope independently of the model\.
The common invariant is therefore not that every repair is atom\-local\. Rather, the framework adapter determines the writable scope, while the granularity of that scope follows the native representation\. SkillOpt and SkillEvolBench use region\-level editing, ContinualSkillBench MODIFY uses bounded windows, and ContinualSkillBench CREATE and CCV repair may rewrite complete skills only within explicitly authorized paths\.
This separation is deliberate: prompt instructions specify*how*the model should transform authorized content, whereas programmatic scope enforcement determines*what*content is permitted to change\.
## Appendix GExtended Related Work
##### Learning from agent experience\.
A broad line of research studies how language\-model agents can improve through interaction without modifying model parameters\. Reflexion stores verbal reflections derived from task feedback in episodic memory and reuses them in subsequent attempts\([Shinn et al\., 2023](https://arxiv.org/html/2609.30861#bib.bib2)\), while ExpeL extracts reusable natural\-language insights from collections of agent experiences\([Zhao et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib3)\)\. Agent\-Pro iteratively refines a language\-level behavioral policy through reflection and search\([Zhang et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib5)\)\. Voyager maintains an expanding library of executable code skills that can be retrieved and composed for later tasks\([Wang et al\., 2024a](https://arxiv.org/html/2609.30861#bib.bib4)\)\. More recent memory systems explicitly organize reusable experience: Agent Workflow Memory induces recurring workflows from successful trajectories\([Wang et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib6)\), and A\-MEM continuously restructures an interconnected memory network as new experiences arrive\([Xu et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib7)\)\. These methods establish reusable external state as an effective mechanism for non\-parametric adaptation; our work focuses on regularizing repeated revision of such procedural state\.
##### Skill generation, optimization, and maintenance\.
SkillOpt formulates a natural\-language skill as trainable external state and optimizes it through bounded textual edits, validation\-gated acceptance, and checkpoint selection\([Yang et al\., 2026b](https://arxiv.org/html/2609.30861#bib.bib8)\)\. Trace2Skill aggregates trajectory\-local lessons before hierarchically consolidating them into transferable skills\([Ni et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib9)\), while CoEvoSkills jointly evolves structured skill packages and a surrogate verifier\([Zhang et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib10)\)\. SkillEvolBench evaluates transfer from episodic experience to procedural skills under context shifts, adversarial shortcuts, and skill composition\([Lei et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib11)\); ContinualSkillBench studies skill creation and modification across sequential tasks and highlights the difficulty of consolidating experience into compact reusable skills\([Guan et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib12)\)\. Closest to our reliability motivation, GSE uses a global skill\-relation graph, cross\-task consolidation, and replay\-driven verification\([Yang et al\., 2026a](https://arxiv.org/html/2609.30861#bib.bib13)\); SkillCommit validates broader abstractions before committing them\([He and Yang, 2026](https://arxiv.org/html/2609.30861#bib.bib14)\); and SkillAdam uses optimization history and adaptive edit budgets to stabilize iterative skill optimization\([Li et al\., 2026](https://arxiv.org/html/2609.30861#bib.bib15)\)\. These methods provide safeguards within particular evolution algorithms, whereasSkillEvoRegasks whether complementary anti\-overfitting principles can transfer across different update semantics\.
##### Optimization of prompts and external language\-model state\.
ProTeGi uses textual feedback analogous to gradients together with beam search to improve prompts\([Pryzant et al\., 2023](https://arxiv.org/html/2609.30861#bib.bib16)\); OPRO treats an LLM itself as an optimizer over natural\-language solutions\([Yang et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib17)\); EvoPrompt and Promptbreeder use evolutionary search over prompts\([Guo et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib18);[Fernando et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib20)\); and PromptAgent formulates prompt optimization as strategic search\([Wang et al\., 2024b](https://arxiv.org/html/2609.30861#bib.bib19)\)\. DSPy extends optimization to modular language\-model programs\([Khattab et al\., 2024](https://arxiv.org/html/2609.30861#bib.bib21)\), while TextGrad propagates language\-model feedback through compound AI systems in analogy to automatic differentiation\([Yuksekgonul et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib22)\)\. These works primarily address how to search for higher\-performing external state\.SkillEvoReginstead regularizes the transitions proposed by an existing updater\.
##### Regularization and counterexample\-based validation\.
Classical learning combines multiple mechanisms to control overfitting: dropout reduces feature co\-adaptation\([Srivastava et al\., 2014](https://arxiv.org/html/2609.30861#bib.bib23)\), weight decay constrains unnecessary capacity\([Krogh and Hertz, 1991](https://arxiv.org/html/2609.30861#bib.bib24)\), and data augmentation or adversarial training exposes models to informative perturbations\([Goodfellow et al\., 2015](https://arxiv.org/html/2609.30861#bib.bib25);[Zhang et al\., 2018](https://arxiv.org/html/2609.30861#bib.bib26)\)\. Validation\-based early stopping provides an additional safeguard against late\-stage deterioration\([Prechelt, 1998](https://arxiv.org/html/2609.30861#bib.bib27)\)\.SkillEvoRegtransfers the principles behind these techniques to discrete skill evolution rather than reproducing their parametric implementations\. CCV is also related to counterexample\-guided synthesis, differential testing, and program repair, where carefully chosen failures reveal weaknesses in candidate programs\. Counterexample\-guided program repair, for example, combines failing examples with fault localization to direct corrections\([Orvalho et al\., 2025](https://arxiv.org/html/2609.30861#bib.bib28)\)\. In our setting, the validated object is a skill transition: incumbent and candidate are first qualified on the original source and then compared on the same candidate\-conditioned attacked case, turning the counterexample into an update\-level behavioral check\.相似文章
重新思考自我进化:一种受约束的探索-利用过程以缓解技能过拟合
本文提出SkillBoost,一种三阶段的受约束探索-利用框架,用于缓解LLM智能体自我进化中的技能过拟合。它在23个模型-基准配置上达到了最先进的性能,并证明了优化后的技能可以迁移到其他智能体。
SkillEvo:基于多轮交互反馈的自我更新进化梯度
SkillEvo 引入了一种通过使用多轮交互反馈和治理层来持续改进 AI 代理技能的方法,以保持进化梯度,超越了自反思和单轮问答驱动的方法。
@Yif_Yang: 介绍 SkillOpt — 一个面向智能体技能的优化器。不再微调模型权重,而是将自然语言…
介绍 SkillOpt,一个将自然语言技能视为可训练外部参数而非微调模型权重的优化器。它通过有界编辑和验证门控实现稳定、可控的技能更新,在 7 个模型的 6 个基准测试的 52 个设置中取得最佳或并列最佳结果。
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
This paper introduces SkillMisevo-Gym and SkillMisevo-Bench to study how self-improving LLM agents can evolve unsafe skills from compromised experience, plus SafeEvolve as a mitigation wrapper. Experiments across 25 agent-method configurations show skill misevolution is widespread and can persist across sessions, though SafeEvolve reduces fresh-session harm significantly.
GraphSkillEvo:图结构代理技能的进化优化
GraphSkillEvo是一个进化优化框架,它将代理技能表示为图结构工件,以提高大型语言模型(LLM)的性能,在多个基准测试中超越基线。