Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

arXiv cs.AI Papers

Summary

This paper proposes SkillBoost, a three-stage constrained exploration-exploitation framework to mitigate skill overfitting in LLM agent self-evolution. It achieves state-of-the-art performance across 23 model-benchmark configurations and demonstrates that optimized skills transfer to other agents.

arXiv:2607.26643v1 Announce Type: new Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconstrained exploration causes regression on previously solved cases. This tension motivates a constrained search view of skill self-evolution, governed by an exploration--exploitation trade-off. We propose SkillBoost, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound. Experiments across 23 model--benchmark configurations show that SkillBoost achieves state-of-the-art performance while mitigating overfitting, outperforming both human-crafted and LLM-generated skills. Transfer experiments further show that optimized skills can be reused by other agents on similar tasks.
Original Article
View Cached Full Text

Cached at: 07/31/26, 04:01 AM

# Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Source: [https://arxiv.org/abs/2607.26643](https://arxiv.org/abs/2607.26643)
[View PDF](https://arxiv.org/pdf/2607.26643)

> Abstract:Enabling large language model \(LLM\) agents to accumulate and reuse experience from past interactions remains a central challenge in real\-world applications\. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training\. However, data\-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments\. Overexploiting these trajectories overfits the current batch, while unconstrained exploration causes regression on previously solved cases\. This tension motivates a constrained search view of skill self\-evolution, governed by an exploration\-\-exploitation trade\-off\. We propose SkillBoost, a three\-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior\-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound\. Experiments across 23 model\-\-benchmark configurations show that SkillBoost achieves state\-of\-the\-art performance while mitigating overfitting, outperforming both human\-crafted and LLM\-generated skills\. Transfer experiments further show that optimized skills can be reused by other agents on similar tasks\.

## Submission history

From: Hongqiang Lin \[[view email](https://arxiv.org/show-email/f92c40ed/2607.26643)\] **\[v1\]**Wed, 29 Jul 2026 09:05:40 UTC \(663 KB\)

Similar Articles

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

arXiv cs.AI

This paper introduces SkillMisevo-Gym and SkillMisevo-Bench to study how self-improving LLM agents can evolve unsafe skills from compromised experience, plus SafeEvolve as a mitigation wrapper. Experiments across 25 agent-method configurations show skill misevolution is widespread and can persist across sessions, though SafeEvolve reduces fresh-session harm significantly.

OpenSkill: Open-World Self-Evolution for LLM Agents

Hugging Face Daily Papers

OpenSkill is a framework for LLM agents to self-evolve skills and verification signals from open-world resources without target-task supervision, achieving high performance across benchmarks.

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Hugging Face Daily Papers

SkillOpt introduces a systematic text-space optimizer for agent skills that trains skills as external agent state with stable updates and zero deployment inference overhead, achieving superior performance across multiple benchmarks and execution environments.

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe

Hugging Face Daily Papers

SkillOpt-Lite proposes a minimal viable pipeline for skill optimization in autonomous agents, achieving better and faster self-evolution by treating all components as editable code and integrating into production coding agents. It formalizes skill optimization via Zeroth-Order optimization and outperforms prior methods on benchmarks.