From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
Summary
This paper systematically evaluates model-generated skills for language agents across the full lifecycle of experience generation, extraction, and consumption, finding that skills are beneficial on average but exhibit non-trivial negative transfer, leading to a meta-skill that improves skill quality.
View Cached Full Text
Cached at: 05/25/26, 02:35 AM
Paper page - From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
Source: https://huggingface.co/papers/2605.23899 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Language agents benefit from reusable skills that encode domain-specific procedures, but their effectiveness varies significantly across different extraction and consumption scenarios, requiring careful evaluation and meta-skill guidance to optimize performance.
Language agentsincreasingly improve by reusingskills-- structured procedural artifacts distilled from past experience. In particular, domain-level andmodel-generated skillsare especially promising. They offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle -- experience generation,skill extraction, andskill consumption-- to ask whether suchskillsactually work, when they work, and what makes them succeed or fail. To close this gap, we build a utility-grounded evaluation framework that provides systematic experimental results across extractors and target agents, covering five diverse agentic task domains. We find thatmodel-generated skillsare beneficial on average but exhibit non-trivialnegative transfer, and that neither extractors nor targets behave uniformly. A model can be a strong extractor yet a weak consumer, or vice versa, with skill utility independent of model scale or baseline task strength. To explain these patterns, we then dissect each lifecycle stage in depth, analyzing how experience composition shapes skill quality, what properties characterize usefulskills, and how the same skill transfers across different consumers. Finally, we translate these findings into a concretemeta-skillthat guidesskill extractiontoward the features tied to actual utility, which consistently improves skill quality across domains and substantially reducesnegative transfer.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.23899
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.23899 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.23899 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.23899 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@dair_ai: Great paper demystifying agent skills.
A paper demystifies agent skills by analyzing 8,135 normalized trials, challenging the assumption that skills primarily inject knowledge into models.
@dair_ai: Finally, a good paper testing whether Agent Skills actually help. Worth reading if you are maintaining a skill library …
A benchmark study shows that injecting Agent Skills in Web Development tasks often reduces performance and increases token cost, with failure modes like length-distracted and content-misled models, highlighting the need for per-deployment evaluation.
Demystifying Agent Skills: Why They Work-Until They Don't
This paper investigates the conditions under which skills enhance LLM agent performance, analyzing success and failure factors through controlled experiments and proposing a taxonomy of skill-use modes.
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
This paper proposes SkillBoost, a three-stage constrained exploration-exploitation framework to mitigate skill overfitting in LLM agent self-evolution. It achieves state-of-the-art performance across 23 model-benchmark configurations and demonstrates that optimized skills transfer to other agents.
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
RESOURCE2SKILL is a framework that distills executable agent skills from multimodal resources like tutorial videos, code repositories, articles, and artifacts into a hierarchical SkillWiki, improving agent performance by 11.9 percentage points over no-skill agents.