SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
Summary
SkillGym proposes an automatic pipeline that crawls reproducible skills from the internet, builds verifiable difficulty-controlled environments, and collects 19k verified trajectories to train skill-use agents. Fine-tuning Qwen3.5 models (2B to 122B) on these trajectories improves performance across four skill-use benchmarks, with the 9B SFT model outperforming a 397B untrained model on two of them.
View Cached Full Text
Cached at: 10/01/26, 04:24 PM
Paper page - SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
Source: https://huggingface.co/papers/2609.37539
Abstract
SkillsequipLLMagentswithprofessionalknowledgeandguidancetocompletelong-horizonandcomplextasks.Althoughskillshavebeenwidelyadoptedinrecentagentparadigmsandharnesses,howtosynthesizereliabletrainingdataandhowtotrainagentsforskilluseremainunderexplored.Inthiswork,weproposeSkillGym,anautomaticpipelinetobuildverifiableenvironments,collecttrajectories,andtrainskill-useagents.SkillGymfirstcrawlsalargevolumeofskillsfromtheinternet,thenkeepsthosewhoseworkflowscanrunreproduciblyoffline.Abuilder-reviewerpipelineisusedtoconstructdifficulty-controlledtasks,spanningfourtasktypes,eachwithareferencesolutionandanexecutableverifier.Withthispipeline,webuild6.8kenvironmentsandcollect19kverifiedsuccessfultrajectoriesforsupervisedfinetuning.FinetuningonthesetrajectoriesimprovesLLMsofdifferentfamiliesandsizes,from2Bto122Bparametersacrossfourskill-usebenchmarks;OurQwen3.5-9BSFTmodeloutperformsthe397Buntrainedmodelontwoofthem.Furtheranalysisshowsthattrainingteachesagentstoinvokeskills,raisingtherateofreadingtherelevantskillfrom28%to96%,andthatthegainsholdacrossreasoningstructures,extendingtotasktypesthatformaminorityofthetrainingdataandtoskillsheldoutfromtraining
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.37539
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper5
#### reasonwang/SkillGym-Qwen3.5-9B Text Generation• 9B• Updatedabout 4 hours ago • 32
#### reasonwang/SkillGym-Qwen3.5-4B Text Generation• 5B• Updatedabout 4 hours ago • 2
#### reasonwang/SkillGym-Qwen3.5-27B Text Generation• 3.05M• Updatedabout 4 hours ago • 2
#### reasonwang/SkillGym-Qwen3.5-122B-A10B Text Generation• 125B• Updatedabout 4 hours ago • 11
Browse 5 models citing this paper## Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.37539 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.37539 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
@dair_ai: An interesting idea is to continually train specialized models on skills. This paper explores that idea. They propose S…
SkillGym is a framework that converts human-written skills into training environments for LLMs, enabling fine-tuning that boosts performance on benchmarks like Terminal-Bench and SkillsBench, surpassing scores from models such as Claude Sonnet 4.6 and GPT-5.4 Mini.
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
CUA-Gym introduces a scalable pipeline for generating verifiable training environments and tasks for computer-use agents, addressing data scarcity. The resulting dataset and models achieve strong performance on benchmarks like OSWorld-Verified and WebArena.
SkillGen: Verified Inference-Time Agent Skill Synthesis
This article introduces SkillGen, a multi-agent framework that synthesizes and verifies reusable inference-time skills for LLM agents by contrasting successful and failed trajectories. The method ensures skills are auditable and empirically verified for their net positive impact on agent performance.
SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving
SkillGym transforms human-written agent skills into executable training environments for LLMs, enabling supervised fine-tuning and reinforcement learning to enhance real-world problem-solving capabilities and performance on benchmarks.
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Introduces SKT, a verified data synthesis pipeline for skill-use training of language model agents, producing 4,000 task packages and 27,164 verified trajectories from 2,000 public skills. Supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance across models and agent harnesses.