SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Hugging Face Daily Papers Papers

Summary

SkillGym proposes an automatic pipeline that crawls reproducible skills from the internet, builds verifiable difficulty-controlled environments, and collects 19k verified trajectories to train skill-use agents. Fine-tuning Qwen3.5 models (2B to 122B) on these trajectories improves performance across four skill-use benchmarks, with the 9B SFT model outperforming a 397B untrained model on two of them.

Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:24 PM

Paper page - SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Source: https://huggingface.co/papers/2609.37539

Abstract

SkillsequipLLMagentswithprofessionalknowledgeandguidancetocompletelong-horizonandcomplextasks.Althoughskillshavebeenwidelyadoptedinrecentagentparadigmsandharnesses,howtosynthesizereliabletrainingdataandhowtotrainagentsforskilluseremainunderexplored.Inthiswork,weproposeSkillGym,anautomaticpipelinetobuildverifiableenvironments,collecttrajectories,andtrainskill-useagents.SkillGymfirstcrawlsalargevolumeofskillsfromtheinternet,thenkeepsthosewhoseworkflowscanrunreproduciblyoffline.Abuilder-reviewerpipelineisusedtoconstructdifficulty-controlledtasks,spanningfourtasktypes,eachwithareferencesolutionandanexecutableverifier.Withthispipeline,webuild6.8kenvironmentsandcollect19kverifiedsuccessfultrajectoriesforsupervisedfinetuning.FinetuningonthesetrajectoriesimprovesLLMsofdifferentfamiliesandsizes,from2Bto122Bparametersacrossfourskill-usebenchmarks;OurQwen3.5-9BSFTmodeloutperformsthe397Buntrainedmodelontwoofthem.Furtheranalysisshowsthattrainingteachesagentstoinvokeskills,raisingtherateofreadingtherelevantskillfrom28%to96%,andthatthegainsholdacrossreasoningstructures,extendingtotasktypesthatformaminorityofthetrainingdataandtoskillsheldoutfromtraining

View arXiv pageView PDFGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.37539

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper5

#### reasonwang/SkillGym-Qwen3.5-9B Text Generation• 9B• Updatedabout 4 hours ago • 32 #### reasonwang/SkillGym-Qwen3.5-4B Text Generation• 5B• Updatedabout 4 hours ago • 2 #### reasonwang/SkillGym-Qwen3.5-27B Text Generation• 3.05M• Updatedabout 4 hours ago • 2 #### reasonwang/SkillGym-Qwen3.5-122B-A10B Text Generation• 125B• Updatedabout 4 hours ago • 11 Browse 5 models citing this paper## Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37539 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37539 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

SkillGen: Verified Inference-Time Agent Skill Synthesis

arXiv cs.LG

This article introduces SkillGen, a multi-agent framework that synthesizes and verifies reusable inference-time skills for LLM agents by contrasting successful and failed trajectories. The method ensures skills are auditable and empirically verified for their net positive impact on agent performance.

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Hugging Face Daily Papers

Introduces SKT, a verified data synthesis pipeline for skill-use training of language model agents, producing 4,000 task packages and 27,164 verified trajectories from 2,000 public skills. Supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance across models and agent harnesses.