Tag
SkillGym proposes an automatic pipeline that crawls reproducible skills from the internet, builds verifiable difficulty-controlled environments, and collects 19k verified trajectories to train skill-use agents. Fine-tuning Qwen3.5 models (2B to 122B) on these trajectories improves performance across four skill-use benchmarks, with the 9B SFT model outperforming a 397B untrained model on two of them.
Introduces SKT, a verified data synthesis pipeline for skill-use training of language model agents, producing 4,000 task packages and 27,164 verified trajectories from 2,000 public skills. Supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance across models and agent harnesses.