Tag
This paper introduces VHD-Play, a pipeline that generates diverse agentic reinforcement learning environments by first solving mathematical models, significantly improving training for language-model agents like Qwen3.6-35B-A3B at low cost and extending to external benchmarks.