Tag
This paper introduces VHD-Play, a pipeline that generates diverse agentic reinforcement learning environments by first solving mathematical models, significantly improving training for language-model agents like Qwen3.6-35B-A3B at low cost and extending to external benchmarks.
Vincent Weisser announces the release of over 365,000 open and agentic reinforcement learning environments for software engineering, terminal, and search agents.
Prime Intellect publishes 365,000+ tasks for reinforcement learning agents, covering SWE, terminal, and search tasks, accessible via a single API and command.
This paper introduces RLVP (Reward the Outcome, Penalize the Path), a reinforcement learning method that uses a verifiable penalty for path violations and outcome reward to achieve near-zero constraint violations with high task success, improving sample efficiency in real-world agentic environments.
Qwen-AgentWorld introduces language world models for agentic environments, covering seven domains with long chain-of-thought reasoning. The work includes a new benchmark, AgentWorldBench, and shows that world modeling improves downstream agent performance.
A comprehensive survey on agentic environment engineering for LLMs, covering environment modeling, synthesis, evaluation, and application, with a focus on agent-environment co-evolution.
OpenEnv, a framework for creating and deploying isolated execution environments for agentic RL training, has moved to Hugging Face and is now governed by a committee including Meta-PyTorch, NVIDIA, and others.