Tag
Vincent Weisser announces the release of over 365,000 open and agentic reinforcement learning environments for software engineering, terminal, and search agents.
Prime Intellect publishes 365,000+ tasks for reinforcement learning agents, covering SWE, terminal, and search tasks, accessible via a single API and command.
This paper introduces RLVP (Reward the Outcome, Penalize the Path), a reinforcement learning method that uses a verifiable penalty for path violations and outcome reward to achieve near-zero constraint violations with high task success, improving sample efficiency in real-world agentic environments.
Qwen-AgentWorld introduces language world models for agentic environments, covering seven domains with long chain-of-thought reasoning. The work includes a new benchmark, AgentWorldBench, and shows that world modeling improves downstream agent performance.
A comprehensive survey on agentic environment engineering for LLMs, covering environment modeling, synthesis, evaluation, and application, with a focus on agent-environment co-evolution.
OpenEnv, a framework for creating and deploying isolated execution environments for agentic RL training, has moved to Hugging Face and is now governed by a committee including Meta-PyTorch, NVIDIA, and others.