Tag
Xiaomi released an open-source repository of reinforcement learning environments on Hugging Face, offering cost savings for AI task acquisition compared to purchasing similar resources.
Xiaomi has open-sourced approximately 7,000 reinforcement learning environments used to train its MiMo model on HuggingFace, covering domains like code, cybersecurity, general, music, and web development.
Release of SmolDataEnvs, a collection of 5,000 verifiable RL environment tasks for training small models in code and data science, fully open source.
This paper from Salesforce audits public RL environments for terminal agents, finding significant defects, and introduces RIVER, a training recipe that filters defective environments and penalizes repetitive behavior to enhance model performance.
CodeMidas is an agentic pipeline that creates reinforcement learning environments from source code, scaling the training of coding agents and showing performance improvements on benchmarks like issue repair and program construction.
Jerry Liu shares thoughts on how forward-deployed engineer (FDE) work will shift toward defining goals, environments, and evals while automated optimization processes handle tactical implementation.
The author argues that RL environments serve as the essential data for building AI agents, enabling systematic training, prompt optimization, and evaluation rather than manual iteration.
Standard Machines launches reinforcement-learning environments for chip design, aiming to accelerate tape-out and push frontier AI capabilities.
A technical tutorial on building a reinforcement learning environment for LLMs using the open-source Verifiers library, with Othello as a working example.
OpenReward environments now integrate directly into TRL's GRPOTrainer via a single OpenRewardSpec, allowing zero-glue-code training against a catalog of RL environments. The integration is experimental and part of a broader effort to make environment and agent RL first-class in TRL.
This article recommends a website called Sophon, which aggregates AI papers, models, benchmarks, leaderboards, and reinforcement learning environments. It provides real-time rankings, comparisons, and subscription features, and is hailed as the Bloomberg terminal for AI research.
Hugging Face Hub has surpassed 4,000 public reinforcement learning environments, positioning itself as a potentially largest platform for RL environments.
Researchers release Terminal Wrench, a dataset of 331 reward-hackable terminal environments with 3,632 exploit trajectories spanning sysadmin, ML, and security tasks.