@kubasienki: Whoa, often such releases seed progress.
Summary
Vincent Weisser announces the release of over 365,000 open and agentic reinforcement learning environments for software engineering, terminal, and search agents.
View Cached Full Text
Cached at: 07/24/26, 09:07 AM
Whoa, often such releases seed progress.
Vincent Weisser (@vincentweisser): We’re publishing over 365,000 open and agentic RL Environments for SWE, terminal, and search agents
The open research ecosystem has produced many great datasets for the three main agentic domains - software engineering, terminal use, and web research - but every one of them
Similar Articles
@KaiZhang_CS: Check out one of the best open-source search agents trained by @jianxie_ !! glad to see early experience methods work o…
Yu Su's team trained a frontier Deep Research Agent on an academic budget using 8K synthetic samples and RL, releasing fully open training infrastructure and models from 2B to 35B parameters.
@svpino: This is huge for anyone who wants to see open models succeed! First time I've seen open models release checkpoint-level…
The post highlights the release of K2 Horizon, a connected fleet of six open foundation models from 0.9B to 375B parameters, with checkpoint-level logs, training code, and evaluations, marking a significant step for open AI development.
@eliebakouch: a LOT of RL environments (365k tasks) curated in one place, one command
Prime Intellect publishes 365,000+ tasks for reinforcement learning agents, covering SWE, terminal, and search tasks, accessible via a single API and command.
@ClementDelangue: Super happy to release SmolDataEnvs: 5,000 verifiable RL environment tasks for hill-climbing small models in code and d…
Release of SmolDataEnvs, a collection of 5,000 verifiable RL environment tasks for training small models in code and data science, fully open source.
@rohanpaul_ai: Another fantastic open source release. DeepReinforce just dropped Ornith-1.0, an MIT-licensed open-source family of age…
DeepReinforce releases Ornith-1.0, an MIT-licensed open-source family of agentic coding LLMs including a 397B MoE model that surpasses Claude Opus 4.7 on SWE-Bench and Terminal-Bench, using a novel self-improving training strategy.