@samsja19: We are also releasing prime-rl 0.7.0 which has full support for verifiers v1 and bring your own harness for training. W…
Summary
Prime Intellect released verifiers v1 and prime-rl 0.7.0, an RL training tool with full support for verifiers, multiple algorithms like GRPO and OPD, and performance improvements.
View Cached Full Text
Cached at: 07/14/26, 04:26 AM
We are also releasing prime-rl 0.7.0 which has full support for verifiers v1 and bring your own harness for training.
We have been battle tested it in prof for weeks now, on the menu:
- Verifiers v1 integration
- Full Algorithm layers (GRPO, OPD, OPSD, SFT, ECHO
- Bunch of performance improvement and saner default
in prod*
Similar Articles
Today, we are releasing verifiers v1 (3 minute read)
Prime Intellect Lab is out of beta, offering a platform to train models with support for various architectures and modalities, enabling self-improving agents.
@samsja19: prime-rl can now train 1T parameters MoE blazingly fast, under 5 minutes per step, or 1k steps in ~3 days To achieve th…
Prime Intellect released prime-rl v0.6.0, enabling reinforcement learning at trillion-parameter MoE scale with sub-5-minute step times and optimized inference, training, and rollout.
@samsja19: We spend a lot of time designing an elegant algorithm api in prime rl that expressive and extensible but doesn't sacrif…
Prime-rl adds a first-class algorithms layer with six built-in RL algorithms (GRPO, MaxRL, OPD, OPSD, SFT, ECHO), making it easier to implement custom algorithms with a single file.
@samsja19: https://x.com/samsja19/status/2076846033922035818
PRIME-RL is a framework for large-scale asynchronous reinforcement learning, designed to be hackable and scale to 1000+ GPUs with support for various models and environments.
@eliebakouch: every infra piece you need to know to do RL on GLM-5 https://primeintellect.ai/blog/rl-at-1t-scale…
Prime Intellect releases prime-rl v0.6.0, enabling efficient reinforcement learning at trillion-parameter scale on large Mixture-of-Experts models, with sub-5-minute step times and optimizations for asynchronous RL.