@MiniMax_AI: Congrats to our long-term partner SGLang/RadixArk on the launch of Miles v0.1! From M-Series to H3 and Music 3, we’ve b…
Summary
RadixArk launches Miles v0.1, an open-source reinforcement learning framework for large language models and multimodal models, aimed at simplifying and scaling RL training.
View Cached Full Text
Cached at: 08/18/26, 08:33 PM
Congrats to our long-term partner SGLang/RadixArk on the launch of Miles v0.1!
From M-Series to H3 and Music 3, we’ve been closely building and pushing the OSS ecosystem forward together.
Nothing beats the feeling of seeing things we built together keep growing @radixark @MiniMax_AI
RadixArk (@radixark): Today we’re launching Miles v0.1, an open-source RL framework for LLMs and multimodal models.
RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale.
Over the past 9 months, 72
Similar Articles
@sgl_project: SGLang is proud to be the native rollout engine for Miles. We're here to keep the tokens flowing and the GPUs busy Grea…
SGLang is announced as the native rollout engine for Miles v0.1, an open-source reinforcement learning framework for LLMs and multimodal models, aimed at improving throughput, cache efficiency, and stability in RL training at scale.
Miles: A PyTorch-Native Stack for Large-Scale LLM RL Post-Training (14 minute read)
Miles is an open-source PyTorch-native framework from RadixArk for large-scale LLM reinforcement learning post-training, integrating SGLang, Megatron-LM, and Ray for high-throughput rollout and distributed training.
@PyTorch: Built on PyTorch, Ray, SGLang, and NVIDIA Megatron-LM, Miles is an open source framework from RadixArk for large-scale …
Miles is an open source framework from RadixArk for large-scale LLM reinforcement learning post-training, integrating PyTorch, Ray, SGLang, and NVIDIA Megatron-LM with support for MoE, low-precision, and fault tolerance.
Miles v0.1: Production-level Post-training (20 minute read)
Miles v0.1 is a production-ready system for frontier post-training, optimizing reinforcement learning loops with fully asynchronous RL and agentic rollouts via SGLang.
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
The MiniMax-M2 series introduces Mixture-of-Experts language models that achieve high performance on agentic tasks with minimal activated parameters (9.8B per token out of 229.9B total), leveraging agent-driven data pipelines, a scalable RL system called Forge, and a checkpoint that takes early steps toward self-evolution.