Tag
The paper 'Ring-Zero' presents a stable training pipeline for scaling zero reinforcement learning (without supervised fine-tuning) to 1 trillion parameters, achieving emergent reasoning capabilities such as self-verification and structured formatting, and demonstrating strong performance on mathematical benchmarks.
Ring-Zero scales reinforcement learning with verifiable rewards to trillion-parameter models, demonstrating emergent reasoning behaviors such as self-verification and structured formatting without human-annotated data.
Oyster-II proposes a reinforcement learning framework for constructive safety alignment in LLMs, overcoming limitations of prior SFT-based methods via a multi-stage Zero-RL paradigm, achieving state-of-the-art safety performance while preserving general capabilities.