zero-rl

Tag

Cards List
#zero-rl

@rosinality: https://arxiv.org/abs/2607.12395 RL without SFT. What could be an interesting point is that they were able to make reas…

X AI KOLs Timeline · 5d ago Cached

The paper 'Ring-Zero' presents a stable training pipeline for scaling zero reinforcement learning (without supervised fine-tuning) to 1 trillion parameters, achieving emergent reasoning capabilities such as self-verification and structured formatting, and demonstrating strong performance on mathematical benchmarks.

0 favorites 0 likes
#zero-rl

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

arXiv cs.CL · 6d ago Cached

Ring-Zero scales reinforcement learning with verifiable rewards to trillion-parameter models, demonstrating emergent reasoning behaviors such as self-verification and structured formatting without human-annotated data.

0 favorites 0 likes
#zero-rl

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models

arXiv cs.AI · 2026-07-07 Cached

Oyster-II proposes a reinforcement learning framework for constructive safety alignment in LLMs, overcoming limitations of prior SFT-based methods via a multi-stage Zero-RL paradigm, achieving state-of-the-art safety performance while preserving general capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback