monotonic-policy-improvement

Tag

Cards List
#monotonic-policy-improvement

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Hugging Face Daily Papers · 2026-06-28 Cached

We introduce MIPI (Monotonic Inference Policy Improvement) and its instantiation MIPU, a two-step RL framework for LLMs that addresses the training-inference mismatch by explicitly aligning optimization with inference-policy improvement. Under FP8-quantized rollout, MIPU achieves improved reasoning performance and training stability across Qwen3-1.7B and Qwen3-4B models.

0 favorites 0 likes
← Back to home

Submit Feedback