Tag
This paper introduces calibrated importance sampling to address the training-inference mismatch in reinforcement learning for large language models, improving policy updates and performance on mathematical reasoning benchmarks.