calibrated-importance-sampling

Tag

Cards List
#calibrated-importance-sampling

Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper introduces calibrated importance sampling to address the training-inference mismatch in reinforcement learning for large language models, improving policy updates and performance on mathematical reasoning benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback