Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

Hugging Face Daily Papers Papers

Summary

This paper introduces calibrated importance sampling to address the training-inference mismatch in reinforcement learning for large language models, improving policy updates and performance on mathematical reasoning benchmarks.

We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To account for this discrepancy in policy updates, we introduce calibrated importance sampling (CIS). CIS is motivated by an empirically supported logit-displacement characterization that expresses the mismatch as an additive displacement varepsilon_t in log-odds, determined by the per-logit perturbation before the softmax, whose distribution is approximately invariant to token confidence. This characterization motivates a confidence-aware truncation: large positive displacements are truncated at a single constant threshold, which maps back to an importance-ratio cap that tightens as token confidence increases. Theoretically, we show that CIS replaces the unbounded second moment that governs the error of exact importance sampling with a term bounded by a constant, at the cost of a bias controlled by the truncated excess. In evaluation across three mixture-of-experts models and five mathematical reasoning benchmarks, CIS achieves the highest five-benchmark average on all three models among the evaluated baselines. Diagnostic analyses show that CIS places less truncation bias on low-confidence tokens than truncated importance sampling, while upward clipping of small importance weights reduces held-out accuracy.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:13 AM

Paper page - Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It

Source: https://huggingface.co/papers/2609.32444

Abstract

Westudytraining-inferencemismatchinreinforcementlearningwithverifiablerewards(RLVR)forlargelanguagemodels,whererolloutsaresampledbyaninferenceenginewhilegradientsarecomputedbyatrainingengine,andthetwoenginesassigndifferentprobabilitiestothesametokens.Toaccountforthisdiscrepancyinpolicyupdates,weintroducecalibratedimportancesampling(CIS).CISismotivatedbyanempiricallysupportedlogit-displacementcharacterizationthatexpressesthemismatchasanadditivedisplacementvarepsilon_tinlog-odds,determinedbytheper-logitperturbationbeforethesoftmax,whosedistributionisapproximatelyinvarianttotokenconfidence.Thischaracterizationmotivatesaconfidence-awaretruncation:largepositivedisplacementsaretruncatedatasingleconstantthreshold,whichmapsbacktoanimportance-ratiocapthattightensastokenconfidenceincreases.Theoretically,weshowthatCISreplacestheunboundedsecondmomentthatgovernstheerrorofexactimportancesamplingwithatermboundedbyaconstant,atthecostofabiascontrolledbythetruncatedexcess.Inevaluationacrossthreemixture-of-expertsmodelsandfivemathematicalreasoningbenchmarks,CISachievesthehighestfive-benchmarkaverageonallthreemodelsamongtheevaluatedbaselines.DiagnosticanalysesshowthatCISplaceslesstruncationbiasonlow-confidencetokensthantruncatedimportancesampling,whileupwardclippingofsmallimportanceweightsreducesheld-outaccuracy.

View arXiv pageView PDFProject pageGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.32444

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.32444 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.32444 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.32444 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Hugging Face Daily Papers

We introduce MIPI (Monotonic Inference Policy Improvement) and its instantiation MIPU, a two-step RL framework for LLMs that addresses the training-inference mismatch by explicitly aligning optimization with inference-policy improvement. Under FP8-quantized rollout, MIPU achieves improved reasoning performance and training stability across Qwen3-1.7B and Qwen3-4B models.