Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
Summary
This paper introduces calibrated importance sampling to address the training-inference mismatch in reinforcement learning for large language models, improving policy updates and performance on mathematical reasoning benchmarks.
View Cached Full Text
Cached at: 09/29/26, 04:13 AM
Paper page - Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
Source: https://huggingface.co/papers/2609.32444
Abstract
Westudytraining-inferencemismatchinreinforcementlearningwithverifiablerewards(RLVR)forlargelanguagemodels,whererolloutsaresampledbyaninferenceenginewhilegradientsarecomputedbyatrainingengine,andthetwoenginesassigndifferentprobabilitiestothesametokens.Toaccountforthisdiscrepancyinpolicyupdates,weintroducecalibratedimportancesampling(CIS).CISismotivatedbyanempiricallysupportedlogit-displacementcharacterizationthatexpressesthemismatchasanadditivedisplacementvarepsilon_tinlog-odds,determinedbytheper-logitperturbationbeforethesoftmax,whosedistributionisapproximatelyinvarianttotokenconfidence.Thischaracterizationmotivatesaconfidence-awaretruncation:largepositivedisplacementsaretruncatedatasingleconstantthreshold,whichmapsbacktoanimportance-ratiocapthattightensastokenconfidenceincreases.Theoretically,weshowthatCISreplacestheunboundedsecondmomentthatgovernstheerrorofexactimportancesamplingwithatermboundedbyaconstant,atthecostofabiascontrolledbythetruncatedexcess.Inevaluationacrossthreemixture-of-expertsmodelsandfivemathematicalreasoningbenchmarks,CISachievesthehighestfive-benchmarkaverageonallthreemodelsamongtheevaluatedbaselines.DiagnosticanalysesshowthatCISplaceslesstruncationbiasonlow-confidencetokensthantruncatedimportancesampling,whileupwardclippingofsmallimportanceweightsreducesheld-outaccuracy.
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.32444
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.32444 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.32444 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.32444 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
This paper diagnoses Training-Inference Mismatch (TIM) in LLM reinforcement learning, showing that small numerical disagreements between training and inference token probabilities can cause training collapse, and proposes remedies.
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
This paper introduces reinforcement learning with metacognitive feedback (RLMF) and metacognitive data selection to improve large language model calibration, enabling faithful expression of intrinsic uncertainty and surpassing standard RL by up to 63%.
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
We introduce MIPI (Monotonic Inference Policy Improvement) and its instantiation MIPU, a two-step RL framework for LLMs that addresses the training-inference mismatch by explicitly aligning optimization with inference-policy improvement. Under FP8-quantized rollout, MIPU achieves improved reasoning performance and training stability across Qwen3-1.7B and Qwen3-4B models.
Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
This paper studies stable miscalibration in large language models, where high-confidence errors remain locally stable under perturbations, using diagnostics like audit scores and probes to assess calibration and internal sensitivity.
Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models
This paper identifies Calibration Drift Under Reasoning (CDUR), where increasing chain-of-thought reasoning budgets causes LLMs to become systematically overconfident in incorrect answers, and proposes a Hypothesis Lock-In model and a calibration-aware stopping rule (CABStop) to mitigate the issue.