Improving LLMs via Validator-to-Generator Alignment
Summary
A new method, FLORA (Frequency-corrected Learning of Ordered Rank Alignment), improves LLMs by aligning generator and validator modes using a principled frequency correction. Experiments show substantial gains in G-V consistency and generator performance on benchmarks like IFEval and HumanEval.
View Cached Full Text
Cached at: 07/07/26, 04:36 AM
# Improving LLMs via Validator-to-Generator Alignment
Source: [https://arxiv.org/html/2607.02668](https://arxiv.org/html/2607.02668)
Juan Diego Rodriguez♠Jocelyn Zhang♠Katrin Erk♢Greg Durrett♣
♠Department of Computer Science, The University of Texas at Austin ♢Departments of Linguistics and Computer Science, University of Massachusetts Amherst ♣Department of Computer Science & Center for Data Science, New York University \{juand\-r, jocelynzhang\}@utexas\.edukerk@umass\.edugdurrett@nyu\.edu
###### Abstract
Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs\. The generator\-validator \(G\-V\) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re\-queried to validate them\. In this work, we introduce a new formulation of G\-V consistency that involves a principled correction for utterance frequency\. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G\-V consistency unworkable\. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency\-corrected generator score emerges naturally\. Our method,*Frequency\-corrected Learning of Ordered Rank Alignment*\(FLORA\), is a training objective implementing frequency\-corrected G\-V consistency for real\-world LLMs\. Our experimental results show that training with FLORA substantially improves both G\-V consistency and generator performance over prior methods, with gains of up to\+27\+27pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks\.111Code and data available at[https://github\.com/juand\-r/flora](https://github.com/juand-r/flora)
Improving LLMs via Validator\-to\-Generator Alignment
Juan Diego Rodriguez♠Jocelyn Zhang♠Katrin Erk♢Greg Durrett♣♠Department of Computer Science, The University of Texas at Austin♢Departments of Linguistics and Computer Science, University of Massachusetts Amherst♣Department of Computer Science & Center for Data Science, New York University\{juand\-r, jocelynzhang\}@utexas\.edukerk@umass\.edugdurrett@nyu\.edu
## 1Introduction
Large language models \(LLMs\) can be used in two complementary modes: as*generators*that produce candidate responses, and as*validators*that assess response correctness or felicity\. These two modes are crucial for usages of LLMs such as self\-refinement\(Madaan et al\.,[2023](https://arxiv.org/html/2607.02668#bib.bib14)\)and backtracking during chain\-of\-thought\(Yao et al\.,[2023](https://arxiv.org/html/2607.02668#bib.bib29)\)\. However, these two modes exhibit divergent behavior, reflecting an underlying inconsistency: even frontier LLMs may generate responses with high probability but judge them to be incorrect, or vice\-versa\. Accordingly, past work has examined training LMs to be explicitly generator\-validator \(G\-V\) consistent in the hope that this will also make them more accurate\(West et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib26); Li et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib11); Rodriguez et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib19)\)\. But formalizing this turns out to be surprisingly difficult\. Past approaches close the gap on G\-V correlation\(Rodriguez et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib19)\), but these approaches risk contributing to pathological behaviors like suppression of correct responses or contributing to further inconsistency through contradictory training signals\.
Figure 1:An LLM may generate outputs inconsistent with how it validates them, since low probability options may simply be unlikely to say\. FLORA provides a principled correction for generator and allows for training of G\-V consistent models\.This work aims to improve G\-V consistency through two contributions\. First, we axiomatize G\-V consistency based on a model of a rational probabilistic agent responding to prompts\. Although LLMs are*not*rational agents, this model allows us to examine what the relationship between a generator and a validator should be in cases where the generator may prefer some possible responses over others, but a validator finds them all to be likely\. We derive a theoretical relationship between generator and validator probabilities, which implies adjusting generator scores by subtracting off*incorrect probabilities*, or how likely a response is to be sampled when an*incorrect*response is requested\. Since this value reflects the base likelihood of the response, we call the final quantity the*frequency\-corrected generator score*\.
Figure 2:Generator\-validator gap on a long\-form generation task from IFEval\. A short email with few placeholders \(red\) does not follow the instruction and fails validation but has high generator score\. Another email \(green\) has lower generator score but correctly follows the instruction and is validated as correct\.Second, we use this relationship in a training objective calledFrequency\-corrected Learning of Ordered Rank Alignment\(FLORA\)\. This objective encourages rank\-alignment of validator scores with frequency\-corrected generator scores, representing an approximation of our theoretical G\-V consistency for rational agents\. We fine\-tune LLMs using this objective function in combination with standard losses to ensure generator and validator correctness\.
Our experiments focus on tasks where validators outperform generators\. Our goal is to observe that the generator improves without degradation of the validator\. We particularly focus on*generator AUROC*, or ensuring that the frequency\-corrected generator distribution correctly ranks the set of positive responses above the set of negative responses, even down into the tail of the distribution, unlike past work which focuses on the head\(Li et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib11)\)\. This is a strong check of consistency and of generator correctness\.
We evaluate on three tasks: instruction\-following, coding, and eliciting taxonomic knowledge\. Two of these tasks are long\-form \(several sentences, for IFEval, or Python code\), unlike previous work which focused on short answers of one or a couple words\(Rodriguez et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib19); Li et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib11)\)Across these tasks and three LLMs, we find that our method closes the G\-V gap more strongly than previous work, and leads to improved discriminability of correct and incorrect completions\. Concretely, FLORA improves generator AUROC by up to\+7\.3\+7\.3pp and generator\-validator correlation by up to\+27\+27pp over the strongest prior alignment method, while preserving or improving validator quality\.
## 2Background and Motivation
The extent to which LLMs have consistent internal processing is an important scientific question\(Pres et al\.,[2026](https://arxiv.org/html/2607.02668#bib.bib16)\), and the mismatch between generation and validation is an important kind of inconsistency to measure and repair\. Much of what a model “knows” never shows up in its sampled outputs, because it lives in discriminative judgments rather than in fluent continuations\(Gekhman et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib7)\); this is problematic for evaluations targeting model knowledge\(Wang et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib25); Biderman et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib2)\)\. Better aligning generators with an LLM’s own validation capability is also important for reward modeling\(Yuan et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib30)\)and reranking applications such as mathematical reasoning\(Cobbe et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib6); Lightman et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib12)\), code generation\(Chen et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib5)\), and factual question answering\.
Rodriguez et al\. \([2025](https://arxiv.org/html/2607.02668#bib.bib19)\)assert that the log\-odds of the generator and validator should correlate\. However, there are cases where this relationship cannot hold, arising from two main sources: \(a\) the role of response frequency and \(b\) aleatoric uncertainty \(multiple correct responses\)\.
#### Motivating Example
A well\-known issue with using LLMs to generate responses is surface form ambiguityHoltzman et al\. \([2021](https://arxiv.org/html/2607.02668#bib.bib9)\)\. Given the question*“What do you call it when blood flow to the heart is suddenly blocked?”*,*heart attack*and*myocardial infarction*are both correct answers\. A good validator should score them both highly, yet it would be odd to expect a generator to score both highly, since*heart attack*is the more common expression\.
This becomes more complex when questions have multiple possible correct answers beyond paraphrases\. Consider a question that admits several valid answers, e\.g\.,*“What’s an example of a noble gas?”*\(Figure[1](https://arxiv.org/html/2607.02668#S1.F1)\)\. There are 7 possible correct answers to this question, plus potential different surface forms of elements \(e\.g\.,*“Helium”*vs\.*“He”*\)\. It is unlikely that a generator will actually generate*Helium*and*Oganesson*with equal probability when given this question, despite both being correct\. Generator scores conflate the correctness of an utterance in response to a prompt with the*frequency*of that utterance\.
Figure 3:Model of agent response generation\. We represent a generator𝒑𝑮\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\}as a mixture distribution over latent validity vectors𝐯\\mathbf\{v\}, which indicate the subset of responses a model believes to be the correct response\.π\\piplaces a distribution overyyconditioned on each𝐯\\mathbf\{v\}\.There is no single “perfect” generator in this setting\. One view is that a model should tend towards a uniform distribution over possible optionsZhang et al\. \([2024](https://arxiv.org/html/2607.02668#bib.bib32)\), while another view is that the model’s distribution over valid answers should match the estimated corpus\-frequency distribution\(Tomov et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib22)\)\. A perfect validator, in contrast, should certainly say “Yes” with probability near 1 to each correct option and “No” to incorrect options\.
Figure[2](https://arxiv.org/html/2607.02668#S1.F2)shows this effect in practice on a long\-form generation task from IFEval\. The validator is imperfect, but most correct responses \(green\) outscore most incorrect responses \(red\) on the validator log\-odds \(y\-axis\)\. However, there are incorrect responses that have much higher generator score than correct responses, partially due to being short and generic and sticking close to the prompt\.
This example shows that generator and validator scores should not generally be expected to match, or even to correlate linearly, without controlling for competition and frequency effects\. Next we derive a relationship between them which accounts for both frequency effects and aleatoric uncertainty\.
## 3Problem Formulation
#### Problem Setup
We assume a promptxxand a set of possible responses𝒴\\mathcal\{Y\}\.222Our analysis will require𝒴\\mathcal\{Y\}to be finite in size, but it can be very large:ΣN\\Sigma^\{N\}for a token vocabularyΣ\\Sigmaand someNN\. We do not assume that𝒴\\mathcal\{Y\}is tractable\.For eachy∈𝒴y\\in\\mathcal\{Y\}, we assume that there is a validity labelv\(x,y\)∈\{0,1\}v\(x,y\)\\in\\\{0,1\\\}\. We assume that validity is binary and unambiguous\. We consider an agent, which is an LLM for the purposes of this work, operating in two modes: a generator𝒑𝑮\(y∣x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}, and a validator𝒑𝑽\(v=1∣x,y\)\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(v\{=\}1\\mid x,y\)\}\. For an LLM,pVp\_\{V\}andpGp\_\{G\}are implemented with different prompts\.
#### Properties of Consistency
We would like to choosepVp\_\{V\}andpGp\_\{G\}to be consistent, under the intuition that improving consistency will also improve accuracy\. We can straightforwardly understandpVp\_\{V\}as the LM’s current “belief” thatyyis a correct response toxx\. However, since there are in general many correct responses,pGp\_\{G\}additionally has to model competition between these to act as a generator\.
We define𝐯=\{0,1\}\|𝒴\|\\mathbf\{v\}=\\\{0,1\\\}^\{\|\\mathcal\{Y\}\|\}as a vector of validities associated with each possible response string, assuming that every possible response is either correct or incorrect, even though the agent may be uncertain as to which\. Let𝐯y∈\{0,1\}\\mathbf\{v\}\_\{y\}\\in\\\{0,1\\\}denote the validity assigned to a particular stringyy\. In the context of Figure[1](https://arxiv.org/html/2607.02668#S1.F1), a correct𝐯\\mathbf\{v\}should assign 1 to the 7 noble gases and all ways of writing them down and 0 to all other strings\. Figure[3](https://arxiv.org/html/2607.02668#S2.F3)shows possible𝐯\\mathbf\{v\}values in green\.
Assume that for a promptxx, we have a known and fixed𝐯\\mathbf\{v\}\. We define an agent to be*consistent*with its beliefs𝐯\\mathbf\{v\}if \(1\) its validator returns exactly the responses in𝐯\\mathbf\{v\}; \(2\) its generator only places mass on valid responses:𝒑𝑮\(y∣x\)\>0\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\>0iff𝐯y=1\\mathbf\{v\}\_\{y\}=1\. We call this asupport constraintmotivated by the maxim of quality\(Grice,[1975](https://arxiv.org/html/2607.02668#bib.bib8)\): conditional on its latent belief state𝐯\\mathbf\{v\}, the agent never generates a response it believes to be incorrect\. However, some options \(e\.g\., xenon and krypton\) may be assigned very low probability\.
In practice, agents have uncertainty over𝐯\\mathbf\{v\}\. We model this as a distributionp\(𝐯∣x\)\>0p\(\\mathbf\{v\}\\mid x\)\>0reflecting the agent’s belief over possible response sets givenxx\. Like in real LLMs, we assume that every𝐯\\mathbf\{v\}has nonzero \(but possibly tiny\) probability mass\. \(Note that this distribution is a theoretical construct and cannot be materialized in practice\.\) We can now define*consistency*with respect to this distributionp\(𝐯∣x\)p\(\\mathbf\{v\}\\mid x\)\. A validator is consistent if it satisfies the following relationship:
𝒑𝑽\(y∣x\):=p\(vy=1∣x\)=∑𝐯:vy=1p\(𝐯∣x\)\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}:=p\(v\_\{y\}=1\\mid x\)=\\sum\_\{\\mathbf\{v\}:v\_\{y\}=1\}p\(\\mathbf\{v\}\\mid x\)\(1\)
We represent generation through the use of a distributionπ\(y∣𝐯,x\)\\pi\(y\\mid\\mathbf\{v\},x\)that conditions on𝐯\\mathbf\{v\}\.π\\pi, although intractable to fully represent in practice, effectively describes how mass should be distributed among the correct options of𝐯\\mathbf\{v\}\. To obey the support constraint,π\\pionly assigns nonzero probability to itemsiiwhere𝐯i=1\\mathbf\{v\}\_\{i\}=1\. But within these items,π\\pigoverns which surface form to use for a given meaning, splitting probability over surface form realizations\(Zhang et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib32); Holtzman et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib9); Zhang et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib31)\), and which response to express when multiple are deemed correct\. We define the generator as a mixture model over the𝐯\\mathbf\{v\}:
𝒑𝑮\(y∣x\)\\displaystyle\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}:=∑𝐯p\(𝐯∣x\)π\(y∣𝐯,x\)\\displaystyle=\\sum\_\{\\mathbf\{v\}\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi\(y\\mid\\mathbf\{v\},x\)\(2\)=∑𝐯:vy=1p\(𝐯∣x\)π\(y∣𝐯,x\)\\displaystyle\\phantom\{:\}=\\sum\_\{\\mathbf\{v\}:v\_\{y\}=1\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi\(y\\mid\\mathbf\{v\},x\)This process is depicted in Figure[3](https://arxiv.org/html/2607.02668#S2.F3)\.
### 3\.1Deriving a G\-V relationship
We can now relate the generator and the validator we have defined so far\.
We first define one more quantityπ′\(y∣𝐯,x\)\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\. This is a policy for selecting a response when asked for an*incorrect*one, with the related constraint that it never generates an answer it believes to be correct, i\.e\.,π′\(y∣𝐯,x\)=0\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)=0whenva=1v\_\{a\}=1\. From that, we can define
𝒑𝑮′\(y∣x\)\\displaystyle\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}:=∑𝐯P\(𝐯∣x\)π′\(y∣𝐯,x\)\\displaystyle=\\sum\_\{\\mathbf\{v\}\}P\(\\mathbf\{v\}\\mid x\)\\,\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\(3\)=∑𝐯:vy=0p\(𝐯∣x\)π′\(y∣𝐯,x\)\\displaystyle\\phantom\{:\}=\\sum\_\{\\mathbf\{v\}:v\_\{y\}=0\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)
###### Theorem 1\(Main\)\.
Assume a problemxx, solution space𝒴\\mathcal\{Y\}, and generator𝐩𝐆\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\}and validator𝐩𝐕\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\}as defined previously\. Further assume that0<𝐩𝐕\(y∣x\)<1\{0<\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}<1\}for ally∈𝒴y\\in\\mathcal\{Y\}\. Then
𝒑𝑽\(y∣x\)1−𝒑𝑽\(y∣x\)=𝒑𝑮\(y∣x\)𝒑𝑮′\(y∣x\)⋅r\(y,x\),\\frac\{\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\{1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}=\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\}\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\}\\cdot r\(y,x\),\(4\)where
r\(y,x\):=𝔼\[π′\(y∣𝐯,x\)∣vy=0,x\]𝔼\[π\(y∣𝐯,x\)∣vy=1,x\]\.r\(y,x\):=\\frac\{\\mathbb\{E\}\[\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=0,x\]\}\{\\mathbb\{E\}\[\\pi\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=1,x\]\}\.
#### Proof Sketch
We give intuition for the proof and a full proof in Appendix[A](https://arxiv.org/html/2607.02668#A1)\.
The conditional expectation of the policy for picking a correct response from among the set of correct answers given thataais valid is:
𝔼\[π\(y∣𝐯,x\)∣vy=1,x\]\\displaystyle\\mathbb\{E\}\\big\[\\pi\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=1,x\\big\]\(5\)=∑𝐯:vy=1P\(𝐯∣x\)π\(y∣𝐯,x\)P\(vy=1∣x\)\\displaystyle\\quad=\\frac\{\\sum\_\{\\mathbf\{v\}:v\_\{y\}=1\}P\(\\mathbf\{v\}\\mid x\)\\,\\pi\(y\\mid\\mathbf\{v\},x\)\}\{P\(v\_\{y\}=1\\mid x\)\}=𝒑𝑮\(y∣x\)𝒑𝑽\(y∣x\)\\displaystyle\\quad=\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\}\{\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}
Similarly,
𝔼\[π′\(y∣𝐯,x\)∣vy=0,x\]\\displaystyle\\mathbb\{E\}\\big\[\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=0,x\\big\]\(6\)=∑𝐯:vy=0p\(𝐯∣x\)π′\(y∣𝐯,x\)P\(vy=0∣x\)\\displaystyle\\quad=\\frac\{\\sum\_\{\\mathbf\{v\}:v\_\{y\}=0\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\}\{P\(v\_\{y\}=0\\mid x\)\}=𝒑𝑮′\(y∣x\)1−𝒑𝑽\(y∣x\)\\displaystyle\\quad=\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\}\{1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}
Since0<𝒑𝑽\(y∣x\)<10<\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}<1the main result follows from Eqn[5](https://arxiv.org/html/2607.02668#S3.E5)and[6](https://arxiv.org/html/2607.02668#S3.E6)\.
#### Intuition
First, we note that0<𝒑𝑽\(y∣x\)<10<\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}<1is a mild condition in practice\. We might reasonably expect that an agent backed by a Transformer produces a distributionp\(𝐯∣x\)p\(\\mathbf\{v\}\\mid x\)with support everywhere due to the nature of softmax; no set of answers is ruled out structurally\. Our needed condition follows from this\.
In the limit of zero epistemic uncertainty, these probabilities become very small andppplaces all probability mass on one configuration𝐯⋆∈\{0,1\}\|A\|\\mathbf\{v\}^\{\\star\}\\in\\\{0,1\\\}^\{\|A\|\}\. In this case, the validator probability𝒑𝑽\(y∣x\)\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}will approach 0 or 1, and the generator probabilities become the policies at configuration𝐯⋆\\mathbf\{v\}^\{\\star\}:𝒑𝑮\(y∣x\)=π\(y∣𝐯⋆,x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}=\\pi\(y\\mid\\mathbf\{v\}^\{\\star\},x\)and𝒑𝑮′\(y∣x\)=π′\(y∣𝐯⋆,x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}=\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\}^\{\\star\},x\)\.
Finally, the value ofr\(y,x\)r\(y,x\)is very important to reasoning about the generator\-validator relationship\. The numerator ofrris the expected value of the probability of a response being generated as a*incorrect*response conditioned on it being marked as incorrect in𝐯\\mathbf\{v\}\. This involves marginalizing over all possible configurations of𝐯\\mathbf\{v\}wherevy=0v\_\{y\}=0\. We can think of this as saying: in aggregate, what fraction of theπ′\\pi^\{\\prime\}mass is assigned toyywhen it’s a valid option? Intuitively, this is related to how frequent we are to sayyy: a string that’s simply more frequent will be uttered with higher probability in a larger number of contexts\. The denominator ofrris similar, but for the case of a correct response\.
### 3\.2Mapping Rational Agent Behavior to LLM Behavior
LLMs are not rational agents of the sort in the previous section\. However, we can nevertheless seek to impose the regularity we derived in Equation[4](https://arxiv.org/html/2607.02668#S3.E4)to their behavior if we can approximate𝒑𝑽\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\},𝒑𝑮\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\}, and𝒑𝑮′\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\}\. To do so,𝒑𝑽\(y∣x\)\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\},𝒑𝑮\(y∣x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}and𝒑𝑮′\(y∣x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}can be elicited from an LLM as follows\. Given template functionTV:x,y→promptV\(x,y\)T\_\{V\}:x,y\\rightarrow\\text\{prompt\}\_\{V\}\(x,y\)mapping questions and answers to prompts, and template functionsTG:x→promptG\(x\)T\_\{G\}:x\\rightarrow\\text\{prompt\}\_\{G\}\(x\)andTG′:x→promptG′\(x\)T^\{\\prime\}\_\{G\}:x\\rightarrow\\text\{prompt\}\_\{G\}^\{\\prime\}\(x\), asking for correct and incorrect answers to a given question, respectively:
𝒑𝑽\(y∣x\)\\displaystyle\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}≈PLLM\(Yes∣TV\(x,y\)\),\\displaystyle\\approx P\_\{\\mathrm\{LLM\}\}\\bigl\(\\text\{Yes\}\\mid T\_\{V\}\(x,y\)\\bigr\),𝒑𝑮\(y∣x\)\\displaystyle\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}≈PLLM\(y∣TG\(x\)\),\\displaystyle\\approx P\_\{\\mathrm\{LLM\}\}\\bigl\(y\\mid T\_\{G\}\(x\)\\bigr\),𝒑𝑮′\(y∣x\)\\displaystyle\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}≈PLLM\(y∣TG′\(x\)\)\.\\displaystyle\\approx P\_\{\\mathrm\{LLM\}\}\\bigl\(y\\mid T^\{\\prime\}\_\{G\}\(x\)\\bigr\)\.
There are two sources of error in these approximations\. First, LLMs do not have an internally consistent beliefp\(𝐯∣x\)p\(\\mathbf\{v\}\\mid x\)about the answers to a problem\. Second, even if they did, these prompts do not necessarily elicit rational responses to these\.
Nevertheless, we argue that these approximations are useful\. Empirically, most of the mass for LLMs when prompted for Yes/No answers lies on the*Yes*and*No*tokens, so at least LLMs can follow these instructions\.333In our experiments we aggregate over the tokensYes,yes,YES,␣Yesand␣yesfor the positive class, andNo,no,NO,␣Noand␣nofor the negative class\.\.𝒑𝑮\(y∣x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}follows closely from how LLMs are trained to answer questions or follow instructions\. On the other hand,𝒑𝑮′\(y∣x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}is likely less well\-approximated, since asking for incorrect responses is not where post\-training focuses its efforts; the fidelity of this proxy is an empirical question, especially for smaller or non\-instruction\-tuned models\.
LettingsV\(y∣x\):=log𝒑𝑽\(y∣x\)1−𝒑𝑽\(y∣x\)s\_\{V\}\(y\\mid x\):=\\log\\frac\{\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\{1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}be the validator log\-odds, we then have
sV\(y∣x\)=log𝒑𝑮\(y∣x\)−log𝒑𝑮′\(y∣x\)\+logr\(y∣x\)\\begin\{split\}s\_\{V\}\(y\\mid x\)=\{\}&\\log\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\-\\log\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\\\\ &\{\}\+\\log r\(y\\mid x\)\\end\{split\}\(7\)
Unfortunately,logr\(y∣x\)\\log r\(y\\mid x\)depends on conditional expectations over quantities that we cannot directly observe\. Intuitively, what this term measures is the ratio of the probability of generatingyywhen prompted for a wrong response \(and the model thinksyyis incorrect\) and the probability of generatingyywhen prompted for a correct response \(and the model thinksyyis correct\)\. Assuming thatπ\\piandπ′\\pi^\{\\prime\}are both capturing the frequency ofyy, and this frequency is not biased by response correctness, this ratio can be approximated as 1\.
As a result, disregardingrr, Eqn[7](https://arxiv.org/html/2607.02668#S3.E7)motivates the following adjusted generator score:
sAdj\(y∣x\):=log𝒑𝑮\(y∣x\)−log𝒑𝑮′\(y∣x\)\.s\_\{\\text\{Adj\}\}\(y\\mid x\)\\;:=\\;\\log\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\-\\log\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\.\(8\)By construction,sAdjs\_\{\\text\{Adj\}\}tracks the validator log\-oddssVs\_\{V\}up to the residuallogr\\log r, and so should correlate better withsVs\_\{V\}than the raw generator scorelogpG\\log p\_\{G\}does\. In the next section, we operationalizesAdjs\_\{\\text\{Adj\}\}as a training objective for our model\.
## 4Training for Frequency Correction
#### Data Condition
We consider a training setting where we have a dataset𝒟=\{\(xi,\{\(yij,vij\)\}\)\}\\mathcal\{D\}=\\\{\(x\_\{i\},\\\{\(y\_\{ij\},v\_\{ij\}\)\\\}\)\\\}\. That is, eachxix\_\{i\}is paired with a set of outputs\{\(yij,vij\)\}\\\{\(y\_\{ij\},v\_\{ij\}\)\\\}where theyijy\_\{ij\}are the outputs themselves and thevijv\_\{ij\}are correctness labels\. We operate in a semi\-supervised setting wherevijv\_\{ij\}may not be available for all\(xi,yij\)\(x\_\{i\},y\_\{ij\}\)pairs\.
As our training involves imposing consistency, we operate over pairs of points \(e\.g\., similar to*hydrogen*and*Rn*in Figure[1](https://arxiv.org/html/2607.02668#S1.F1)\)\. We define a labeled set𝒯=\{\(\(xi,yi,vi\),\(xj,yj,vj\)\)\}\\mathcal\{T\}=\\\{\(\(x\_\{i\},y\_\{i\},v\_\{i\}\),\(x\_\{j\},y\_\{j\},v\_\{j\}\)\)\\\}of pairs with veracity labels\. We define an unlabeled set𝒰=\{\(\(xi,yi\),\(xj,yj\)\)\}\\mathcal\{U\}=\\\{\(\(x\_\{i\},y\_\{i\}\),\(x\_\{j\},y\_\{j\}\)\)\\\}of pairs without labels\.
Furthermore, throughout this work we requirexi=xjx\_\{i\}=x\_\{j\}; we sample points*within the same prompt*\. We also require\|pV\(yj∣xj\)−pV\(yi∣xi\)\|\>δ\|p\_\{V\}\(y\_\{j\}\\mid x\_\{j\}\)\-p\_\{V\}\(y\_\{i\}\\mid x\_\{i\}\)\|\>\\deltato avoid noise from too\-close validator pairs, following past work\(Gekhman et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib7); Rodriguez et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib19)\)\.
#### Training Objective
We use a loss designed to do two things: \(1\) teach the generator to rank pairs in the same way as the validator, and \(2\) use the labeled examples to improve the discriminability of generator and validator scores\. We train using the following loss, where we letyj=y\+y\_\{j\}=y^\{\+\}andyi=y−y\_\{i\}=y^\{\-\}denote the higher\- and lower\-ranked completions under the validator:
ℒ=\\displaystyle\\mathcal\{L\}=ℒpref\(xi,yi,xj,yj\)\+\\displaystyle\\mathcal\{L\}\_\{\\text\{pref\}\}\(x\_\{i\},y\_\{i\},x\_\{j\},y\_\{j\}\)\+λG\(viℒNLL\(yi∣xi\)\+vjℒNLL\(yj∣xj\)\)\+\\displaystyle\\lambda\_\{G\}\(v\_\{i\}\\mathcal\{L\}\_\{\\text\{NLL\}\}\(y\_\{i\}\\mid x\_\{i\}\)\+v\_\{j\}\\mathcal\{L\}\_\{\\text\{NLL\}\}\(y\_\{j\}\\mid x\_\{j\}\)\)\+λV\(ℒNLL\(vi∣xi,yi\)\+ℒNLL\(vj∣xj,yj\)\)\\displaystyle\\lambda\_\{V\}\(\\mathcal\{L\}\_\{\\text\{NLL\}\}\(v\_\{i\}\\mid x\_\{i\},y\_\{i\}\)\+\\mathcal\{L\}\_\{\\text\{NLL\}\}\(v\_\{j\}\\mid x\_\{j\},y\_\{j\}\)\)\(9\)
Theℒpref\\mathcal\{L\}\_\{\\text\{pref\}\}loss is always active, whileℒNLL\\mathcal\{L\}\_\{\\text\{NLL\}\}terms are only active for labeled examples\. The generator terms only apply to positive examples \(i\.e\.,vi=1v\_\{i\}=1orvj=1v\_\{j\}=1\) and the validator terms are only present for labeled examples, which have ground truth Yes/No completions\.
#### Preference Loss
We use a pairwise logistic preference loss given by
ℒpref\(x,y−,y\+\)=−logσ\(s\(y\+\|x\)−s\(y−\|x\)\)\\mathcal\{L\}\_\{\\text\{pref\}\}\(x,y^\{\-\},y^\{\+\}\)=\-\\log\\sigma\\\!\\left\(s\(y^\{\+\}\|x\)\-s\(y^\{\-\}\|x\)\\right\)\(10\)whereσ\\sigmais the logistic sigmoid andsG\(y∣x\)s\_\{G\}\(y\\mid x\)is the generator’s log\-likelihood of completionyygiven promptxx,
s\(y∣x\)=∑t=1\|y\|logPLLM\(yt∣x,y<t\)\.s\(y\\mid x\)=\\sum\_\{t=1\}^\{\|y\|\}\\log P\_\{\\text\{LLM\}\}\(y\_\{t\}\\mid x,y\_\{<t\}\)\.\(11\)
To account for completion frequency, we consider two empirical estimators of the adjusted scoresAdjs\_\{\\text\{Adj\}\}from Eqn[8](https://arxiv.org/html/2607.02668#S3.E8)\(§[3\.1](https://arxiv.org/html/2607.02668#S3.SS1)\)\. The first,NegFC, directly substitutes a negative\-prompt elicitation forpG′p\_\{G^\{\\prime\}\}:
sNeg\(y∣x\)=sG\(y∣x\)−logPLLM\(y∣TG′\(x\)\)\.s\_\{\\text\{Neg\}\}\(y\\mid x\)=s\_\{G\}\(y\\mid x\)\-\\log P\_\{\\text\{LLM\}\}\(y\\mid T^\{\\prime\}\_\{G\}\(x\)\)\.\(12\)The second,Unconditional FC, uses the unconditional log probability\(Meister et al\.,[2023](https://arxiv.org/html/2607.02668#bib.bib15)\)as a cheaper proxy,
sPMI\(y∣x\)=sG\(y∣x\)−logPLLM\(y\),s\_\{\\text\{PMI\}\}\(y\\mid x\)=s\_\{G\}\(y\\mid x\)\-\\log P\_\{\\text\{LLM\}\}\(y\),\(13\)which captures a similar frequency\-penalization effect without requiring an elicitation promptTG′T^\{\\prime\}\_\{G\}for incorrect answers\.
#### Sampling strategy
Finally, we ensure that the preference losses are consistent with the labels\. For any pair\(x,y−\),\(x,y\+\)\(x,y^\{\-\}\),\(x,y^\{\+\}\)wheresV\(y\+∣x\)\>sV\(y−\|x\)\+δs\_\{V\}\(y^\{\+\}\\mid x\)\>s\_\{V\}\(y^\{\-\}\|x\)\+\\delta, the preference loss will push the generator score of\(x,y−\)\(x,y^\{\-\}\)down and\(x,y\+\)\(x,y^\{\+\}\)up; this is incorrect ify−y^\{\-\}is a positive example ory\+y^\{\+\}is a negative example\. In addition, ify−y^\{\-\}is positive the preference loss and generatorℒNLL\\mathcal\{L\}\_\{\\text\{NLL\}\}will compete in opposite directions\. To avoid these problems, we only keep a labeled\(x,y−\)\(x,y^\{\-\}\)when it is negative, and similarly we only keep a labeled\(x,y\+\)\(x,y^\{\+\}\)when it is positive\.
## 5Experimental Setup
### 5\.1Tasks and Datasets
We evaluate on three datasets which we adapted from previous work to test our method on instruction following, coding, and conceptual knowledge\. We use IFEval\(Zhou et al\.,[2023](https://arxiv.org/html/2607.02668#bib.bib33)\)for instruction following, HumanEvalChen et al\. \([2021](https://arxiv.org/html/2607.02668#bib.bib5)\)for Python coding, and a mix of datasets to probe for knowledge of taxonomic category relations \(hyponymy\)\(Rosch,[1975](https://arxiv.org/html/2607.02668#bib.bib20); Banks and Connell,[2023](https://arxiv.org/html/2607.02668#bib.bib1); Stoinski et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib21); Castro et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib4); Van Overschelde et al\.,[2004](https://arxiv.org/html/2607.02668#bib.bib24); Uyeda and Mandler,[1980](https://arxiv.org/html/2607.02668#bib.bib23)\)\. Each dataset has a different notion of correctness, detailed below\. These datasets are described here, with further details on datasets given in Appendix[C](https://arxiv.org/html/2607.02668#A3)\. Prompt templates for each of these tasks are shown in Appendix[D](https://arxiv.org/html/2607.02668#A4)\.
#### Instruction Following
This task uses prompts from IFEval\(Zhou et al\.,[2023](https://arxiv.org/html/2607.02668#bib.bib33)\), e\.g\., “*Write a poem that’s at least 350 words about the beauty of eucalyptus trees and their many uses\.*” which specify explicit content and format constraints\. A correct response should follow all constraints in the prompt, while an incorrect response violates one or more\. We generate, using GPT\-4o mini, correct responses via paraphrases of the original prompt and incorrect responses by introducing violations \(e\.g\., missing required elements or breaking constraints\)\. Labels are assigned using rule\-based checks for verifiable constraints together with LLM\-as\-a\-judge evaluations for content correctness where deterministic verification is not possible\. We use 79 prompts for the training set and 20 for the test set\.
#### Coding
This task uses prompts from the HumanEval\(Chen et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib5)\)benchmark of Python programming exercises\. Each prompt asks to complete a Python function to solve a given problem\. Here correct responses are those which pass all the unit tests, while incorrect ones fail one or more\. We use various LLMs to generate correct and incorrect responses\. To stress\-test our method with examples that are correct but unlikely, we systematically convert the variables of all the correct solutions to uppercase\.
#### Hyponymy
This task consists of giving exemplars to categories, i\.e\., answering “An example of a X is a \_\_” for various nouns X\. We evaluated on the ten categories used in the original experiments of\(Rosch,[1975](https://arxiv.org/html/2607.02668#bib.bib20)\), supplemented with GPT\-5\-generated incorrect completions in order to obtain a balanced test set\. We decided on a cutoff between items that are members of the category while those that are not\. While there is some subjectivity in this decision, our cutoff is validated by the fact that LLMs with validator prompts can successfully distinguish between the positive and negative classes \(Table[5](https://arxiv.org/html/2607.02668#A2.T5)\)\.
For all three tasks, we evaluate on a held\-out set of prompts\. For IFEval and HumanEval, we randomly split the data\. For Hyponymy we train on a different set of categories than the ten in\(Rosch,[1975](https://arxiv.org/html/2607.02668#bib.bib20)\)\. Further details of the dataset composition and construction process is given in Appendix[C](https://arxiv.org/html/2607.02668#A3)\.
### 5\.2Target Models for G\-V Alignment
As we defined the dataset at the beginning of Section[4](https://arxiv.org/html/2607.02668#S4), our target models for G\-V consistency do not necessarily have to be those from which responses were generated\. In fact, having a range of responses \(as we generated for each dataset\) that we require G\-V consistency on allows us to stress\-test each model beyond the mode of the distribution that it generates\.
We use a range of models including those in the Gemma family, namely Gemma\-2\-9b\-it \(G2\-9b\-it\) and Gemma\-4\-31b\-it \(G4\-31b\-it\), and Qwen family, namely Qwen\-3\.5\-9B \(Q3\.5\-9b\)\. For each task, we select models in an appropriate performance range: for instance, on IFEval, we only use larger instruction\-tuned models because smaller models are generally not capable enough at the task\. For each task and model combination, we only proceed with models where validator AUROC is over 65%\); this eliminates Gemma\-2\-9b\-it from the HumanEval experiments\.
### 5\.3Evaluation
For all methods, we evaluate both using the raw generator scoressG\(y∣x\)=logPLLM\(y∣x\)s\_\{G\}\(y\\mid x\)=\\log P\_\{LLM\}\(y\\mid x\)as well as the two corrected forms: frequency\-corrected scoressPMI=sG\(y∣x\)−logPLLM\(y\)s\_\{\\text\{PMI\}\}=s\_\{G\}\(y\\mid x\)\-\\log P\_\{LLM\}\(y\), and negative prompting\-corrected scoressNeg=sG\(y∣x\)−sG′\(y∣x\)s\_\{\\text\{Neg\}\}=s\_\{G\}\(y\\mid x\)\-s\_\{G\}^\{\\prime\}\(y\\mid x\)\. All metrics are computed independently for each prompt, and averaged over prompts\.
#### Accuracy
We evaluate the validator and generator performance at discriminating correct and incorrect answers through AUROC scores of the associated classifiers:𝒑𝑮\(y∣x\)\>τG\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\>\\tau\_\{G\}andsV\(y∣x\)\>τVs\_\{V\}\(y\\mid x\)\>\\tau\_\{V\}for thresholdsτG\\tau\_\{G\}andτV\\tau\_\{V\}\. Each defines an ROC curve and we refer to the corresponding AUROC scores as𝐑𝐎𝐂𝐆\\mathbf\{ROC\_\{G\}\}and𝐑𝐎𝐂𝐕\\mathbf\{ROC\_\{V\}\}\. For the generator, we use one of several different possible generator scores:𝒑𝑮\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\},sPMIs\_\{\\text\{PMI\}\}orsNegs\_\{\\text\{Neg\}\}\. We additionally measure thevalidator accuracyatτV=0\\tau\_\{V\}=0in order to measure how well\-calibrated the validator is\.
We focus primarily on𝐑𝐎𝐂𝐆\\mathbf\{ROC\_\{G\}\}; our focus in this work is on tasks where the validator is stronger than the generator and can be used to improve the generator’s performance\.
#### Correlations
FollowingRodriguez et al\. \([2025](https://arxiv.org/html/2607.02668#bib.bib19)\), we measure the Pearson correlationρ\\rhobetween the validator and \(possibly\-corrected\) generator scores over the full set of candidate answers\.
Table 1:Eval\-time frequency correction applied to Gemma\-2\-9b\-it, except HumanEval, which uses Gemma\-4\-31b\-it\. Both NegFC and Unconditional FC improve over the raw generator score, but neither dominates across all tasks\. Subsequent sections apply the per\-task best eval\-time correction to all methods for a fair comparison\.
### 5\.4Baselines
We compare our method against the following baselines in addition to theBase\(untrained\) model:
#### SFT
We sample pairs of completions as in RankAlign, but use a Supervised Fine\-Tuning \(SFT\) loss, i\.e\., only the negative log\-likelihood terms from Eqn\.[9](https://arxiv.org/html/2607.02668#S4.E9)\.
#### Consistency FT
Li et al\. \([2024](https://arxiv.org/html/2607.02668#bib.bib11)\)SFT only on the subset of examples where generator and validator agree\. We use the version fromRodriguez et al\. \([2025](https://arxiv.org/html/2607.02668#bib.bib19)\), which measures agreement via𝟙lG\(z,yA\)\>tG=𝟙lV\(z,yA\)\>tV\\mathds\{1\}\_\{l\_\{G\}\(z,y\_\{A\}\)\>t\_\{G\}\}=\\mathds\{1\}\_\{l\_\{V\}\(z,y\_\{A\}\)\>t\_\{V\}\}wheretGt\_\{G\}andtVt\_\{V\}are thresholds set as the average generator and validator scores over the dataset\.
#### RankAlign
We use the preference loss of RankAlignRodriguez et al\. \([2025](https://arxiv.org/html/2607.02668#bib.bib19)\)\. Compared to our method this \(1\) does not apply any frequency correction during training; \(2\) does not enforce sampled pairs to have the same prompt during training, \(3\) does not use NLL terms\.
## 6Results
### 6\.1Frequency Correction Helps at Test Time
Frequency correction improves generator–validator consistency even before training\.
We first ask whether frequency correction, applied at test time, improves the generator’s ability to discriminate correct from incorrect answers and its consistency with the model’s own validator\. Table[1](https://arxiv.org/html/2607.02668#S5.T1)compares the raw generator scoresGs\_\{G\}with the two corrected variants introduced in Section[4](https://arxiv.org/html/2607.02668#S4): NegTC \(sNegs\_\{\\text\{Neg\}\}\), which subtracts the log probability when using the negative prompt, and Unconditional FC \(sPMIs\_\{\\text\{PMI\}\}\), which subtracts the unconditional log probability of the completion\.
Both corrections improve ROCGandρ\\rhoover the raw generator score on all tasks\. This indicates that some amount of apparent G\-V gap goes away when using these corrected terms, indicating that they form a better starting point for improving G\-V consistency and generator discriminability\.
### 6\.2Main Results
FLORA outperforms prior consistency alignment methods on both ROCGandρ\\rho\.
Table[2](https://arxiv.org/html/2607.02668#S6.T2)compares methods on ROCGandρ\\rhofor IFEval, HumanEval, and Hyponymy datasets\. All methods are evaluated with the per\-task best eval\-time correction identified in Section[6\.1](https://arxiv.org/html/2607.02668#S6.SS1)\. Full results with standard errors are given in Table[4](https://arxiv.org/html/2607.02668#A2.T4)in Appendix[B](https://arxiv.org/html/2607.02668#A2)\.
Table 2:Main results across tasks and models\.ROCG\(generator AUROC\) measures generator discriminability, andρ\\rhomeasures generator–validator Pearson correlation\)\. All methods use the per\-task best eval\-time correction\.FLORA achieves a higherROCG\\text\{ROC\}\_\{G\}than the other baselines in five out of six settings \(tying RankAlign on Hyponymy with Qwen\-3\.5\-9B\), with gains ranging from \+0\.9 to \+7\.3pp over the next\-best method\. G\-V consistency improves by a wider margin, withρ\\rho\+27pp over the next\-best method on both IFEval with Gemma\-2\-9b\-it \(FLORA\-PMI: 62\.8 vs\. 35\.5 for Base\) and HumanEval \(FLORA\-Neg: 76\.3 vs\. 49\.0 for Consistency FT\)\. On the two Hyponymy settings, FLORA is competitive on both metrics, trailing RankAlign by at most 1\.4pp onρ\\rho\.
#### FLORA Preserves Validator Quality
Improvements in G\-V consistency do not come at the cost of validator quality\.
A standard concern with preference\-style training is*likelihood displacement*\(Razin et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib18); Xu et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib28)\): the model may appear better aligned only because its validator has degraded\. To rule this out, we measure the validator accuracy and AUROC\. Table[5](https://arxiv.org/html/2607.02668#A2.T5)shows the validator accuracies per\-method \(ROCV\\text\{ROC\}\_\{V\}andAccV\\text\{Acc\}\_\{V\}\)\.
Figure 4:Change in validator accuracy \(ΔAccV\\Delta\\,\\text\{Acc\}\_\{V\}, Method−\-Base, in percentage points\) after FLORA training\.Our primary aim is for validator accuracy not to degrade after FLORA training\.However, FLORA actually preserves or improves validator accuracy across all five settings, with the largest gains on HumanEval \(\+14–17pp\) and IFEval/G2\-9b\-it \(\+4–8pp\)\.
#### The effect of frequency correction
Table[3](https://arxiv.org/html/2607.02668#S6.T3)ablates the frequency correction term applied*during training*\. Here the “w/o typ\. corr\.” variant trains on the raw generator log\-likelihood\. Both variants apply the same eval\-time correction, so the comparison isolates the effect of frequency correction during training\. Removing it hurts both ROCGandρ\\rhoon every task, with the largest drops on HumanEval \(−6\.3\-6\.3pp ROCG,−16\.1\-16\.1ppρ\\rho\)\.
Table 3:Removing frequency correction during training \(shown here: Gemma\-2\-9b\-it\) has a detrimental effect on both G\-V consistency and generator performance\. Frequency correction is applied at test time in all cases\.
### 6\.3Qualitative Analysis of Alignment
Figure 5:Generator score versus validator log\-odds for candidate responses to a single IFEval prompt \(base Gemma\-2\-9b\-it; green: correct, red: incorrect\)\. Each panel plots that model’s own validator on theyy\-axis\.*Left:*the raw generator scorelogpG\\log p\_\{G\}is uncorrelated with correctness and does not separate correct from incorrect completions \(Gen AUROC43\.143\.1,ρ=−\.24\\rho\{=\}\{\-\}\.24\)\.*Middle:*applying the frequency \(PMI\) correction at test time already recovers most of the alignment \(AUROC85\.385\.3,ρ=\.64\\rho\{=\}\.64\)\.*Right:*after FLORA\-PMI training, the corrected generator score ranks correct above incorrect completions and tracks the validator \(AUROC90\.890\.8,ρ=\.75\\rho\{=\}\.75\)\.Across many prompts, we observe cleaner consistency between generator scores and validator scores after applying our method\. Figure[5](https://arxiv.org/html/2607.02668#S6.F5)illustrates a case of this for a prompt in IFEval and a set of possible completions\. X\-axes in the left panel are the raw generator scores, while x\-axes in the right two panel show the corrected generator scoressPMIs\_\{\\mathrm\{PMI\}\}\. For the base model \(left\), the relationship between generator scores and validator scores is noisy: some high\-likelihood generations receive low validator scores and vice versa, indicating a misalignment between what the generator prefers and what the validator deems correct\. Using frequency correction for the generator scores at test time improves both generator\-validator correlation and generator ROC \(middle panel\)\. After training, \(right panel\) the relationship between generator and validator scores becomes even more structured, with a clearer trend in which higher generator scores correspond to higher validator scores\.
## 7Related Work
#### Generator\-Validator Gap\.
Generator\-validator gaps have been discussed in a number of contexts, including\(West et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib26)\),\(Li et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib11)\), andRodriguez et al\. \([2025](https://arxiv.org/html/2607.02668#bib.bib19)\)\. Our work builds on these notions and is the first to derive a principled relationship between generator probabilities and validator log\-odds, which\(Rodriguez et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib19)\)treated heuristically\.
#### Preference Learning\.
The ranking\-based preference objective we use is similar to DPORafailov et al\. \([2023](https://arxiv.org/html/2607.02668#bib.bib17)\)\. In our case, the validator produces preferences for the generator, broadly related to self\-rewarding LMsYuan et al\. \([2024](https://arxiv.org/html/2607.02668#bib.bib30)\)\. However, rather than a general post\-training method, we view FLORA as a mechanism to achieve*consistency*in an LLMs’ predictions\.
#### Frequency correction\.
Past work has explored similar ideas in frequency correction\.Holtzman et al\. \([2021](https://arxiv.org/html/2607.02668#bib.bib9)\)correct for frequency in multiple\-choice settings\. Their domain\-conditional PMI is similar to our PMI objective, but it is motivated as a heuristic\.Liu et al\. \([2021](https://arxiv.org/html/2607.02668#bib.bib13)\)uses the idea of “anti\-experts” for controlled text generation, which gives rise to a similar form \(subtracting log probabilities from the anti\-expert\)\.Li et al\. \([2023](https://arxiv.org/html/2607.02668#bib.bib10)\)then extend this idea to contrasting strong and weak language models\. However, none of these approaches targets consistency per se\.
#### Theoretical Formulation\.
We are not the first to analyze LLM behavior by decomposing output distributions over a latent variable\.Xie et al\. \([2022](https://arxiv.org/html/2607.02668#bib.bib27)\)formulate in\-context learning as posterior inference over a latent document concept, andBigelow et al\. \([2025](https://arxiv.org/html/2607.02668#bib.bib3)\)similarly adopt a latent variable formulation to unify in\-context learning and activation steering\. Our derivation in §[3\.1](https://arxiv.org/html/2607.02668#S3.SS1)uses a similar idea, but applies it to a different question: the discrepancy between a single model’s generator and validator distributions on a given prompt\.
## 8Conclusion
In this paper, we presented FLORA, a method for improving LLMs via validator\-to\-generator alignment\. We derive a relationship between generator probabilities and validator log\-odds based on a model of how a rational agent might generate a response\. This relationship suggests a frequency\-based correction factor\. When trained to align validator and generator with this factor, we see improved results across taxonomic categorization \(Hyponymy\) and two long\-form generation problems, IFEval and HumanEval\.
## Limitations
One limitation of this work is a gap between the theoretical motivation and practical implementation\. We rely on an assumption that the intractablerrterm in Equation[4](https://arxiv.org/html/2607.02668#S3.E4)is close to 1\. Although we believe this is a reasonable assumption as discussed in the text, it is difficult to validate this asrris intractable to estimate even on relatively simple examples\. Nevertheless, we see that our theoretically\-derived training objective performs well empirically, lending credence to our method\.
A second limitation is the focus on evaluating fixed sets of outputs\. We focus on evaluating the tails of the model’s distribution: samples which are potentially low likelihood but which we still believe the model should rank correctly\. We believe this setting is relevant, as LLM judges are frequently asked to reckon about fixed responses from other models, and various kinds of self\-consistency or best\-of\-N approaches use a related kind of scoring\. However, we do not demonstrate an impact on the 1\-best answers produced by greedy decoding\.
Furthermore, our framework assumes that correctness is a binary notion\. Relaxing this to accommodate partially\-correct answers would be an interesting direction for future\.
Finally, we note that evaluating within\-prompt rather than across a diverse set of prompts prevents us from evaluating directly on the datasets used to measure the G\-V gap in\(Rodriguez et al\.,[2025](https://arxiv.org/html/2607.02668#bib.bib19)\), because most prompts in that work only have a small set of completions\. To still enable a comparison, we re\-implement RankAlign as a baseline and evaluate it on our datasets alongside FLORA \(Tables[4](https://arxiv.org/html/2607.02668#A2.T4),[5](https://arxiv.org/html/2607.02668#A2.T5)\)\.
## Acknowledgments
We thank Jacob Andreas for insightful discussion around this work\. This work was supported by NSF CAREER Award IIS\-2145280, NSF grant IIS\-2433071, the NSF AI Institute for Foundations of Machine Learning \(IFML\), and the NSF under Cooperative Agreement 2421782 and the Simons Foundation grant MPS\-AI\-00010515 awarded to the NSF\-Simons AI Institute for Cosmic Origins — CosmicAI,[https://www\.cosmicai\.org/](https://www.cosmicai.org/)\. We also thank members of the TAUR Lab for helpful feedback\.
## References
- Banks and Connell \(2023\)Briony Banks and Louise Connell\. 2023\.Category production norms for 117 concrete and abstract categories\.*Behavior Research Methods*, 55\(3\):1292–1313\.
- Biderman et al\. \(2024\)Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Jeffrey Hsu, Mimansa Jaiswal, Wilson Y\. Lee, Haonan Li, and 11 others\. 2024\.[Lessons from the Trenches on Reproducible Evaluation of Language Models](https://doi.org/10.48550/ARXIV.2405.14782)\.*CoRR*, abs/2405\.14782\.
- Bigelow et al\. \(2025\)Eric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman, Tomer Ullman, Hidenori Tanaka, and Ekdeep Singh Lubana\. 2025\.Belief dynamics reveal the dual nature of in\-context learning and activation steering\.*arXiv preprint arXiv:2511\.00617*\.
- Castro et al\. \(2021\)Nichol Castro, Taylor Curley, and Christopher Hertzog\. 2021\.Category norms with a cross\-sectional sample of adults in the united states: Consideration of cohort, age, and historical effects on semantic categories\.*Behavior research methods*, 53\(2\):898–917\.
- Chen et al\. \(2021\)Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others\. 2021\.[Evaluating Large Language Models Trained on Code](https://arxiv.org/abs/2107.03374)\.*CoRR*, abs/2107\.03374\.
- Cobbe et al\. \(2021\)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\. 2021\.[Training verifiers to solve math word problems](https://arxiv.org/abs/2110.14168)\.*CoRR*, abs/2110\.14168\.
- Gekhman et al\. \(2025\)Zorik Gekhman, Eyal Ben\-David, Hadas Orgad, Eran Ofek, Yonatan Belinkov, Idan Szpektor, Jonathan Herzig, and Roi Reichart\. 2025\.[Inside\-out: Hidden factual knowledge in llms](https://doi.org/10.48550/ARXIV.2503.15299)\.In*Proceedings of the Conference on Language Modeling \(COLM\)*\.
- Grice \(1975\)H\. P\. Grice\. 1975\.[Logic and conversation](http://www.ucl.ac.uk/ls/studypacks/Grice-Logic.pdf)\.In Peter Cole and Jerry L\. Morgan, editors,*Syntax and Semantics: Vol\. 3: Speech Acts*, pages 41–58\. Academic Press, New York\.
- Holtzman et al\. \(2021\)Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer\. 2021\.[Surface form competition: Why the highest probability answer isn’t always right](https://doi.org/10.18653/v1/2021.emnlp-main.564)\.In*Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing*, pages 7038–7051, Online and Punta Cana, Dominican Republic\. Association for Computational Linguistics\.
- Li et al\. \(2023\)Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis\. 2023\.[Contrastive Decoding: Open\-ended Text Generation as Optimization](https://doi.org/10.18653/v1/2023.acl-long.687)\.In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 12286–12312, Toronto, Canada\. Association for Computational Linguistics\.
- Li et al\. \(2024\)Xiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto, and Percy Liang\. 2024\.[Benchmarking and improving generator\-validator consistency of language models](https://openreview.net/forum?id=phBS6YpTzC)\.In*The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024*\. OpenReview\.net\.
- Lightman et al\. \(2024\)Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\. 2024\.[Let’s verify step by step](https://openreview.net/forum?id=v8L0pN6EOi)\.In*The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024*\. OpenReview\.net\.
- Liu et al\. \(2021\)Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A\. Smith, and Yejin Choi\. 2021\.[DExperts: Decoding\-Time Controlled Text Generation with Experts and Anti\-Experts](https://doi.org/10.18653/v1/2021.acl-long.522)\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 6691–6706, Online\. Association for Computational Linguistics\.
- Madaan et al\. \(2023\)Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark\. 2023\.[Self\-refine: Iterative refinement with self\-feedback](http://papers.nips.cc/paper_files/paper/2023/hash/91edff07232fb1b55a505a9e9f6c0ff3-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023*\.
- Meister et al\. \(2023\)Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell\. 2023\.[Locally typical sampling](https://doi.org/10.1162/TACL_A_00536)\.*Trans\. Assoc\. Comput\. Linguistics*, 11:102–121\.
- Pres et al\. \(2026\)Itamar Pres, Belinda Z Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, and Jacob Andreas\. 2026\.[Position: It’s time to optimize for self\-consistency](https://time-for-consistency.github.io/assets/pdfs/Consistency_Pos_Paper.pdf)\.
- Rafailov et al\. \(2023\)Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D\. Manning, Stefano Ermon, and Chelsea Finn\. 2023\.[Direct preference optimization: Your language model is secretly a reward model](http://papers.nips.cc/paper_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023*\.
- Razin et al\. \(2025\)Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen, Sanjeev Arora, and Boris Hanin\. 2025\.[Unintentional unalignment: Likelihood displacement in direct preference optimization](https://openreview.net/forum?id=uaMSBJDnRv)\.In*The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025*\. OpenReview\.net\.
- Rodriguez et al\. \(2025\)Juan Diego Rodriguez, Wenxuan Ding, Katrin Erk, and Greg Durrett\. 2025\.[RankAlign: A Ranking View of the Generator\-Validator Gap in Large Language Models](https://arxiv.org/abs/2504.11381)\.In*Proceedings of the Conference on Language Modeling \(COLM\)*\.
- Rosch \(1975\)Eleanor Rosch\. 1975\.Cognitive representations of semantic categories\.*Journal of experimental psychology: General*, 104\(3\):192\.
- Stoinski et al\. \(2024\)Laura M Stoinski, Jonas Perkuhn, and Martin N Hebart\. 2024\.Thingsplus: New norms and metadata for the things database of 1854 object concepts and 26,107 natural object images\.*Behavior Research Methods*, 56\(3\):1583–1603\.
- Tomov et al\. \(2025\)Tim Tomov, Dominik Fuchsgruber, Tom Wollschläger, and Stephan Günnemann\. 2025\.[The Illusion of Certainty: Uncertainty quantification for LLMs fails under ambiguity](https://doi.org/10.48550/ARXIV.2511.04418)\.*CoRR*, abs/2511\.04418\.
- Uyeda and Mandler \(1980\)Katherine M Uyeda and George Mandler\. 1980\.Prototypicality norms for 28 semantic categories\.*Behavior Research Methods & Instrumentation*, 12\(6\):587–595\.
- Van Overschelde et al\. \(2004\)James P Van Overschelde, Katherine A Rawson, and John Dunlosky\. 2004\.Category norms: An updated and expanded version of the norms\.*Journal of memory and language*, 50\(3\):289–335\.
- Wang et al\. \(2024\)Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber\-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank\. 2024\.[“My Answer is C”: First\-Token Probabilities Do Not Match Text Answers in Instruction\-Tuned Language Models](https://doi.org/10.18653/v1/2024.findings-acl.441)\.In*Findings of the Association for Computational Linguistics: ACL 2024*, pages 7407–7416, Bangkok, Thailand\. Association for Computational Linguistics\.
- West et al\. \(2024\)Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D\. Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Raghavi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, and Yejin Choi\. 2024\.[The Generative AI Paradox: "What It Can Create, It May Not Understand"](https://openreview.net/forum?id=CF8H8MS5P8)\.In*The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024*\. OpenReview\.net\.
- Xie et al\. \(2022\)Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma\. 2022\.[An explanation of in\-context learning as implicit bayesian inference](https://openreview.net/forum?id=RdJVFCHjUMI)\.In*The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\-29, 2022*\. OpenReview\.net\.
- Xu et al\. \(2024\)Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim\. 2024\.[Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation](https://proceedings.mlr.press/v235/xu24t.html)\.In*Forty\-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21\-27, 2024*, Proceedings of Machine Learning Research, pages 55204–55224\. PMLR / OpenReview\.net\.
- Yao et al\. \(2023\)Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan\. 2023\.[Tree of thoughts: Deliberate problem solving with large language models](http://papers.nips.cc/paper_files/paper/2023/hash/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html)\.In*Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023*\.
- Yuan et al\. \(2024\)Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston\. 2024\.[Self\-rewarding language models](https://proceedings.mlr.press/v235/yuan24d.html)\.In*Forty\-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21\-27, 2024*, Proceedings of Machine Learning Research, pages 57905–57923\. PMLR / OpenReview\.net\.
- Zhang et al\. \(2025\)Yiming Zhang, Harshita Diddee, Susan Holm, Hanchen Liu, Xinyue Liu, Vinay Samuel, Barry Wang, and Daphne Ippolito\. 2025\.[NoveltyBench: Evaluating Language Models for Humanlike Diversity](https://doi.org/10.48550/arXiv.2504.05228)\.In*Proceedings of the Conference on Language Modeling \(COLM\)*\.
- Zhang et al\. \(2024\)Yiming Zhang, Avi Schwarzschild, Nicholas Carlini, Zico Kolter, and Daphne Ippolito\. 2024\.[Forcing Diffuse Distributions out of Language Models](https://arxiv.org/abs/2404.10859)\.In*Proceedings of the Conference on Language Modeling \(COLM\)*\.
- Zhou et al\. \(2023\)Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou\. 2023\.[Instruction\-following evaluation for large language models](https://doi.org/10.48550/ARXIV.2311.07911)\.*CoRR*, abs/2311\.07911\.
## Appendix AProof of Theorem[1](https://arxiv.org/html/2607.02668#Thmtheorem1)
#### Step 1: RewritepGp\_\{G\}andpG′p\_\{G^\{\\prime\}\}\.
By definition,
𝒑𝑮\(y∣x\)=∑𝐯p\(𝐯∣x\)π\(y∣𝐯,x\)\.\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}=\\sum\_\{\\mathbf\{v\}\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi\(y\\mid\\mathbf\{v\},x\)\.Split the sum into\{𝐯:vy=1\}\\\{\\mathbf\{v\}:v\_\{y\}=1\\\}and\{𝐯:vy=0\}\\\{\\mathbf\{v\}:v\_\{y\}=0\\\}\. By the support constraint onπ\\pi, every term in the second sum vanishes, so
𝒑𝑮\(y∣x\)=∑𝐯:vy=1p\(𝐯∣x\)π\(y∣𝐯,x\)\.\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}=\\sum\_\{\\mathbf\{v\}:v\_\{y\}=1\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi\(y\\mid\\mathbf\{v\},x\)\.\(14\)By the same reasoning,
𝒑𝑮′\(y∣x\)=∑𝐯:vy=0p\(𝐯∣x\)π′\(y∣𝐯,x\)\.\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}=\\sum\_\{\\mathbf\{v\}:v\_\{y\}=0\}p\(\\mathbf\{v\}\\mid x\)\\,\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\.\(15\)
#### Step 2: Express with conditional expectation\.
For any functionf\(𝐯\)f\(\\mathbf\{v\}\)and setA⊆\{0,1\}\|𝒴\|A\\subseteq\\\{0,1\\\}^\{\|\\mathcal\{Y\}\|\}withP\(𝐯∈A\)\>0P\(\\mathbf\{v\}\\in A\)\>0,
𝔼\[f\(𝐯\)∣𝐯∈A\]=∑𝐯∈Ap\(𝐯∣x\)f\(𝐯\)∑𝐯∈Ap\(𝐯∣x\)\.\\mathbb\{E\}\\\!\\left\[f\(\\mathbf\{v\}\)\\mid\\mathbf\{v\}\\in A\\right\]=\\frac\{\\sum\_\{\\mathbf\{v\}\\in A\}p\(\\mathbf\{v\}\\mid x\)\\,f\(\\mathbf\{v\}\)\}\{\\sum\_\{\\mathbf\{v\}\\in A\}p\(\\mathbf\{v\}\\mid x\)\}\.Applying this withf\(𝐯\)=π\(y∣𝐯,x\)f\(\\mathbf\{v\}\)=\\pi\(y\\mid\\mathbf\{v\},x\)andA=\{𝐯:vy=1\}A=\\\{\\mathbf\{v\}:v\_\{y\}=1\\\}, the fact that𝒑𝑽\(y∣x\)\>0\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\>0ensures that the denominator is greater than 0, as the denominator equals𝒑𝑽\(y∣x\)\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}by \([1](https://arxiv.org/html/2607.02668#S3.E1)\)\. The numerator equals𝒑𝑮\(y∣x\)\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}by \([14](https://arxiv.org/html/2607.02668#A1.E14)\)\. Hence
𝔼\[π\(y∣𝐯,x\)∣vy=1,x\]=𝒑𝑮\(y∣x\)𝒑𝑽\(y∣x\)\.\\mathbb\{E\}\\\!\\left\[\\pi\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=1,x\\right\]=\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\}\{\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\.\(16\)The analogous calculation withf\(𝐯\)=π′\(y∣𝐯,x\)f\(\\mathbf\{v\}\)=\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)andA=\{𝐯:vy=0\}A=\\\{\\mathbf\{v\}:v\_\{y\}=0\\\}, using𝒑𝑽\(y∣x\)<1\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}<1to ensure1−𝒑𝑽\(y∣x\)\>01\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\>0, yields
𝔼\[π′\(y∣𝐯,x\)∣vy=0,x\]=𝒑𝑮′\(y∣x\)1−𝒑𝑽\(y∣x\)\.\\mathbb\{E\}\\\!\\left\[\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=0,x\\right\]=\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\}\{1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\.\(17\)
#### Step 3: Take the ratio\.
The denominators on the right\-hand sides of \([16](https://arxiv.org/html/2607.02668#A1.E16)\) and \([17](https://arxiv.org/html/2607.02668#A1.E17)\) are strictly positive by the hypothesis0<𝒑𝑽\(y∣x\)<10<\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}<1\. The numerators are nonnegative as sums of nonnegative terms\.
Dividing \([17](https://arxiv.org/html/2607.02668#A1.E17)\) by \([16](https://arxiv.org/html/2607.02668#A1.E16)\),
𝒑𝑮′\(y∣x\)/\(1−𝒑𝑽\(y∣x\)\)𝒑𝑮\(y∣x\)/𝒑𝑽\(y∣x\)=𝔼\[π′\(y∣𝐯,x\)∣vy=0,x\]𝔼\[π\(y∣𝐯,x\)∣vy=1,x\]\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}/\(1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\)\}\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}/\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}=\\frac\{\\mathbb\{E\}\\\!\\left\[\\pi^\{\\prime\}\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=0,x\\right\]\}\{\\mathbb\{E\}\\\!\\left\[\\pi\(y\\mid\\mathbf\{v\},x\)\\mid v\_\{y\}=1,x\\right\]\}The left\-hand side simplifies to
𝒑𝑽\(y∣x\)1−𝒑𝑽\(y∣x\)⋅𝒑𝑮′\(y∣x\)𝒑𝑮\(y∣x\),\\frac\{\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\{1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\\cdot\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\}\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\},and rearranging gives
𝒑𝑽\(y∣x\)1−𝒑𝑽\(y∣x\)=𝒑𝑮\(y∣x\)𝒑𝑮′\(y∣x\)⋅r\(y,x\)\.\\frac\{\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}\{1\-\{\\color\[rgb\]\{0\.1328125,0\.546875,0\.1328125\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1328125,0\.546875,0\.1328125\}\\boldsymbol\{p\_\{V\}\}\(y\\mid x\)\}\}=\\frac\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G\}\}\(y\\mid x\)\}\}\{\{\\color\[rgb\]\{0\.1171875,0\.390625,0\.78515625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.1171875,0\.390625,0\.78515625\}\\boldsymbol\{p\_\{G^\{\\prime\}\}\}\(y\\mid x\)\}\}\\cdot r\(y,x\)\.
## Appendix BAdditional Results
Full results on generator ROC \(ROCG\) and Pearson correlation \(ρ\\rho\) are given in Table[4](https://arxiv.org/html/2607.02668#A2.T4)\. Each entry is the average score across prompts, while±\\pmindicates the standard error\.
Table 4:Main results across tasks and models\. ROCGmeasures generator discriminability, andρ\\rhomeasures generator–validator correlation\. All methods use the per\-task best eval\-time correction\. Models: G2\-9b\-it= Gemma\-2\-9b\-it, Q3\.5\-9b = Qwen\-3\.5\-9B, G4\-31b\-it = Gemma\-4\-31b\-it\.The full results on validator accuracy \(AccV\) and validator AUROC \(ROCV\) are given in Table[5](https://arxiv.org/html/2607.02668#A2.T5)\.
Table 5:Validator performance across tasks and models\. ROCVmeasures validator ROC\-AUC and AccVmeasures validator accuracy at threshold 0\. Both are independent of the eval\-time frequency correction \(they only use the validator log\-odds against the gold label\)\. FLORA matches or exceeds the base model on most tasks, indicating no validator\-side likelihood displacement\.
## Appendix CDataset construction Details
### C\.1IFEval
IFEvalZhou et al\. \([2023](https://arxiv.org/html/2607.02668#bib.bib33)\)is an instruction\-following benchmark designed to evaluate whether model outputs satisfy explicit, verifiable constraints specified in a prompt\. To construct our dataset, we sample and filter prompts from IFEval to ensure they are well\-formed, English\-only, and support the generation of sufficiently diverse outputs\. For each prompt, we generate candidate responses by introducing prompt variations, including equivalent rephrasings \(to produce positive samples\) and violations of content or format constraints \(to produce negative samples\)\. This yields a set of responses with nontrivial separability\. Ground\-truth labels are assigned using a combination of IFEval\-style rule\-based verification for format constraints \(e\.g\., keyword counts, structural requirements\) and prompt\-induced correctness assumptions for content constraints \(e\.g\., relevance or factual consistency\)\. A response is labeled as correct only if it satisfies both criteria\. We further divide the dataset into in\-domain and out\-of\-domain splits, where the former evaluates performance on seen prompts with held\-out responses, and the latter measures generalization to entirely unseen prompts\.
### C\.2Hyponymy
Hyponymy is the relationship between a higher\-level category and it’s exemplars \(e\.g\., \(furniture, table\)\)\. We adopted the data used by\(Rosch,[1975](https://arxiv.org/html/2607.02668#bib.bib20)\), consisting of ten classes, each with 30 example items\. These examples ranged from very prototypical \(e\.g\., table\) to less usual \(e\.g\., lamp\), with the least prototypical not arguably not really belonging to the category at all\. Since most of the examples were positive \(actual examples of a category\), we used GPT\-5 to sample additional examples\.
In order to train with a disjoint set of categories, we combined the datasets from\(Banks and Connell,[2023](https://arxiv.org/html/2607.02668#bib.bib1); Stoinski et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib21); Castro et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib4); Van Overschelde et al\.,[2004](https://arxiv.org/html/2607.02668#bib.bib24); Uyeda and Mandler,[1980](https://arxiv.org/html/2607.02668#bib.bib23)\), in order to have a diverse set\. Then we manually removed the categories which overlapped with the ten in\(Rosch,[1975](https://arxiv.org/html/2607.02668#bib.bib20)\)\.
### C\.3HumanEval
We generated solutions to the 164 Python problems in the HumanEval benchmark, requiring at least 10 passing and 10 failing solutions per problem\. We used a diverse set of LLMs to generate solutions from: gpt\-3\.5\-turbo, gpt\-4, gpt\-4o, gpt\-4o\-mini, gpt\-4\.1, gpt\-5, and gpt\-5\.5, as well as Llama\-3\.1\-8B\-Instruct, deepseek\-coder\-1\.3b\-instruct, Phi\-3\-mini\-4k\-instruct, Mistral\-7B\-Instruct\-v0\.3, and OLMo\-2\-0425\-1B\-Instruct\. Each generation was prompted under one of three difference persona strategies: normal, beginner, unusual, refactorable\.
Finally, we filtered the solutions that were too short or too long \(less than 10 characters or more than 900 characters\)\. 82 problems were randomly sampled to be the held\-out test set\.
## Appendix DPrompts
IFEval: Example Prompt and Candidate ResponsesIFEval Prompt
Write a song about a man who rents a room in a house with a bunch of other people, and he absolutely hated it\. Your song should contain at least 10 words in all capital letters that are adjectives or verbs\. Commas are not allowed in the song\.Positive Candidate Response
\*\*Title: Cramped and Cursed\*\*Waking up in this WORN\-OUT placeRoommates scatter like RATS in a raceSocks and dishes piling HIGHVoices echo and SHOUT from the skyEvery night’s a MESS that never endsPanic and noise I CAN’T defendNo space to breathe it FEELS so tightIn this gloomy box my dreams take flightLife here is FROZEN like a bad mistakeI search for peace but ALL I do is acheSlamming doors and DRIFTING dreamsI’m haunted by all these NOISY screamsA hopeful heart that’s TRAPPED insideYearning to escape I can’t CONFIDELonging for laughter not DISTRACTIONBut I keep STUMBLING on this chain reactionFingers crossed I’ll BREAK these chainsFrom this cursed house with all its painsSomeday I’ll find that FREEDOM so trueBut for now it’s just ME and this messy crewNegative Candidate Response
\*\*Room to Love\*\*In a HOUSE filled with laughterShe found a SPACE so brightA heart full of wonderIn the soft morning LIGHTDREAMS linger on the wallsWith FRIENDS who understandThey SHOUT and dance in the hallsCreating memories hand in handThe KITCHEN filled with chatterCups of coffee in the airSharing STORIES and secretsShowing love everywhereEvery NOOK is a treasureEach moment feels so freeA PLACE where her heart soarsAs happy as can beLiving in a FLAIR so vividWhere life’s a grand DISPLAYTogether they SHINE like starsIn their beautiful balletFigure 6:Example prompt and candidate responses fromIFEvalHumanEval: Example Prompt and Candidate ResponsesHumanEval Prompt
```
def string_to_md5(text):
"""
Given a string ’text’, return its md5 hash equivalent string.
If ’text’ is an empty string, return None.
>>> string_to_md5(’Hello world’) == ’3e25960a79dbc69b674cd4ec67a72c62’
"""
```
Positive Candidate Response
```
if not text:
return None
import hashlib
HASH_OBJECT = hashlib.md5(text.encode())
return HASH_OBJECT.hexdigest()
```
Negative Candidate Response
```
if not text:
return None
m = hashlib.md5()
m.update(text.encode(’utf-8’))
return m.hexdigest()
```
Figure 7:Example prompt and candidate responses fromHumanEvalHyponym: Example Hypernym and Candidate ResponsesHypernym
clothingPositive Candidate Response
leotardNegative Candidate Response
purseFigure 8:Example hypernym and candidate responses \(hyponyms\) fromHyponymPrompts: IFEval TemplatesGenerator Template<ifeval\_prompt\>Validator TemplateYou are a strict binary validator\.You will be given a prompt and a response\. Determine if the response follows ALL instructions in the prompt exactly\.Rules:\- Every constraint must be satisfied perfectly\.\- If ANY constraint is violated, output NO\.\- Do not give partial credit\.\- Do not infer or assume missing requirements\.\- If you are uncertain, output NO\.Prompt:<ifeval\_prompt\>Response:<candidate\_response\>Output exactly one word: YES or NO\.Figure 9:Prompt templates forIFEval\.Prompts: HumanEval TemplatesGenerator TemplateComplete the following Python function:<humaneval\_prompt\>
Solution:Validator TemplateIs this a correct solution to the programming problem?Problem:<humaneval\_prompt\>
Solution:<candidate\_response\>
Answer Yes or No:Figure 10:Prompt templates forHumanEval\.Prompts: Hyponym TemplatesGenerator TemplateComplete the sentence: An example of a<hypernym\>is aValidator TemplateAnswer whether the first word is a type/kind of the second word\.Is a dog a kind of animal? Answer: YesIs a car a kind of fruit? Answer: NoIs an apple a kind of fruit? Answer: YesIs a chair a kind of vehicle? Answer: NoIs a<candidate\_response\>a kind of<hypernym\>? Answer:Figure 11:Prompt templates forHyponym\.
## Appendix EDataset Licenses
We summarize the licenses of the datasets used in this work and confirm that our use is consistent with their intended terms\.
#### HumanEval\.
HumanEval\(Chen et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib5)\)is released by OpenAI under the MIT License \(see[https://github\.com/openai/human\-eval](https://github.com/openai/human-eval)\)\. The MIT License permits use, modification, and redistribution, including for research purposes such as ours, provided the original copyright and license notice are retained\. Our derived dataset of model\-generated solutions is built on top of these problems and is used solely for non\-commercial research\.
#### IFEval\.
#### Hyponymy stimuli \(Rosch, 1975\)\.
The category–exemplar lists we use for the Hyponymy task originate from the published norms inRosch \([1975](https://arxiv.org/html/2607.02668#bib.bib20)\), which appear as tables of example items within the article itself\. We are not aware of an explicit data license accompanying the original paper; our use of these short lists of category exemplars for non\-commercial academic research, with full citation, is consistent with fair use as commonly applied to stimulus materials reported in published psychology research\. The additional category datasets used to construct disjoint training categories \(Banks and Connell,[2023](https://arxiv.org/html/2607.02668#bib.bib1); Stoinski et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib21); Castro et al\.,[2021](https://arxiv.org/html/2607.02668#bib.bib4); Van Overschelde et al\.,[2004](https://arxiv.org/html/2607.02668#bib.bib24); Uyeda and Mandler,[1980](https://arxiv.org/html/2607.02668#bib.bib23)\) are likewise drawn from items reported in published academic articles or accompanying supplementary materials; THINGSplus\(Stoinski et al\.,[2024](https://arxiv.org/html/2607.02668#bib.bib21)\)in particular is released under CC BY 4\.0\. We use only the textual category–exemplar pairs and cite each source\.Similar Articles
Reliability-Aware LLM Alignment from Inconsistent Human Feedback
Proposes Reliability-Guided Preference Optimization (RGPO) to handle inconsistent human feedback in LLM alignment by estimating annotator reliability and dynamically modulating training based on consensus, achieving superior performance over standard RLHF methods.
Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
The paper proposes position-selective self-distillation for training LLM judges from natural language feedback, using per-position entropy shifts to mask memorization-prone tokens and improve out-of-distribution generalization over outcome-supervised RL methods like GRPO by 2–9 points on subjective tasks.
LLM-as-an-Improver: Turning Verification into Better Candidates
This paper introduces LLM-as-an-Improver, a method that uses verification feedback to generate improved candidate solutions for LLMs, enhancing performance beyond initial candidate pools.
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
WildFeedback is a novel framework that leverages in-situ user feedback from actual LLM conversations to automatically create preference datasets for aligning language models with human preferences, addressing scalability and bias issues in traditional annotation-based alignment methods.
Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking
This paper presents a method to fine-tune LLaMA 3 8B as an efficient reranker for Retrieval-Augmented Generation using knowledge distillation and 4-bit quantization, achieving 14-21% gains in retrieval metrics over cross-encoder baselines with reduced inference cost.