Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
Summary
This paper introduces LLM-NRM, an option-level psychometric framework for multiple-choice benchmarks that models the full distribution over answer choices rather than binary correctness, showing that incorrect responses carry useful measurement information and improving ability estimation and benchmarking efficiency.
View Cached Full Text
Cached at: 08/05/26, 07:42 AM
# Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
Source: [https://arxiv.org/html/2608.02966](https://arxiv.org/html/2608.02966)
## Every Wrong Answer Counts: Option\-Level Psychometrics for LLM Multiple\-Choice Benchmarks
###### Abstract
Most multiple\-choice question \(MCQ\) benchmarks evaluate Large Language Models \(LLMs\) only by whether they select the correct answers\. This binary scoring treats all incorrect responses alike, even though an LLM’s preferences among incorrect options may contain systematic and useful information about its behavior and ability\. We introduce the LLM Nominal Response Model \(LLM\-NRM\), an option\-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option\-level item characteristics, while separating model\-specific response calibration sharpness, positional preference, and difficulty\-dependent fallback behavior\. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM\-NRM predicts held\-out LLM\-item interactions more accurately than binary Item Response models and conventional nominal\-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0\.920 with the external human\-preference Arena\.ai Elo leaderboard\. Distractor identity contributes \+101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full\-information ability estimates with Spearman 0\.943\. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full\-bank ranking with Kendall’s correlation 0\.85, corresponding to a 770 times reduction\. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes\.
## Introduction
Figure 1:A real exampleof an ARC\-Challenge MCQ item with item characteristic curves from our LLM\-NRM fitting\. \(a\) In this item, "Saturn" is a plausible distractor that reflects partial knowledge of LLM examinees, whereas "the Sun" is the correct answer\. \(b\) Binary scoring collapses all three choices into the same "incorrect" outcome, but option\-level modeling estimates a separate response curve instead for each option across LLM ability spectrumθ\\theta\. \(c\) Because these distractor preferences vary with ability, their identities provide Fisher information that correctness does not retain\.Large Language Model \(LLM\) benchmarking commonly evaluates performance on standard multiple\-choice question \(MCQ\) benchmarks with accuracy scores\(Wanget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib22); Reinet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib23)\)\. However, accuracy is a limited measure of model capability, because it is tied to a particular benchmark: raw scores depend on its item composition, while empirical item difficulty depends on the evaluated model population\. As a result, model rankings can shift with changes in the benchmark, scoring protocol, or comparison set\(Laloret al\.[2016](https://arxiv.org/html/2608.02966#bib.bib59); Perlitzet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib60); Alzahraniet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib61)\)\. A single accuracy value also collapses every response to a binary outcome, hiding model confidence and systematic error patterns\(Zhuet al\.[2023](https://arxiv.org/html/2608.02966#bib.bib6)\)\. It does not by itself quantify measurement uncertainty\(Perlitzet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib60)\), and it loses discriminative power as the strongest models approach benchmark saturation\(Wanget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib22)\)\.
Item Response Theory \(IRT\) provides a principled framework for addressing some of these limitations by modeling interactions between latent ability and item characteristics\(Rasch[1960](https://arxiv.org/html/2608.02966#bib.bib37); Lord and Novick[1968](https://arxiv.org/html/2608.02966#bib.bib38); Birnbaum[1968](https://arxiv.org/html/2608.02966#bib.bib16)\)\. Recent studies apply psychometric models to LLM evaluation for benchmark analysis, adaptive testing, difficulty estimation, and multilingual evaluation\(Zhouet al\.[2026](https://arxiv.org/html/2608.02966#bib.bib45); Land and Bikel[2026](https://arxiv.org/html/2608.02966#bib.bib51); Zhuanget al\.[2025](https://arxiv.org/html/2608.02966#bib.bib42); Liet al\.[2025](https://arxiv.org/html/2608.02966#bib.bib44); Zhuet al\.[2025](https://arxiv.org/html/2608.02966#bib.bib46); Lioret al\.[2026](https://arxiv.org/html/2608.02966#bib.bib43); Zhanget al\.[2026](https://arxiv.org/html/2608.02966#bib.bib58)\)\. These approaches reveal benchmark saturation, item quality issues, and differences between models with similar accuracy scores\(Vaniaet al\.[2021](https://arxiv.org/html/2608.02966#bib.bib41); Zhouet al\.[2026](https://arxiv.org/html/2608.02966#bib.bib45); Land and Bikel[2026](https://arxiv.org/html/2608.02966#bib.bib51)\)\. However, they generally reduce responses to binary correctness alone, discarding the option\-level information contained in multiple\-choice predictions\.
In the meantime, LLM responses contain richer behavioral signals beyond correctness, including option\-selection biases, positional preferences, self\-evaluation, and response consistency patterns\(Zhenget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib55); Yanget al\.[2025](https://arxiv.org/html/2608.02966#bib.bib56); Kadavathet al\.[2022](https://arxiv.org/html/2608.02966#bib.bib47); Wuet al\.[2025](https://arxiv.org/html/2608.02966#bib.bib52); Chaudhuryet al\.[2026](https://arxiv.org/html/2608.02966#bib.bib53)\)\. Yet, these signals are typically analyzed separately from latent ability and item characteristics, leaving the full response structure underutilized\.
In this paper, we demonstrate that the probability distribution over multiple\-choice options produced by an LLM is itself a psychometric response\. This naturally connects to the human\-testing Nominal Response Model \(NRM\), a categorical extension of IRT that models relationships between latent ability and response option preferences\(Bock[1972](https://arxiv.org/html/2608.02966#bib.bib1)\)\. Unlike binary IRT, which scores each response only by correctness, the NRM gives every option its own discrimination and attractiveness, capturing how each distractor’s appeal varies with ability and thereby extracting more measurement information per item than correctness alone\.
We propose the LLM Nominal Response Model \(LLM\-NRM\), which adapts NRM for LLM evaluation by modeling latent ability, item characteristics, and LLM\-specific behavioral factors\. Our contributions are as follows:
- •To our knowledge, this work presents the first application of the Nominal Response Model \(NRM\) to LLMs, revealing latent response preferences across benchmark MCQ answer options and providing a richer view of model behavior beyond answer correctness;
- •We propose an LLM\-adapted NRM formulation that incorporates behavioral factors such as calibration sharpness, ability\-gated guessing, and positional bias, moving beyond traditional item\-response modeling toward a richer representation of LLM decision processes;
- •Through large\-scale experiments on 189 LLMs and 31,554 MCQ items across 14 benchmarks, we show that LLM\-NRM improves held\-out response prediction, yields ability estimates that best match external human\-preference rankings, and establishes that distractor choices are informative enough to recover ability and compress benchmarks by orders of magnitude\.
## LLM Nominal Response Model
We introduce the LLM Nominal Response Model \(LLM\-NRM\), an option\-aware extension of Bock’s NRM for analyzing LLM responses to multiple\-choice questions\. Unlike conventional evaluation, which retains only the correctness of the most probable option, LLM\-NRM models the complete response distribution over all valid options\. This preserves richer information about distractor preferences and how option attractiveness varies across models of different abilities\.
LLM\-NRM retains the original NRM parameterization for MCQ items\(Penfield[2014](https://arxiv.org/html/2608.02966#bib.bib5)\)and instead augments the respondent model with additional latent response processes motivated by empirical characteristics commonly observed in LLMs: \(i\) response calibration sharpness, \(ii\) positional preference, and \(iii\) difficulty\-dependent fallback behavior\. These LLM\-specific extensions separately capture systematic response variation that would otherwise be absorbed into examinee ability and item modeling\. Unlike human testing that typically records one categorical response per item\(Thissen and Steinberg[1984](https://arxiv.org/html/2608.02966#bib.bib2); Suh and Bolt[2010](https://arxiv.org/html/2608.02966#bib.bib3)\), LLM evaluation provides full option distributions across large item banks, enabling these respondent\-specific processes to be estimated\.
### Problem Formulation
Let𝒥\\mathcal\{J\}denote a collection of LLMs as examinees andℐ\\mathcal\{I\}denote a collection of MCQ items\. Each MCQ itemi∈ℐi\\in\\mathcal\{I\}presentsKiK\_\{i\}answer options indexed by𝒦i=\{1,…,Ki\}\\mathcal\{K\}\_\{i\}=\\\{1,\\ldots,K\_\{i\}\\\}, corresponding to the displayed answer positions such as A, B, C and D in the prompt\. Letgi∈𝒦ig\_\{i\}\\in\\mathcal\{K\}\_\{i\}denote the index of the keyed correct option, and defineKmax=maxi∈ℐKiK\_\{\\max\}=\\max\_\{i\\in\\mathcal\{I\}\}K\_\{i\}as the largest number of options among all items\.
An LLM’s response to an MCQ item is inherently distributional, as each autoregressive step produces a probability over next tokens\. So when an LLMj∈𝒥j\\in\\mathcal\{J\}answers an MCQ itemi∈ℐi\\in\\mathcal\{I\}, it expresses a degree of preference for every available option\. We represent this observed preference𝐩jiobs\\mathbf\{p\}^\{\\mathrm\{obs\}\}\_\{ji\}as:
𝐩jiobs=\(pji1obs,…,pjiKiobs\)∈ΔKi−1,\\mathbf\{p\}^\{\\mathrm\{obs\}\}\_\{ji\}=\\left\(p^\{\\mathrm\{obs\}\}\_\{ji1\},\\ldots,p^\{\\mathrm\{obs\}\}\_\{jiK\_\{i\}\}\\right\)\\in\\Delta^\{K\_\{i\}\-1\},\(1\)withpjikobsp^\{\\mathrm\{obs\}\}\_\{jik\}the normalized preference for optionkkandΔKi−1\\Delta^\{K\_\{i\}\-1\}the probability simplex\.
LLM\-NRM treats𝐩jiobs\\mathbf\{p\}^\{\\mathrm\{obs\}\}\_\{ji\}as the observed response, rather than reducing it to a binary indicator of whether the most probable option equalsgig\_\{i\}\.
### Bock’s NRM
Bock’s NRM\(Bock[1972](https://arxiv.org/html/2608.02966#bib.bib1)\)models unordered response options by assigning each examineej∈𝒥j\\in\\mathcal\{J\}a latent abilityθj∈ℝ\\theta\_\{j\}\\in\\mathbb\{R\}, and each optionk∈𝒦ik\\in\\mathcal\{K\}\_\{i\}a discrimination parameteraik∈ℝa\_\{ik\}\\in\\mathbb\{R\}plus an interceptcik∈ℝc\_\{ik\}\\in\\mathbb\{R\}\. The discrimination controls how the option’s relative attractiveness varies with ability, while the intercept captures its baseline attractiveness\. The resulting response distribution is
pjikNRM=exp\(aikθj\+cik\)∑k′∈𝒦iexp\(aik′θj\+cik′\)\.p^\{\\mathrm\{NRM\}\}\_\{jik\}=\\frac\{\\exp\\left\(a\_\{ik\}\\theta\_\{j\}\+c\_\{ik\}\\right\)\}\{\\sum\_\{k^\{\\prime\}\\in\\mathcal\{K\}\_\{i\}\}\\exp\\left\(a\_\{ik^\{\\prime\}\}\\theta\_\{j\}\+c\_\{ik^\{\\prime\}\}\\right\)\}\.\(2\)
Because the softmax is invariant to item\-wise shifts in its utilities, the identification constraints are imposed
∑k∈𝒦iaik=0,∑k∈𝒦icik=0,\\sum\_\{k\\in\\mathcal\{K\}\_\{i\}\}a\_\{ik\}=0,\\qquad\\sum\_\{k\\in\\mathcal\{K\}\_\{i\}\}c\_\{ik\}=0,\(3\)and anchor the ability scale with the priorθj∼𝒩\(0,1\)\\theta\_\{j\}\\sim\\mathcal\{N\}\(0,1\)\.
The classical NRM is powerful for human testing, but LLMs violate its assumptions in three systematic ways:
- •differences in distribution sharpness that may reflect calibration rather than knowledge\(Kadavathet al\.[2022](https://arxiv.org/html/2608.02966#bib.bib47); Zhuet al\.[2023](https://arxiv.org/html/2608.02966#bib.bib6)\),
- •content\-independent preferences for displayed positions\(Zhenget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib55); Pezeshkpour and Hruschka[2024](https://arxiv.org/html/2608.02966#bib.bib7); Attali and Bar\-Hillel[2003](https://arxiv.org/html/2608.02966#bib.bib8)\),
- •and fallback guessing patterns on items that are difficult relative to the model’s ability\(Ivgiet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib9)\)\.
Fitting NRM directly would absorb these non\-ability effects intoθj\\theta\_\{j\}and the option parameters, potentially biasing both\.
### Response Sharpness and Positional Bias
We first augment the NRM ability\-driven utilities with global response sharpness and content\-independent positional preference\.
##### Response Sharpness\.
Two LLMs with identical preference orderings over options may nevertheless produce distributions with substantially different concentrations because of differences in calibration rather than knowledge\(Kadavathet al\.[2022](https://arxiv.org/html/2608.02966#bib.bib47); Zhuet al\.[2023](https://arxiv.org/html/2608.02966#bib.bib6)\)\. Classical NRM has no examinee\-specific parameter that changes concentration while preserving option ordering, so this variation may instead be absorbed intoθj\\theta\_\{j\}and the item discriminations\.
We capture the global component of this variation with a positive sharpness parameter for each LLM:
sj∈ℝ\+\.s\_\{j\}\\in\\mathbb\{R\}^\{\+\}\.\(4\)
Mathematically,sjs\_\{j\}is a per\-LLM inverse temperature that changes distribution concentration without altering option ordering:sj\>1s\_\{j\}\>1sharpens the distribution, whereassj<1s\_\{j\}<1flattens it\. Thus,sjs\_\{j\}captures model\-level calibration, or to what degree of confidence an LLM expresses its preferences, separately from item\-specific knowledge\.
##### Positional Bias\.
LLMs can exhibit systematic preferences for displayed answer positions or their labels independently of option content, and reordering the same options can substantially change measured accuracy\(Pezeshkpour and Hruschka[2024](https://arxiv.org/html/2608.02966#bib.bib7)\)\. Because the NRM interceptscikc\_\{ik\}are shared across LLMs, classical NRM cannot separate these model\-specific positional effects from option attractiveness\.
Rather than controlling for positional bias through computationally expensive permutation protocols\(Zhenget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib55)\), we estimate it jointly with ability:
𝜹j=\(δj1,…,δjKmax\)∈ℝKmax,\\boldsymbol\{\\delta\}\_\{j\}=\\left\(\\delta\_\{j1\},\\ldots,\\delta\_\{jK\_\{\\max\}\}\\right\)\\in\\mathbb\{R\}^\{K\_\{\\max\}\},\(5\)whereδjk\\delta\_\{jk\}represents LLMjj’s content\-independent preference for displayed positionkk\.
Incorporating both model\-specific sharpness and positional bias, we define the ability\-driven response logit as
ujik=sj\(aikθj\+cik\+δjk\)\.u\_\{jik\}=s\_\{j\}\\left\(a\_\{ik\}\\theta\_\{j\}\+c\_\{ik\}\+\\delta\_\{jk\}\\right\)\.\(6\)
### Difficulty\-Gated Guessing Fallback
Classical NRM represents all responses through the same ability\-driven option utilities\. When an item is difficult relative to an LLM’s ability, however, its option preferences may increasingly reflect model\-specific fallback strategies, such as label priors or option heuristics, rather than the item’s content\(Ivgiet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib9); Zhenget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib55); Balepuret al\.[2024](https://arxiv.org/html/2608.02966#bib.bib15)\)\. Binary 3PL\-IRT introduces a per\-item lower asymptote\(Birnbaum[1968](https://arxiv.org/html/2608.02966#bib.bib16)\), but it neither models a complete fallback distribution over options nor allows fallback behavior to vary across LLMs, and its guessing parameters are usually poorly identified\(Barton and Lord[1981](https://arxiv.org/html/2608.02966#bib.bib19); Maris and Bechger[2009](https://arxiv.org/html/2608.02966#bib.bib17); San Martínet al\.[2015](https://arxiv.org/html/2608.02966#bib.bib18)\)\.
We therefore define the guessing fallback gate instead:
Gji=σ\[w0−κ\(θj−b~i\)\],G\_\{ji\}=\\sigma\\\!\\left\[w\_\{0\}\-\\kappa\\left\(\\theta\_\{j\}\-\\widetilde\{b\}\_\{i\}\\right\)\\right\],\(7\)wherew0∈ℝw\_\{0\}\\in\\mathbb\{R\}is a global intercept,κ≥0\\kappa\\geq 0is a global slope, andb~i\\widetilde\{b\}\_\{i\}is the item\-difficulty index defined below\. The logistic linkσ\\sigmamaps the ability\-difficulty gap to a fallback interpolation weight in\(0,1\)\(0,1\)\. The interceptw0w\_\{0\}sets the transition location, whileκ\\kappacontrols its sharpness\. Consequently, fallback becomes more prominent as item difficulty increases relative to LLM ability\.
#### An Approximation of Difficulty Index\.
The fallback gate requires a scalar measure of item difficultyb~i\\tilde\{b\}\_\{i\}derived from the NRM option parameters\. In 2PL\-IRT, item difficulty is defined as the ability at which the correct\-response probability equals12\\tfrac\{1\}\{2\}\. NRM has no explicit scalar difficulty parameter, but an analogous difficulty threshold can be expressed through the correct\-versus\-incorrect log\-odds margin
fi\(θ\)=aigiθ\+cigi−log∑k≠giexp\(aikθ\+cik\)\.f\_\{i\}\(\\theta\)=a\_\{ig\_\{i\}\}\\theta\+c\_\{ig\_\{i\}\}\-\\log\\\!\\\!\\sum\_\{k\\neq g\_\{i\}\}\\\!\\exp\\left\(a\_\{ik\}\\theta\+c\_\{ik\}\\right\)\.\(8\)
Solving for this threshold numerically during optimization would be costly\. We therefore construct a local and differentiable index by linearizingfif\_\{i\}aroundθ=0\\theta=0, the center of the ability prior:
fi\(θ\)≈fi\(0\)\+fi′\(0\)⋅θf\_\{i\}\(\\theta\)\\approx f\_\{i\}\(0\)\+f\_\{i\}^\{\\prime\}\(0\)\\cdot\\theta\(9\)Thus,
b~i=−fi\(0\)fi′\(0\)=log∑k≠giexp\(cik\)−cigiaigi−∑k≠giwikaik,\\widetilde\{b\}\_\{i\}=\-\\frac\{f\_\{i\}\(0\)\}\{f\_\{i\}^\{\\prime\}\(0\)\}=\\frac\{\\log\\sum\_\{k\\neq g\_\{i\}\}\\exp\(c\_\{ik\}\)\-c\_\{ig\_\{i\}\}\}\{a\_\{ig\_\{i\}\}\-\\sum\_\{k\\neq g\_\{i\}\}w\_\{ik\}a\_\{ik\}\},\(10\)where
wik=exp\(cik\)∑k′≠giexp\(cik′\)\.w\_\{ik\}=\\frac\{\\exp\(c\_\{ik\}\)\}\{\\sum\_\{k^\{\\prime\}\\neq g\_\{i\}\}\\exp\(c\_\{ik^\{\\prime\}\}\)\}\.\(11\)Here,wikw\_\{ik\}is the distractor softmax weight atθ=0\\theta=0\. Equivalently,b~i\\widetilde\{b\}\_\{i\}is the local threshold estimate obtained by taking one Newton step from the center of the ability prior\. This construction introduces no additional free item parameter and remains differentiable during optimization\.
#### Fallback pattern\.
The scalar gateGjiG\_\{ji\}determines how strongly fallback behavior contributes to a response\. Inspired by respondent\-specific bias vectors in human annotation models\(Dawid and Skene[1979](https://arxiv.org/html/2608.02966#bib.bib20); Welinderet al\.[2010](https://arxiv.org/html/2608.02966#bib.bib21)\), we assign each LLM a vector of fallback logits to specify the distribution over options within that regime
𝝆j=\(ρj1,…,ρjKmax\)∈ℝKmax,\\boldsymbol\{\\rho\}\_\{j\}=\\left\(\\rho\_\{j1\},\\ldots,\\rho\_\{jK\_\{\\max\}\}\\right\)\\in\\mathbb\{R\}^\{K\_\{\\max\}\},\(12\)whereρjk\\rho\_\{jk\}is LLMjj’s fallback logit for displayed positionkk\.
Notably, the positional\-bias vector𝜹j\\boldsymbol\{\\delta\}\_\{j\}and fallback vector𝝆j\\boldsymbol\{\\rho\}\_\{j\}play distinct roles\. The former captures a persistent positional effect in the ability\-driven utilities, whereas the latter defines the response pattern that receives increasing weight asGjiG\_\{ji\}grows\.
### Full LLM\-NRM Response Distribution
Altogether, we define the response preference distribution under LLM\-NRM as an interpolation combining the ability\-driven and fallback processes:
pjikLLM\-NRM=exp\[\(1−Gji\)ujik\+Gjiρjk\]∑k′∈𝒦iexp\[\(1−Gji\)ujik′\+Gjiρjk′\]\.\\boxed\{p^\{\\mathrm\{LLM\\text\{\-\}NRM\}\}\_\{jik\}=\\frac\{\\exp\\left\[\(1\-G\_\{ji\}\)\\,u\_\{jik\}\+G\_\{ji\}\\,\\rho\_\{jk\}\\right\]\}\{\\sum\_\{k^\{\\prime\}\\in\\mathcal\{K\}\_\{i\}\}\\exp\\left\[\(1\-G\_\{ji\}\)\\,u\_\{jik^\{\\prime\}\}\+G\_\{ji\}\\,\\rho\_\{jk^\{\\prime\}\}\\right\]\}\}\.\(13\)
Notably, withsj=1s\_\{j\}=1,𝜹j=𝝆j=𝟎\\boldsymbol\{\\delta\}\_\{j\}=\\boldsymbol\{\\rho\}\_\{j\}=\\mathbf\{0\}, andGji=0G\_\{ji\}=0, our modeling reduces exactly to the classical NRM\.
Because the softmax is invariant to adding the same constant to all logits, we impose the LLM\-level identification constraints:
∑k≤Kmaxδjk=0,and∑k≤Kmaxρjk=0,\\sum\_\{k\\leq K\_\{\\max\}\}\\delta\_\{jk\}=0,\\text\{ and \}\\sum\_\{k\\leq K\_\{\\max\}\}\\rho\_\{jk\}=0,\(14\)and positivity are enforced through unconstrained reparameterizations:sj=softplus\(s¯j\)s\_\{j\}=\\operatorname\{softplus\}\(\\bar\{s\}\_\{j\}\),κ=softplus\(κ¯\)\\kappa=\\operatorname\{softplus\}\(\\bar\{\\kappa\}\), withs¯j,κ¯∈ℝ\\bar\{s\}\_\{j\},\\bar\{\\kappa\}\\in\\mathbb\{R\}\.
### Fitting
In total, LLM\-NRM estimates2\(Ki−1\)2\(K\_\{i\}\-1\)free parameters per MCQ item, plus2Kmax2K\_\{\\max\}free scalars per LLM along with two more global scalars\(w0,κ\)\(w\_\{0\},\\kappa\)\. Given a group of LLMs𝒥\\mathcal\{J\}and their observed preference𝐩jiobs\\mathbf\{p\}^\{\\mathrm\{obs\}\}\_\{ji\}on a set of MCQ itemsℐ\\mathcal\{I\}, we fit all parameters jointly by maximum a posteriori \(MAP\) estimation with Adam optimizer on the soft cross\-entropy loss:
ℒ=−∑i∈ℐ,j∈𝒥,k∈𝒦ipjikobslogpjikLLM\-NRM\+12∑j∈𝒥θj2,\\mathcal\{L\}=\-\\sum\_\{i\\in\\mathcal\{I\},j\\in\\mathcal\{J\},k\\in\\mathcal\{K\}\_\{i\}\}p^\{\\mathrm\{obs\}\}\_\{jik\}\\,\\log p^\{\\mathrm\{LLM\\text\{\-\}NRM\}\}\_\{jik\}\+\\frac\{1\}\{2\}\\sum\_\{j\\in\\mathcal\{J\}\}\\theta\_\{j\}^\{2\},\(15\)where12θj2\\frac\{1\}\{2\}\\theta\_\{j\}^\{2\}is the regularization term of the priorθ∼𝒩\(0,1\)\\theta\\sim\\mathcal\{N\}\(0,1\)\.
## Experimental Setup
##### MCQ item bank\.
We assemble a large and diverse evaluation set comprising*31,554 MCQ items*from 14 benchmarks spanning factual knowledge, reasoning, and commonsense: MMLU\-Pro\(Wanget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib22)\), GPQA\-Diamond\(Reinet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib23)\), ARC\-Challenge\(Clarket al\.[2018](https://arxiv.org/html/2608.02966#bib.bib24)\), AGIEval\(Zhonget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib25)\), CommonsenseQA\(Talmoret al\.[2019](https://arxiv.org/html/2608.02966#bib.bib26)\), LogiQA 2\.0\(Liuet al\.[2023](https://arxiv.org/html/2608.02966#bib.bib27)\), OpenBookQA\(Mihaylovet al\.[2018](https://arxiv.org/html/2608.02966#bib.bib28)\), TruthfulQA\(Linet al\.[2022](https://arxiv.org/html/2608.02966#bib.bib29)\), MedQA\-USMLE\(Jinet al\.[2021](https://arxiv.org/html/2608.02966#bib.bib30)\), RACE\(Laiet al\.[2017](https://arxiv.org/html/2608.02966#bib.bib31)\), SocialIQA\(Sapet al\.[2019](https://arxiv.org/html/2608.02966#bib.bib32)\), QASC\(Khotet al\.[2020](https://arxiv.org/html/2608.02966#bib.bib33)\), ReClor\(Yuet al\.[2020](https://arxiv.org/html/2608.02966#bib.bib34)\), and MMLU\(Hendryckset al\.[2021](https://arxiv.org/html/2608.02966#bib.bib35)\)\. Because MMLU and MMLU\-Pro partially overlap, we remove from MMLU every MCQ item that also appears in MMLU\-Pro\. The resulting collection contains between 2 and 10 answer options per MCQ item, with a mean of 5\.82\. This variation allows us to evaluate LLM\-NRM across both conventional and large\-option MCQ formats\.
##### LLM fleet\.
Our examinee fleet contains*189 LLMs*, accessed through either publicly released weights or official APIs\. We intentionally cover a broad range of LLM ability, parameter scale, training recipe, architecture, and development period\. The collection includes 64 open\-weight LLM families, including Qwen, Llama, Gemma, Phi, Mistral, Falcon, and Pythia, with parameter counts ranging from 70 million to 235 billion and release dates from February 2019 to May 2026\. Different model architectures are also covered, including dense Transformers, mixture\-of\-experts models, and state\-space models\. API\-accessible systems including DeepSeek\-V4 Pro, GLM\-5\.1, and Kimi\-K2\.6 extend the fleet to recent frontier\-scale LLMs\.
##### Answer\-choice distributions\.
For each LLM–MCQ item pair, we extract the log\-probability assigned to every valid answer option at the first token position immediately following the question and answer instruction\. All experiments apply common settings in LLM MCQ benchmarking, using zero\-shot prompting, temperatureT=1T=1, with chain\-of\-thought disabled\. We compute option scores from the full\-vocabulary log\-softmax whenever complete logits are available\. Some APIs expose only the top\-20 next\-token logits where we retain valid option tokens present in the returned set, and assign zero observed mass to unreported options\. The global answer\-instruction prompt is designed to concentrate next\-token mass on option tokens, with only 0\.56% letter\-mass leakage reported across the entire fleet\.
##### Baselines\.
We compare LLM\-NRM with classical binary 1PL–4PL IRT\(Barton and Lord[1981](https://arxiv.org/html/2608.02966#bib.bib19)\), Bock’s NRM\(Bock[1972](https://arxiv.org/html/2608.02966#bib.bib1)\), as well as prior state\-of\-the\-art approaches including Deep\-IRT\(Yeung[2019](https://arxiv.org/html/2608.02966#bib.bib39)\),β3\\beta^\{3\}\-IRT\(Chenet al\.[2019](https://arxiv.org/html/2608.02966#bib.bib40)\), PSN\-IRT\(Zhouet al\.[2026](https://arxiv.org/html/2608.02966#bib.bib45)\), and SD\-IR\(Wanget al\.[2026](https://arxiv.org/html/2608.02966#bib.bib36)\)\. Binary models natively predict only correctness, and whenever they are evaluated on complete responses, their predicted incorrect mass is distributed using an MCQ\-item\-specific wrong\-option distribution estimated exclusively from training cells\. All models are trained with Adam optimizer with learning rate 0\.05 for 2,500 steps, under which a stable convergence is always reached\.
## A Better Modeling of LLM Responses
We organize our evaluation around three questions: \(i\) whether LLM\-NRM better models LLM responses, \(ii\) how much information incorrect\-option identity contributes beyond correctness, and \(iii\) whether this additional information enables more efficient benchmarking\.
##### Held\-out Response Prediction\.
Table 1:Held\-out response prediction under 5\-fold cross\-validation\. Values report the mean and standard deviation across five disjoint test folds\. We report top\-1 accuracy \(Acc\) and log\-loss \(LL\) for both response channels, together with the Brier score \(Brier\) for binary correctness\. Arrows indicate the preferred direction\.We evaluate held\-out prediction of the LLM–item response matrix using 5\-fold cell\-level cross\-validation\. Table[1](https://arxiv.org/html/2608.02966#Sx4.T1)reports performance in two channels: binary correctness, where option\-aware models predict the probability of the keyed answer; and answer\-option prediction, where binary models distribute non\-key probability mass across distractors using training\-set frequencies\.
The binary channel favors binary IRT baselines, yet LLM\-NRM achieves the highest correctness accuracy among competing models \(0\.83950\.8395vs\.0\.82320\.8232\)\. In the answer\-option channel, which directly evaluates option\-level modeling, LLM\-NRM provides larger gains, achieving0\.70700\.7070accuracy and0\.81810\.8181log\-loss compared with0\.64270\.6427and0\.98790\.9879for the strongest binary baselines, and0\.66980\.6698and0\.91710\.9171for hard\-channel NRM\. Ablation results further show that all proposed components contribute in a complementary way, as removing all of them reduces LLM\-NRM \(full\) to NRM \(soft channel\) with significant performance gap\.
##### External Validity of the Ability Scale\.
Table 2:Rank agreement between each measurement and Arena\.ai text Elo scores\. Values report Spearman correlation with bootstrap standard deviation\. LLM\-NRM achieves the highest observed Spearman correlation\.Predictive performance alone does not guarantee that estimated LLM abilities capture meaningful model differences\. We therefore compare LLM\-NRM ability rankings with independent Arena\.ai Text Elo scores, derived from crowdsourced pairwise human preferences on open\-ended prompts\(Chianget al\.[2024](https://arxiv.org/html/2608.02966#bib.bib57)\)\. This evaluates whether the learned scale aligns with external human judgments rather than benchmark accuracy alone\. We match 48 LLMs with a July 2026 leaderboard snapshot111Arena\.ai Text leaderboard:https://arena\.ai/leaderboard/text\.and compute Spearman rank correlation, with uncertainty estimated by bootstrap resampling of the matched LLMs\.
As shown in Table[2](https://arxiv.org/html/2608.02966#Sx4.T2), LLM\-NRM achieves the strongest agreement with Arena Elo \(ρ=0\.920\\rho=0\.920\), outperforming raw accuracy \(0\.8660\.866\) and the strongest competing latent ability estimate \(0\.8850\.885\)\. Since Arena outcomes are not used during training, this result indicates that option\-level response patterns capture model differences aligned with human preferences beyond aggregate MCQ correctness\. The automatically estimated scale provides a reproducible and lower\-cost proxy for Arena\-style ranking\.
## Information Beyond Binary Correctness

\(a\)Information retained by the response channels

\(b\)Informative distractors across benchmark dates
Figure 2:Ability information beyond binary correctness\.\(a\)Mean item\-level Fisher information in binary correctness and the additional information contributed by distractor identity\. The full stacked length represents the information in the complete option response; percentages denote its increase over binary scoring\.\(b\)Average number of detected informative distractors per item versus benchmark publication date\. More recent benchmarks tend to expose more distinguishable distractor structure\.We next examine the source of this predictive advantage of LLM\-NRM: binary scoring observes only*whether*an LLM is wrong, whereas option\-level scoring also observes*which*distractor it selects\. We next demonstrate this additional information both theoretically and empirically\.
##### Fisher Information from Distractor Identity\.
For the nominal\-response component, as commonly defined in prior works\(Bock[1972](https://arxiv.org/html/2608.02966#bib.bib1); Suh and Bolt[2010](https://arxiv.org/html/2608.02966#bib.bib3); Garcia\-Perez[2014](https://arxiv.org/html/2608.02966#bib.bib64)\), the Fisher Information in the full option identity decomposes into the information retained by binary correctness and an additional nonnegative term contributed by distractor identity:
IiNRM\(θ\)=Iibinary\(θ\)\+Iiwrong\(θ\),I^\{\\mathrm\{NRM\}\}\_\{i\}\(\\theta\)=I^\{\\mathrm\{binary\}\}\_\{i\}\(\\theta\)\+I^\{\\mathrm\{wrong\}\}\_\{i\}\(\\theta\),\(16\)whereIiwrong\(θ\)≥0\.I^\{\\mathrm\{wrong\}\}\_\{i\}\(\\theta\)\\geq 0\.
Figure[2](https://arxiv.org/html/2608.02966#Sx5.F2)\(a\) reports the item\-level Fisher information retained by coarsened binary correctness and the additional contribution of distractor identity, averaged over the ability distribution and over items within each benchmark\. The full option response provides\+101%\+101\\%more Fisher information per item than binary scoring on average, and the gain is positive for every benchmark, showing that ability\-dependent distractor preferences are a systematic property of LLM responses rather than an isolated feature of a few datasets\.
##### Informativeness across benchmarks\.
Figure[2](https://arxiv.org/html/2608.02966#Sx5.F2)\(b\) characterizes how the available distractor signal varies across benchmark designs, as we define an incorrect option as informative distractor if it’s nevertheless the single most likely response at some ability level\. More recently released benchmarks tend to contain more informative distractors per item, and many recent benchmarks have more than 1 informative distractor per item\. It indicates that the response structure discarded by binary scoring remains substantial, and may become increasingly consequential as MCQ benchmarks incorporate richer sets of distractors\.
##### Ability Estimation from Incorrect Responses\.
We next isolate the contribution of distractor identity with another 5\-fold cross\-validation over the LLM fleet\. In each fold, MCQ item parameters are fit using the training LLMs and then held fixed while estimatingθ\\thetafor the held\-out LLMs under two conditions: \(i\) Binary\-Only, where the estimator observes only correctness signals; and \(ii\) Incorrect\-Only, where it observes only the categorical signals from incorrectly answered MCQ items\. For each held\-out LLM, we compare the estimatedθ\\thetawith the reference estimate obtained using the full response information\. Table[3](https://arxiv.org/html/2608.02966#Sx5.T3)reports the resulting correlations across all folds, quantifying the information contributed by each response signal\.
Table 3:Correlation between held\-outθ\\thetaestimates under each observation scenario and the full\-fleet reference estimate \(mean±\\pmsd over 5\-fold cross\-validation\)\.The results confirm that distractor identity is a meaningful source of information for ability estimation\. Even without observing whether an answer is correct, the pattern of selected distractors yields a strong estimate ofθ\\theta, indicating that error structure alone is highly informative but not a random effect\. Neither scenario leads to a perfect correlation, suggesting that the two signals provide complementary information\.
## Data Efficient LLM Benchmarking

\(a\)Ranking new LLMs with fewer MCQ items

\(b\)Calibrating an MCQ item bank with fewer LLMs
Figure 3:Data\-efficient LLM evaluation and benchmark calibration\.\(a\) Kendall correlation between rankings from information\-selected subsets and the full item bank\. LLM\-NRM performs best, especially at small item budgets\. \(b\) Held\-out correctness log\-loss after calibration with fewer LLMs\. LLM\-NRM produces useful item parameters with only 5 LLMs and approaches full\-fleet performance near 40 MCQ items\. Bands show variation across splits in \(a\) and calibration subsets in \(b\)\.As a natural consequence, we examine how LLM\-NRM’s richer response modeling enables more efficient LLM MCQ benchmarking\. Prior compression methods are binary\-correctness\-based, typically need around 100 anchor items, and calibrate on thousands of models\(Maia Poloet al\.[2024](https://arxiv.org/html/2608.02966#bib.bib65); Kipniset al\.[2025](https://arxiv.org/html/2608.02966#bib.bib66)\)\. We therefore investigate two questions as two typical costs of benchmarking on our approach: \(i\) how many MCQ items are required to reliably rank new LLMs? and \(ii\) how many LLMs are needed to calibrate a reusable MCQ item bank?
##### Compression of MCQ Benchmarks\.
We first vary the number of MCQ items administered to each held\-out LLM\. Each model uses its calibrated MCQ item parameters to select an information\-rich subset of the specified size, then estimates LLM ability from only the selected responses\. We evaluate the resulting leaderboard using Kendall’sτ\\tauwith the ranking obtained from the complete MCQ item bank\.
Figure[3](https://arxiv.org/html/2608.02966#Sx6.F3)\(a\) shows that LLM\-NRM remains effective even with a very small MCQ item budget\. Using only the 41 most informative MCQ items, the reduced benchmark achieves a Kendall correlation of0\.850\.85, corresponding to a×770\\times 770compression rate\. Overall, LLM\-NRM performs best in the low\-budget regime, while PSN\-IRT becomes competitive only when several hundred MCQ items are available\. These results indicate that a new LLM can be faithfully calibrated by evaluating it on only a small set of MCQ items\.
##### Calibration with Fewer LLMs\.
We next consider the reverse calibration problem\. Unlike conventional psychometric IRT, which assumes a large population of examinees calibrating a relatively small item bank, LLM benchmarking often enters the opposite regime: tens of thousands of MCQ items but only tens or hundreds of LLMs for calibration\. Flexible binary IRT models must therefore estimate item parameters from limited responses\. We test whether LLM\-NRM alleviates this bottleneck by randomly sampling calibration subsets, estimating item parameters, and evaluating the resulting MCQ banks on held\-out LLMs with binary log loss on correctness for fair comparison\.
Figure[3](https://arxiv.org/html/2608.02966#Sx6.F3)\(b\) shows that LLM\-NRM consistently achieves the lowest held\-out correctness loss across all calibration fleet sizes\. Unlike binary IRT models, which become unstable with small calibration fleets due to limited information in binary responses, LLM\-NRM exploits option\-level response distributions to obtain richer and more data\-efficient calibration signals\. Compared with soft\-channel NRM, LLM\-NRM’s modeling of LLM\-specific response sharpness, positional bias, and guessing behavior further improves generalization to unseen LLMs\. These results demonstrate that LLM\-NRM reduces both the number of MCQ items required to evaluate new LLMs and the number of LLMs needed for benchmark calibration, lowering evaluation costs\.
## Discussion and Conclusion
We introduced LLM\-NRM, an option\-aware psychometric model that recovers information discarded by binary MCQ scoring\. By modeling option distributions, LLM\-NRM jointly estimates LLM ability, item characteristics, and behavioral factors such as response calibration sharpness, positional preference, and difficulty\-dependent fallback behavior\. Across 189 LLMs and 31,554 items from 14 benchmarks, it improves held\-out response prediction, aligns ability estimates with human\-preference rankings, and reveals that distractor choices provide substantial information beyond correctness\. These signals also improve evaluation efficiency by preserving benchmark rankings with fewer items and reducing the number of LLMs required for calibration\.
On the other hand, the current formulation assumes access to comparable option probabilities and models ability with a single latent dimension\. Future work can explore partial probability observations, multidimensional abilities, and extensions beyond fixed\-option questions\. After all, our results demonstrate that LLM errors are informative signals that reveal their underlying capabilities and decision patterns, enabling more reliable evaluation beyond accuracy\-based rankings\.
## References
- N\. Alzahrani, H\. Alyahya, Y\. Alnumay, S\. AlRashed, S\. Alsubaie, Y\. Almushayqih, F\. Mirza, N\. Alotaibi, N\. Al\-Twairesh, A\. Alowisheq, M\. S\. Bari, and H\. Khan \(2024\)When benchmarks are targets: revealing the sensitivity of large language model leaderboards\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13787–13805\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.744),[Link](https://aclanthology.org/2024.acl-long.744/)Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p1.1)\.
- Y\. Attali and M\. Bar\-Hillel \(2003\)Guess where: the position of correct answers in multiple\-choice test items as a psychometric variable\.Journal of Educational Measurement40\(2\),pp\. 109–128\.Cited by:[2nd item](https://arxiv.org/html/2608.02966#Sx2.I2.i2.p1.1)\.
- N\. Balepur, A\. Ravichander, and R\. Rudinger \(2024\)Artifacts or abduction: how do LLMs answer multiple\-choice questions without the question?\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 10308–10330\.Cited by:[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1)\.
- M\. A\. Barton and F\. M\. Lord \(1981\)An upper asymptote for the three\-parameter logistic item\-response model\.Technical reportTechnical ReportRR\-81\-20,Educational Testing Service\.Cited by:[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1),[Baselines\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.25.25.6)\.
- A\. Birnbaum \(1968\)Some latent trait models and their use in inferring an examinee’s ability\.InStatistical Theories of Mental Test Scores,F\. M\. Lord and M\. R\. Novick \(Eds\.\),pp\. 397–479\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1),[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.20.20.6)\.
- R\. D\. Bock \(1972\)Estimating item parameters and latent ability when responses are scored in two or more nominal categories\.Psychometrika37\(1\),pp\. 29–51\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p4.1),[Bock’s NRM](https://arxiv.org/html/2608.02966#Sx2.SSx2.p1.5),[Baselines\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.56.56.6),[Fisher Information from Distractor Identity\.](https://arxiv.org/html/2608.02966#Sx5.SSx6.SSSx2.Px1.p1.1)\.
- B\. Chaudhury, M\. F\. Wang, H\. H\. Park, R\. Ghosh, S\. Hong, and J\. O\. Woo \(2026\)Quantifying consistency in llm logical reasoning via structural uncertainty\.arXiv preprint arXiv:2606\.17312\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p3.1)\.
- Y\. Chen, T\. Silva Filho, R\. B\. Prudencio, T\. Diethe, and P\. Flach \(2019\)β3\\beta^\{3\}\-IRT: a new item response model and its applications\.InThe 22nd international conference on artificial intelligence and statistics,pp\. 1013–1021\.Cited by:[Baselines\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.31.31.1)\.
- W\. Chiang, L\. Zheng, Y\. Sheng, A\. N\. Angelopoulos, T\. Li, D\. Li, H\. Zhang, B\. Zhu, M\. Jordan, J\. E\. Gonzalez,et al\.\(2024\)Chatbot arena: an open platform for evaluating llms by human preference\.arXiv preprint arXiv:2403\.04132\.Cited by:[External Validity of the Ability Scale\.](https://arxiv.org/html/2608.02966#Sx4.SSx6.SSSx2.Px2.p1.1)\.
- P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. Tafjord \(2018\)Think you have solved question answering? Try ARC, the AI2 reasoning challenge\.arXiv preprint arXiv:1803\.05457\.Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- A\. P\. Dawid and A\. M\. Skene \(1979\)Maximum likelihood estimation of observer error\-rates using the EM algorithm\.Journal of the Royal Statistical Society: Series C \(Applied Statistics\)28\(1\),pp\. 20–28\.Cited by:[Fallback pattern\.](https://arxiv.org/html/2608.02966#Sx2.SSx4.SSSx2.p1.1)\.
- M\. A\. Garcia\-Perez \(2014\)Multiple\-choice tests: polytomous irt models misestimate item information\.The Spanish Journal of Psychology17,pp\. E88\.Cited by:[Fisher Information from Distractor Identity\.](https://arxiv.org/html/2608.02966#Sx5.SSx6.SSSx2.Px1.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- M\. Ivgi, O\. Yoran, J\. Berant, and M\. Geva \(2024\)From loops to oops: fallback behaviors of language models under uncertainty\.arXiv preprint arXiv:2407\.06071\.Cited by:[3rd item](https://arxiv.org/html/2608.02966#Sx2.I2.i3.p1.1),[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1)\.
- D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. Szolovits \(2021\)What disease does this patient have? A large\-scale open domain question answering dataset from medical exams\.Applied Sciences11\(14\),pp\. 6421\.Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- S\. Kadavath, T\. Conerly, A\. Askell, T\. Henighan, D\. Drain, E\. Perez, N\. Schiefer, Z\. Hatfield\-Dodds, N\. DasSarma, E\. Tran\-Johnson,et al\.\(2022\)Language models \(mostly\) know what they know\.arXiv preprint arXiv:2207\.05221\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p3.1),[1st item](https://arxiv.org/html/2608.02966#Sx2.I2.i1.p1.1),[Response Sharpness\.](https://arxiv.org/html/2608.02966#Sx2.SSx3.SSS0.Px1.p1.1)\.
- T\. Khot, P\. Clark, M\. Guerquin, P\. Jansen, and A\. Sabharwal \(2020\)QASC: a dataset for question answering via sentence composition\.InProceedings of the 34th AAAI Conference on Artificial Intelligence,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- A\. Kipnis, K\. Voudouris, L\. Schulze Buschoff, and E\. Schulz \(2025\)Metabench\-a sparse benchmark of reasoning and knowledge in large language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 31734–31770\.Cited by:[Data Efficient LLM Benchmarking](https://arxiv.org/html/2608.02966#Sx6.p1.1)\.
- G\. Lai, Q\. Xie, H\. Liu, Y\. Yang, and E\. Hovy \(2017\)RACE: large\-scale reading comprehension dataset from examinations\.InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- J\. P\. Lalor, H\. Wu, and H\. Yu \(2016\)Building an evaluation scale using item response theory\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,pp\. 648–657\.External Links:[Document](https://dx.doi.org/10.18653/v1/D16-1062),[Link](https://aclanthology.org/D16-1062/)Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p1.1)\.
- S\. Land and D\. M\. Bikel \(2026\)Auditing llm benchmarks with item response theory\.arXiv preprint arXiv:2605\.30504\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.
- P\. Li, X\. Tang, S\. Chen, Y\. Cheng, R\. Metoyer, T\. Hua, and N\. V\. Chawla \(2025\)Adaptive testing for llm evaluation: a psychometric alternative to static benchmarks\.arXiv preprint arXiv:2511\.04689\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022\)TruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- G\. Lior, T\. Frostig, G\. Stanovsky, and M\. Eyal \(2026\)Extending item response theory for efficient and meaningful multilingual evaluation\.arXiv preprint arXiv:2606\.15643\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.
- H\. Liu, J\. Liu, L\. Cui, Z\. Teng, N\. Duan, M\. Zhou, and Y\. Zhang \(2023\)Logiqa 2\.0—an improved dataset for logical reasoning in natural language understanding\.IEEE/ACM Transactions on Audio, Speech, and Language Processing31,pp\. 2947–2962\.Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- F\. M\. Lord and M\. R\. Novick \(1968\)Statistical theories of mental test scores\.Addison\-Wesley\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.15.15.6)\.
- F\. Maia Polo, L\. Weber, L\. Choshen, Y\. Sun, G\. Xu, and M\. Yurochkin \(2024\)TinyBenchmarks: evaluating llms with fewer examples\.arXiv preprint arXiv:2402\.14992\.Cited by:[Data Efficient LLM Benchmarking](https://arxiv.org/html/2608.02966#Sx6.p1.1)\.
- G\. Maris and T\. Bechger \(2009\)On interpreting the model parameters for the three parameter logistic model\.Measurement: Interdisciplinary Research and Perspectives7\(2\),pp\. 75–88\.Cited by:[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1)\.
- T\. Mihaylov, P\. Clark, T\. Khot, and A\. Sabharwal \(2018\)Can a suit of armor conduct electricity? A new dataset for open book question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- R\. D\. Penfield \(2014\)An NCME instructional module on polytomous item response theory models\.Educational Measurement: Issues and Practice33\(1\),pp\. 36–48\.Cited by:[LLM Nominal Response Model](https://arxiv.org/html/2608.02966#Sx2.p2.1)\.
- Y\. Perlitz, E\. Bandel, A\. Gera, O\. Arviv, L\. Ein\-Dor, E\. Shnarch, N\. Slonim, M\. Shmueli\-Scheuer, and L\. Choshen \(2024\)Efficient benchmarking \(of language models\)\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 2519–2536\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.139),[Link](https://aclanthology.org/2024.naacl-long.139/)Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p1.1)\.
- P\. Pezeshkpour and E\. Hruschka \(2024\)Large language models sensitivity to the order of options in multiple\-choice questions\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 2006–2017\.Cited by:[2nd item](https://arxiv.org/html/2608.02966#Sx2.I2.i2.p1.1),[Positional Bias\.](https://arxiv.org/html/2608.02966#Sx2.SSx3.SSS0.Px2.p1.1)\.
- G\. Rasch \(1960\)Studies in mathematical psychology: i\. probabilistic models for some intelligence and attainment tests\.\.Copenhagen: Danmarks Pædagogiske Institut\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.10.10.6)\.
- D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman \(2024\)GPQA: a graduate\-level Google\-proof Q&A benchmark\.InFirst Conference on Language Modeling,Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p1.1),[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- E\. San Martín, J\. González, and F\. Tuerlinckx \(2015\)On the unidentifiability of the fixed\-effects 3PL model\.Psychometrika80\(2\),pp\. 450–467\.Cited by:[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1)\.
- M\. Sap, H\. Rashkin, D\. Chen, R\. Le Bras, and Y\. Choi \(2019\)Social iqa: commonsense reasoning about social interactions\.InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing \(EMNLP\-IJCNLP\),pp\. 4463–4473\.Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- Y\. Suh and D\. M\. Bolt \(2010\)Nested logit models for multiple\-choice item response data\.Psychometrika75\(3\),pp\. 454–473\.Cited by:[LLM Nominal Response Model](https://arxiv.org/html/2608.02966#Sx2.p2.1),[Fisher Information from Distractor Identity\.](https://arxiv.org/html/2608.02966#Sx5.SSx6.SSSx2.Px1.p1.1)\.
- A\. Talmor, J\. Herzig, N\. Lourie, and J\. Berant \(2019\)CommonsenseQA: a question answering challenge targeting commonsense knowledge\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- D\. Thissen and L\. Steinberg \(1984\)A response model for multiple choice items\.Psychometrika49\(4\),pp\. 501–519\.Cited by:[LLM Nominal Response Model](https://arxiv.org/html/2608.02966#Sx2.p2.1)\.
- C\. Vania, P\. M\. Htut, W\. Huang, D\. Mungra, R\. Yuanzhe Pang, J\. Phang, H\. Liu, K\. Cho, and S\. R\. Bowman \(2021\)Comparing test sets with item response theory\.InAnnual Meeting of the Association for Computational Linguistics,Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.
- Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.\(2024\)MMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.InAdvances in Neural Information Processing Systems 37, Datasets and Benchmarks Track,Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p1.1),[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- Z\. Wang, W\. Wu, G\. Wang, G\. Ye, and Z\. Cheng \(2026\)MetaEval: measuring the discrimination of benchmarks for efficient llm evaluation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 33773–33781\.Cited by:[Baselines\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.46.46.6)\.
- P\. Welinder, S\. Branson, S\. Belongie, and P\. Perona \(2010\)The multidimensional wisdom of crowds\.InAdvances in Neural Information Processing Systems 23,Cited by:[Fallback pattern\.](https://arxiv.org/html/2608.02966#Sx2.SSx4.SSSx2.p1.1)\.
- X\. Wu, W\. Lin, O\. Akgul, and L\. Bauer \(2025\)Estimating llm consistency: a user baseline vs surrogate metrics\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 30518–30532\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p3.1)\.
- Z\. Yang, P\. Jian, and C\. Li \(2025\)Option symbol matters: investigating and mitigating multiple\-choice option symbol bias of large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 1902–1917\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p3.1)\.
- C\. Yeung \(2019\)Deep\-IRT: make deep learning based knowledge tracing explainable using item response theory\.InProceedings of the 12th International Conference on Educational Data Mining,Note:arXiv:1904\.11738Cited by:[Baselines\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.30.30.6)\.
- W\. Yu, Z\. Jiang, Y\. Dong, and J\. Feng \(2020\)ReClor: a reading comprehension dataset requiring logical reasoning\.InInternational Conference on Learning Representations,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- Y\. Zhang, X\. Fei, A\. Mohamed, S\. A\. Carneiro, M\. Konomi, M\. Geng, A\. Asaad, G\. Shang, and M\. Vazirgiannis \(2026\)The masked advantage: uncovering local\-language access to cultural knowledge in llms\.arXiv preprint arXiv:2606\.07422\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.
- C\. Zheng, H\. Zhou, F\. Meng, J\. Zhou, and M\. Huang \(2024\)Large language models are not robust multiple choice selectors\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 19426–19454\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p3.1),[2nd item](https://arxiv.org/html/2608.02966#Sx2.I2.i2.p1.1),[Positional Bias\.](https://arxiv.org/html/2608.02966#Sx2.SSx3.SSS0.Px2.p2.4),[Difficulty\-Gated Guessing Fallback](https://arxiv.org/html/2608.02966#Sx2.SSx4.p1.1)\.
- W\. Zhong, R\. Cui, Y\. Guo, Y\. Liang, S\. Lu, Y\. Wang, A\. Saied, W\. Chen, and N\. Duan \(2024\)AGIEval: a human\-centric benchmark for evaluating foundation models\.InFindings of the Association for Computational Linguistics: NAACL 2024,Cited by:[MCQ item bank\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px1.p1.1)\.
- H\. Zhou, H\. Huang, Z\. Zhao, L\. Han, H\. Wang, K\. Chen, M\. Yang, W\. Bao, J\. Dong, B\. Xu,et al\.\(2026\)Lost in benchmarks? rethinking large language model benchmarking with item response theory\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 35085–35093\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1),[Baselines\.](https://arxiv.org/html/2608.02966#Sx3.SSx6.SSSx2.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.02966#Sx4.T1.41.41.6)\.
- C\. Zhu, B\. Xu, Q\. Wang, Y\. Zhang, and Z\. Mao \(2023\)On the calibration of large language models and alignment\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 9778–9795\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p1.1),[1st item](https://arxiv.org/html/2608.02966#Sx2.I2.i1.p1.1),[Response Sharpness\.](https://arxiv.org/html/2608.02966#Sx2.SSx3.SSS0.Px1.p1.1)\.
- Y\. Zhu, D\. Liu, Z\. Lin, W\. Tong, S\. Zhong, and J\. Shao \(2025\)The llm already knows: estimating llm\-perceived question difficulty via hidden representations\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 1160–1176\.Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.
- Y\. Zhuang, Q\. Liu, Z\. A\. Pardos, P\. C\. Kyllonen, J\. Zu, Z\. Huang, S\. Wang, and E\. Chen \(2025\)Position: ai evaluation should learn from how we test humans\.External Links:2306\.10512,[Link](https://arxiv.org/abs/2306.10512)Cited by:[Introduction](https://arxiv.org/html/2608.02966#Sx1.p2.1)\.Similar Articles
Auditing LLM Benchmarks with Item Response Theory
This paper introduces an Item Response Theory-based method to detect mislabeled examples in LLM benchmarks at 95% precision, tracing errors to labeling heuristics and annotation issues.
Benchmarking LLM Competence on Logical Inference over Probability Operators
This paper introduces a benchmark of 14,320 procedurally-generated prompts for evaluating LLMs on logical inference over probability operators like 'probably', 'might', and 'must'. Testing 29 models, the authors find systematic answer biases and show that only 9 exceed random chance.
BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.
Benchmarking Different Methods of LLM Confidence Estimation
This article benchmarks various blackbox and whitebox methods for LLM confidence estimation, including verbalized confidence, linguistic uncertainty, reasoning-length, P(Answer), P(True), and self-consistency, comparing their effectiveness for tasks like active learning and safety classification.
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark
This paper introduces ProverbIT, a novel Italian benchmark of 100 multiple-choice questions to test LLMs' ability to complete proverbs. Evaluating 13 models, it finds that performance drops significantly in multiple-choice formats without correct answers, suggesting reliance on memorized patterns rather than deep semantic understanding.