超越分数对齐的 LLM-as-a-Judge 评估:残差评分难度的心理测量学分析

arXiv cs.CL 论文

摘要

该论文从心理测量学视角评估 LLM-as-a-Judge,使用 Many-Facet Rasch 模型将评分分解为潜在质量、评分者严苛度与残差难度,发现人类与 LLM 在汇总对齐之外仍存在明显的残差难度结构错配。

arXiv:2610.02877v1 Announce Type: new Abstract: Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.
查看原文
查看缓存全文

缓存时间: 2026/10/05 10:03

# Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty
Source: [https://arxiv.org/html/2610.02877](https://arxiv.org/html/2610.02877)
Longwei CongAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationEmail:[l\.cong@dipf\.de](mailto:[email protected])Sonja HahnAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationEmail:[s\.hahn@dipf\.de](mailto:[email protected])Sebastian GombertAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationEmail:[s\.gombert@dipf\.de](mailto:[email protected])Leon CamusAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationEmail:[l\.camus@dipf\.de](mailto:[email protected])Fabian ZehnerAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationAffiliation:Centre for International Student Assessment \(ZIB\)Email:[f\.zehner@dipf\.de](mailto:[email protected])Hendrik DrachslerAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationAffiliation:Faculty of Computer Science, Goethe University FrankfurtEmail:[h\.drachsler@dipf\.de](mailto:[email protected])Ulf KroehneAffiliation:DIPF \| Leibniz Institute for Research and Information in EducationAffiliation:Chemnitz University of TechnologyEmail:[u\.kroehne@dipf\.de](mailto:[email protected])

###### Abstract

Large language models \(LLMs\) are widely used as automatic judges, with validity typically assessed via alignment with human scores\. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult\. In this paper, we study this problem in summarization evaluation from a psychometric perspective\. We fit Many\-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating\-scale thresholds\. Building on this decomposition, we define residual hardness as a model\-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure\.

Across 17 open\-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness\. Human and LLM judges differ in which summary–dimension units remain difficult, and this mismatch is strongly dimension\-dependent\. Consistency shows a pronounced LLM\-hard shift, whereas coherence shows a human\-hard shift\. We further show that human\-easy but LLM\-hard cases are partially predictable from observable source–summary properties\. These findings suggest that aggregate human alignment reflects only part of LLM\-as\-a\-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human–LLM collaboration\.

## 1Introduction

Large language models \(LLMs\) have attracted substantial attention for their strong performance across a wide range of natural language processing tasks[Naveed et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib29)\. Their advanced language understanding and reasoning capabilities have also enabled their use as automatic judges for complex tasks, offering a scalable alternative or complement to traditional expert\-driven evaluation[Gu et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib4)\. LLM\-as\-a\-judge methods have been applied to social intelligence evaluation[Wang et al\. \(2024b\)](https://arxiv.org/html/2610.02877#bib.bib24), multimodal evaluation[Chen et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib25), and domain\-specific settings such as finance[Yu et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib26), law[Cheong et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib27), and education[Cong et al\. \(2026b\)](https://arxiv.org/html/2610.02877#bib.bib5)\. In this setting, an LLM is prompted to assign scores, rankings, labels, or pairwise preferences according to task\-specific criteria or natural\-language rubrics\. The appeal is that LLMs could approximate aspects of human judgment while overcoming the cost, time, and scalability bottlenecks inherent in human evaluation[Zheng et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib14)\. Consequently, prior work typically validates LLM judges by measuring their alignment with human judgments, most often through correlations with human ratings or agreement with human preference labels[Gu et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib4)\.

However, agreement on average quality provides only a coarse view of judging behavior and reveals little about the difficulty structure of evaluation\([Rodriguez et al\., 2021](https://arxiv.org/html/2610.02877#bib.bib8);[Cong et al\., 2026a](https://arxiv.org/html/2610.02877#bib.bib28)\)\. Two judge groups may assign similar overall scores while struggling with different evaluation instances\. This distinction is important because an automatic judge should not only reproduce aggregate human scores on benchmarks, but also behave reliably on individual cases in real\-world evaluation settings[Li et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib10);[Zhou et al\. \(2026\)](https://arxiv.org/html/2610.02877#bib.bib3)\. If humans and LLMs disagree about which cases are difficult to judge, aggregate alignment metrics may overstate the reliability of LLM\-as\-a\-judge systems\.

We argue that a shift toward a more measurement\-oriented perspective is essential for a deeper understanding of LLM judging behavior\. Rather than evaluating only whether LLM judges reproduce human scores on average, we need diagnostics that distinguish alignment in output quality from alignment in judging difficulty\.

To address this gap, we draw on Item Response Theory \(IRT\), a psychometric framework for modeling observed ratings in terms of latent traits and systematic measurement facets\. Specifically, we use a rating\-scale Many\-Facet Rasch Model \(MFRM\)[Eckes \(2015\)](https://arxiv.org/html/2610.02877#bib.bib18)because it separates latent item quality from systematic rating effects such as rater severity, dimension\-specific severity, and rating\-scale thresholds\. Building on this model, we define residual hardness as the average absolute residual between observed ratings and the ratings expected by the fitted MFRM\. Conceptually, this metric provides a post\-fit diagnostic of local judging instability after accounting for the systematic facets of the evaluation process\. It therefore distinguishes low estimated summary quality from large residual deviations in the rating process, without treating residual hardness as an intrinsic latent measure of item difficulty\.

We study this diagnostic in a SummEval case study of summarization evaluation, a setting in which each system summary is rated along multiple dimensions, such as coherence, consistency, fluency, and relevance, that place different demands on judges\. This makes summarization a useful testbed for analyzing whether human–LLM alignment extends beyond aggregate quality scores to the structure of judging difficulty[Fabbri et al\. \(2021\)](https://arxiv.org/html/2610.02877#bib.bib21);[Liu et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib12)\.

Our analysis is guided by three research questions\.

##### RQ1

Does alignment between human raters and LLM judges in latent summary quality also imply alignment in residual hardness?

##### RQ2

How does residual\-hardness mismatch vary across evaluation dimensions?

##### RQ3

Can textual features predict which summaries are easy for humans but difficult for LLMs to evaluate?

These questions are practically important for human–LLM collaboration scenarios\. Cases that are likely to be human\-easy but LLM\-hard can be routed to human review, additional source\-grounding checks, or stronger judge models\([Bondi et al\., 2022](https://arxiv.org/html/2610.02877#bib.bib22);[Wang et al\., 2024c](https://arxiv.org/html/2610.02877#bib.bib23)\)\. A difficulty\-aware view therefore supports more targeted human–LLM collaboration in evaluation\.

The contributions of this paper are as follows:

- •We propose a psychometrically grounded diagnostic that separates latent instance quality from residual judging difficulty, providing a more fine\-grained view of LLM\-as\-a\-judge behavior\.
- •We show that human–LLM alignment in latent instance quality does not imply alignment in residual hardness, revealing substantial differences in the residual\-hardness patterns of human and LLM judge panels\.
- •We demonstrate that human\-easy but LLM\-hard cases are partially predictable from observable summary\-, source\-, and system\-level features, supporting difficulty\-aware human–LLM evaluation pipelines\.

## 2Background

### 2\.1LLM\-as\-a\-judge and Human Alignment

As complex evaluation tasks become increasingly difficult to assess with fixed references or surface\-form metrics, recent work has turned to large language models as automatic judges[Gu et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib4)\. LLM judges have been used for general\-purpose evaluation of model outputs, including summarization, dialogue, and instruction following[Liu et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib12);[Fu et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib13);[Zheng et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib14), as well as for data annotation[Tan et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib11)and domain\-specific educational assessment[Gombert et al\. \(2026\)](https://arxiv.org/html/2610.02877#bib.bib9)\.

The validity of LLM\-as\-a\-judge systems is typically assessed by measuring alignment with human judgments\. Prior work reports correlations with human quality ratings[Liu et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib12);[Fu et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib13), rank correlations between systems[Liu et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib12);[Fu et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib13), agreement with pairwise human preferences[Zheng et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib14);[Li et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib16);[Wang et al\. \(2024d\)](https://arxiv.org/html/2610.02877#bib.bib15), or agreement with aggregated human labels[Bavaresco et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib17)\. These studies show that LLMs can serve as useful evaluators in some settings, but also that their reliability varies substantially across datasets, evaluated properties, judge models, and forms of human annotation[Bavaresco et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib17)\.

This literature establishes human alignment as a central criterion for evaluating LLM judges\. However, most existing alignment metrics focus on whether LLMs reproduce human scores, rankings, preferences, or labels\. They provide less information about whether human and LLM judges share the same structure of judging difficulty\. A judge may correlate with humans in aggregate quality assessment while still diverging on which evaluation instances are unstable, ambiguous, or difficult to rate\. This distinction is important for model training[Yu et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib30), benchmark construction[Zhou et al\. \(2026\)](https://arxiv.org/html/2610.02877#bib.bib3), and human–AI collaboration[Pan et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib31), where the reliability of individual judgments matters as much as aggregate agreement\. We therefore adopt a psychometric view of LLM\-as\-a\-judge alignment that separates latent quality and systematic rating facets from post\-fit local residual variation\.

### 2\.2Item Response Theory

Item Response Theory \(IRT\) provides a measurement framework for modeling observed responses in terms of latent respondent characteristics and item properties[van der Linden \(2016\)](https://arxiv.org/html/2610.02877#bib.bib6)\. The Rasch model is a foundational IRT model in which the probability of a correct response depends on the difference between a respondent’s latent ability and an item’s difficulty:

P⁡\(ym​n=1∣θm,bn\)=σ⁡\(θm−bn\),P\(y\_\{mn\}=1\\mid\\theta\_\{m\},b\_\{n\}\)=\\sigma\(\\theta\_\{m\}\-b\_\{n\}\),whereym​n∈\{0,1\}y\_\{mn\}\\in\\\{0,1\\\}denotes the observed response of respondentmmto itemnn,θm\\theta\_\{m\}denotes the latent ability of respondentmm,bnb\_\{n\}denotes the difficulty of itemnn, andσ⁡\(⋅\)\\sigma\(\\cdot\)is the logistic sigmoid\.

The Many\-Facet Rasch Model \(MFRM\)[Eckes \(2015\)](https://arxiv.org/html/2610.02877#bib.bib18)extends the Rasch model by incorporating additional facets that may systematically influence observed scores\. It allows the measurement model to include further sources of variation, such as rater severity, task difficulty, scoring dimensions, or category thresholds\.

Recent work has begun to apply IRT to LLM\-as\-a\-judge evaluation in several ways[Ye et al\. \(2026\)](https://arxiv.org/html/2610.02877#bib.bib32)\. Choi et al\. use a graded\-response IRT model to diagnose LLM\-as\-a\-judge reliability, formalizing reliability in terms of prompt\-level intrinsic consistency and alignment with human quality assessments[Choi et al\. \(2026\)](https://arxiv.org/html/2610.02877#bib.bib7)\. Recent educational assessment work has also applied many\-facet Rasch measurement to evaluate whether LLM raters exhibit acceptable severity, consistency, and rater effects in automated scoring\([Jiao et al\., 2025](https://arxiv.org/html/2610.02877#bib.bib19);[Wang et al\., 2025](https://arxiv.org/html/2610.02877#bib.bib20)\)\. These studies establish psychometric models as useful tools for assessing LLM raters\. However, they primarily ask whether LLM raters are reliable and human\-aligned at the score or rater level\. In contrast, we analyze whether human annotators and LLM judges exhibit similar post\-fit residual\-hardness patterns\.

## 3Method

We first obtain ratings from a panel of LLM judges on the summaries and dimensions rated by human annotators\. We then analyze these ratings with MFRM, modeling summary quality, rater severity, dimension effects, and rating\-scale thresholds\. Based on the fitted model, we compare human and LLM judges in latent measurement space and introduce a residual\-based measure of judging hardness\.

### 3\.1Dataset and LLMs

We use SummEval[Fabbri et al\. \(2021\)](https://arxiv.org/html/2610.02877#bib.bib21), a multi\-rater benchmark for evaluating generated summaries\. SummEval was constructed in the context of news summarization evaluation and contains summaries generated for source articles from the CNN/DailyMail dataset\. The human\-annotated subset used in our study consists of 1,600 system summaries, obtained from 100 source documents and 16 summarization systems\. Each evaluated unit corresponds to a system summary for a source document\. The dataset includes summaries produced by both extractive and abstractive summarization systems\. Human ratings cover four evaluation dimensions: coherence, consistency, fluency, and relevance\. The original annotations include three expert annotators and five crowd annotators, with anonymized rater IDs\. The primary human MFRM uses all eight human raters, comprising the three expert annotators and five crowd annotators\.

To construct LLM\-as\-a\-judge annotations, we prompt 17 open\-weight language models to rate the same summary–dimension units on the same 1–5 scale\. The models span multiple developers and model families, including Google[Team et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib42), Qwen[Qwen Team \(2026\)](https://arxiv.org/html/2610.02877#bib.bib43), Microsoft[Abdin et al\. \(2024b\)](https://arxiv.org/html/2610.02877#bib.bib44);[Abdin et al\. \(2024a\)](https://arxiv.org/html/2610.02877#bib.bib48), Meta\-Llama[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2610.02877#bib.bib45), MistralAI[Liu et al\. \(2026\)](https://arxiv.org/html/2610.02877#bib.bib46), and OpenChat[Wang et al\. \(2024a\)](https://arxiv.org/html/2610.02877#bib.bib47), with parameter sizes ranging from 0\.8B to 14B\. We choose this panel to obtain a diverse, reproducible, and locally runnable set of LLM judges\. Decoding is performed greedily withdo\_sample=Falseand temperature set to 0\. Under this setting, token sampling is disabled, so changing the sampling seed does not produce independent generations\. Details of the model set and the prompt are provided in Appendix[A\.1](https://arxiv.org/html/2610.02877#A1.SS1)\.

### 3\.2Many\-Facet Rasch Measurement

We fit rating\-scale MFRM models separately for human raters and LLM raters\. For a ratingyr​i​dy\_\{rid\}assigned by raterrrto summaryiion dimensiondd, the adjacent\-category logit is defined as:

log⁡P⁡\(yr​i​d=c\+1\)P⁡\(yr​i​d=c\)=θi−ρr−δd−τc,c=1,…,4\.\\begin\{split\}\\log\\frac\{P\(y\_\{rid\}=c\+1\)\}\{P\(y\_\{rid\}=c\)\}&=\\theta\_\{i\}\-\\rho\_\{r\}\-\\delta\_\{d\}\-\\tau\_\{c\},\\\\ &\\qquad c=1,\\ldots,4\.\\end\{split\}\(1\)
Here,θi\\theta\_\{i\}denotes the latent quality of summaryii,ρr\\rho\_\{r\}denotes the severity of raterrr,δd\\delta\_\{d\}denotes the severity associated with evaluation dimensiondd, andτc\\tau\_\{c\}denotes the threshold between adjacent rating categoriesccandc\+1c\+1\. Higherθi\\theta\_\{i\}indicates higher estimated summary quality\. Higherρr\\rho\_\{r\}indicates a stricter rater, higherδd\\delta\_\{d\}indicates a dimension on which high scores are harder to obtain after controlling for summary quality and rater severity, and higherτc\\tau\_\{c\}indicates a more difficult transition to the next higher rating category\. Details of the estimation procedure are provided in Appendix[A\.2\.1](https://arxiv.org/html/2610.02877#A1.SS2.SSS1)\.

We assess the adequacy of the MFRM estimation using standard psychometric diagnostics\. First, we examine standardized residuals and report MAE and RMSE as descriptive summaries of rating\-level prediction error\. Second, we report item separation reliability, which indicates whether the calibrated measures distinguish summaries with sufficient precision relative to measurement error\. We additionally verify that all five response categories are used by both human and LLM raters\.

### 3\.3Residual Hardness

Classical agreement metrics measure whether two groups assign similar scores, but they do not directly capture whether the same items exhibit larger local deviations after accounting for the main structure of the rating process\. We use the term residual hardness as a post\-fit diagnostic of local judging instability\. In the MFRM, expected ratings are determined by summary quality, rater severity, dimension severity, and rating\-scale thresholds\. If ratings for a summary–dimension unit remain far from their expected values after these facets have been accounted for, the unit exhibits high residual hardness\.

For each observed rating, the fitted MFRM yields an expected rating:

y^r​i​d=E\[yr​i​d∣θ^i,ρ^r,δ^d,τ^\]\.\\hat\{y\}\_\{rid\}=E\[y\_\{rid\}\\mid\\hat\{\\theta\}\_\{i\},\\hat\{\\rho\}\_\{r\},\\hat\{\\delta\}\_\{d\},\\hat\{\\tau\}\]\.\(2\)
The response residual is:

er​i​d=yr​i​d−y^r​i​d\.e\_\{rid\}=y\_\{rid\}\-\\hat\{y\}\_\{rid\}\.\(3\)
For a summary\-dimension unit\(i,d\)\(i,d\), residual hardness is defined as the mean absolute residual over raters in groupgg:

Hi​d\(g\)=1\|Rg\|​∑r∈Rg\|er​i​d\|\.H^\{\(g\)\}\_\{id\}=\\frac\{1\}\{\|R\_\{g\}\|\}\\sum\_\{r\\in R\_\{g\}\}\|e\_\{rid\}\|\.\(4\)
Residual hardness is not low summary quality or a conventional latent item\-difficulty parameter\. Rather, it summarizes the magnitude of local deviations from the fitted MFRM after accounting for the modeled facets\. We therefore interpret it as a post\-fit diagnostic of local judging instability rather than as an intrinsic latent measure of item difficulty, ambiguity, or uncertainty\. Larger values may reflect local rater disagreement, model misfit, or remaining unmodeled structure\. The human and LLM analyses use the same MFRM form but separately estimated parameters, and we compare whether the resulting panels exhibit similar residual\-hardness patterns\. Appendix[A\.2\.2](https://arxiv.org/html/2610.02877#A1.SS2.SSS2)relates residual hardness to local rater agreement, and Appendix[A\.2\.6](https://arxiv.org/html/2610.02877#A1.SS2.SSS6)examines sensitivity to the use of separate calibrations\.

### 3\.4Predicting Human\-Easy and LLM\-Hard Cases

To examine whether human–LLM mismatch is predictable from observable features, we formulate a binary classification task at the summary–dimension level\. The positive class consists of units in the human\-easy/LLM\-hard quadrant, and all other units are treated as negative\. We select features that capture observable properties of the source–summary pair\. TF–IDF features capture lexical and topical content, length and compression features reflect information density and omission, and sentence counts provide coarse structural cues\. Digit and negation counts approximate factual\-detail and verification demands, while specific system IDs are excluded to avoid label leakage\. We train class\-balanced logistic regression models using these features\. A detailed list of input features is provided in Appendix[A\.3\.1](https://arxiv.org/html/2610.02877#A1.SS3.SSS1)\.

We evaluate both a pooled model over all dimensions and dimension\-specific models using 5\-fold stratified cross\-validation\. In each fold, 80% of the data are used for training and 20% for testing while preserving the class ratio\. Performance is reported using AUROC and AUPRC\. Because the task is imbalanced, we interpret AUPRC relative to the positive\-class rate\.

## 4Results

### 4\.1RQ1

##### Model adequacy\.

Before comparing human and LLM judging structures, we verify that the MFRM estimations provide usable measurement models\. Table[1](https://arxiv.org/html/2610.02877#S4.T1)summarizes the main fit diagnostics, including MAE and RMSE as descriptive summaries of rating\-level prediction error, mean and standard deviation of standardized residuals, and item separation reliability\. Across both estimations, standardized residuals are centered near zero and have standard deviations close to one\. Item separation reliability is high in both estimations, indicating that the model distinguishes latent summary quality with sufficient precision\. These diagnostics support using the fitted MFRM estimates to compare latent quality and residual hardness across human and LLM judge groups\. All five response categories are used by both human and LLM raters, although their usage distributions differ\.

Table 1:MFRM fit diagnostics\. MAE and RMSE are descriptive summaries of rating\-level prediction error\. Res\. M and Res\. SD denote the mean and standard deviation of standardized residuals\. Item rel\. denotes item separation reliability\.
##### Quality alignment does not imply hardness alignment\.

At the summary level, LLM judges show moderate alignment with human latent summary quality\. The correlation between human\- and LLM MFRM quality estimates is Pearsonr=0\.372r=0\.372and Spearmanρ=0\.371\\rho=0\.371over 1,600 summaries\. This indicates that LLM judges capture part of the human ranking of summary quality\.

However, this alignment does not transfer to residual hardness\. When residual hardness is aggregated at the same summary level, the human–LLM correlation drops tor=0\.283r=0\.283andρ=0\.176\\rho=0\.176\. The gap is even sharper at the summary\-dimension level, where each point corresponds to a summary evaluated on one dimension\. There, human and LLM residual hardness correlate only weakly, withr=0\.145r=0\.145andρ=0\.089\\rho=0\.089over 6,400 summary\-dimension units\. This pattern is also visible in Figure[1](https://arxiv.org/html/2610.02877#S4.F1)\. The quality estimates form a clear positive trend, whereas the hardness estimates are substantially more diffuse, especially at the summary–dimension level\.

![Refer to caption](https://arxiv.org/html/2610.02877v1/plots/human_llm_quality_hardness_scatter_combined_3.png)Figure 1:Human–LLM alignment in latent summary quality \(left\), summary\-level residual hardness \(middle\), and summary–dimension residual hardness \(right\)\.
##### Robustness to judge capability and human\-panel composition\.

The weak residual\-hardness alignment is not driven solely by the smallest LLM judges\. When we refit the LLM MFRM using only the seven judges with at least 7B parameters, quality alignment remains moderate \(ρ=0\.408\\rho=0\.408\), whereas summary–dimension residual\-hardness alignment remains near zero \(ρ=0\.039\\rho=0\.039\)\. We additionally refit the human MFRM using only the three expert annotators and compare it with the independently fitted≥7\\geq 7B LLM panel\. In this setting, quality alignment is substantially stronger \(ρ=0\.688\\rho=0\.688\), while residual\-hardness alignment remains weak \(ρ=0\.092\\rho=0\.092\)\. These analyses indicate that the contrast between quality alignment and residual\-hardness alignment is not explained solely by the smallest LLM judges or by pooling expert and crowd human ratings\.

These results answer RQ1 negatively\. LLMs partially align with humans in latent summary quality, but this quality alignment does not carry over to the structure of residual hardness\. In other words, LLMs can rank summaries in a moderately human\-aligned way while exhibiting substantially different residual\-hardness patterns\.

### 4\.2RQ2

We next examine whether the weak overall residual\-hardness alignment is accompanied by systematic differences across evaluation dimensions\. We classified each summary–dimension unit as easy or hard within each judge group by thresholding residual hardness at the group’s overall median\. This median split is an operational definition rather than a theoretically privileged cutoff\. Appendix[A\.3\.2](https://arxiv.org/html/2610.02877#A1.SS3.SSS2)examines the sensitivity of the downstream RQ3 prediction analysis to a stricter quartile\-based quadrant definition\.

Table[2](https://arxiv.org/html/2610.02877#S4.T2)reports the percentage of units classified as hard for each judge group\. Consistency shows the largest LLM\-hard shift, with 74\.7% of units classified as LLM\-hard compared with 50\.8% human\-hard\. Coherence shows the opposite pattern, with 57\.3% human\-hard but only 34\.8% LLM\-hard\. Fluency and relevance are closer to balanced, with much smaller hard\-rate gaps\. This structure is also visible in Figure[2](https://arxiv.org/html/2610.02877#S4.F2), where consistency shows a pronounced concentration of points in the human\-easy/LLM\-hard region, whereas coherence shows the opposite tendency\. This pattern also persists under the expert\-only versus≥7\\geq 7B LLM comparison, where a\+26\.2\+26\.2percentage point consistency gap remains\.

Table 2:Dimension\-level residual hardness gaps\. Values are percentages of summary–dimension units classified as hard within each judge group\. Gap is LLM minus human hard\-rate\.At the system\-type level, the mismatch is weaker than at the dimension level\. As reported in Table[3](https://arxiv.org/html/2610.02877#S4.T3), extractive systems show a modest LLM\-hard shift, whereas abstractive systems show a small human\-hard shift\.

We additionally examined rater\-level fit and misfit sensitivity in Appendix[A\.2\.5](https://arxiv.org/html/2610.02877#A1.SS2.SSS5)\. The main residual\-hardness findings remain stable after excluding raters flagged by descriptive infit or outfit criteria\. To assess whether the conclusions depend on direct comparability of raw residual magnitudes across separate calibrations, we also report normalization\- and calibration\-based sensitivity checks in Appendix[A\.2\.6](https://arxiv.org/html/2610.02877#A1.SS2.SSS6)\. The weak residual\-hardness alignment and the dominant consistency/coherence contrast remain stable across these alternative specifications\.

Table 3:System\-level residual hardness gaps\. Values are percentages of summary–dimension units classified as hard\. Gap is LLM minus human hard\-rate\.![Refer to caption](https://arxiv.org/html/2610.02877v1/plots/human_llm_hardness_scatter_4dimension_panels.png)Figure 2:Human–LLM residual hardness by evaluation dimension\. Dashed lines indicate the overall median residual hardness within each judge group, partitioning the space into human\-easy/LLM\-easy, human\-easy/LLM\-hard, human\-hard/LLM\-easy, and human\-hard/LLM\-hard regions\.
### 4\.3RQ3

For RQ3, we analyzed whether this mismatch can be predicted directly from observable properties of the source document and summary\. Table[4](https://arxiv.org/html/2610.02877#S4.T4)shows that human\-easy/LLM\-hard cases are partially predictable from observable features\. The pooled all\-dimension model reaches AUROC 0\.631 and AUPRC 0\.341, compared with a positive rate of 23\.3%\. Relevance is the most predictable task, with AUROC 0\.717, followed by consistency, fluency, and coherence\. These results suggest that the human\-easy/LLM\-hard quadrant is not arbitrary\. It is associated with observable properties of the evaluated summary and source document\. This predictive signal is robust to feature ablations and stricter grouped split strategies\. Additional ablations, confidence intervals, grouped\-split robustness checks and quadrant\-definition sensitivity checks are reported in Appendix[A\.3\.2](https://arxiv.org/html/2610.02877#A1.SS3.SSS2)\.

Table 4:Prediction of human\-easy/LLM\-hard cases from all summary\-dimension units\. The positive class is the human\-easy/LLM\-hard quadrant; all other quadrants are negative\.Figure[3](https://arxiv.org/html/2610.02877#S4.F3)interprets the dimension\-specific classifiers using signed SHAP values[Lundberg and Lee \(2017\)](https://arxiv.org/html/2610.02877#bib.bib37)\. Because several surface features are correlated, particularly source length, summary length, and compression ratio, we interpret the SHAP values descriptively as model associations rather than as independent or causal feature effects\. For relevance, higher compression ratio is associated with SHAP contributions toward human\-easy/LLM\-hard classification, whereas longer summaries tend to contribute in the opposite direction\. Coherence shows a similar pattern, where longer source documents and higher compression ratios are associated with contributions toward the human\-easy/LLM\-hard quadrant, while longer summaries tend to contribute in the opposite direction\.

![Refer to caption](https://arxiv.org/html/2610.02877v1/plots/shap_beeswarm_top5_by_dimension.png)Figure 3:SHAP beeswarm plots for human\-easy/LLM\-hard classification by evaluation dimension\. Positive SHAP values indicate features that push predictions toward the human\-easy/LLM\-hard quadrant, while negative values indicate the opposite\. Color denotes feature value from low to high\.Consistency exhibits a different profile\. Its most influential predictors include digit count, sentence count, source length, negation count, and coarse system type, indicating that mismatch in this dimension is tied more strongly to factual\-detail and verification\-related cues than to compression alone\. Fluency shows a weaker but still interpretable pattern, in which longer summaries contribute most strongly to human\-easy/LLM\-hard classification, with smaller contributions from system type, digit count, and compression ratio\.

## 5Discussion

Our results show that human alignment in LLM\-as\-judge evaluation is not a single property\. Although LLM judges moderately align with human latent summary quality, this agreement does not imply that they share the same residual hardness structure\. The weak correlation between human and LLM residual hardness indicates that the cases that remain difficult after controlling for summary quality, rater severity, dimension severity, and rating thresholds are only partially overlapping across judge groups\. This distinction matters because aggregate quality alignment can make an evaluator appear reliable while masking item\-level instability[Lalor et al\. \(2016\)](https://arxiv.org/html/2610.02877#bib.bib1);[Lalor et al\. \(2018\)](https://arxiv.org/html/2610.02877#bib.bib2)\.

The mismatch structure appears at the dimension level\. Consistency shows a large LLM\-hard shift, whereas coherence shows the opposite human\-hard shift\. Consistency judgments require checking whether summary claims are supported by the source document[Fabbri et al\. \(2021\)](https://arxiv.org/html/2610.02877#bib.bib21);[Kryscinski et al\. \(2020\)](https://arxiv.org/html/2610.02877#bib.bib33);[Wang et al\. \(2020\)](https://arxiv.org/html/2610.02877#bib.bib34), which may impose a factual\-verification burden on LLM judges\. In contrast, coherence judgments depend more on discourse organization, readability, and holistic interpretation[Zhao et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib35);[Barzilay and Lapata \(2005\)](https://arxiv.org/html/2610.02877#bib.bib36);[Fabbri et al\. \(2021\)](https://arxiv.org/html/2610.02877#bib.bib21), where human annotators may apply more nuanced or variable standards\. Fluency and relevance show smaller gaps, suggesting that the divergence is not a general failure of the LLM judges but a dimension\-specific difference in residual judging behavior\.

The prediction results further indicate that part of the human–LLM residual\-hardness mismatch is associated with observable source–summary properties, consistent with prior evidence that LLM\-as\-judge behavior exhibits persistent biases and dimension\-dependent unreliability[Ye et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib38);[Shen et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib39)\. In our setting, the class\-balanced logistic regression model predicts human\-easy/LLM\-hard cases above the positive\-class baseline, providing modest evidence that part of the observed mismatch is associated with observable source–summary properties\. SHAP analysis reveals two feature profiles\. For relevance and coherence, compression\-related features dominate, providing modest evidence that some of the observed mismatch is predictable from observable source–summary properties\. For consistency, digit count, sentence count, source length, negation count, and system type point to factual\-detail and verification\-related cues, consistent with prior work framing factual consistency evaluation as verification\-intensive[Chen et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib40);[Gabriel et al\. \(2021\)](https://arxiv.org/html/2610.02877#bib.bib41)\. Overall, part of the residual\-hardness mismatch is measurable and partially predictable from observable text properties\.

These findings have methodological implications for evaluating LLM judges\. Existing validation practices[Liu et al\. \(2023\)](https://arxiv.org/html/2610.02877#bib.bib12);[Gu et al\. \(2025\)](https://arxiv.org/html/2610.02877#bib.bib4)often treat human alignment as a score\-level property, measured through correlations, rank agreement, or preference agreement\. Our results show that such metrics are incomplete\. Two judge groups may agree on the latent quality ordering of outputs while disagreeing on which individual cases are difficult to judge\. Psychometric residual diagnostics therefore provide a complementary view of evaluator reliability\. In particular, residual hardness and dimension\-specific hardness gaps can reveal failure modes that are invisible to aggregate agreement metrics\.

Practically, this suggests that LLM judges should not be deployed as uniform replacements for human annotators\. Their reliability depends on the evaluation dimension and on properties of the evaluated text\. A difficulty\-aware evaluation pipeline could use such diagnostics to route cases selectively[Bondi et al\. \(2022\)](https://arxiv.org/html/2610.02877#bib.bib22);[Wang et al\. \(2024c\)](https://arxiv.org/html/2610.02877#bib.bib23)\. Cases predicted to be human\-easy but LLM\-hard may be prioritized for human review, additional source\-grounding checks, or stronger judge models\. In this sense, these diagnostics may support more targeted human–LLM collaboration rather than simply accepting or rejecting LLM judges at the aggregate level\.

## 6Conclusion

We presented a psychometric analysis of human–LLM alignment in summarization evaluation\. Moving beyond aggregate agreement metrics, we asked whether human and LLM judges share the same structure of residual judging difficulty after controlling for latent summary quality, rater severity, dimension severity, and rating\-scale thresholds\. Using a Many\-Facet Rasch Model, we showed that moderate human–LLM alignment in latent summary quality does not imply alignment in residual hardness\. In particular, human and LLM judges differ substantially in their summary–dimension residual\-hardness patterns, with especially pronounced mismatch for consistency and coherence\.

We further showed that human\-easy but LLM\-hard cases are partially predictable from observable properties of the source–summary pair, with different predictive patterns across evaluation dimensions\. These findings suggest that score\-level human alignment provides only a partial view of LLM\-as\-judge reliability\. Psychometric residual diagnostics can reveal failure modes that remain hidden under aggregate agreement metrics\.

More broadly, our results argue for a difficulty\-aware view of LLM\-as\-judge evaluation\. Rather than treating LLM judges as uniform replacements for human annotators, future evaluation pipelines should account for which cases are likely to remain difficult for LLMs even when overall score alignment appears acceptable\. This perspective can support more targeted human–LLM collaboration and more informative evaluation of automatic judges\. The empirical findings are specific to the SummEval setting studied here, and establishing whether similar patterns generalize across tasks, domains, and stronger judge populations remains an important direction for future work\.

## Limitations

Our study has several limitations\. First, we examine residual hardness mismatch in summarization evaluation using a single multi\-rater benchmark, which limits the generalizability of the findings to other LLM\-as\-judge settings\. However, this benchmark is well suited to our research question because it provides repeated ratings across multiple evaluation dimensions and includes anonymized rater identifiers required for the MFRM\. More broadly, public summarization benchmarks with repeated human ratings, identifiable raters, and summary\-level multidimensional quality scores are extremely limited, and SummEval is, to our knowledge, the most suitable publicly available benchmark for this analysis\.

Second, our results are based on a fixed panel of open\-weight LLM judges under a single prompting setup\. Different model families, prompting strategies, or evaluation protocols may yield different residual hardness structures\.

Third, residual hardness should not be interpreted as a standalone latent construct of judging difficulty or ambiguity\. Rather, it is a post\-fit diagnostic that operationalizes local judging difficulty through residual deviations from the fitted MFRM\. Its magnitude may reflect local rater disagreement, model misspecification, or remaining structure not captured by the model\. Accordingly, our analyses establish differences in residual\-hardness patterns between the evaluated human and LLM panels, but do not assume that these residuals have identical substantive meaning across panels\.

Fourth, the RQ3 prediction analysis identifies correlates of human–LLM mismatch from observable features, but does not establish the causal mechanisms underlying these patterns\.

## Acknowledgments

This research was conducted within the project “Assessment for Learning with AI \(ALwAI\)” funded by the Leibniz Association under the Leibniz Competition \(project no\. T163/2024\)\.

## References

- M\. Abdin, J\. Aneja, H\. Awadalla, A\. Awadallah, A\. A\. Awan, N\. Bach, A\. Bahree, A\. Bakhtiari, J\. Bao, H\. Behl, A\. Benhaim, M\. Bilenko, J\. Bjorck, S\. Bubeck, M\. Cai, Q\. Cai, V\. Chaudhary, D\. Chen, D\. Chen, W\. Chen, Y\. Chen, Y\. Chen, H\. Cheng, P\. Chopra, X\. Dai, M\. Dixon, R\. Eldan, V\. Fragoso, J\. Gao, M\. Gao, M\. Gao, A\. Garg, A\. D\. Giorno, A\. Goswami, S\. Gunasekar, E\. Haider, J\. Hao, R\. J\. Hewett, W\. Hu, J\. Huynh, D\. Iter, S\. A\. Jacobs, M\. Javaheripi, X\. Jin, N\. Karampatziakis, P\. Kauffmann, M\. Khademi, D\. Kim, Y\. J\. Kim, L\. Kurilenko, J\. R\. Lee, Y\. T\. Lee, Y\. Li, Y\. Li, C\. Liang, L\. Liden, X\. Lin, Z\. Lin, C\. Liu, L\. Liu, M\. Liu, W\. Liu, X\. Liu, C\. Luo, P\. Madan, A\. Mahmoudzadeh, D\. Majercak, M\. Mazzola, C\. C\. T\. Mendes, A\. Mitra, H\. Modi, A\. Nguyen, B\. Norick, B\. Patra, D\. Perez\-Becker, T\. Portet, R\. Pryzant, H\. Qin, M\. Radmilac, L\. Ren, G\. de Rosa, C\. Rosset, S\. Roy, O\. Ruwase, O\. Saarikivi, A\. Saied, A\. Salim, M\. Santacroce, S\. Shah, N\. Shang, H\. Sharma, Y\. Shen, S\. Shukla, X\. Song, M\. Tanaka, A\. Tupini, P\. Vaddamanu, C\. Wang, G\. Wang, L\. Wang, S\. Wang, X\. Wang, Y\. Wang, R\. Ward, W\. Wen, P\. Witte, H\. Wu, X\. Wu, M\. Wyatt, B\. Xiao, C\. Xu, J\. Xu, W\. Xu, J\. Xue, S\. Yadav, F\. Yang, J\. Yang, Y\. Yang, Z\. Yang, D\. Yu, L\. Yuan, C\. Zhang, C\. Zhang, J\. Zhang, L\. L\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, and X\. ZhouPhi\-3 technical report: a highly capable language model locally on your phone\.External Links:2404\.14219,[Link](https://arxiv.org/abs/2404.14219)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- Abdinet al\.\(2024b\)M\. Abdin, J\. Aneja, H\. Behl, S\. Bubeck, R\. Eldan, S\. Gunasekar, M\. Harrison, R\. J\. Hewett, M\. Javaheripi, P\. Kauffmann, J\. R\. Lee, Y\. T\. Lee, Y\. Li, W\. Liu, C\. C\. T\. Mendes, A\. Nguyen, E\. Price, G\. de Rosa, O\. Saarikivi, A\. Salim, S\. Shah, X\. Wang, R\. Ward, Y\. Wu, D\. Yu, C\. Zhang, and Y\. ZhangPhi\-4 technical report\.External Links:2412\.08905,[Link](https://arxiv.org/abs/2412.08905)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- Barzilay and Lapata \(2005\)R\. Barzilay and M\. LapataModeling local coherence: an entity\-based approach\.InProceedings of the 43rd Annual Meeting of the Association for Computational Linguistics \(ACL’05\),K\. Knight, H\. T\. Ng, and K\. Oflazer \(Eds\.\),Ann Arbor, Michigan,pp\. 141–148\.External Links:[Link](https://aclanthology.org/P05-1018/),[Document](https://dx.doi.org/10.3115/1219840.1219858)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p2.1)\.
- Bavarescoet al\.\(2025\)A\. Bavaresco, R\. Bernardi, L\. Bertolazzi, D\. Elliott, R\. Fernández, A\. Gatt, E\. Ghaleb, M\. Giulianelli, M\. Hanna, A\. Koller, A\. Martins, P\. Mondorf, V\. Neplenbroek, S\. Pezzelle, B\. Plank, D\. Schlangen, A\. Suglia, A\. K\. Surikuchi, E\. Takmaz, and A\. TestoniLLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 238–255\.External Links:[Link](https://aclanthology.org/2025.acl-short.20/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-short.20),ISBN 979\-8\-89176\-252\-7Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p2.1)\.
- Bondiet al\.\(2022\)E\. Bondi, R\. Koster, H\. Sheahan, M\. Chadwick, Y\. Bachrach, T\. Cemgil, U\. Paquet, and K\. DvijothamRole of human\-ai interaction in selective prediction\.Proceedings of the AAAI Conference on Artificial Intelligence36\(5\),pp\. 5286–5294\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/20465),[Document](https://dx.doi.org/10.1609/aaai.v36i5.20465)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.SS0.SSS0.Px3.p2.1),[§5](https://arxiv.org/html/2610.02877#S5.p5.1)\.
- Chenet al\.\(2024\)D\. Chen, R\. Chen, S\. Zhang, Y\. Wang, Y\. Liu, H\. Zhou, Q\. Zhang, Y\. Wan, P\. Zhou, and L\. SunMLLM\-as\-a\-judge: assessing multimodal llm\-as\-a\-judge with vision\-language benchmark\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1)\.
- Chenet al\.\(2023\)S\. Chen, S\. Gao, and J\. HeEvaluating factual consistency of summaries with large language models\.External Links:2305\.14069,[Link](https://arxiv.org/abs/2305.14069)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p3.1)\.
- Cheonget al\.\(2024\)I\. Cheong, K\. Xia, K\. J\. K\. Feng, Q\. Z\. Chen, and A\. X\. Zhang\(A\)i am not a lawyer, but…: engaging legal experts towards responsible llm policies for legal advice\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’24,New York, NY, USA,pp\. 2454–2469\.External Links:ISBN 9798400704505,[Link](https://doi.org/10.1145/3630106.3659048),[Document](https://dx.doi.org/10.1145/3630106.3659048)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1)\.
- Choiet al\.\(2026\)J\. Choi, S\. Park, C\. Cho, H\. Park, and B\. KimDiagnosing the reliability of llm\-as\-a\-judge via item response theory\.arXiv preprint arXiv:2602\.00521\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.00521)Cited by:[§2\.2](https://arxiv.org/html/2610.02877#S2.SS2.p3.1)\.
- Conget al\.\(2026a\)L\. Cong, S\. Hahn, S\. Gombert, L\. Camus, H\. Drachsler, and U\. KroehneEstimating LLM grading ability and response difficulty in automatic short answer grading via item response theory\.InProceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications \(BEA 2026\),E\. Kochmar, B\. Alhafni, S\. Bannò, M\. Bexte, J\. Burstein, A\. Horbach, R\. Laarmann\-Quante, A\. Tack, V\. Yaneva, and Z\. Yuan \(Eds\.\),San Diego, California, USA,pp\. 259–271\.External Links:[Link](https://aclanthology.org/2026.bea-1.19/),[Document](https://dx.doi.org/10.18653/v1/2026.bea-1.19),ISBN 979\-8\-89176\-409\-5Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p2.1)\.
- Conget al\.\(2026b\)L\. Cong, L\. Hammerla, S\. Hahn, S\. Gombert, H\. Drachsler, and U\. KroehneAutomatic short answer grading with LLMs: from memorization to reasoning\.InProceedings of the 16th International Learning Analytics and Knowledge Conference,New York, NY, USA\.External Links:[Document](https://dx.doi.org/10.1145/3785022.3785031)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1)\.
- Eckes \(2015\)T\. EckesIntroduction to many\-facet rasch measurement\.Peter Lang Verlag,Berlin, Deutschland\.External Links:[Document](https://dx.doi.org/10.3726/978-3-653-04844-5),[Link](https://www.peterlang.com/document/1045610)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p4.1),[§2\.2](https://arxiv.org/html/2610.02877#S2.SS2.p2.1)\.
- Fabbriet al\.\(2021\)A\. R\. Fabbri, W\. Kryściński, B\. McCann, C\. Xiong, R\. Socher, and D\. RadevSummEval: re\-evaluating summarization evaluation\.Transactions of the Association for Computational Linguistics9,pp\. 391–409\.External Links:[Link](https://aclanthology.org/2021.tacl-1.24/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00373)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p5.1),[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p1.1),[§5](https://arxiv.org/html/2610.02877#S5.p2.1)\.
- Fuet al\.\(2024\)J\. Fu, S\. Ng, Z\. Jiang, and P\. LiuGPTScore: evaluate as you desire\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 6556–6576\.External Links:[Link](https://aclanthology.org/2024.naacl-long.365/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.365)Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p2.1)\.
- Gabrielet al\.\(2021\)S\. Gabriel, A\. Celikyilmaz, R\. Jha, Y\. Choi, and J\. GaoGO FIGURE: a meta evaluation of factuality in summarization\.InFindings of the Association for Computational Linguistics: ACL\-IJCNLP 2021,C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 478–487\.External Links:[Link](https://aclanthology.org/2021.findings-acl.42/),[Document](https://dx.doi.org/10.18653/v1/2021.findings-acl.42)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p3.1)\.
- Gombertet al\.\(2026\)S\. Gombert, Z\. Sun, F\. Zehner, J\. Lossjew, T\. Wyrwich, B\. K\. Czinczel, D\. Bednorz, M\. Kubsch, D\. Di Mitri, K\. Neumann, and H\. DrachslerAre rubrics all you need? towards rubric\-based automatic short answer scoring via guided rubric\-answer alignment\.InProceedings of the LAK26: 16th International Learning Analytics and Knowledge Conference,LAK ’26,New York, NY, USA,pp\. 272–282\.External Links:ISBN 9798400720666,[Link](https://doi.org/10.1145/3785022.3785064),[Document](https://dx.doi.org/10.1145/3785022.3785064)Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- Guet al\.\(2025\)J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu, S\. Wang, K\. Zhang, Y\. Wang, W\. Gao, L\. Ni, and J\. GuoA survey on llm\-as\-a\-judge\.External Links:2411\.15594,[Link](https://arxiv.org/abs/2411.15594)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p1.1),[§5](https://arxiv.org/html/2610.02877#S5.p4.1)\.
- Jiaoet al\.\(2025\)H\. Jiao, D\. Song, and W\. LeeComparing human and ai rater effects using the many\-facet rasch model\.External Links:2505\.18486,[Link](https://arxiv.org/abs/2505.18486)Cited by:[§2\.2](https://arxiv.org/html/2610.02877#S2.SS2.p3.1)\.
- Kryscinskiet al\.\(2020\)W\. Kryscinski, B\. McCann, C\. Xiong, and R\. SocherEvaluating the factual consistency of abstractive text summarization\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 9332–9346\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.750/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.750)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p2.1)\.
- Laloret al\.\(2018\)J\. P\. Lalor, H\. Wu, T\. Munkhdalai, and H\. YuUnderstanding deep learning performance through an examination of test set difficulty: a psychometric case study\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 4711–4716\.External Links:[Link](https://aclanthology.org/D18-1500/),[Document](https://dx.doi.org/10.18653/v1/D18-1500)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p1.1)\.
- Laloret al\.\(2016\)J\. P\. Lalor, H\. Wu, and H\. YuBuilding an evaluation scale using item response theory\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,J\. Su, K\. Duh, and X\. Carreras \(Eds\.\),Austin, Texas,pp\. 648–657\.External Links:[Link](https://aclanthology.org/D16-1062/),[Document](https://dx.doi.org/10.18653/v1/D16-1062)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p1.1)\.
- Liet al\.\(2025\)D\. Li, B\. Jiang, L\. Huang, A\. Beigi, C\. Zhao, Z\. Tan, A\. Bhattacharjee, Y\. Jiang, C\. Chen, T\. Wu, K\. Shu, L\. Cheng, and H\. LiuFrom generation to judgment: opportunities and challenges of LLM\-as\-a\-judge\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 2757–2791\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.138/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.138),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p2.1)\.
- Liet al\.\(2023\)X\. Li, T\. Zhang, Y\. Dubois, R\. Taori, I\. Gulrajani, C\. Guestrin, P\. Liang, and T\. B\. HashimotoAlpacaEval: an automatic evaluator of instruction\-following models\.GitHub\.Note:[https://github\.com/tatsu\-lab/alpaca\_eval](https://github.com/tatsu-lab/alpaca_eval)Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p2.1)\.
- Liuet al\.\(2026\)A\. H\. Liu, K\. Khandelwal, S\. Subramanian, V\. Jouault, A\. Rastogi, A\. Sadé, A\. Jeffares, A\. Jiang, A\. Cahill, A\. Gavaudan, A\. Sablayrolles, A\. Héliou, A\. You, A\. Ehrenberg, A\. Lo, A\. Eliseev, A\. Calvi, A\. Sooriyarachchi, B\. Bout, B\. Rozière, B\. D\. Monicault, C\. Lanfranchi, C\. Barreau, C\. Courtot, D\. Grattarola, D\. Dabert, D\. de las Casas, E\. Chane\-Sane, F\. Ahmed, G\. Berrada, G\. Ecrepont, G\. Guinet, G\. Novikov, G\. Kunsch, G\. Lample, G\. Martin, G\. Gupta, J\. Ludziejewski, J\. Rute, J\. Studnia, J\. Amar, J\. Delas, J\. S\. Roberts, K\. Yadav, K\. Chandu, K\. Jain, L\. Aitchison, L\. Fainsin, L\. Blier, L\. Zhao, L\. Martin, L\. Saulnier, L\. Gao, M\. Buyl, M\. Jennings, M\. Pellat, M\. Prins, M\. Poirée, M\. Guillaumin, M\. Dinot, M\. Futeral, M\. Darrin, M\. Augustin, M\. Chiquier, M\. Schimpf, N\. Grinsztajn, N\. Gupta, N\. Raghuraman, O\. Bousquet, O\. Duchenne, P\. Wang, P\. von Platen, P\. Jacob, P\. Wambergue, P\. Kurylowicz, P\. R\. Muddireddy, P\. Chagniot, P\. Stock, P\. Agrawal, Q\. Torroba, R\. Sauvestre, R\. Soletskyi, R\. Menneer, S\. Vaze, S\. Barry, S\. Gandhi, S\. Waghjale, S\. Gandhi, S\. Ghosh, S\. Mishra, S\. Aithal, S\. Antoniak, T\. L\. Scao, T\. Cachet, T\. S\. Sorg, T\. Lavril, T\. N\. Saada, T\. Chabal, T\. Foubert, T\. Robert, T\. Wang, T\. Lawson, T\. Bewley, T\. Bewley, T\. Edwards, U\. Jamil, U\. Tomasini, V\. Nemychnikova, V\. Phung, V\. Maladière, V\. Richard, W\. Bouaziz, W\. Li, W\. Marshall, X\. Li, X\. Yang, Y\. E\. Ouahidi, Y\. Wang, Y\. Tang, and Z\. RamziMinistral 3\.External Links:2601\.08584,[Link](https://arxiv.org/abs/2601.08584)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- Liuet al\.\(2023\)Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. ZhuG\-eval: NLG evaluation using gpt\-4 with better human alignment\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 2511–2522\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.153/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.153)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p5.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p2.1),[§5](https://arxiv.org/html/2610.02877#S5.p4.1)\.
- Lundberg and Lee \(2017\)S\. M\. Lundberg and S\. LeeA unified approach to interpreting model predictions\.InProceedings of the 31st International Conference on Neural Information Processing Systems,NIPS’17,Red Hook, NY, USA,pp\. 4768–4777\.External Links:ISBN 9781510860964Cited by:[§4\.3](https://arxiv.org/html/2610.02877#S4.SS3.p2.1)\.
- Naveedet al\.\(2025\)H\. Naveed, A\. U\. Khan, S\. Qiu, M\. Saqib, S\. Anwar, M\. Usman, N\. Akhtar, N\. Barnes, and A\. MianA comprehensive overview of large language models\.ACM Trans\. Intell\. Syst\. Technol\.16\(5\)\.External Links:ISSN 2157\-6904,[Link](https://doi.org/10.1145/3744746),[Document](https://dx.doi.org/10.1145/3744746)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1)\.
- Panet al\.\(2024\)Q\. Pan, Z\. Ashktorab, M\. Desmond, M\. Santillán Cooper, J\. Johnson, R\. Nair, E\. Daly, and W\. GeyerHuman\-centered design recommendations for LLM\-as\-a\-judge\.InProceedings of the 1st Human\-Centered Large Language Modeling Workshop,N\. Soni, L\. Flek, A\. Sharma, D\. Yang, S\. Hooker, and H\. A\. Schwartz \(Eds\.\),pp\. 16–29\.External Links:[Link](https://aclanthology.org/2024.hucllm-1.2/),[Document](https://dx.doi.org/10.18653/v1/2024.hucllm-1.2)Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p3.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.5: towards native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- Rodriguezet al\.\(2021\)P\. Rodriguez, J\. Barrow, A\. Hoyle, J\. P\. Lalor, R\. Jia, and J\. Boyd\-GraberEvaluation examples are not equally informative: how should that change NLP leaderboards?\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 4486–4503\.External Links:[Link](https://aclanthology.org/2021.acl-long.346/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.346)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p2.1)\.
- Shenet al\.\(2023\)C\. Shen, L\. Cheng, X\. Nguyen, Y\. You, and L\. BingLarge language models are not yet human\-level evaluators for abstractive summarization\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 4215–4233\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.278/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.278)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p3.1)\.
- Tanet al\.\(2024\)Z\. Tan, D\. Li, S\. Wang, A\. Beigi, B\. Jiang, A\. Bhattacharjee, M\. Karami, J\. Li, L\. Cheng, and H\. LiuLarge language models for data annotation and synthesis: a survey\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 930–957\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.54/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.54)Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p1.1)\.
- Teamet al\.\(2025\)G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- W\.J\. van der Linden \(Ed\.\) \(2016\)W\.J\. van der Linden \(Ed\.\)Handbook of item response theory: volume 1: models\.Behavioral Sciences,Chapman and Hall/CRC\.External Links:ISBN 9781315374512,[Document](https://dx.doi.org/10.1201/9781315374512)Cited by:[§2\.2](https://arxiv.org/html/2610.02877#S2.SS2.p1.1)\.
- Wanget al\.\(2020\)A\. Wang, K\. Cho, and M\. LewisAsking and answering questions to evaluate the factual consistency of summaries\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 5008–5020\.External Links:[Link](https://aclanthology.org/2020.acl-main.450/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.450)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p2.1)\.
- Wanget al\.\(2024a\)G\. Wang, S\. Cheng, X\. Zhan, X\. Li, S\. Song, and Y\. LiuOpenChat: advancing open\-source language models with mixed\-quality data\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 57021–57040\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/fc8781fb328fb1fd069584a4519a2709-Paper-Conference.pdf)Cited by:[§3\.1](https://arxiv.org/html/2610.02877#S3.SS1.p2.1)\.
- Wanget al\.\(2024b\)R\. Wang, H\. Yu, W\. Zhang, Z\. Qi, M\. Sap, Y\. Bisk, G\. Neubig, and H\. ZhuSOTOPIA\-π\\pi: interactive learning of socially intelligent language agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12912–12940\.External Links:[Link](https://aclanthology.org/2024.acl-long.698/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.698)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1)\.
- Wanget al\.\(2024c\)X\. Wang, H\. Kim, S\. Rahman, K\. Mitra, and Z\. MiaoHuman\-llm collaborative annotation through effective verification of llm labels\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,CHI ’24,New York, NY, USA\.External Links:ISBN 9798400703300,[Link](https://doi.org/10.1145/3613904.3641960),[Document](https://dx.doi.org/10.1145/3613904.3641960)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.SS0.SSS0.Px3.p2.1),[§5](https://arxiv.org/html/2610.02877#S5.p5.1)\.
- Wanget al\.\(2024d\)Y\. Wang, Z\. Yu, W\. Yao, Z\. Zeng, L\. Yang, C\. Wang, H\. Chen, C\. Jiang, R\. Xie, J\. Wang, X\. Xie, W\. Ye, S\. Zhang, and Y\. ZhangPandaLM: an automatic evaluation benchmark for llm instruction tuning optimization\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 43573–43593\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/be3b0d51a2b86cb4ffe50f13480217e0-Paper-Conference.pdf)Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p2.1)\.
- Wanget al\.\(2025\)Y\. Wang, J\. Huang, L\. Du, Y\. Guo, Y\. Liu, and R\. WangEvaluating large language models as raters in large\-scale writing assessments: a psychometric framework for reliability and validity\.Computers and Education: Artificial Intelligence9,pp\. 100481\.External Links:ISSN 2666\-920X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.caeai.2025.100481),[Link](https://www.sciencedirect.com/science/article/pii/S2666920X25001213)Cited by:[§2\.2](https://arxiv.org/html/2610.02877#S2.SS2.p3.1)\.
- Yeet al\.\(2026\)H\. Ye, J\. Jin, Y\. Xie, X\. Zhang, and G\. SongLarge language model psychometrics: a systematic review of evaluation, validation, and enhancement\.External Links:2505\.08245,[Link](https://arxiv.org/abs/2505.08245)Cited by:[§2\.2](https://arxiv.org/html/2610.02877#S2.SS2.p3.1)\.
- Yeet al\.\(2025\)J\. Ye, Y\. Wang, Y\. Huang, D\. Chen, Q\. Zhang, N\. Moniz, T\. Gao, W\. Geyer, C\. Huang, P\. Chen, N\. Chawla, and X\. ZhangJustice or prejudice? quantifying biases in llm\-as\-a\-judge\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 102351–102390\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/fdca08d371e4b6c031397909e20043bd-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p3.1)\.
- Yuet al\.\(2025\)J\. Yu, S\. Sun, X\. Hu, J\. Yan, K\. Yu, and X\. LiImprove LLM\-as\-a\-judge ability as a general ability\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 14099–14115\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.712/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.712),ISBN 979\-8\-89176\-332\-6Cited by:[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p3.1)\.
- Yuet al\.\(2024\)Y\. Yu, Z\. Yao, H\. Li, Z\. Deng, Y\. Jiang, Y\. Cao, Z\. Chen, J\. W\. Suchow, Z\. Cui, R\. Liu, Z\. Xu, D\. Zhang, K\. Subbalakshmi, G\. Xiong, Y\. He, J\. Huang, D\. Li, and Q\. XieFINCON: a synthesized llm multi\-agent system with conceptual verbal reinforcement for enhanced financial decision making\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1)\.
- Zhaoet al\.\(2023\)W\. Zhao, M\. Strube, and S\. EgerDiscoScore: evaluating text generation with BERT and discourse coherence\.InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics,A\. Vlachos and I\. Augenstein \(Eds\.\),Dubrovnik, Croatia,pp\. 3865–3883\.External Links:[Link](https://aclanthology.org/2023.eacl-main.278/),[Document](https://dx.doi.org/10.18653/v1/2023.eacl-main.278)Cited by:[§5](https://arxiv.org/html/2610.02877#S5.p2.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p1.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p2.1)\.
- Zhouet al\.\(2026\)H\. Zhou, H\. Huang, Z\. Zhao, L\. Han, H\. Wang, K\. Chen, M\. Yang, W\. Bao, J\. Dong, B\. Xu, C\. Zhu, H\. Cao, and T\. ZhaoLost in benchmarks? rethinking large language model benchmarking with item response theory\.InProceedings of the Fortieth AAAI Conference on Artificial Intelligence and Thirty\-Eighth Conference on Innovative Applications of Artificial Intelligence and Sixteenth Symposium on Educational Advances in Artificial Intelligence,AAAI’26/IAAI’26/EAAI’26\.External Links:ISBN 978\-1\-57735\-906\-7,[Link](https://doi.org/10.1609/aaai.v40i41.40814),[Document](https://dx.doi.org/10.1609/aaai.v40i41.40814)Cited by:[§1](https://arxiv.org/html/2610.02877#S1.p2.1),[§2\.1](https://arxiv.org/html/2610.02877#S2.SS1.p3.1)\.

## Appendix AAppendix

### A\.1LLM Judge Prompt

We use the same task instruction for all open\-weight LLM judges\. Table[5](https://arxiv.org/html/2610.02877#A1.T5)lists the 17 open\-weight LLM judges used in our main analysis\. For models whose tokenizer supports a system role, the prompt is formatted as a chat conversation using the model’s native chat template\. If the model does not support a system role, the system instruction is prepended to the user message\. Decoding is deterministic\.

> System:You are a careful evaluator of news summaries\. User:Rate the summary for the requested evaluation dimension\. Use a 1–5 integer scale where 1 is very poor and 5 is excellent\. Dimension: \{dimension\} Criterion: \{rubric\} Source document: \{source document\} Summary: \{summary\} Return ONLY one number: 1, 2, 3, 4, or 5\. Answer:

The dimension field is one of coherence, consistency, fluency, or relevance\. The criterion field contains the corresponding SummEval rubric\. We parse the first valid integer in\{1,…,5\}\\\{1,\\ldots,5\\\}from the model output\. If no valid score is found, we retry once with a longer generation budget and apply the same parser\. All prompts included the full source document without explicit truncation\. As a length check, the longest rendered prompt contained 694 whitespace\-delimited words, with maximum source and summary lengths of 540 and 133 words, respectively\. These lengths are far below the context limits of the local judges\.

Table 5:17 open\-weight LLM judges used for SummEval annotation in our analysis\.
### A\.2MFRM

#### A\.2\.1MFRM Estimation Details

We estimate the MFRM parameters by maximum likelihood\. For each observed rater–summary–dimension rating, we computeηr​i​d=θi−ρr−δd\\eta\_\{rid\}=\\theta\_\{i\}\-\\rho\_\{r\}\-\\delta\_\{d\}, construct the five category logits implied by the adjacent\-category formulation, and normalize them with a softmax\. The objective is the negative log\-likelihood of the observed rating categories with weakL2L\_\{2\}regularization\. We optimize this objective in PyTorch using AdamW with learning rate0\.050\.05and weight decay10−410^\{\-4\}\. For identification, summary, rater, and dimension parameters are mean\-centered during optimization\. The step thresholds are constrained to be ordered using an ordered parameterization and are initialized as evenly spaced values from−1\.5\-1\.5to1\.51\.5\. All other facet parameters are initialized at zero\. Optimization uses early stopping when the loss does not improve by10−710^\{\-7\}for 300 epochs\.

#### A\.2\.2Agreement\-based Sanity Check for Residual Hardness

As a sanity check, we examine whether residual hardness covaries with local rater agreement\. For each summary–dimension unitii, we compute the observed pairwise agreement

Pi=∑cni​c​\(ni​c−1\)ni​\(ni−1\),P\_\{i\}=\\frac\{\\sum\_\{c\}n\_\{ic\}\(n\_\{ic\}\-1\)\}\{n\_\{i\}\(n\_\{i\}\-1\)\},\(5\)whereni​cn\_\{ic\}denotes the number of raters assigning categoryccto unitii, andnin\_\{i\}is the total number of ratings for that unit\. We then define the local chance\-corrected agreement contribution as

κi=Pi−Pe1−Pe,\\kappa\_\{i\}=\\frac\{P\_\{i\}\-P\_\{e\}\}\{1\-P\_\{e\}\},\(6\)wherePeP\_\{e\}is estimated separately for the corresponding judge panel and evaluation dimension\. This quantity is used as a local agreement diagnostic rather than as the standard global kappa statistic\.

Table 6:Spearman correlations between residual hardness and summary–dimension\-level agreement contribution\. Negative correlations indicate that units with higher residual hardness contribute less to rater agreement\.Residual hardness is negatively associated with the local agreement diagnostics across all dimensions in both judge panels\. This relationship is intended only as a sanity check showing that larger residual deviations tend to coincide with greater local rater disagreement\. It does not independently validate residual hardness as a distinct construct or establish that residual hardness has identical substantive meaning across the human and LLM panels\. Indeed, the associations are substantially stronger for human ratings than for LLM ratings\. Unlike the agreement diagnostic, residual hardness is conditional on the fitted MFRM facets\.

#### A\.2\.3Judge\-Capability and Human\-Panel Sensitivity

Table 7:Sensitivity of human–LLM alignment to LLM judge capability and human\-panel composition\. Residual\-hardness correlations are computed at the summary–dimension level\.Restricting the LLM panel to the seven judges with at least 7B parameters does not remove the contrast between quality alignment and residual\-hardness alignment\. The contrast is also preserved when the human MFRM is estimated using only the three expert annotators\. In the expert\-only versus≥7\\geq 7B comparison, the consistency hard\-rate gap remains\+26\.2\+26\.2percentage points\. These results indicate that the main RQ1 and RQ2 patterns are not driven solely by the smallest LLM judges or by expert–crowd disagreement in the human panel\.

#### A\.2\.4Rater\-by\-Dimension Interaction Sensitivity

To examine whether the weak human–LLM residual\-hardness alignment primarily reflects an omitted rater\-by\-dimension interaction, we refit the MFRM with an additional rater–dimension term\. Under this richer specification, human–LLM residual\-hardness alignment remains weak \(r=0\.277r=0\.277,ρ=0\.122\\rho=0\.122\)\. Thus, the main RQ1 pattern is not solely attributable to omission of rater\-specific dimension effects\.

#### A\.2\.5Rater\-Misfit Sensitivity

To examine whether the main residual\-hardness findings are driven by poorly fitting raters, we conducted a post\-hoc sensitivity analysis after excluding raters with infit or outfit mean\-square outside the conventional descriptive range of 0\.5–1\.5\. The original MFRM calibrations were kept fixed, and residual\-hardness summaries were recomputed using only the remaining raters\.

Table[8](https://arxiv.org/html/2610.02877#A1.T8)shows that the central RQ1 conclusion remains unchanged\. Human–LLM alignment in latent quality remains moderate, whereas residual\-hardness alignment remains weak at both the summary level and the summary–dimension level\. If anything, excluding flagged raters further reduces the residual\-hardness correlations, suggesting that the weak human–LLM hardness correspondence is not an artifact of poorly fitting raters\.

Table 8:RQ1 sensitivity after excluding raters with infit or outfit mean\-square outside the descriptive 0\.5–1\.5 range\.Table[9](https://arxiv.org/html/2610.02877#A1.T9)shows that the dimension\-level pattern is also stable\. Consistency remains LLM\-hard relative to humans, whereas coherence remains human\-hard relative to LLMs\. The smaller gaps for fluency and relevance vary more after exclusion, which is expected because their original gaps are close to zero\.

Table 9:RQ2 dimension hard\-rate gap sensitivity after excluding flagged raters\. Values are percentage\-point gaps, computed as LLM hard\-rate minus human hard\-rate\.
#### A\.2\.6MFRM Calibration Comparability Sensitivity

Because the main analyses compare residual hardness derived from separately calibrated human\-only and LLM\-only MFRMs, we examine whether the conclusions depend on direct comparability of raw residual magnitudes across calibrations\. We recompute the residual\-hardness comparisons under three alternatives: within\-groupzz\-scored residual hardness, percentile\-rank residual hardness, and residuals from a combined MFRM fit to pooled human and LLM ratings\. For the combined\-model analysis, absolute residuals are aggregated separately for human and LLM raters after fitting the shared model\.

Table 10:Residual\-hardness alignment under alternative calibration and normalization choices\.The main conclusion does not rely on direct comparability of raw residual magnitudes across separately calibrated models\. As expected, within\-groupzz\-scoring leaves the correlations unchanged because it preserves within\-group ordering\. Rank\-based normalization yields slightly smaller correlations, but the overall pattern remains the same\. Under the combined MFRM, residual\-hardness alignment is somewhat higher, yet it remains clearly weaker than the corresponding quality alignment and remains weak in absolute terms at the summary–dimension level\. The dominant dimension pattern is also stable across specifications\. Consistency remains LLM\-hard relative to humans, whereas coherence remains human\-hard relative to LLMs\.

### A\.3Additional Information and Experiments for RQ3

#### A\.3\.1Input Features for RQ3

Table[11](https://arxiv.org/html/2610.02877#A1.T11)summarizes the input features used for the RQ3 human\-easy/LLM\-hard prediction experiments\. We include these features as interpretable, observable proxies for factors that may affect judging difficulty\. TF–IDF features capture lexical and topical content, length and compression features approximate information density and omission, and sentence, digit, and negation counts provide coarse structural and factual\-detail cues\. Coarse system type captures broad extractive versus abstractive differences, while specific system IDs are excluded to reduce label leakage\.

Table 11:Input features used for RQ3 prediction of human\-easy/LLM\-hard cases\. TF–IDF features usemin\_df=2\\texttt\{min\\\_df\}=2, unigram/bigram ranges, and up to 6,000 features each for the summary and source fields\.
#### A\.3\.2RQ3 Prediction Robustness and Ablations

We run additional robustness checks for the RQ3 human\-easy/LLM\-hard prediction task\. All analyses use the pooled summary–dimension dataset \(N=6,400N=6\{,\}400, positive class=1,494=1\{,\}494, positive rate=23\.3%=23\.3\\%\)\. As in the main analysis, we retain coarse system type but exclude specific system IDs\. Confidence intervals are percentile bootstrap intervals computed from out\-of\-fold predictions\.

Table 12:RQ3 feature ablation for predicting human\-easy/LLM\-hard summary–dimension units\. Brackets denote 95% bootstrap confidence intervals from out\-of\-fold predictions\. Metadata includes coarse system type only; specific system IDs are excluded\.Table 13:RQ3 split robustness for the full model\. Brackets denote 95% bootstrap confidence intervals from out\-of\-fold predictions\. Grouped splits hold out all dimensions of the same summary, all summaries from the same source document, or all summaries from the same generation system, respectively\.These checks show that the RQ3 signal is not driven by a single feature family or solely by random\-fold leakage across repeated summaries\. TF–IDF features provide most of the predictive signal, surface features retain above\-chance predictive performance, and the full model consistently improves over the ablations\. Performance decreases under stricter grouped splits, especially when holding out entire source documents, but predictive signal remains in all settings\. For the surface\-feature model, we additionally perform a compression\-drop ablation under summary\-grouped 5\-fold cross\-validation\. Removing compression ratio changes AUROC only from0\.5710\.571to0\.5660\.566and AUPRC from0\.2930\.293to0\.2860\.286\. Overall, these results support partial predictability of human\-easy/LLM\-hard cases that persists under alternative feature sets and split strategies rather than high\-accuracy classification\.

Table 14:RQ3 quartile\-threshold sensitivity for predicting human\-easy/LLM\-hard cases from all summary–dimension units\.We also test a stricter quadrant definition based on task\-specific quartiles\. For each prediction task, human\-easy units are those whose human residual hardness falls in the lower quartile of that task’s evaluation group, and LLM\-hard units are those whose LLM residual hardness falls in the upper quartile of the same group\. Table[14](https://arxiv.org/html/2610.02877#A1.T14)shows that the pooled predictive signal remains above chance under this more selective definition \(AUROC=0\.676=0\.676\), although dimension\-specific robustness is uneven\. In particular, coherence decreases to AUROC=0\.578=0\.578\. We therefore interpret this analysis as supporting pooled predictability rather than uniformly strong robustness across dimensions\.

相似文章

LLM-as-Judge的几何学:为何LLM间共识并非人类对齐

arXiv cs.CL

本文从几何角度分析了为何作为裁判的LLM彼此之间高度一致,但与人类仅弱相关,发现LLM间共识在主观评分标准上反映的是坍塌子空间,而非真正的人类对齐。基于人类数据的后验校准提高了对齐,但即使经过校准的LLM也未达到人类的可靠性。

换一个裁判,分数就变了:审计大模型作为裁判的可靠性

arXiv cs.CL

本文审计了大模型作为裁判(LLM-as-judge)评估的可靠性,表明即使候选回复固定不变,更换评估模型也可能改变评分。论文考察了Qwen3和MiniMax模型的扩展与升级路径,得出结论:裁判升级不可互换,并提出了最佳报告实践。