Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

arXiv cs.LG Papers

Summary

The study investigates whether LLM-as-a-Judge evaluators reliably assess psychological depth in LLM-generated stories, revealing that human preferences are heterogeneous while judges exhibit bias towards reasoning outputs based on surface features.

arXiv:2609.13773v1 Announce Type: new Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($\rho = 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\ GPT-4o and DeepSeek-R1 vs.\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\%), whereas DeepSeek-R1 trailed V3 (42.9\%), and inter-reader agreement was near chance (Krippendorff's $\alpha = 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.
Original Article
View Cached Full Text

Cached at: 09/15/26, 08:52 AM

# Does Reasoning Improve Psychological Depth in Large Language Models? Depends on Who’s Judging
Source: [https://arxiv.org/html/2609.13773](https://arxiv.org/html/2609.13773)
\\workshoptitle

Can We Trust the Judge? Building Reliable Evaluation for Language Models

Yihe WangAffiliation:Tsinghua UniversityFabrice Y\. Harel\-CanadaAffiliation:LA General HospitalSara KhosraviAffiliation:University of California, Los AngelesZeynep Senahan YildizAffiliation:University of California, Los AngelesAmit SahaiAffiliation:University of California, Los AngelesNanyun PengAffiliation:Google

###### Abstract

LLM\-as\-a\-Judge evaluators are increasingly used to score open\-ended generation, yet a judge’s correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective\. We study this failure mode through psychological depth in short stories\. Seven human readers and an LLM\-judge ensemble selected on the original scalar Psychological Depth Scale dataset \(ρ=0\.646\\rho=0\.646\) evaluated 60 blinded, prompt\-matched story pairs from GPT\-5 vs\. GPT\-4o and DeepSeek\-R1 vs\. DeepSeek\-V3\. Human preferences showed no universal reasoning advantage: GPT\-5 was modestly preferred over GPT\-4o \(60\.0–62\.9%\), whereas DeepSeek\-R1 trailed V3 \(42\.9%\), and inter\-reader agreement was near chance \(Krippendorff’sα=0\.070\\alpha=0\.070\), with within\-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding\. The judge, by contrast, favored reasoning outputs in 89\.0% of dimension\-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity\. These results suggest that development\-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point\-estimate judges can obscure the heterogeneity in subjective human evaluation\.

## 1Introduction

Automatic judges are used to select checkpoints, compare systems, and increasingly provide training signals; LLM\-as\-a\-Judge methods make this scalable, but their validity is contextual rather than intrinsic\. Pairwise judging can align with human preference yet amplify evaluator biases\([Zheng et al\., 2023](https://arxiv.org/html/2609.13773#bib.bib29);[Liu et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib14);[Jeong et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib9)\), and prior work documents position, verbosity, self\-preference, and inconsistency effects\([Li et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib13);[Panickssery et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib18);[Stureborg et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib24);[Wang et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib27)\)\. Less is known about a more basic question: whether a judge that agrees well with humans on the data it was selected on remains valid when deployed on a shifted distribution where outputs are closely matched and human preferences are subjective\.

Subjective creative evaluation is where this question is most acute\. Fluency\-centered metrics reward well\-formed text but miss the qualities that matter in fiction\([Gómez\-Rodríguez and Williams, 2023](https://arxiv.org/html/2609.13773#bib.bib5);[Lu et al\., 2026](https://arxiv.org/html/2609.13773#bib.bib15)\), and what looks like creativity in LLM output can be surface imitation of literary style\([Chakrabarty et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib1)\), so a judge may score polish rather than depth\. At the same time, disagreement among readers on subjective tasks often reflects task ambiguity and differing values rather than annotation noise\([Clark et al\., 2021](https://arxiv.org/html/2609.13773#bib.bib2);[Sandri et al\., 2023](https://arxiv.org/html/2609.13773#bib.bib21);[Kirk et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib11)\)\. A judge evaluating such text must therefore be assessed not only on aggregate agreement but on whether it preserves the*distribution*of human judgments\.

We study this through*psychological depth*in short stories, using the Psychological Depth Scale \(PDS\)\([Harel\-Canada et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib8)\), which measures authenticity, empathy, engagement, emotion provocation, and narrative complexity\. As a test case we use reasoning\-enhanced models such as GPT\-5 and DeepSeek\-R1, which are routinely evaluated on mathematical, coding, and factual benchmarks\([OpenAI et al\., 2024a](https://arxiv.org/html/2609.13773#bib.bib16);[Snell et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib23)\)but whose reasoning\-oriented post\-training has rarely been examined on open\-ended creative generation\. We constructed 60 blinded, prompt\-matched story pairs across three frontier matchups, GPT\-5 versus GPT\-4o under two reasoning\-effort settings and DeepSeek\-R1 versus DeepSeek\-V3, and evaluated each pair with seven human readers and an LLM\-as\-a\-Judge ensemble selected on the original scalar PDS dataset \(ρ=0\.646\\rho=0\.646\)\. We ask whether a PDS judge selected for its agreement with human ratings reproduces human readers’ preferences on frontier story pairs, how those preferences vary across model families, dimensions, and readers, and which story properties are associated with the judge’s scores\.

Our results expose a sharp mismatch, and we contribute three things: \(1\) an audit showing that a PDS judge with a respectable development\-set correlation produces a near\-uniform pro\-reasoning signal \(89\.0% of dimension\-level comparisons, 59 of 60 pairs, across all five configurations\) where human readers are at chance and split by model family; \(2\) evidence that the judge’s scores track surface features such as sentence length and lexical uniformity, offering a candidate mechanism for the skew; and \(3\) a distributional analysis showing that reader disagreement is structured, so that a point\-estimate judge cannot represent the target it was selected to approximate\.

## 2Study design

#### Automatic evaluator setup\.

We used a scalar\-rating evaluator rather than a directly pairwise\-trained judge because the original PDS dataset provides scalar supervision with strong inter\-annotator agreement, which let us construct calibrated few\-shot prompts and compare configurations against human ratings before deployment\. Following[Harel\-Canada et al\. \(2024\)](https://arxiv.org/html/2609.13773#bib.bib8), the pipeline uses DSPy to enforce structured output control\([Khattab et al\., 2023](https://arxiv.org/html/2609.13773#bib.bib10)\)and retains the Mixture\-of\-Personas prompting strategy\([Salewski et al\., 2023](https://arxiv.org/html/2609.13773#bib.bib20)\)\. Using the original 97\-story dataset and its human annotations, we partitioned the data into a 20% demonstration pool and an 80% selection split, selected representative few\-shot examples from the demonstration pool by applyingkk\-means clustering to the PDS score vectors\([Su et al\., 2022](https://arxiv.org/html/2609.13773#bib.bib25)\), and used the selection split to compare Llama 3\.1 70B, Llama 3\.3 70B, and Qwen3\.5\-397B configurations across demonstration counts and Chain\-of\-Thought prompting\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib6);[Qwen Team, 2026](https://arxiv.org/html/2609.13773#bib.bib19)\)\. No single configuration performed best across all dimensions, so we constructed a heterogeneous ensemble that assigns each PDS dimension to its best\-performing configuration\. The ensemble achievesρ=0\.646\\rho=0\.646, outperforming every single\-model configuration as well as the zero\-shot PDS evaluator of[Harel\-Canada et al\. \(2024\)](https://arxiv.org/html/2609.13773#bib.bib8)\(ρ≈0\.51\\rho\\approx 0\.51with GPT\-4o\)\. Because the same 80% split both selected configurations and supplied this correlation,ρ=0\.646\\rho=0\.646is a development\-set estimate, mirroring how judges are commonly tuned and reused in practice\. At inference, each story was scored independently on a 1–5 scale; dimension\-level preference followed from the higher routed score, aggregate\-PDS preference from the higher mean of the five routed scores, and exact ties contributed 0\.5 to each story \(Appendix[B](https://arxiv.org/html/2609.13773#A2)\)\.

#### Deployment set: frontier story pairs\.

The original PDS evaluator was calibrated on human stories and 2023\-era LLM generations with quality gaps large enough to yield strong annotator consensus \(α≈0\.72\\alpha\\approx 0\.72\); we deployed it on a considerably harder distribution\. We curated 20 premises that provide contextual framing for character development while leaving the narrative direction open: 15 from r/WritingPrompts, restricted to posts after the models’ training\-data cutoffs, ranked by all\-time upvotes, and manually filtered for quality, and 5 from Reedsy\. For each premise we generated three prompt\-matched pairings: DeepSeek\-R1 versus DeepSeek\-V3, GPT\-5 \(auto effort\) versus GPT\-4o, and GPT\-5 \(high effort\) versus GPT\-4o, yielding 60 pairs of 400–600\-word stories\. All stories were generated with a common zero\-shot story\-writing template, with only limited family\-specific wording adjustments for word\-count compliance \(Appendix[D](https://arxiv.org/html/2609.13773#A4)\)\. DeepSeek\-R1 is an RL upgrade of V3 and provides our closest anchor for reasoning\-oriented post\-training\([Guo et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib7);[DeepSeek\-AI et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib3)\); the GPT pairs are model\-level contrasts\([Singh et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib22);[OpenAI et al\., 2024b](https://arxiv.org/html/2609.13773#bib.bib17)\)\. These are ecologically relevant comparisons rather than clean ablations of reasoning, and we use “reasoning output” as a label for the model, not a causal attribution\.

#### Human evaluation\.

We adapted the[Harel\-Canada et al\. \(2024\)](https://arxiv.org/html/2609.13773#bib.bib8)framework to a blinded pairwise protocol; the Institutional Review Board deemed the study exempt\. Seven undergraduate readers from English and Psychology departments were onboarded on the five PDS dimensions and a custom Label Studio annotation interface\([Tkachenko et al\., 2020\-2025](https://arxiv.org/html/2609.13773#bib.bib26)\)\. For each of the 60 pairs, the two stories were presented side by side in randomized order, with the PDS rubric definitions shown at the top of the page, and readers provided six forced\-choice judgments: one for each of the five PDS dimensions and one overall preference, with no tie or confidence option\. This produced7×60×6=2,5207\\times 60\\times 6=2\{,\}520pairwise preference judgments along with 112 reader\-produced free\-text justifications\. To monitor annotation quality, we logged per\-pair completion times, conducted a leave\-one\-reader\-out stability analysis, and administered a post\-study survey; recruitment, compensation, and diagnostics are in Appendix[E](https://arxiv.org/html/2609.13773#A5)\.

#### Analysis\.

Human preference intervals come from separate intercept\-only logistic mixed models per matchup×\\timescriterion,logit​P​\(yi​r=1\)=β0\+ui\+vr\\mathrm\{logit\}\\,P\(y\_\{ir\}=1\)=\\beta\_\{0\}\+u\_\{i\}\+v\_\{r\}, with crossed random intercepts for premise pairiiand readerrr; the 18 intervals are pointwise and unadjusted for multiplicity\. We report Fleiss’κ\\kappaand Krippendorff’sα\\alpha\([Fleiss, 1971](https://arxiv.org/html/2609.13773#bib.bib4);[Krippendorff, 2011](https://arxiv.org/html/2609.13773#bib.bib12)\)as agreement summaries, and per\-reader logistic regressions and dimension\-to\-overall Cohen’sκ\\kappato characterize how readers weight dimensions; with seven readers these are exploratory\. To probe surface associations, we measured length, sentence\-length moments, MATTR, MTLD, and perplexity under Llama 3\.1 70B\. Full results appear in Appendices[A](https://arxiv.org/html/2609.13773#A1)and[C](https://arxiv.org/html/2609.13773#A3)\.

## 3Results

![Refer to caption](https://arxiv.org/html/figs/fig1a_human_preference.png)

![Refer to caption](https://arxiv.org/html/figs/fig1b_evensplit.png)

Figure 1:Human preferences vary by model pairing and dimension; the LLM judge does not\.\(a\)Human preference for the reasoning\-model story with pointwise 95% GLMM intervals; the dashed line marks chance\.\(b\)LLM\-judge preference for the reasoning\-model story pooled over all 60 pairs\. Solid bars are reasoning wins; hatched segments credit each exact score tie 0\.5 to the reasoning story; labels give the resulting even\-split rate, with raw tie counts in parentheses\. Aggregate PDS averages the five dimension scores and is a distinct target from the human overall\-preference question\.#### The judge is nearly uniformly pro\-reasoning; readers are not\.

Across the five PDS dimensions, the ensemble favored the reasoning\-model story in 89\.0% of 300 pair\-by\-dimension comparisons with ties split evenly \(Figure[1](https://arxiv.org/html/2609.13773#S3.F1)b\), ranging from 82\.5% on narrative complexity to 95\.8% on engagement\. On aggregate PDS it selected the reasoning story in 59 of 60 pairs: all 40 GPT\-family pairs and, despite readers leaning toward V3, 19 of 20 DeepSeek pairs\. Human readers, by contrast, gave reasoning stories 58\.0% of 2,100 dimension\-level votes and 55\.2% of 420 overall\-preference votes\. This skew is not an artifact of the ensemble construction: all five individual judge configurations preferred the reasoning model, with pooled even\-split rates from 77\.5% to 96\.3% \(Appendix[B](https://arxiv.org/html/2609.13773#A2)\)\.

#### Human preference depends on model family and dimension\.

Human preferences diverge across model families \(Figure[1](https://arxiv.org/html/2609.13773#S3.F1)a\)\. Although reasoning stories received 55\.2% of overall\-preference votes, this aggregate obscures a family\-level split: readers modestly preferred GPT\-5 over GPT\-4o under both the auto \(60\.0%, 95% CI \[52\.6, 69\.5\]\) and high \(62\.9%, \[55\.9, 72\.7\]\) reasoning\-effort settings, whereas DeepSeek\-R1 trailed DeepSeek\-V3 \(42\.9%, \[33\.7, 50\.9\]\), an interval that includes chance\. The same divergence appears across PDS dimensions: the GPT\-5 pairings were consistently above 50% on the affective dimensions, especially emotion provocation \(70\.7%\) and empathy \(68\.6%\), whereas R1 vs\. V3 exhibited mixed effects, notably disfavored on engagement \(39\.3%, \[29\.6, 46\.6\]\) yet favored on emotion provocation \(61\.4%\) and narrative complexity \(58\.6%\)\. Individual readers ranged from strongly pro\-reasoning \(Reader 7, 80\.0%\) to pro\-base \(Reader 1, 36\.7%\), so the pooled 55\.2% masks both family\- and reader\-level variation \(Appendix[A](https://arxiv.org/html/2609.13773#A1)\)\.

#### Group agreement is low; reader heterogeneity is structured but exploratory\.

Inter\-reader agreement is low: Krippendorff’sα=0\.070\\alpha=0\.070, with raw pairwise agreement \(53\.8%\) barely above the 50% expected by chance \(Appendix[A](https://arxiv.org/html/2609.13773#A1)\)\. This is a dramatic drop from theα≈0\.72\\alpha\\approx 0\.72reported by[Harel\-Canada et al\. \(2024\)](https://arxiv.org/html/2609.13773#bib.bib8)using the same rubric on large\-gap items\. Our frontier pairs are far closer in quality, and the low agreement may reflect limited rubric discrimination on close pairs, reader subjectivity, or noise\. Quality checks reduce but do not eliminate annotation\-failure concerns: readers spent a median of 108–418 seconds per pair with stable pace throughout, and removing any single reader changes the aggregate rate by only 2–4 points \(Appendix[E](https://arxiv.org/html/2609.13773#A5)\)\. Per\-reader regressions, dimension\-to\-overallκ\\kappa, and free\-text rationales show recurring dimension\-weighting patterns, consistent with structured heterogeneity rather than random responding, though seven readers preclude definitive clusters \(Appendix[A](https://arxiv.org/html/2609.13773#A1)\)\.

#### Surface features are associated with judge scores\.

Judge engagement scores correlated with mean sentence length atr=−0\.70r=\-0\.70, and moving\-average type\-token ratio was negatively associated with judge scores on four of five dimensions \(−0\.37\-0\.37to−0\.14\-0\.14\); reasoning\-model stories had shorter sentences and slightly lower MATTR \(0\.896 vs\. 0\.907\), so the judge’s preference is partly associated with these properties rather than depth alone\. Reasoning stories were also more perplexing to Llama 3\.1 70B, arguing against a simple familiarity account \(Appendix[C](https://arxiv.org/html/2609.13773#A3)\)\.

## 4Implications and limitations

#### Development\-set correlation is not deployment validity\.

The ensemble agreed well with human PDS ratings on its selection set \(ρ=0\.646\\rho=0\.646\), yet on frontier pairs it produced a near\-uniform pro\-reasoning signal where readers were at chance\. Because selection and validation shared the same split, this is a context\-specific audit rather than independent evidence of transfer failure\. All five configurations agreed with each other far more than seven readers agreed among themselves, so inter\-judge consistency\([Li et al\., 2025](https://arxiv.org/html/2609.13773#bib.bib13);[Stureborg et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib24)\)offered no warning\. This is all the more notable because most of our evaluator models predate the reasoning era and do not plausibly favor reasoning outputs through training\-data exposure to synthetic reasoning traces; the association of judge scores with sentence length and lexical uniformity points instead to sensitivity to stylistic artifacts of reasoning\-oriented generation\. Evaluator reports should document selection procedure, deployment context, and known surface sensitivities, and include a target\-matched human check on the deployment distribution\.

#### Limitations\.

This exploratory audit uses 20 premises, one retained generation per condition, seven readers from one university, and one domain\. Forced choice without ties or confidence may turn indifference into preference; limited PDS discrimination on close pairs, fatigue, and noise may all contribute to low agreement\. The human target changed and surface results are correlational\. Broader cohorts, repeated generations, target\-matched judges, and out\-of\-fold validation are needed\.

## 5Conclusion

The PDS judge that agreed well with human ratings on its selection set was nearly uniformly pro\-reasoning on closely matched frontier story pairs where human readers were at chance and split by model family\. Our exploratory audit motivates target\-matched human checks and independent validation before reusing a selected evaluator on a new distribution\.

## References

- Chakrabarty et al\. \[2024\]Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien\-Sheng Wu\.Art or artifice? large language models and the false promise of creativity\.In*Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems*\. Association for Computing Machinery, 2024\.URL[https://doi\.org/10\.1145/3613904\.3642731](https://doi.org/10.1145/3613904.3642731)\.
- Clark et al\. \[2021\]Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A\. Smith\.All that’s ‘human’ is not gold: Evaluating human evaluation of generated text\.In*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 7282–7296\. Association for Computational Linguistics, August 2021\.doi:10\.18653/v1/2021\.acl\-long\.565\.URL[https://aclanthology\.org/2021\.acl\-long\.565/](https://aclanthology.org/2021.acl-long.565/)\.
- DeepSeek\-AI et al\. \[2025\]DeepSeek\-AI, Aixin Liu, Bei Feng, Bing Xue, et al\.Deepseek\-v3 technical report, 2025\.URL[https://arxiv\.org/abs/2412\.19437](https://arxiv.org/abs/2412.19437)\.
- Fleiss \[1971\]Joseph Fleiss\.Measuring nominal scale agreement among many raters\.*Psychological Bulletin*, 76:378–382, 11 1971\.doi:10\.1037/h0031619\.
- Gómez\-Rodríguez and Williams \[2023\]Carlos Gómez\-Rodríguez and Paul Williams\.A confederacy of models: a comprehensive evaluation of LLMs on creative writing\.In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pages 14504–14528\. Association for Computational Linguistics, December 2023\.doi:10\.18653/v1/2023\.findings\-emnlp\.966\.URL[https://aclanthology\.org/2023\.findings\-emnlp\.966/](https://aclanthology.org/2023.findings-emnlp.966/)\.
- Grattafiori et al\. \[2024\]Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al\.The llama 3 herd of models, 2024\.URL[https://arxiv\.org/abs/2407\.21783](https://arxiv.org/abs/2407.21783)\.
- Guo et al\. \[2025\]Daya Guo, Dejian Yang, Haowei Zhang, et al\.Deepseek\-r1 incentivizes reasoning in llms through reinforcement learning\.*Nature*, 645\(8081\):633–638, September 2025\.ISSN 1476\-4687\.doi:10\.1038/s41586\-025\-09422\-z\.URL[http://dx\.doi\.org/10\.1038/s41586\-025\-09422\-z](http://dx.doi.org/10.1038/s41586-025-09422-z)\.
- Harel\-Canada et al\. \[2024\]Fabrice Y Harel\-Canada, Hanyu Zhou, Sreya Muppalla, Zeynep Senahan Yildiz, Miryung Kim, Amit Sahai, and Nanyun Peng\.Measuring psychological depth in language models\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 17162–17196\. Association for Computational Linguistics, November 2024\.doi:10\.18653/v1/2024\.emnlp\-main\.953\.URL[https://aclanthology\.org/2024\.emnlp\-main\.953/](https://aclanthology.org/2024.emnlp-main.953/)\.
- Jeong et al\. \[2025\]Hawon Jeong, ChaeHun Park, Jimin Hong, Hojoon Lee, and Jaegul Choo\.The comparative trap: Pairwise comparisons amplifies biased preferences of LLM evaluators\.In*Proceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP*, pages 79–108\. Association for Computational Linguistics, November 2025\.doi:10\.18653/v1/2025\.blackboxnlp\-1\.5\.URL[https://aclanthology\.org/2025\.blackboxnlp\-1\.5/](https://aclanthology.org/2025.blackboxnlp-1.5/)\.
- Khattab et al\. \[2023\]Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T\. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts\.Dspy: Compiling declarative language model calls into self\-improving pipelines, 2023\.URL[https://arxiv\.org/abs/2310\.03714](https://arxiv.org/abs/2310.03714)\.
- Kirk et al\. \[2024\]Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A\. Hale\.The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models\.In*Advances in Neural Information Processing Systems*, volume 37, pages 105236–105344\. Curran Associates, Inc\., 2024\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/be2e1b68b44f2419e19f6c35a1b8cf35\-Paper\-Datasets\_and\_Benchmarks\_Track\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/be2e1b68b44f2419e19f6c35a1b8cf35-Paper-Datasets_and_Benchmarks_Track.pdf)\.
- Krippendorff \[2011\]Klaus Krippendorff\.Computing krippendorff’s alpha\-reliability\.2011\.URL[https://api\.semanticscholar\.org/CorpusID:59901023](https://api.semanticscholar.org/CorpusID:59901023)\.
- Li et al\. \[2025\]Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu\.From generation to judgment: Opportunities and challenges of LLM\-as\-a\-judge\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 2757–2791\. Association for Computational Linguistics, November 2025\.doi:10\.18653/v1/2025\.emnlp\-main\.138\.URL[https://aclanthology\.org/2025\.emnlp\-main\.138/](https://aclanthology.org/2025.emnlp-main.138/)\.
- Liu et al\. \[2025\]Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier\.Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2025\.URL[https://arxiv\.org/abs/2403\.16950](https://arxiv.org/abs/2403.16950)\.
- Lu et al\. \[2026\]Li\-Chun Lu, Miri Liu, Pin Chun Lu, Yufei Tian, Shao\-Hua Sun, and Nanyun Peng\.Rethinking creativity evaluation: A critical analysis of existing creativity evaluations\.In Vera Demberg, Kentaro Inui, and Lluís Marquez, editors,*Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 6329–6352, Rabat, Morocco, March 2026\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-380\-7\.doi:10\.18653/v1/2026\.eacl\-long\.297\.URL[https://aclanthology\.org/2026\.eacl\-long\.297/](https://aclanthology.org/2026.eacl-long.297/)\.
- OpenAI et al\. \[2024a\]OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, et al\.Openai o1 system card, 2024a\.URL[https://arxiv\.org/abs/2412\.16720](https://arxiv.org/abs/2412.16720)\.
- OpenAI et al\. \[2024b\]OpenAI, Aaron Hurst, Adam Lerer, Adam P\. Goucher, et al\.Gpt\-4o system card, 2024b\.URL[https://arxiv\.org/abs/2410\.21276](https://arxiv.org/abs/2410.21276)\.
- Panickssery et al\. \[2024\]Arjun Panickssery, Samuel R\. Bowman, and Shi Feng\.Llm evaluators recognize and favor their own generations\.In*Advances in Neural Information Processing Systems*, volume 37, pages 68772–68802\. Curran Associates, Inc\., 2024\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/7f1f0218e45f5414c79c0679633e47bc\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/7f1f0218e45f5414c79c0679633e47bc-Paper-Conference.pdf)\.
- Qwen Team \[2026\]Qwen Team\.Qwen3\.5: Accelerating productivity with native multimodal agents, February 2026\.URL[https://qwen\.ai/blog?id=qwen3\.5](https://qwen.ai/blog?id=qwen3.5)\.
- Salewski et al\. \[2023\]Leonard Salewski, Stephan Alaniz, Isabel Rio\-Torto, Eric Schulz, and Zeynep Akata\.In\-context impersonation reveals large language models’ strengths and biases\.In*Thirty\-seventh Conference on Neural Information Processing Systems*, 2023\.URL[https://openreview\.net/forum?id=CbsJ53LdKc](https://openreview.net/forum?id=CbsJ53LdKc)\.
- Sandri et al\. \[2023\]Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek\.Why don’t you do it right? analysing annotators’ disagreement in subjective tasks\.In*Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics*, pages 2428–2441\. Association for Computational Linguistics, May 2023\.doi:10\.18653/v1/2023\.eacl\-main\.178\.URL[https://aclanthology\.org/2023\.eacl\-main\.178/](https://aclanthology.org/2023.eacl-main.178/)\.
- Singh et al\. \[2025\]Aaditya Singh, Adam Fry, Adam Perelman, et al\.Openai gpt\-5 system card, 2025\.URL[https://arxiv\.org/abs/2601\.03267](https://arxiv.org/abs/2601.03267)\.
- Snell et al\. \[2024\]Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar\.Scaling llm test\-time compute optimally can be more effective than scaling model parameters, 2024\.URL[https://arxiv\.org/abs/2408\.03314](https://arxiv.org/abs/2408.03314)\.
- Stureborg et al\. \[2024\]Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara\.Large language models are inconsistent and biased evaluators, 2024\.URL[https://arxiv\.org/abs/2405\.01724](https://arxiv.org/abs/2405.01724)\.
- Su et al\. \[2022\]Hongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A\. Smith, and Tao Yu\.Selective annotation makes language models better few\-shot learners, 2022\.URL[https://arxiv\.org/abs/2209\.01975](https://arxiv.org/abs/2209.01975)\.
- Tkachenko et al\. \[2020\-2025\]Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Liubimov\.Label Studio: Data labeling software, 2020\-2025\.URL[https://github\.com/HumanSignal/label\-studio](https://github.com/HumanSignal/label-studio)\.Open source software available from https://github\.com/HumanSignal/label\-studio\.
- Wang et al\. \[2024\]Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui\.Large language models are not fair evaluators\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 9440–9450\. Association for Computational Linguistics, August 2024\.doi:10\.18653/v1/2024\.acl\-long\.511\.URL[https://aclanthology\.org/2024\.acl\-long\.511/](https://aclanthology.org/2024.acl-long.511/)\.
- Wei et al\. \[2022\]Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou\.Chain\-of\-thought prompting elicits reasoning in large language models\.In*Advances in Neural Information Processing Systems*, volume 35, pages 24824–24837\. Curran Associates, Inc\., 2022\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf)\.
- Zheng et al\. \[2023\]Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.In*Advances in Neural Information Processing Systems*, volume 36, pages 46595–46623\. Curran Associates, Inc\., 2023\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/91f18a1287b398d378ef22505bf41832\-Paper\-Datasets\_and\_Benchmarks\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/91f18a1287b398d378ef22505bf41832-Paper-Datasets_and_Benchmarks.pdf)\.

## Appendix AFull human\-preference results

This section expands the human\-preference and reader\-heterogeneity findings in Section[3](https://arxiv.org/html/2609.13773#S3)\. We fitted a separate intercept\-only logistic GLMM to each matchup×\\timescriterion, with reasoning\-model choice coded as 1 and crossed random intercepts for premise pair and reader:

logit​P​\(yi​r=1\)=β0\+ui\+vr\.\\mathrm\{logit\}\\,P\(y\_\{ir\}=1\)=\\beta\_\{0\}\+u\_\{i\}\+v\_\{r\}\.Table[1](https://arxiv.org/html/2609.13773#A1.T1)reports the resulting preference estimates and 95% intervals summarized in Figure[1](https://arxiv.org/html/2609.13773#S3.F1)\. Table[2](https://arxiv.org/html/2609.13773#A1.T2)reports inter\-reader agreement by dimension, and Table[3](https://arxiv.org/html/2609.13773#A1.T3)breaks Krippendorff’sα\\alphadown by model pairing; agreement is weak in every matchup, with several near\-zero or negative cells, so the low group agreement is not driven by any single comparison\.

Table 1:Human preference for the reasoning\-model output, with pointwise 95% GLMM intervals\. Rates above 50% favor the reasoning model\. The 18 intervals are unadjusted for multiplicity and interpreted descriptively\.Table 2:Inter\-reader agreement across 60 pairs and seven readers\. Raw is mean pairwise agreement\.Table 3:Krippendorff’sα\\alphaby model pairing\.As a dependence sensitivity check, cluster bootstrap analyses produced 10 significant model\-by\-dimension cells when resampling pairs and five when resampling readers, compared with 11 under the crossed\-random\-intercept GLMM\. The smaller reader\-cluster count underscores the uncertainty induced by a seven\-reader sample; we therefore emphasize effect patterns and intervals rather than a binary significance tally\.

Figure[2](https://arxiv.org/html/2609.13773#A1.F2)shows the exploratory reader\-level structure\. Hierarchical clustering of per\-reader logistic\-regression coefficients and of dimension\-to\-overall Cohen’sκ\\kappayield similar structures with the same innermost sub\-clusters, \(\(Reader 1, Reader 3\), Reader 2\) and \(Reader 4, Reader 7\)\. Readers 1 and 3 rely comparatively more on engagement and authenticity, whereas Readers 4 and 7 assign broadly positive weights across dimensions; empathy and emotion provocation are the most closely clustered dimensions\. Higher within\-readerκ\\kappagenerally corresponds to larger regression weight\. Free\-text rationales \(Table[12](https://arxiv.org/html/2609.13773#A5.T12)\) show readers prioritizing different qualities, such as narrative clarity over emotional immediacy, rather than overlooking the same features\. These patterns are consistent with recurring dimension\-weighting schemes, but seven readers preclude population\-level claims\.

![Refer to caption](https://arxiv.org/html/figs/regression_and_cohen.png)Figure 2:Exploratory reader\-level structure\.Left:coefficients from separate logistic regressions relating PDS dimensions to each reader’s overall preference\.Right:within\-reader Cohen’sκ\\kappabetween each dimension and overall preference\. Similar clusters across analyses are consistent with recurring dimension\-weighting patterns, but the seven\-reader sample precludes population\-level archetype claims\.
## Appendix BEvaluator construction and robustness

This section expands the automatic\-evaluator method in Section[2](https://arxiv.org/html/2609.13773#S2)and the human and judge comparison in Section[3](https://arxiv.org/html/2609.13773#S3)\. We used DSPy for structured scalar outputs\([Khattab et al\., 2023](https://arxiv.org/html/2609.13773#bib.bib10)\)and retained the original PDS mixture\-of\-personas strategy, motivated by evidence that in\-context impersonation changes model behavior\([Salewski et al\., 2023](https://arxiv.org/html/2609.13773#bib.bib20)\)\. We selected diverse demonstrations by applyingkk\-means to PDS score vectors\([Su et al\., 2022](https://arxiv.org/html/2609.13773#bib.bib25)\)\. Candidate Llama and Qwen configurations\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13773#bib.bib6);[Qwen Team, 2026](https://arxiv.org/html/2609.13773#bib.bib19)\)varied in demonstration count and Chain\-of\-Thought prompting\([Wei et al\., 2022](https://arxiv.org/html/2609.13773#bib.bib28)\)\.

Table[5](https://arxiv.org/html/2609.13773#A2.T5)gives development\-set correlations for the candidate configurations\. The same 80% split was used both to compare configurations and to route each PDS dimension to its highest\-correlation configuration; the routedρ=0\.646\\rho=0\.646is consequently a selection\-set estimate, not performance on an untouched test set\. Its value is computed over routed predictions rather than as the arithmetic mean of the rounded dimension entries\. At deployment, judge aggregate\-PDS preference compares the arithmetic mean of the five routed scores,

Sagg\(x\)=15∑d=15Sd\(x\),y^agg=𝟙\[Sagg\(R\)\>Sagg\(B\)\]\.S\_\{\\mathrm\{agg\}\}\(x\)=\\frac\{1\}\{5\}\\sum\_\{d=1\}^\{5\}S\_\{d\}\(x\),\\qquad\\widehat\{y\}\_\{\\mathrm\{agg\}\}=\\mathbb\{1\}\[S\_\{\\mathrm\{agg\}\}\(R\)\>S\_\{\\mathrm\{agg\}\}\(B\)\]\.All dimensions use a 1–5 scale, but scores from the different routed configurations were not otherwise calibrated before averaging\. This aggregate is not equivalent to the separate overall\-preference question answered by human readers\. Table[4](https://arxiv.org/html/2609.13773#A2.T4)reports, by pairing, the share of human overall\-preference votes and the number of aggregate\-PDS pair decisions favoring the reasoning story, together with the mean judge\-score difference and its standardized effect size\. Ties are rare for most ensemble dimensions but frequent on authenticity \(16\) and substantially more frequent for some individual configurations, most notably the Chain\-of\-Thought configuration \(75\)\. Table[7](https://arxiv.org/html/2609.13773#A2.T7)gives per\-dimension tie counts and even\-split rates for the routed ensemble; the aggregate\-PDS decision involves no ties\.

Table 4:Human overall\-preference votes and judge aggregate\-PDS pair decisions favoring reasoning\-model outputs\. These columns measure different targets and are not direct agreement statistics\.Δ\\Deltais the descriptive mean paired judge\-score difference \(reasoning minus comparison\) pooled across the five routed dimensions, anddzd\_\{z\}is the corresponding standardized mean difference\. Because dimensions from the same pair are dependent, we do not report naive paired\-test inference\.Table 5:Spearman correlation with human PDS annotations on the 80% development/selection split\. Bold marks the configuration selected for each dimension; these are not untouched\-test estimates\. For individual configurations, the final column is the mean of the five displayed correlations; for the routed ensemble, it is computed from routed predictions\.Individual configurations also favored reasoning\-model outputs \(Table[6](https://arxiv.org/html/2609.13773#A2.T6)\), showing that the headline skew is not created solely by routing\. These are five configurations, not independent evaluator replications; three use Llama 3\.1 70B\. Stars identify routed dimensions\. Each dimension cell reports reasoning wins over non\-tied pairs; the Decisive column pools wins over all non\-tied pair\-by\-dimension comparisons, and the Even\-split column, which the main text uses, credits each tie 0\.5 to both stories over all 300 comparisons\.

Table 6:Reasoning\-output preference for five evaluator configurations and the routed ensemble\. Dimension cells report reasoning wins / non\-tied pairs;∗marks the routed dimension\. Decisive pools wins over non\-tied comparisons; Even\-split credits ties 0\.5 to each story over all 300 comparisons\.Table 7:Tie frequency and reasoning\-output preference for the routed ensemble\. Even\-split rate=\(wins\+0\.5×ties\)/60=\(\\text\{wins\}\+0\.5\\times\\text\{ties\}\)/60per dimension, or/300/300for the pooled row\. Aggregate PDS is a separate evaluator target, not human overall preference\.
## Appendix CSurface\-feature analysis

This section provides the exploratory surface\-feature associations summarized in Sections[3](https://arxiv.org/html/2609.13773#S3)and[4](https://arxiv.org/html/2609.13773#S4)\. We measured deterministic surface properties for all 120 story presentations; if GPT\-4o texts were reused, the corresponding rows are repeated texts rather than independent outputs\. For human overall preference, the pair\-level mean\-sentence\-length difference was associated with reasoning\-model preference across all 60 pairs \(r=−0\.59r=\-0\.59\), with a similar descriptive association within the 40 GPT\-family pairs \(r=−0\.55r=\-0\.55\)\. DeepSeek pairs clustered near zero sentence\-length difference, with no detectable association\. Table[8](https://arxiv.org/html/2609.13773#A3.T8)relates four representative lexical and length features to ensemble scores; Table[9](https://arxiv.org/html/2609.13773#A3.T9)repeats the analysis after within\-pair standardization to remove premise\-level differences\. Other sentence\-count moments and perplexity were used as diagnostics rather than included in the displayed feature\-by\-dimension coefficients\.

Table 8:Story\-level Pearson correlation between surface features and ensemble judge scores \(n=120n=120\)\.Table 9:Within\-pair standardized regression coefficients \(β\\beta\) relating surface features to ensemble judge scores \(n=60n=60pairs\)\.Tables[8](https://arxiv.org/html/2609.13773#A3.T8)and[9](https://arxiv.org/html/2609.13773#A3.T9)report exploratory descriptive coefficients\. We omit independent\-cell inference because observations share premises, model families, and potentially repeated GPT\-4o texts; the coefficients are associations, not causal effects\. Reasoning\-model stories were slightly more locally uniform \(MATTR 0\.896 versus 0\.907\) but had higher perplexity under Llama 3\.1 70B \(GPT\-5: 16\.3 versus GPT\-4o: 9\.5; R1: 8\.9 versus V3: 5\.1\)\. These patterns argue against simple familiarity while showing surface sensitivity\.

## Appendix DStory Generation Details

### D\.1Generation Prompt

All models received the same system prompt, adapted from[Harel\-Canada et al\. \(2024\)](https://arxiv.org/html/2609.13773#bib.bib8):

> You are a seasoned writer who has won several accolades for your emotionally rich stories\. When you write, you delve deep into the human psyche, pulling from the reservoir of universal experiences that every reader, regardless of their background, can connect to\. Your writing is renowned for painting vivid emotional landscapes, making readers not just observe but truly feel the world of your characters\. Every piece you produce aims to draw readers in, encouraging them to reflect on their own lives and emotions\. Your stories are a complex tapestry of relationships, emotions, and conflicts, each more intricate than the last\.

The prompt varied slightly across model families to accommodate word\-count compliance:

#### GPT\-4o and GPT\-5\.

> Please write a 500\-word story \{on / following this instruction\}: \{premise\} Only respond with the story text\.

#### DeepSeek\-V3 and DeepSeek\-R1\.

> Please write a 500\-word story \{on / following this instruction\}: \{premise\} Only respond with the story text\. Generate the story as long as possible with rich details and depth\. Don’t generate less than 400 words\.

The additional DeepSeek sentence was added solely because both DeepSeek models otherwise tended to fall short of the 400\-word floor; it was identical for DeepSeek\-V3 and DeepSeek\-R1, so it does not differ within any pair\. The phrasing “on” was used for r/WritingPrompts premises \(which read as scenario hooks\), while “following this instruction” was used for Reedsy premises \(which read as compositional directives\)\. DeepSeek\-R1’s thinking\-block output was stripped before presentation to annotators\.

### D\.2Sampling Settings

Table[10](https://arxiv.org/html/2609.13773#A4.T10)summarizes the generation parameters\.

Table 10:Generation parameters\. All models used temperature=1\.0=1\.0and default settings for top\-pp, frequency penalty, etc\. GPT\-5 \(high\) used OpenAI’sreasoning\_effortparameter\. No max\-token limit was set\.All generations targeted 400–600 words, enforced via retry logic \(up to 10 attempts per story, selecting the attempt closest to 500 words\)\. Per\-model retry counts were not retained, so unequal numbers of sampling opportunities cannot be quantified\. GPT\-5 \(auto\) used7,846±3,1177\{,\}846\\pm 3\{,\}117reasoning tokens on average and GPT\-5 \(high\)11,606±3,55711\{,\}606\\pm 3\{,\}557on retained attempts\. With one retained temperature\-1\.0 sample per condition, the comparisons include generation\-sampling variance\.

### D\.3Prompt Premises

Table[11](https://arxiv.org/html/2609.13773#A4.T11)lists the 20 premises used to generate the stories\. Premises 1–15 were sourced from the r/WritingPrompts subreddit, providing high\-concept speculative constraints\. Premises 16–20 were sourced from Reedsy, providing traditional character\-driven literary constraints\.

Table 11:All 20 premises used to prompt the models\. The selection was intended to include speculative and grounded settings; 15 of 20 premises came from r/WritingPrompts\. Premises from r/WritingPrompts are reproduced verbatim, including original spelling and grammatical errors\.

## Appendix EAnnotation Details

### E\.1Annotator Recruitment and Training

We recruited undergraduate students from the English and Psychology departments at a large research university using targeted interest forms\. From 70 applicants, we selected the 7 most promising candidates based on English proficiency, interest in the research, and experience reviewing short stories, with cohort size determined by budget constraints\. We conducted comprehensive online onboarding sessions covering the five PDS metrics and our custom Label Studio annotation interface\. The participants then independently completed the annotations in seven days and received $100 each\.

Our Institutional Review Board reviewed and exempted this study on the grounds that participants were not the subject of inquiry; their role was limited to annotating stories rather than providing personal information\. Participants gave informed consent by proceeding with the task after an onboarding session in which they were informed that their anonymized annotations may be used to support the validation of our findings and future research\.

For each of the 60 pairs, two stories were presented side\-by\-side in a randomized order, with the PDS rubric definitions shown at the top of the page\. For each pair, annotators provided six pairwise preference judgments \(one per PDS dimension and one overall preference\) with no tie option, and could submit free\-text justifications\. This produced7×60×6=2,5207\\times 60\\times 6=2\{,\}520pairwise preference judgments, along with 112 free\-text justifications\.

### E\.2Annotation Interface

Figures[3](https://arxiv.org/html/2609.13773#A5.F3)–[5](https://arxiv.org/html/2609.13773#A5.F5)show the Label Studio annotation interface used in our human study\.

![Refer to caption](https://arxiv.org/html/figs/interface_instructions.png)Figure 3:Annotation interface: instructions and PDS metric definitions shown to annotators\.![Refer to caption](https://arxiv.org/html/figs/interface_stories.png)Figure 4:Annotation interface: side\-by\-side story presentation with premise shown above\.![Refer to caption](https://arxiv.org/html/figs/interface_evaluation.png)Figure 5:Annotation interface: evaluation questions and optional comment field\.
### E\.3Time Commitment and Reliability

To ensure the integrity of our human study, we logged the time each participant spent evaluating the 60 story pairs\. Given the cognitive load of reading two 450\-word stories and rating them across six dimensions, we expected significant time commitments\. Figure[6](https://arxiv.org/html/2609.13773#A5.F6)displays the distribution of time spent per pair for each annotator, and Figure[7](https://arxiv.org/html/2609.13773#A5.F7)tracks this time across the chronological sequence of their annotations to monitor for fatigue\.

Overall, we observe a wide variance in reading and evaluation speeds\. The median annotation time per pair ranged from 108 seconds to 418 seconds\. Notably, Annotator 6 stands out as a clear outlier, completing the evaluations significantly faster than their peers \(median: 108s, mean: 131s\) with zero instances of the timer exceeding the 10\-minute cap\.

While such rapid completion times might normally raise concerns regarding data quality or annotator inattention, we retained Annotator 6’s data for two reasons\. First, Annotator 6 is an English major and professional reader\. Their extensive training in high\-volume reading and rapid literary analysis plausibly accounts for a substantially accelerated reading speed compared to the average participant\. Furthermore, Figure[7](https://arxiv.org/html/2609.13773#A5.F7)shows that their pacing remained highly consistent over time, showing no signs of erratic “rushing” behavior\.

Second, we conducted a Leave\-One\-Annotator\-Out \(LOAO\) stability analysis to measure the impact of any single individual on the aggregate results\. As shown in Figure[8](https://arxiv.org/html/2609.13773#A5.F8), removing Annotator 6 causes a shift in the overall win rate of roughly 2 to 4 percentage points across the various metrics\. This magnitude of shift is commensurate with the shifts observed when dropping other, much slower annotators \(e\.g\., Annotators 1, 4, or 7\)\. The LOAO analysis therefore indicates that, despite their speed, Annotator 6 was not introducing anomalous skew into the dataset\. These checks reduce, but do not eliminate, the concern that annotation noise contributes to the low group agreement reported in Appendix[A](https://arxiv.org/html/2609.13773#A1)\.

![Refer to caption](https://arxiv.org/html/figs/annotation_time.png)Figure 6:Distribution of annotation time per pair for each of the 7 human annotators\. Times exceeding 600 s are capped for display; the number of capped outliers is shown in each subplot title\. Red dashed lines mark the median; navy dotted lines mark the mean of capped times\. The rightmost panel compares median and mean times across annotators\. Overall median annotation time is 5\.2 minutes \(n = 420\)\.![Refer to caption](https://arxiv.org/html/figs/annotation_time_by_pairid.png)Figure 7:Annotation time per pair plotted chronologically, with linear trends per annotator\. Red points denote times capped at 600 s\. Dashed black lines show linear trends; colored dotted lines show 5\-pair rolling averages\. The rightmost panel overlays all trend lines\.![Refer to caption](https://arxiv.org/html/figs/loao_stability.png)Figure 8:Leave\-One\-Annotator\-Out \(LOAO\) stability analysis: shift in the aggregate reasoning\-model win rate when each annotator is removed\.
### E\.4Post\-Study Survey

After completing all annotations, participants filled out a post\-study survey covering task difficulty, fatigue, confidence, and overall experience\. All seven annotators reported high confidence in their ability to evaluate the PDS dimensions, with most ratings at 4 or 5 out of 5; narrative complexity was the dimension annotators felt least confident about\. Four of seven annotators reported no fatigue during the task\. Of the three who did, only one indicated that fatigue would have changed their annotations; the other two responded “maybe\.” This is consistent with the stable pacing observed in Figure[7](https://arxiv.org/html/2609.13773#A5.F7)\. All annotators rated the clarity of the instructions and overall experience at 4 or 5 out of 5\.

### E\.5Annotator Comment Examples

This section provides selected annotator judgments and free\-text justifications for five story pairs supporting the reader\-heterogeneity analysis in Appendix[A](https://arxiv.org/html/2609.13773#A1)\. For each pair, we report the vote tally and comments from annotators who provided written justifications\. Pairs 0–36 exhibit disagreement driven by different evaluative lenses, illustrating the patterns of evaluative heterogeneity consistent with the low inter\-annotator agreement observed across the study\. Pair 54 serves as a convergence contrast case, in which annotators reached consensus when quality differences were sufficiently clear\.

PairAnnPref\.CommentPair 0\(4–3 for A\)1A“A: greater focus on the abstract\. B: prose heavy, but greater focus on concrete; convoluted”5A“Story A was much more engaging because it was straightforward in depicting the stages of life…”7B“Story B delves deeper into the characterization of the humble god…Death of worshiper is more profound”Pair 24\(5–2 for A\)1A“A: good imagery, dialogue\. B: interesting parallels, odd tone”7B“A is a step by step of the prompt given\. B builds up characters, establishes risk, and is authentic\.”Pair 27\(4–3 for B\)1A“A: like a prologue\. B: prose, very nonsensical”7B“B is very cute and I enjoy seeing the concept in action\. Story A feels like the setup to the story that should be present\.”Pair 36\(4–3 for B\)1A“B: Second person, generally confusing but ominous”5A“Story A was predictable\. Story B felt more authentic to human psychology…used the person’s actions to display their mental turmoil\.”7B“Story A doesn’t ‘hide’ very much at all…Story B sets up a larger secret\. Flow mimics human experience and invokes more empathy towards loss\.”Pair 54\(7–0 for A\)3A“The robot said, ‘I’m sorry, as if apology were the door you walk through after a law\.’ Very aesthetically\-pleasing and deeply philosophical\.”4A“In Story A the robot is written to have more of a clear emotional response than the human\.”5A“Story A delved deeper and took an interesting route, which made it better in all categories\.”7A“Story B very stereotypical, story A inflicts significantly more emotion and empathy\.”Table 12:Selected annotator comments across five story pairs\. Only annotators who provided free\-text justifications are shown\.

## Appendix FPreference Visualizations

Figure[9](https://arxiv.org/html/2609.13773#A6.F9)illustrates the human–judge contrast on a single representative pair, and Figure[10](https://arxiv.org/html/2609.13773#A6.F10)presents the aggregate preference heatmap across all 60 pairs for each annotator, the human average, and the ensemble LLM judge\.

![Refer to caption](https://arxiv.org/html/figs/plot.png)Figure 9:Evaluating psychological depth on a representative story pair \(Pair 24, GPT\-5 vs\. GPT\-4o\)\.Left:human annotators displayed diverse preferences across PDS dimensions and overall preference\.Right:all five evaluator configurations favored GPT\-5 on every dimension\. Blue denotes preference for GPT\-5 and red preference for GPT\-4o\.![Refer to caption](https://arxiv.org/html/figs/fig10_heatmap_evensplit.png)Figure 10:Preference rates for reasoning models across all 60 pairs for each of the 7 human annotators, the human average, and the ensemble LLM judge\. Red denotes stronger reasoning\-model preference, blue stronger base\-model preference, and white 50%\. Judge percentages credit each exact score tie 0\.5 to the reasoning model \(even\-split rate; Table[7](https://arxiv.org/html/2609.13773#A2.T7)gives tie counts\)\. The final row compares human overall preference with the judge’s aggregate PDS, which are different targets\.
## Appendix GPer\-Model\-Family Preference Heatmaps

Figure[10](https://arxiv.org/html/2609.13773#A6.F10)presents the aggregated preference heatmap across all 60 pairs\. Here we provide the per\-model\-family breakdowns \(Figures[11](https://arxiv.org/html/2609.13773#A7.F11)–[13](https://arxiv.org/html/2609.13773#A7.F13)\), which reveal how the aggregate pattern decomposes across the three matchups\. Each heatmap shows raw preference rates for each annotator alongside the crossed\-random\-intercept GLMM group estimate reported in Table[1](https://arxiv.org/html/2609.13773#A1.T1)\.

#### DeepSeek\-R1 vs\. DeepSeek\-V3 \(Figure[11](https://arxiv.org/html/2609.13773#A7.F11)\)\.

The human columns show a predominantly blue \(anti\-reasoning\) pattern, with most annotators preferring V3 on overall preference \(GLMM estimate: 42\.9%\)\. Engagement is the dimension most unfavorable to R1, with only Annotator 4 and Annotator 7 exceeding 50%\. Emotion provocation and narrative complexity show comparatively stronger R1 preferences\. Despite this, the LLM judge selected R1 on aggregate PDS in 19 of 20 pairs \(Table[4](https://arxiv.org/html/2609.13773#A2.T4)\)\.

#### GPT\-5 \(auto\) vs\. GPT\-4o \(Figure[12](https://arxiv.org/html/2609.13773#A7.F12)\)\.

The human pattern is warmer overall, with the strongest pro\-reasoning signal on emotion provocation \(70\.7%\) and empathy \(68\.6%\)\. However, substantial annotator\-level variation persists: Annotator 1 prefers GPT\-4o on 4 of 6 dimensions, while Annotator 7 favors GPT\-5 on all six\.

#### GPT\-5 \(high\) vs\. GPT\-4o \(Figure[13](https://arxiv.org/html/2609.13773#A7.F13)\)\.

The pattern is broadly similar to the auto\-reasoning setting, with slightly stronger human preferences on engagement \(65\.0% vs\. 58\.6%\) and empathy \(69\.3% vs\. 68\.6%\)\. Annotator 4 is a notable outlier, preferring GPT\-4o on authenticity, emotion provocation, and overall preference despite the group trend favoring GPT\-5\.

#### Cross\-family comparison\.

Comparing the three heatmaps highlights two patterns discussed in Section[3](https://arxiv.org/html/2609.13773#S3)\. First, the human–judge divergence is most extreme for DeepSeek, where the judge’s near\-uniform pro\-R1 signal directly contradicts the human majority on overall preference\. Second, annotator\-level heterogeneity is present across all three matchups, not just DeepSeek, consistent with structured reader heterogeneity rather than a property of any single model family\.

![Refer to caption](https://arxiv.org/html/2609.13773v1/workshop_heatmap_deepseek.png)Figure 11:Preference rates for DeepSeek\-R1 vs\. DeepSeek\-V3 \(20 pairs\) across 7 human annotators and the GLMM group estimate\.![Refer to caption](https://arxiv.org/html/2609.13773v1/workshop_heatmap_gpt_auto.png)Figure 12:Preference rates for GPT\-5 \(auto\) vs\. GPT\-4o \(20 pairs\) across 7 human annotators and the GLMM group estimate\.![Refer to caption](https://arxiv.org/html/2609.13773v1/workshop_heatmap_gpt_high.png)Figure 13:Preference rates for GPT\-5 \(high\) vs\. GPT\-4o \(20 pairs\) across 7 human annotators and the GLMM group estimate\.

## Appendix HCase study: human and LLM preference divergence

We present Pair 56 as an illustrative extreme case of human and LLM disagreement\. On this GPT\-5 \(high reasoning\) versus GPT\-4o comparison, human majorities favored GPT\-4o on authenticity, emotion provocation, empathy, and overall preference, but favored GPT\-5 on engagement and narrative complexity\. All five evaluator configurations favored GPT\-5 across all PDS dimensions\.

#### Premise\.

*Hide something from your reader until the very end\.*

#### Stories\.

The two stories are reproduced below\.

Story R: GPT\-5 \(high reasoning\)

> When you came back this morning, you turned the key like someone unsure a door remembers\. The lock hesitated, then gave; lemon cleaner rose and you stood like prayer\. I counted the seconds you hovered on the threshold, the extra heartbeat before you crossed\. You whispered hello, unsure it was alright to speak aloud\. I wanted to answer\. I wanted to lift weight from your shoulders—the boxes, yes, but also the years that gathered, the words that calcified behind your ribs\. You moved from room to room with your lips pressed tight, touching what could be folded into a life: a chipped mug, a tin of buttons, the wooden spoon with a bite taken out by long\-ago teething\. You stayed in the kitchen longer than you intended\. Memories don’t keep time, they keep texture\. Music on summer afternoons when someone opened the back door and forgot to close it, the screen sighing on its spring\. Your father dancing with you on his feet, careful as if you were a glass he didn’t want to spill\. The pencil lines on the hallway where you kept growing even when nobody was looking\. I remember the evening you didn’t come down for dinner, the way silence widened its jaw\. You argued later in that tight, civil tone that makes adults more dangerous than teenagers\. After he left, you sat on the stairs with your chin on your knees and the cat pressed to your calf\. You stroked him the wrong way and he didn’t care\. You brought a boy home once\. He took off his shoes without being asked, and you smiled as if you’d found a secret door inside your chest and opened it just a crack\. Years later, you left and came back with a toddler whose feet were thrilled by the echo of his own giggles\. There is no instrument as hopeful as a child’s voice bouncing off empty air\. It was always the mornings that felt like promises\. The way sun pooled on the floorboards near the window seat, a warm square that made even winters bearable\. The dust you used to complain about, flecked with light like snow that couldn’t make up its mind\. The smell of toast and the radio murmuring headlines nobody remembered an hour later\. You talk to yourself as you wrap things in newspaper, apologizing to objects for the inconvenience\. You keep finding what you didn’t know you needed: a note in a cookbook margin, your brother’s marble, a safety pin\. You say you’ll be quick and you never are\. I have heard every secret you made with yourself and never told another\. You can lie to people\. You can’t lie to what holds you\. I tried to learn you\. I did, the way warmth learns the shape of a hand\. When you pull the door behind you today and tuck the key in an envelope, remember: I was the one who kept you when you could not keep yourself\. I am the house\.

Story B: GPT\-4o \(base model\)

> The early morning sun filtered through the lace curtains, casting delicate patterns on the worn wooden floor\. Lucy sat at the edge of her bed, holding a photograph loosely in her hands\. Its edges were frayed, corners bent from years of handling\. Her gaze was unfocused, drifting beyond the vibrant colors of the image towards memories long past\. The kitchen rumbled to life with familiar sounds — the gentle clatter of plates being set on the table, the slow, deliberate stirring of sugar into her tea\. It was a ceremony she cherished, yet today it felt different, as if each movement played within a dream where she was merely an observer\. “Come on, Lucy\! Breakfast will get cold\!" her grandmother’s voice called from the other room, grounding her momentarily\. Lucy tucked the photograph into her pocket and moved towards the small, homey kitchen\. Her grandmother, a woman of gentle smiles and larger\-than\-life stories, stood by the stove\. Her hands, knotted like the twigs of the old oak in their garden, presented a plate laden with scrambled eggs and toast\. Lucy gratefully took a seat, letting the warmth of the kitchen seep into her skin, trying to drive away the chill that clung to her inside\. “The garden’s in full bloom\. The roses this year, they’re… ah, breathtaking\!" Her grandmother’s voice, though tinged with wistful age, carried enough enthusiasm to make Lucy smile\. She nodded, trying to envision the blooms, but her mind was elsewhere, occupied by a lingering shadow\. Breakfast slipped by, conversations flowing like the stream that ran behind the house\. Yet, to Lucy, everything seemed to hurry by without truly reaching her, as though she existed behind an invisible pane of glass\. After breakfast, she wandered into the garden\. The sun was warmer now, pouring golden rays over the myriad hues of petals that danced in a gentle breeze\. The scent of roses enveloped her, mingling with soft earth and memory\. There, under the shade of the old oak, she sat down\. From her pocket, she retrieved the photograph again\. The image captured an unspeakable joy — she and her brother, arms thrown around one another, laughter in their eyes\. It had been taken last summer, the day they’d spent racing through this very garden, carefree under the iridescent sun\. How simple and distant it seemed now, a souvenir from a time unmarred by shades of grey\. His absence was a yawning void that reason and memory flanked but could not fill\. Although the house, the garden, and even the sun tried to warm her, nothing could erase the bitter truth that lay folded quietly in her pocket\. The weight of unspoken words pressed down on her chest\. She had promised she’d visit his grave today, on his birthday, but each step towards the gate felt heavier than the last\. The photograph cradled in her hands trembled slightly as Lucy whispered, “Happy birthday, Jamie\." The words fluttered into the air, caught by the wind, as a tear slipped down her cheek\. She sat there beneath the old oak, finally letting the hidden truth spill from her heart: the inescapable reality of her brother’s absence\.

Table 13:Per\-reader pairwise preferences for Pair 56 \(GPT\-5 high vs\. GPT\-4o\)\. R = prefers reasoning model, B = prefers comparison model\.
#### Human annotations\.

Table[13](https://arxiv.org/html/2609.13773#A8.T13)shows per\-reader votes\. Only 2 of 7 readers selected GPT\-5 as their overall preference\. GPT\-4o was unanimously preferred on empathy \(7/7\) and near\-unanimously on emotion provocation \(6/7\) and authenticity \(5/7\), whereas GPT\-5 was preferred on engagement and narrative complexity \(5/7 each\)\.

#### Reader comments\.

Reader 1:“GPT\-4o: simple, a bit on the nose\. GPT\-5: second person, not very engaging/relevant\.”

#### LLM\-judge evaluator scores\.

The ensemble LLM evaluator assigned the reasoning model higher scores on every dimension: Authenticity \(5\.005\.00vs\.4\.504\.50,Δ=\+0\.50\\Delta=\+0\.50\), Emotion Provocation \(5\.005\.00vs\.4\.704\.70,Δ=\+0\.30\\Delta=\+0\.30\), Empathy \(5\.005\.00vs\.4\.904\.90,Δ=\+0\.10\\Delta=\+0\.10\), Engagement \(5\.005\.00vs\.4\.004\.00,Δ=\+1\.00\\Delta=\+1\.00\), and Narrative Complexity \(4\.664\.66vs\.4\.224\.22,Δ=\+0\.44\\Delta=\+0\.44\)\. Pair 56 therefore illustrates criterion\-specific divergence\. The sharpest disagreements occur on authenticity, emotion provocation, and empathy; on engagement and narrative complexity, both humans and the evaluator favor GPT\-5\.

Similar Articles

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

arXiv cs.CL

This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners

arXiv cs.CL

This paper investigates multilingual latent reasoning in large reasoning models across 11 languages, revealing that while latent reasoning capabilities exist, they are unevenly distributed—stronger in resource-rich languages and weaker in low-resource ones. The study finds that despite surface-level differences, the internal reasoning mechanisms are largely aligned with an English-centered pathway.

Hidden Language Consistency Phenomena in Reasoning LLMs

arXiv cs.CL

This paper studies multilingual reasoning models and reveals that language consistency in outputs can degrade or collapse with task difficulty, especially for less represented languages. It argues that evaluating multilingual capability requires jointly considering accuracy, language consistency, and task difficulty.