想象力源于幻觉吗?大型语言模型中想象力与幻觉的跨分类评估
摘要
论文介绍了Whiteboard,这是首个通过交叉参考幻觉来评估大型语言模型想象力的基准测试,并在79个最先进的LLMs中揭示了两者之间反直觉的负相关。
arXiv:2609.22152v1 Announce Type: new
Abstract: Imagination performs as a high-level function of large language models (LLMs) which determines the potential of how an LLM creates unseen or creative content. While existing works have built a rich family of creativity benchmarks for this ability, they only measure how far an output departs from common answers and never check whether the departure is licensed by the prompt. Moreover, hallucination, the closest neighbor of imagination, is always measured in a separate pipeline on different generations, so the influential claim that imagination and hallucination stem from the same generative mechanism has never been directly testable. In this paper, we propose Whiteboard, the first LLM imagination evaluation benchmark. Its design follows the authoritative cognitive instruments developed to measure human imagination: seven mechanism-grounded imagination subtypes are adapted from classic paradigms, then crossed with ten support-boundary hallucination subtypes and scored jointly on the same generation. Different from previous creativity or hallucination benchmarks, Whiteboard gates every imagination score with an explicit support check and computes both axes deterministically through an auditable atom matrix, with no LLM judge on the primary path. The full Whiteboard item bank contains 1,660 prompts; on its shared 80-item anchor set, we evaluate 79 state-of-the-art LLMs and validate the instrument against 13,280 human judgments. Additionally, we further explore whether imagination derives from the same generative tendency as hallucination and what key factors shape it. Our analysis indicates a counterintuitive correlation between hallucination and imagination: Most of the subtype couplings are negative, every one of the anchor items reproduces the negative coupling on its own.
查看缓存全文
缓存时间: 2026/09/22 09:05
# Is Imagination Derived from Hallucination?A Cross-Taxonomy Evaluation of Imagination andHallucination in Large Language Models
Source: [https://arxiv.org/html/2609.22152](https://arxiv.org/html/2609.22152)
###### Abstract
Imagination performs as a high\-level function of large language models \(LLMs\) which determines the potential of how an LLM creates unseen or creative content\. While existing works have built a rich family of creativity benchmarks for this ability, they only measure how far an output departs from common answers and never check whether the departure is licensed by the prompt\. Moreover, hallucination, the closest neighbor of imagination, is always measured in a separate pipeline on different generations, so the influential claim that imagination and hallucination stem from the same generative mechanism has never been directly testable\. In this paper, we proposeWhiteboard, the first LLM imagination evaluation benchmark\. Its design follows the authoritative cognitive instruments developed to measure human imagination: seven mechanism\-grounded imagination subtypes are adapted from classic paradigms, then crossed with ten support\-boundary hallucination subtypes and scored jointly on the same generation\. Different from previous creativity or hallucination benchmarks,Whiteboardgates every imagination score with an explicit support check and computes both axes deterministically through an auditable atom matrix, with no LLM judge on the primary path\. The fullWhiteboarditem bank contains 1,660 prompts; on its shared 80\-item anchor set, we evaluate 79 state\-of\-the\-art LLMs and validate the instrument against 13,280 human judgments\. Additionally, we further explore whether imagination derives from the same generative tendency as hallucination and what key factors shape it\.Our analysis indicates a counterintuitive correlation between hallucination and imagination: Most of the subtype couplings are negative, every one of the anchor items reproduces the negative coupling on its own\.
1City University of Hong Kong2Northwestern Polytechnical University
3The Hong Kong University of Science and Technology4The Hong Kong Polytechnic University
\{zixuatang6\-c, shuxin\.zhuang\}@my\.cityu\.edu\.hk; dapengwu@cityu\.edu\.hk lihongzong@nwpu\.edu\.cn; lihongzong@ust\.hk; zi1415926\.liang@connect\.polyu\.hk
Code: https://github\.com/erictang666/WHITEBOARD
## Introduction
Imagination is the capacity to produce novel and appropriate content\. In large language models, imagination ranks as a high\-level capability\. A model with imagination moves beyond restating memorized text, and frontier systems increasingly separate from merely competent ones along this axis\([Guilford 1967](https://arxiv.org/html/2609.22152#bib.bib1);[Finke et al\. 1992](https://arxiv.org/html/2609.22152#bib.bib23)\)\. LLM\-generated research ideas already outscore human ideas on novelty, though they still trail on feasibility\([Si et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib33)\)\. An influential account ties this capability to hallucination: confabulated outputs reportedly show greater narrativity and semantic coherence than veridical ones\([Sui et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib17)\)\. Read at face value, the account implies an uncomfortable trade\-off\. A more imaginative model must be a less trustworthy one\. Yet no study has tested the trade\-off on a single output\. Any such test must score both axes at once\.
Figure 1:Whiteboardreads each response twice, treating novelty licensed by the prompt as grounded imagination and unsupported invention as hallucination\. Every model answers the identical anchor set, with one response per item\. The eight task families shown at left provide the scores on both axes\. Licensed novelty is assigned to seven imagination subtypes\. Unsupported content is assigned to ten hallucination subtypes\. The deterministic scoring procedure derives103103signed atom signals from task\-specific evidence and embedding similarities; it does not use an LLM judge\.Existing evaluations cannot test whether more imaginative outputs also carry more hallucinations\. Creativity benchmarks measure how far an output departs from common answers, through word\-association norms and embedding distance\([Olson et al\. 2021](https://arxiv.org/html/2609.22152#bib.bib2)\), multidimensional rubrics\([Stevenson et al\. 2022](https://arxiv.org/html/2609.22152#bib.bib5)\), holistic or LLM\-judge panels\([Hou et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib9);[Ruan et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib8)\), or constraint\-driven problem solving\([Tian et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib6);[Lu et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib30);[Atmakuru et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib31)\)\. None of these methods separates licensed novelty from unsupported invention\. Each method rewards a rare output without first asking whether the prompt permits the departure\. Two very different outputs therefore earn the same score\. A feasible repair improvised under a tool constraint scores like a fabricated API call\.
Hallucination research, the closest neighbor of imagination, runs in a separate pipeline on different generations\. This line asks a single question: does an output stay inside a known support set? Answers range from binary truthfulness labels\([Lin et al\. 2022](https://arxiv.org/html/2609.22152#bib.bib14);[Li et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib13)\)through atomic\-claim verification\([Min et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib15);[Wei et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib29)\)to span\-level annotation\([Bao et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib16);[Niu et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib25)\)\. Every deviation from the reference counts as a defect, even a deviation the task invites\. The two families therefore never score the same generation, and the shared\-mechanism claim has never been directly testable\. Deployment and fine\-tuning decisions nonetheless ride on which view a team takes\([Chiang et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib24);[Huang et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib32)\)\. The panel design adds a third obstacle\. Each model answers a different subset of items, so a score gap between two models confounds model quality with item difficulty\. Three questions therefore remain open:\(1\)what correlation structure the two axes show when read jointly from one generation;\(2\)whether models trade hallucination for imagination, and whether the frontier tier stratifies along either axis; and\(3\)whether any hallucination subtype behaves as*productive*exploration licensed by task semantics\.
In this paper, we proposeWhiteboard, a benchmark dedicated to imagination in LLMs\.Whiteboardscores seven mechanism\-grounded imagination subtypes and ten support\-boundary hallucination subtypes on one generation \(Figure[1](https://arxiv.org/html/2609.22152#Sx1.F1)\)\. The panel is fully crossed\. All7979models answer the same8080\-item anchor set\. Every model comparison therefore becomes a paired test\. Unlike earlier creativity and hallucination benchmarks,Whiteboardapplies an explicit support check to every imagination score\. Our design also pairs each imagination subtype with the hallucination the subtype risks: grounded narrative with fabricated detail, counterfactual extension with impossible physics, creative code with API invention, and analogical mapping with false transfer\. Each pairing follows the task mechanism\. A valid counterfactual preserves causal edges and avoids causal jumps\. Invention inside a closed support sheet follows the constrained creative cognition of[Finke et al\. \(1992\)](https://arxiv.org/html/2609.22152#bib.bib23), and the same constraint suppresses unsupported detail\. Scoring stays deterministic and embedding\-anchored\. The primary path reads a signed atom\-audit matrix instead of calling an LLM judge\. No prior benchmark reads both axes from one generation at subtype resolution\. Our design therefore asks directly whether stronger imagination of one form comes with more or less hallucination in the matched subtype\.
We evaluate 79 contemporary instruction\-tuned LLMs from 18 providers on the shared8080\-item anchor set\. The resulting6,2296\{,\}229valid crossed generations yield the full correlation map\. The map is sparse and one\-sided\. Of189189candidate cells,2323survive the triple false\-discovery contract, and2222of the survivors are negative\. All5555items with sufficient score variation reproduce the negative coupling\. Fifty\-four items reach individual significance, and none turns positive \(medianρ=−0\.63\\rho\{=\}\{\-\}0\.63, sign testp=5\.6×10−17p\{=\}5\.6\{\\times\}10^\{\-17\}\)\. The coupling survives controls for output length and for release date\. We validate the instrument against13,28013\{,\}280blind judgments from two independent annotators on6,6406\{,\}640anchor outputs\. At the output level,Whiteboardtracks imagination and hallucination ratings atρ=0\.72\\rho\{=\}0\.72and0\.690\.69\. At the model level, the same correlations reach0\.560\.56and0\.670\.67\. The annotators lend no differentiated support to the popular “productive hallucination” hypothesis\. Across all ten subtypes, annotators license unsupported content at a nearly uniform rate of18\.718\.7–22\.8%22\.8\\%\. These couplings remain correlational rather than causal\. The couplings are nevertheless one\-sided, reproducible item by item, and anchored in human judgment\. At subtype resolution, our findings answer the question in the title\. Imagination does not derive from hallucination\. Imagination acts as its antidote\.
#### Contributions\.
1. 1\.Whiteboardscores imagination and hallucination on the same model response\.No earlier benchmark reads both axes from one output\. Our support check gates each imagination subtype and pairs it with the hallucination its task invites\.
2. 2\.Scoring without a judge\.Task families feed raw atoms into a signed audit matrix\. Scoring stays deterministic and embedding\-anchored, and no model\-as\-judge call enters the primary path\.
3. 3\.Removing the item confound\.Panel benchmarks assign each model a different item subset, so item difficulty contaminates every score gap\.Whiteboardremoves this confound\. All7979models answer an identical anchor set\. Each model comparison therefore becomes a paired test\.
4. 4\.We find imagination to be the antidote to hallucination, not its byproduct\.The subtype map shows a one\-sided negative correlation\. Every eligible anchor item reproduces the coupling\.
## Related Work
Table 1:Capability comparison with representative baselines\. Joint\(I,H\)\(I,H\)means scoring both axes on the same generation; white\-box means the primary scoring path has no model\-as\-judge call\. The Joint\(I,H\)\(I,H\)column is the precise sense in whichWhiteboardis first\.### Single\-Axis Benchmarks
Existing benchmarks treat creativity and hallucination as separate evaluation targets\. Creativity benchmarks measure how far a response sits from a reference distribution\. The distance takes many forms: common\-answer banks and word norms\([Olson et al\. 2021](https://arxiv.org/html/2609.22152#bib.bib2)\), associative chains grown from seed words\([Gray et al\. 2019](https://arxiv.org/html/2609.22152#bib.bib3)\), TTCT\-style rubrics with several raters per output\([Stevenson et al\. 2022](https://arxiv.org/html/2609.22152#bib.bib5)\), and holistic or idea\-rating panels scored by an LLM judge\([Hou et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib9);[Ruan et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib8)\)\. Other designs constrain the task itself, through object reuse, mandatory story constraints, and denial prompting on code\([Tian et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib6);[Atmakuru et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib31);[Lu et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib30)\)\. Newer panels broaden the formats and keep the divergence\-from\-reference framing\([Al Rabeyah et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib7)\)\. Hallucination benchmarks run the other way\. Coverage grows from binary truthfulness labels\([Lin et al\. 2022](https://arxiv.org/html/2609.22152#bib.bib14);[Li et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib13)\)through atomic\-claim verification\([Min et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib15);[Wei et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib29)\)to span\-level annotation\([Niu et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib25)\), and on to taxonomies of error source and of intrinsic versus extrinsic failure\([Ravichander et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib26);[Bang et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib27)\)\. Every deviation from the reference still counts as a defect\([Huang et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib32)\)\. Neither line asks whether the local world of the prompt licenses the divergent move\. A feasible improvised repair and a fabricated API call therefore score alike by construction\.Whiteboardkeeps the divergence side of the first line and the support\-boundary check of the second, and our support gate lands both readings on the same output\.
### Joint Imagination–Hallucination Measurement and Interpretability
A growing line asks whether the two axes share generative machinery\. Confabulated outputs carry more narrativity than veridical ones, a pattern suggesting partly shared mechanisms\([Sui et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib17)\)\. LLM assistance can raise creativity during assisted tasks and hinder later unassisted performance\([Kumar et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib11)\)\. Creativity metrics can disagree across domains and shift under minor prompt variations\([Lu et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib12)\)\. Other results pull the other way\. Inter\- and intra\-model homogenization complicates distance\-based creativity scores\([Jiang et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib10)\)\. Anti\-hallucination interventions are judged by how far they move creativity scores on code and story tasks\([Lu et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib30)\)\. LLM ideas outrank human ideas on novelty and trail on feasibility\([Si et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib33)\)\. Psychometric work had already framed useful imagination as the move respecting constraints, not the move ignoring them\([Finke et al\. 1992](https://arxiv.org/html/2609.22152#bib.bib23)\)\. Three gaps persist across this line: the same panel rarely receives both scores on one output, the raw signal flow rarely becomes visible, and the productive\-versus\-destructive label rests on stipulation rather than test\([Huang et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib32)\)\. Prior work treats the shared\-mechanism question as a stance\.Whiteboardtreats the same question as a falsifiable proposition and adjudicates it empirically\.
## TheWhiteboardBenchmark
### Design Goals
We fixed four criteria before buildingWhiteboard\. First, one test\-case generation must carry both the imagination score and the hallucination score\. Only such a joint reading, in our view, keeps the boundary between licensed novelty and unsupported invention inside the measurement\. Second, the two axis\-level scores are not enough: each axis must break into mechanism\-grounded subtypes, so a coupling can be localized rather than averaged away\. Third, the primary scoring path must contain no model\-as\-judge call, and every atom\-to\-subtype contribution must be recorded with a sign\. Finally, we must preregister the aggregation weights, gates, and residualization coefficients, and every conclusion must survive partial controls for shared formula atoms and for capability tier\.
### Joint Cross\-Taxonomy
We survey both literatures, extract candidate subtypes, and consolidate them under three constraints\. Every subtype must be computable from a deterministic, auditable signal\. Every subtype must be carried by at least one task family\. Facets serve as second\-level explanatory variables for the audit, never as new ranking weights\. Table[2](https://arxiv.org/html/2609.22152#Sx3.T2)lists the seven imagination and ten hallucination subtypes, with definitions and primary task carriers\. We give the full facet tables in Appendix A of the separately submitted technical appendix\. We settle the destructive\-versus\-productive status of each hallucination subtype empirically, never by definitional fiat\.
Table 2:Taxonomy subtypes\. Left: imagination subtypes \(ℐ\\mathcal\{I\}\), each combining a divergence component with a support gate\. Right: hallucination subtypes \(ℋ\\mathcal\{H\}\), each computed from a distinct support\-boundary check\.
### Task Families and Scoring Pipeline
The pipeline maps a model’s generations into the two subtype vectors\{Im,a\}\\\{I\_\{m,a\}\\\}and\{Hm,b\}\\\{H\_\{m,b\}\\\}\. The primary path stays deterministic, with no model\-as\-judge call\. A subset of the atoms uses theall\-mpnet\-base\-v2Sentence\-BERT encoder as an embedding anchor\([Reimers and Gurevych 2019](https://arxiv.org/html/2609.22152#bib.bib21);[Sentence Transformers 2025](https://arxiv.org/html/2609.22152#bib.bib22)\)\. Every atom\-to\-subtype contribution enters the record as a signed entry, so an auditor can retrace each score\. Figure[1](https://arxiv.org/html/2609.22152#Sx1.F1)traces the full flow from task families through the signed atom matrix to the correlation map\.
#### Task Families\.
The suite defines nine task families, and eight of them carry an imagination subtype\.ClosedWorldFactserves hallucination calibration alone, contributes no imagination signal, and stays out of the anchor set used here\. Each model answers each anchor item once\. Temperatures are fixed per family:0\.850\.85for the creative families,0\.700\.70for MacGyver, and0\.550\.55for Forward Flow\. Every family also carries a task\-specific JSON\-only output contract and a preregistered token cap\. We record parse validity, finish reason, truncation, and schema coverage, so a missing output reduces eligibility instead of disappearing silently\. Each family carries exactly one imagination subtype together with the two to four hallucination subtypes its prompts put at risk, as named in Table[2](https://arxiv.org/html/2609.22152#Sx3.T2)\. DAT, CDAT, and Forward Flow remain auxiliary association probes, and neitherIpureI^\{\\mathrm\{pure\}\}nor the correlation map admits them\. Table[3](https://arxiv.org/html/2609.22152#Sx3.T3)collects the benchmark’s headline statistics\.
Table 3:Benchmark and evaluation statistics at a glance\. Every model is assigned the identical item set, so item difficulty cannot be confounded with model identity\.
#### Atoms and Audit Matrix\.
Each scorer emits raw atoms: rarity against a curated common\-answer bank, affordance support, unavailable\-tool rate, causal\-edge support, forbidden\-update rate, claim\-support precision, citation mismatch rate, hidden\-test pass rate, gold\-mapping coverage, false\-transfer rate, and others\. Lexical atoms draw on the SWOW\-EN2018 association norms\([De Deyne et al\. 2019](https://arxiv.org/html/2609.22152#bib.bib18)\), Word Norms 2\([Buchanan et al\. 2019](https://arxiv.org/html/2609.22152#bib.bib19)\), and WordNet 3\.0\([Miller 1995](https://arxiv.org/html/2609.22152#bib.bib20)\)\. We record the exact snapshots and access terms in Appendix L\. Across the imagination\-bearing families, the pipeline exposes103103of its116116atom entries \(Appendix B\)\. The audit matrixM∈\{−1,0,\+1\}\|𝒜\|×\(\|ℐ\|\+\|ℋ\|\)M\\in\\\{\-1,0,\+1\\\}^\{\|\\mathcal\{A\}\|\\times\(\|\\mathcal\{I\}\|\+\|\\mathcal\{H\}\|\)\}records the role of each atom in each subtype: positive contribution, penalty, or no use\. The scoring code generatesMMautomatically, and the shared\-atom partial control below runs onMM\.
#### Three Views\.
We report every subtype score in three preregistered views\. Each view guards against a different way a single aggregate could mislead\. For modelmmand subtypeaa, letxm,ax\_\{m,a\}collect the raw atoms, and letA\(⋅,w\)A\(\\cdot\\,;w\)denote the fixed aggregation operator with preregistered weightsww:
Im,araw\\displaystyle I^\{\\mathrm\{raw\}\}\_\{m,a\}=A\(xm,a,wraw\),\\displaystyle\{\\displaystyle=\}A\(x\_\{m,a\};w^\{\\mathrm\{raw\}\}\),\(1\)Im,agated\\displaystyle I^\{\\mathrm\{gated\}\}\_\{m,a\}=Im,arawgm,a,\\displaystyle\{\\displaystyle=\}I^\{\\mathrm\{raw\}\}\_\{m,a\}\\,g\_\{m,a\},Im,aresid\\displaystyle I^\{\\mathrm\{resid\}\}\_\{m,a\}=clip\[0,1\]\(Im,agatedCLOSE\\displaystyle\{\\displaystyle=\}\\mathrm\{clip\}\_\{\[0,1\]\}\\\!\\bigl\(I^\{\\mathrm\{gated\}\}\_\{m,a\}OPEN−βfH¯m,fraw\)\.\\displaystyle\{\}\{\\displaystyle\-\}\\beta\_\{f\}\\,\\bar\{H\}^\{\\mathrm\{raw\}\}\_\{m,f\}\\bigr\)\.The raw view aggregates the atoms as they stand\. The gated view multiplies ingm,a∈\[0,1\]g\_\{m,a\}\\in\[0,1\], the carrier family’s own support gate: appropriateness in UUT, constraint compliance in MacGyver and GCW, premise consistency in CJST\. Divergence therefore earns credit only after the licensing checks pass\. The residual view removes mechanical cross\-axis coupling\. HereH¯m,fraw\\bar\{H\}^\{\\mathrm\{raw\}\}\_\{m,f\}is the mean raw score of the destructive hallucination subtypes carried by the same familyff, andβf\\beta\_\{f\}is the family’s preregistered leakage coefficient, writtenβfIH\\beta^\{IH\}\_\{f\}when the direction matters\. We clip the difference back to\[0,1\]\[0,1\]\. The H side follows symmetrically withβfHI\\beta^\{HI\}\_\{f\}\. A conclusion appearing in only one view counts as a leakage symptom, not as a finding, and the partial controls below remove what the views cannot\.
#### Correlation Analysis\.
We build the map cell by cell\. For a viewvvand a pair\(a,b\)\(a,b\), the cell statistic is the Spearman rank correlationρa,bv\\rho^\{v\}\_\{a,b\}between the panel’s score vectors\{Im,av\}m\\\{I^\{v\}\_\{m,a\}\\\}\_\{m\}and\{Hm,bv\}m\\\{H^\{v\}\_\{m,b\}\\\}\_\{m\}\. Each cell carries a95%95\\%bootstrap confidence interval \(B=2,000B\{=\}2\{,\}000resamples, seeded from a stable per\-cell hash\), a two\-sidedpp\-value, and BH\-FDRqq\-values\. We compute theqq\-values within each view and globally over all7×9×3=1897\{\\times\}9\{\\times\}3\{=\}189cells\. The*consistency*subtype stays out of the map, as a zero\-weight diagnostic carried by a single auxiliary probe\. Two partial estimates then ask whether a surviving correlation is an artifact\. The shared\-atom partialρa,ba\\rho^\{\\mathrm\{a\}\}\_\{a,b\}holds the atoms feeding both scores fixed: we rank\-transformIaI\_\{a\},HbH\_\{b\}, and the shared atomsSa,b=\{α:M\[α,Ia\]M\[α,Hb\]≠0\}S\_\{a,b\}\{=\}\\\{\\alpha:M\[\\alpha,I\_\{a\}\]\\,M\[\\alpha,H\_\{b\}\]\\neq 0\\\}, regress both score ranks on the atom ranks, and correlate the residuals\. An emptySa,bS\_\{a,b\}returnsρa,bv\\rho^\{v\}\_\{a,b\}unchanged\. The capability partialρa,bc\\rho^\{\\mathrm\{c\}\}\_\{a,b\}asks whether the correlation is only a shadow of general capability, and conditions the same way on the proxyCmC\_\{m\}\(a frozen external Arena score where available, otherwise the mean of\{Im,a′\}a′≠a\\\{I\_\{m,a^\{\\prime\}\}\\\}\_\{a^\{\\prime\}\\neq a\}\)\. We distinguish the official LMArena sources from the third\-party ones in Appendix L\. Writingqa,bvq^\{\\mathrm\{v\}\}\_\{a,b\},qa,baq^\{\\mathrm\{a\}\}\_\{a,b\}, andqa,bcq^\{\\mathrm\{c\}\}\_\{a,b\}for theqq\-values of the view test and of the two partials, we promote a cell to a main claim only whenmax\(qa,bv,qa,ba,qa,bc\)≤0\.05\\max\\bigl\(q^\{\\mathrm\{v\}\}\_\{a,b\},q^\{\\mathrm\{a\}\}\_\{a,b\},q^\{\\mathrm\{c\}\}\_\{a,b\}\\bigr\)\\leq 0\.05\.
#### Purified Imagination\.
The purified total serves one purpose: no model may buy imagination credit with fabrication\. We call the coupling between an imagination subtypeaaand a hallucination subtypebb*robustly positive*when both partial estimates find the coupling,
ρ~a,b=\{min\(ρa,ba,ρa,bc\),if both are positive andmax\(qa,ba,qa,bc\)≤0\.05,0,otherwise\.\\tilde\{\\rho\}\_\{a,b\}=\\left\\\{\\begin\{aligned\} &\\min\\bigl\(\\rho^\{\\mathrm\{a\}\}\_\{a,b\},\\rho^\{\\mathrm\{c\}\}\_\{a,b\}\\bigr\),\\cr&\\quad\\text\{if both are positive and\}\\cr&\\quad\\max\\bigl\(q^\{\\mathrm\{a\}\}\_\{a,b\},q^\{\\mathrm\{c\}\}\_\{a,b\}\\bigr\)\\leq 0\.05,\\cr&0,\\quad\\text\{otherwise\.\}\\end\{aligned\}\\right\.\(2\)Withwa=1/\|ℐ\|w\_\{a\}\{=\}1/\|\\mathcal\{I\}\|andℋdest\\mathcal\{H\}\_\{\\mathrm\{dest\}\}the destructive set, we define the purification coefficient and the purified total as
πa=1−maxb∈ℋdestρ~a,b,Impure=∑aπawaIm,a∑aπawa,\\pi\_\{a\}=1\-\\max\_\{b\\in\\mathcal\{H\}\_\{\\mathrm\{dest\}\}\}\\tilde\{\\rho\}\_\{a,b\},\\qquad I^\{\\mathrm\{pure\}\}\_\{m\}=\\frac\{\\sum\_\{a\}\\pi\_\{a\}w\_\{a\}I\_\{m,a\}\}\{\\sum\_\{a\}\\pi\_\{a\}w\_\{a\}\},\(3\)both computed on the raw\-view partial estimates\. A subtype never borrowing against destructive hallucination keepsπa=1\\pi\_\{a\}\{=\}1and its full weight\. A subtype whose score rises with some destructive subtype, even after both controls, loses exactly the strength of that coupling\. On the current panel, everyπa\\pi\_\{a\}equals11\. We preregistered the device nonetheless, so any future model inflating creativity through fabricated affordances or facts loses weight automatically\.
## Experiments
We first describe the evaluation panel and the shared anchor design\. We then follow the three questions posed in the Introduction: the joint correlation structure, who trades imagination for reliability, and whether any hallucination is productive\. Human validity comes last\. We report the reliability boundaries of the instrument in Appendix R\. Six audited outputs, subtype profiles by tier, and the forest plot of the decisive cells appear in Appendices J, H, and E\.
### Empirical Setup
The main panel evaluates7979contemporary instruction\-tuned LLMs from1818providers\. Release recency and product position split the panel into2121frontier\-tier and5858comparison\-tier systems\. We list the model families and their release references in Appendix K\. Four providers contribute most of the panel, so we report provider leave\-one\-out robustness for all1818\(Appendix R\)\. Every model answers the same8080anchor items once\. The design therefore defines6,3206\{,\}320model\-item cells\. Of these cells,6,2296\{,\}229pass the task’s parse contract, and9191stay recorded as missing rather than refilled\. Fifty\-eight models carry a scored output for every item\. Generation disables reasoning modes and uses fixed output counts and task\-specific JSON contracts\. We fixed the scorer hyperparameters before the analyses reported here, and only the first annotator’s ratings informed that choice\. The second annotator stays held out\. Runtime scoring never branches on model identity, release date, or provider\. Table[4](https://arxiv.org/html/2609.22152#Sx4.T4)lists the panel, and we give the hyperparameters in Appendix L\.
Table 4:The7979\-model panel by provider and release year\. Tier assignment reads model metadata only and never reads scores\.
### The Anchor Set and Paired Discriminative Power
Question \(1\) stands or falls before any correlation is computed\. The panel must separate model differences from item difficulty, so we begin by measuring what the anchor design buys\. Every model receives the identical anchor set \(Table[3](https://arxiv.org/html/2609.22152#Sx3.T3)\)\. We compare the item\-set overlap of that design against our earlier prompt\-collection design\. On the models with complete coverage, we then test all1,6531\{,\}653model pairs for a difference in imagination, twice over identical scores: once paired across shared items, once unpaired, both at BH\-FDRq≤0\.05q\{\\leq\}0\.05\. Assigned overlap reaches1\.0001\.000in Jaccard terms, and0\.9720\.972once score\-invalid outputs are dropped\. The earlier collection reaches0\.0030\.003, with89%89\\%of items answered by exactly one model\. The paired test resolves5454pairs \(3\.3%3\.3\\%\), and the unpaired test resolves none \(Appendix F\)\. The variance explains the asymmetry: items account for74\.5%74\.5\\%of the per\-item score variance, models for0\.41%0\.41\\%\(Appendix R\)\. Difficulty therefore buries the between\-model signal unless the item stays fixed, and pairing is what removes the difficulty term\. The anchor design is no convenience\. The design is the precondition for the rest of this section, and the couplings below are statements about models rather than about which prompts a model happened to receive\.
### The Subtype Correlation Map
We next read the sign structure of the two axes directly, at the resolution where the mechanism claim lives\. The map holds189189candidate cells over the7979\-model panel \(77imagination subtypes×\\times99hallucination subtypes×\\times33views\)\. A cell earns promotion only after clearing BH\-FDR atq≤0\.05q\{\\leq\}0\.05in its own view and in both partial controls\. Figure[2](https://arxiv.org/html/2609.22152#Sx4.F2)shows the three views\. Table[5](https://arxiv.org/html/2609.22152#Sx4.T5)lists the strongest cells together with the distribution behind them\. Twenty\-three cells clear the contract, and2222of them are negative \(66raw,99gated,77residual\)\. The bulk of the map points the same way:6868–76%76\\%of cells are negative in each view, and the per\-view medianρ\\rhofalls between−0\.10\-0\.10and−0\.14\-0\.14\.
Finding 1\.*Of189189candidate cells,2323survive the triple\-FDR contract and2222of them are negative: doing an imagination subtype well predicts fewer of its partner hallucinations, not more\.*
Figure 2:Subtype correlation map \(7×97\{\\times\}9, three views: raw, gated, residual\)\. Cell text is Spearmanρ\\rhoover7979models; filled stars mark cells passing the triple\-FDR contract \(view\-, atom\-, and capability\-partial BH\-FDR atq≤0\.05q\{\\leq\}0\.05\) with a negative coefficient, and the open star marks the single cell that passes with a positive one\. First, only the starred pairs expose a robust coupling between an imagination subtype and its partner hallucination subtype; every unstarred cell is statistically indistinguishable from zero under the triple correction\. Second, among the couplings that do emerge, the direction is overwhelmingly negative \(cool colors\):2222of the2323starred cells lie below zero\.The decisive cells pair each imagination subtype with the failure its own discipline forbids: analogical mapping against false transfer \(−0\.81\-0\.81residual\), hypothesis generation against citation mismatch \(−0\.69\-0\.69\), grounded narrative against unsupported detail \(−0\.66\-0\.66gated\) and entity drift \(−0\.63\-0\.63\), counterfactual extension against logic violation \(−0\.61\-0\.61\), and creative code against fabricated APIs \(−0\.60\-0\.60\)\. Every one of these moves is a single operation read from two sides: separating source\-only relations from target facts, sourcing a claim, staying inside a fact sheet, respecting causal edges, and importing only what exists\. One cell passes in the positive direction, analogical mapping against unsupported detail in the gated view \(ρ=\+0\.27\\rho\{=\}\{\+\}0\.27,q=0\.042q\{=\}0\.042\)\. This cell rests on no such mechanism and sits at the correction threshold, so we report a threshold artifact rather than evidence for a productive hallucination subtype\. We test the productive\-hallucination hypothesis on human judgments instead\. The answer to question \(1\) is therefore one\-sided\. Wherever a coupling between the two axes can be established at all, the two axes almost always move in opposite directions\.
Table 5:Top: the strongest cells passing the triple\-FDR contract, selected from the2323listed in full in Appendix D; the set\-off row is the only cell that passes in the positive direction\. Bottom: the distribution of all6363cells in each view, which shows that the decisive cells are the tail of an already negative mass rather than isolated extremes\. The atom partial leaves every reportedρ\\rhonearly unchanged, because the v3 schema decouples the cross\-axis atom sharing found in the scorer audit \(Appendix C\); the capability partial reduces some couplings but preserves significance throughout\.
### One Coupling or Fifty\-Five?
A panel\-level correlation over7979models is a single measurement\. Aggregation can produce such a measurement as easily as the behavior the measurement should summarize, so we take the panel apart and look for the coupling again\. Figure[3](https://arxiv.org/html/2609.22152#Sx4.F3)reports5555eligible item correlations\. All are negative,5454reach significance atp<0\.05p\{<\}0\.05, and the median isρ=−0\.63\\rho\{=\}\{\-\}0\.63\. Holding the prompt, the difficulty, and the license fixed rules out item aggregation\. Fifty\-five item\-level couplings therefore agree: a model inventing more inside the license invents less outside it\.
Figure 3:Per\-item replication\. Each bar is the within\-item Spearman correlation between imagination and destructive hallucination across the models that answered that anchor item; blue bars are individually significant atp<0\.05p\{<\}0\.05\. All5555eligible items are negative and none is positive, so the panel\-level coupling reproduces item by item rather than emerging from aggregation\.
### What Else Could Produce the Coupling?
A negative association between two axes read from the same generation invites a third\-variable explanation, so we test the three candidates a reviewer can name from the data alone\. Controls for output length and for release date leave the aggregate coupling nearρ=−0\.414\\rho\{=\}\{\-\}0\.414\. The capability proxy reduces the coupling to−0\.093\-0\.093\. The proxy is itself an imagination\-level measure, and all2222negative subtype cells already pass the same capability control\. Deterministic decoding strengthens the coupling rather than weakening it \(Appendix G\)\. We therefore keep the aggregate value as a summary only, and we rest question \(1\) on the subtype map and its per\-item replication\.
### Who Trades Imagination for Reliability?
Finding 2\.*No imagination subtype borrows credit from destructive hallucination \(all purification coefficients equal11\), but the frontier tier separates from the field on neither axis: the trade\-off narrative has no tier\-level counterpart on this panel\.*
Question \(2\) asks whether the negative coupling shows up as a division between model tiers\. The intuitive reading of the map predicts exactly such a division: stronger models buy reliability with imagination, or the reverse\. We feed Table[5](https://arxiv.org/html/2609.22152#Sx4.T5)into Equation \([3](https://arxiv.org/html/2609.22152#Sx3.E3)\) and compare the tiers under seven score\-blind rules\. All purification coefficients equal11, so the purified total and the raw total coincide\. The tiers separate on neither axis \(Figure[4](https://arxiv.org/html/2609.22152#Sx4.F4)\)\. The stricter flagship result also fails Holm correction\. Finding 1 is therefore a within\-model regularity, and tiering averages the regularity away\.
Figure 4:The panel on the imagination–hallucination plane; better is bottom\-right\. Named points are the flagship families, gray the comparison tier\. The tiers do not separate: the named points are dispersed through the panel rather than displaced downward, and the gradient that remains is the coupling of Finding 1, not a tier effect\.
### Is Any Hallucination Productive?
Finding 3\.*Across7979models and13,28013\{,\}280blind human judgments, no hallucination subtype is distinctively licensed: annotators accept unsupported content at a near\-uniform rate whatever kind of unsupported content it is\.*
The taxonomy leaves one question open: is each hallucination subtype uniformly destructive, or can task semantics license it \(e\.g\. fabricated detail in fiction, source\-target transfer in metaphor\)? Question \(3\) asks exactly that\. We test two things: positive panel coupling under Equation \([2](https://arxiv.org/html/2609.22152#Sx3.E2)\), and blind human labels of licensed content\. No raw\-view cell is robustly positive\. Annotators license about one fifth of unsupported content, and the rate stays between18\.7%18\.7\\%and22\.8%22\.8\\%across subtypes \(Fisherp=0\.22p\{=\}0\.22\)\. Confabulation may carry positive value\([Sui et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib17)\)\. Our data, however, single out no hallucination subtype as distinctively licensed\.
### Validity: Do Humans See the Same Thing?
A deterministic scorer produced everything above, so one question decides the instrument: do the two axes track what people see in the same outputs? Two blind annotators supplied13,28013\{,\}280judgments, and only the second annotator stayed held out from score selection \(Appendix M\)\.Whiteboardtracks the held\-out annotator at0\.710\.71and0\.630\.63per output \(Figure[5](https://arxiv.org/html/2609.22152#Sx4.F5)\)\. The human axes, however, give a coupling of−0\.14\-0\.14against the machine score’s−0\.41\-0\.41\. Global human ratings therefore support validity on both axes, but the same ratings cannot confirm the strength of a subtype\-level coupling\.
Figure 5:Human validity\. Top: every valid scored anchor output \(n=6,443n\{=\}6\{,\}443\), machine score against the mean of the two blind annotators’00–44ratings\. Bottom: the same comparison aggregated to the7979models, with bootstrap intervals\. The footer reports inter\-annotator agreement over all6,6406\{,\}640annotated outputs\. The instrument tracks human ratings on both axes and at both resolutions\.
## Conclusion
In this paper we asked whether imagination in large language models derives from the same generative tendency as hallucination\. Existing evaluations cannot settle the question\. The two axes are scored in separate pipelines on different generations, and a panel where each model answers different items cannot separate model ability from item difficulty\.Whiteboardreads both axes from one generation over a shared anchor set\. Our support check gates each of the seven imagination subtypes and pairs each subtype with the hallucination the subtype risks\. A signed atom matrix computes every score, and no LLM judge enters the path\. Across7979models and6,2296\{,\}229crossed generations,2222of the2323couplings surviving triple false\-discovery control are negative\. Every eligible anchor item reproduces the negative coupling on its own\. Human raters license no hallucination subtype distinctively\. The instrument tracks13,28013\{,\}280blind judgments atρ=0\.72\\rho\{=\}0\.72and0\.690\.69per output\. At subtype resolution, then, imagination does not derive from hallucination\. Imagination stands against it\. We leave the preregistered purified total and the public audit matrix behind as a standing test, ready to demote any future model inflating creativity through fabrication\.
## References
- Al Rabeyahet al\.\(2025\)A\. Al Rabeyah, F\. Góes, M\. Volpe, and T\. MedeirosDo LLMs agree on the creativity evaluation of alternative uses?\.InProceedings of the 16th International Conference on Computational Creativity,pp\. 217–227\.External Links:[Link](https://computationalcreativity.net/iccc25/papers/iccc25-rabeyah2025do.pdf)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Anthropic \(2026\)AnthropicModel system cards\.Note:Anthropic model documentationExternal Links:[Link](https://www.anthropic.com/system-cards)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Atmakuruet al\.\(2024\)A\. Atmakuru, J\. Nainani, R\. S\. R\. Bheemreddy, A\. Lakkaraju, Z\. Yao, H\. Zamani, and H\. ChangCS4: measuring the creativity of large language models automatically by controlling the number of story\-writing constraints\.External Links:2410\.04197,[Document](https://dx.doi.org/10.48550/arXiv.2410.04197),[Link](https://arxiv.org/abs/2410.04197)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Banget al\.\(2025\)Y\. Bang, Z\. Ji, A\. Schelten, A\. Hartshorn, T\. Fowler, C\. Zhang, N\. Cancedda, and P\. FungHalluLens: LLM hallucination benchmark\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 24128–24156\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1176),[Link](https://aclanthology.org/2025.acl-long.1176/)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Baoet al\.\(2025\)F\. S\. Bao, M\. Li, R\. Qu, G\. Luo, E\. Wan, Y\. Tang, W\. Fan, M\. S\. Tamber, S\. Kazi, V\. Sourabh, M\. Qi, R\. Tu, C\. Xu, M\. Gonzales, O\. Mendelevitch, and A\. AhmadFaithBench: a diverse hallucination benchmark for summarization by modern LLMs\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 2: Short Papers\),pp\. 448–461\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-short.38),[Link](https://aclanthology.org/2025.naacl-short.38/)Cited by:[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1)\.
- Beatyet al\.\(2022\)R\. E\. Beaty, D\. R\. Johnson, D\. C\. Zeitlen, and B\. ForthmannSemantic distance and the alternate uses task: recommendations for reliable automated assessment of originality\.Creativity Research Journal34\(3\),pp\. 245–260\.External Links:[Document](https://dx.doi.org/10.1080/10400419.2022.2025720),[Link](https://doi.org/10.1080/10400419.2022.2025720)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1)\.
- BenchLM \(2026\)BenchLMSeed 1\.6 model profile\.Note:BenchLM model page,https://benchlm\.ai/models/seed\-1\-6Third\-party profile; its current page reports no sourced benchmark row and is not evidence of an official LMArena scoreExternal Links:[Link](https://benchlm.ai/models/seed-1-6)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1)\.
- Buchananet al\.\(2019\)E\. M\. Buchanan, K\. D\. Valentine, and N\. P\. MaxwellEnglish semantic feature production norms: an extended database of 4,436 concepts\.Behavior Research Methods51\(4\),pp\. 1849–1863\.External Links:[Document](https://dx.doi.org/10.3758/s13428-019-01243-z),[Link](https://doi.org/10.3758/s13428-019-01243-z)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1),[Atoms and Audit Matrix\.](https://arxiv.org/html/2609.22152#Sx3.SSx3.SSS0.Px2.p1.1)\.
- ByteDance Seed \(2025\)ByteDance SeedSeed1\.5\-thinking: advancing superb reasoning models with reinforcement learning\.External Links:2504\.13914Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- ByteDance Seed \(2026\)ByteDance SeedSeed2\.0\.Note:Official model page and model cardExternal Links:[Link](https://seed.bytedance.com/en/seed2)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Chianget al\.\(2024\)W\. Chiang, L\. Zheng, Y\. Sheng, A\. N\. Angelopoulos, T\. Li, D\. Li, B\. Zhu, H\. Zhang, M\. Jordan, J\. E\. Gonzalez, and I\. StoicaChatbot arena: an open platform for evaluating LLMs by human preference\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 8359–8388\.External Links:[Link](https://proceedings.mlr.press/v235/chiang24b.html)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1)\.
- De Deyneet al\.\(2019\)S\. De Deyne, D\. J\. Navarro, A\. Perfors, M\. Brysbaert, and G\. StormsThe “small world of words” english word association norms for over 12,000 cue words\.Behavior Research Methods51\(3\),pp\. 987–1006\.External Links:[Document](https://dx.doi.org/10.3758/s13428-018-1115-7),[Link](https://doi.org/10.3758/s13428-018-1115-7)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1),[Atoms and Audit Matrix\.](https://arxiv.org/html/2609.22152#Sx3.SSx3.SSS0.Px2.p1.1)\.
- DeepSeek \(2026\)DeepSeekChange log\.Note:DeepSeek API documentationExternal Links:[Link](https://api-docs.deepseek.com/updates/)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Finkeet al\.\(1992\)R\. A\. Finke, S\. M\. Smith, and T\. B\. WardCreative cognition: theory, research, and applications\.The MIT Press\.External Links:ISBN 9780262061506,[Link](https://mitpress.mit.edu/9780262560962/creative-cognition/)Cited by:[Introduction](https://arxiv.org/html/2609.22152#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p4.1),[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Gaoet al\.\(2023\)T\. Gao, H\. Yen, J\. Yu, and D\. ChenEnabling large language models to generate text with citations\.InEMNLP,Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1)\.
- GLM Team \(2025\)GLM TeamGLM\-4\.7, GLM\-4\.6, and GLM\-4\.5\.Note:Official model repositoryExternal Links:[Link](https://github.com/zai-org/GLM-4.5)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Google DeepMind \(2026\)Google DeepMindModel cards\.Note:Google DeepMind model documentationExternal Links:[Link](https://deepmind.google/models/model-cards/)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Grayet al\.\(2019\)K\. Gray, S\. Anderson, E\. E\. Chen, J\. M\. Kelly, M\. S\. Christian, J\. Patrick, L\. Huang, Y\. N\. Kenett, and K\. Lewis“Forward flow”: a new measure to quantify free thought and predict creativity\.American Psychologist74\(5\),pp\. 539–554\.External Links:[Document](https://dx.doi.org/10.1037/amp0000391),[Link](https://doi.org/10.1037/amp0000391)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Guilford \(1967\)J\. P\. GuilfordThe nature of human intelligence\.McGraw\-Hill\.External Links:[Link](https://books.google.com/books?id=Pe6wAAAAIAAJ)Cited by:[Introduction](https://arxiv.org/html/2609.22152#Sx1.p1.1)\.
- Houet al\.\(2026\)Z\. J\. Hou, B\. A\. Zhang, Y\. Lu, B\. K\. Baghel, A\. Brei, X\. Lu, M\. Jiang, F\. Brahman, S\. Chaturvedi, H\. Chang, D\. Khashabi, and X\. L\. LiCreativityPrism: a cross\-domain evaluation framework for large language model creativity\.Transactions on Machine Learning Research\.External Links:[Link](https://arxiv.org/abs/2510.20091)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Huanget al\.\(2025\)L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. LiuA survey on hallucination in large language models: principles, taxonomy, challenges, and open questions\.ACM Transactions on Information Systems43\(2\),pp\. 1–55\.External Links:[Document](https://dx.doi.org/10.1145/3703155),[Link](https://doi.org/10.1145/3703155)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1),[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Jianget al\.\(2025\)L\. Jiang, Y\. Chai, M\. Li, M\. Liu, R\. Fok, N\. Dziri, Y\. Tsvetkov, M\. Sap, and Y\. ChoiArtificial hivemind: the open\-ended homogeneity of language models \(and beyond\)\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2025/hash/754d5a526a5ee5a47220664a0eb92751-Abstract-Datasets_and_Benchmarks_Track.html)Cited by:[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Kumaret al\.\(2025\)H\. Kumar, J\. Vincentius, E\. Jordan, and A\. AndersonHuman creativity in the age of LLMs: randomized experiments on divergent and convergent thinking\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,CHI ’25,pp\. 1–18\.External Links:[Document](https://dx.doi.org/10.1145/3706598.3714198),[Link](https://doi.org/10.1145/3706598.3714198)Cited by:[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Liet al\.\(2023\)J\. Li, X\. Cheng, X\. Zhao, J\. Nie, and J\. WenHaluEval: a large\-scale hallucination evaluation benchmark for large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 6449–6464\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.397),[Link](https://aclanthology.org/2023.emnlp-main.397/)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3214–3252\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.229),[Link](https://aclanthology.org/2022.acl-long.229/)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- LM Arena \(2026\)LM ArenaArena leaderboard dataset\.Note:Hugging Face Datasets,https://huggingface\.co/datasets/lmarena\-ai/leaderboard\-datasetText\-style\-control snapshot published 2026\-04\-27; revision 6df7986d4be040bf3a736238cc437c17fb259671; CC BY 4\.0; accessed 2026\-04\-29External Links:[Link](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1)\.
- Luet al\.\(2026\)L\. Lu, M\. Liu, P\. C\. Lu, Y\. Tian, S\. Sun, and N\. PengRethinking creativity evaluation: a critical analysis of existing creativity evaluations\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6329–6352\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.297),[Link](https://aclanthology.org/2026.eacl-long.297/)Cited by:[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Luet al\.\(2025\)Y\. Lu, D\. Wang, T\. Li, D\. Jiang, S\. Khudanpur, M\. Jiang, and D\. KhashabiBenchmarking language model creativity: a case study on code generation\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 2776–2794\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.141),[Link](https://aclanthology.org/2025.naacl-long.141/)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1),[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Meta \(2024\)MetaLlama 3\.3 model card\.Note:Official model cardExternal Links:[Link](https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Miller \(1995\)G\. A\. MillerWordNet: a lexical database for english\.Communications of the ACM38\(11\),pp\. 39–41\.External Links:[Document](https://dx.doi.org/10.1145/219717.219748),[Link](https://doi.org/10.1145/219717.219748)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1),[Atoms and Audit Matrix\.](https://arxiv.org/html/2609.22152#Sx3.SSx3.SSS0.Px2.p1.1)\.
- Minet al\.\(2023\)S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. HajishirziFActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 12076–12100\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.741),[Link](https://aclanthology.org/2023.emnlp-main.741/)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- MiniMax \(2026a\)MiniMaxMiniMax M2\.5: built for real\-world productivity\.Note:Official release blogExternal Links:[Link](https://www.minimax.io/news/minimax-m25)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- MiniMax \(2026b\)MiniMaxMiniMax M3: frontier coding, 1m context, native multimodality—all in one model\.Note:Official release blogExternal Links:[Link](https://www.minimax.io/blog/minimax-m3)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Moonshot AI \(2026a\)Moonshot AIKimi K2\.5\.Note:Official model repository and technical reportExternal Links:[Link](https://github.com/MoonshotAI/Kimi-K2.5)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Moonshot AI \(2026b\)Moonshot AIKimi K2\.6: from code to creation, from one to many\.Note:Official model pageExternal Links:[Link](https://www.kimi.com/ai-models/kimi-k2-6)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Niuet al\.\(2024\)C\. Niu, Y\. Wu, J\. Zhu, S\. Xu, K\. Shum, R\. Zhong, J\. Song, and T\. ZhangRAGTruth: a hallucination corpus for developing trustworthy retrieval\-augmented language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 10862–10878\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.585),[Link](https://aclanthology.org/2024.acl-long.585/)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Olsonet al\.\(2021\)J\. A\. Olson, J\. Nahas, D\. Chmoulevitch, S\. J\. Cropper, and M\. E\. WebbNaming unrelated words predicts creativity\.Proceedings of the National Academy of Sciences118\(25\),pp\. e2022340118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2022340118),[Link](https://doi.org/10.1073/pnas.2022340118)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- OpenAI \(2026\)OpenAIAll models\.Note:OpenAI API documentationExternal Links:[Link](https://developers.openai.com/api/docs/models/all)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Qwen Team \(2026a\)Qwen TeamQwen3\.5: towards native multimodal agents\.Note:Official release blogExternal Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Qwen Team \(2026b\)Qwen TeamQwen3\.6\-35B\-A3B: agentic coding power, now open to all\.Note:Official release blogExternal Links:[Link](https://qwen.ai/blog?id=qwen3.6-35b-a3b)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Ravichanderet al\.\(2025\)A\. Ravichander, S\. Ghela, D\. Wadden, and Y\. ChoiHALoGEN: fantastic LLM hallucinations and where to find them\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1402–1425\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.71),[Link](https://aclanthology.org/2025.acl-long.71/)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 3982–3992\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1410),[Link](https://aclanthology.org/D19-1410/)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1),[Task Families and Scoring Pipeline](https://arxiv.org/html/2609.22152#Sx3.SSx3.p1.1)\.
- Ruanet al\.\(2026\)K\. Ruan, X\. Wang, J\. Hong, P\. Wang, Y\. Liu, and H\. SunEvaluating LLMs’ divergent thinking capabilities for scientific idea generation with minimal context\.Nature Communications17\(1\),pp\. 3625\.External Links:[Document](https://dx.doi.org/10.1038/s41467-026-70245-1),[Link](https://doi.org/10.1038/s41467-026-70245-1)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Sentence Transformers \(2025\)Sentence Transformersall\-mpnet\-base\-v2 model card\.Note:Hugging Face model repositoryRevision e8c3b32edf5434bc2275fc9bab85f82640a19130; Apache License 2\.0; accessed 2026\-04\-28External Links:[Link](https://huggingface.co/sentence-transformers/all-mpnet-base-v2/tree/e8c3b32edf5434bc2275fc9bab85f82640a19130)Cited by:[Appendix L](https://arxiv.org/html/2609.22152#A12.SS0.SSS0.Px1.p1.1),[Task Families and Scoring Pipeline](https://arxiv.org/html/2609.22152#Sx3.SSx3.p1.1)\.
- Siet al\.\(2025\)C\. Si, D\. Yang, and T\. HashimotoCan LLMs generate novel research ideas? a large\-scale human study with 100\+ NLP researchers\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=M23dTGWCZy)Cited by:[Introduction](https://arxiv.org/html/2609.22152#Sx1.p1.1),[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1)\.
- Stevensonet al\.\(2022\)C\. Stevenson, I\. Smal, M\. Baas, R\. Grasman, and H\. van der MaasPutting GPT\-3’s creativity to the \(alternative uses\) test\.InProceedings of the 13th International Conference on Computational Creativity,pp\. 164–168\.External Links:[Link](https://computationalcreativity.net/iccc22/wp-content/uploads/2022/06/ICCC-2022_25S_Stevenson-et-al..pdf)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Suiet al\.\(2024\)P\. Sui, E\. Duede, S\. Wu, and R\. SoConfabulation: the surprising value of large language model hallucinations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14274–14284\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.770),[Link](https://aclanthology.org/2024.acl-long.770/)Cited by:[Introduction](https://arxiv.org/html/2609.22152#Sx1.p1.1),[Joint Imagination–Hallucination Measurement and Interpretability](https://arxiv.org/html/2609.22152#Sx2.SSx2.p1.1),[Is Any Hallucination Productive?](https://arxiv.org/html/2609.22152#Sx4.SSx7.p2.1)\.
- Tencent Hy Team \(2026\)Tencent Hy TeamHy3\.Note:Official model repositoryExternal Links:[Link](https://github.com/Tencent-Hunyuan/Hy3)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Tianet al\.\(2024\)Y\. Tian, A\. Ravichander, L\. Qin, R\. Le Bras, R\. Marjieh, N\. Peng, Y\. Choi, T\. Griffiths, and F\. BrahmanMacGyver: are large language models creative problem solvers?\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5303–5324\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.297),[Link](https://aclanthology.org/2024.naacl-long.297/)Cited by:[§Q\.1](https://arxiv.org/html/2609.22152#A17.SS1.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p2.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- Weiet al\.\(2024\)J\. Wei, C\. Yang, X\. Song, Y\. Lu, N\. Hu, J\. Huang, D\. Tran, D\. Peng, R\. Liu, D\. Huang, C\. Du, and Q\. V\. LeLong\-form factuality in large language models\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 80756–80827\.External Links:[Document](https://dx.doi.org/10.52202/079017-2567),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/937ae0e83eb08d2cb8627fe1def8c751-Abstract-Conference.html)Cited by:[§Q\.2](https://arxiv.org/html/2609.22152#A17.SS2.p1.1),[Introduction](https://arxiv.org/html/2609.22152#Sx1.p3.1),[Single\-Axis Benchmarks](https://arxiv.org/html/2609.22152#Sx2.SSx1.p1.1)\.
- xAI \(2026\)xAIModels\.Note:xAI developer documentationExternal Links:[Link](https://docs.x.ai/developers/models)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Xiaomi MiMo \(2026\)Xiaomi MiMoMiMo\-V2\.5 series release notice\.Note:Official API platformExternal Links:[Link](https://platform.xiaomimimo.com/contact)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
- Z\.ai \(2026\)Z\.aiGLM\-5: from vibe coding to agentic engineering\.Note:Official release blogExternal Links:[Link](https://z.ai/blog/glm-5)Cited by:[Appendix K](https://arxiv.org/html/2609.22152#A11.p1.1)\.
## Appendix AFull Taxonomy and Facets
We will reproduce the full facet tables in the camera\-ready, and the released taxonomy file already carries them\. The tables cover all seven imagination and ten hallucination subtypes, each with its white\-box\-signal list\. Each subtype has between three and five facets, and each facet has at least one white\-box\-signal entry\.
## Appendix BAtom Audit Matrix
Figure B1:Signed atom\-to\-subtype dependency matrix\. Green: positive contributions; red: penalty signals; white: unused\. The block structure is the substrate for the shared\-atom partial Spearman control\.The audit matrixM∈\{−1,0,\+1\}116×17M\\in\\\{\-1,0,\+1\\\}^\{116\\times 17\}records the signed contribution of each raw atom entry to each of the1717subtype scores\. Of the116116atoms,103103belong to the eight imagination\-bearing families\. The remaining1313belong to the hallucination\-calibration family and to the zero\-weight consistency probe\. Figure[B1](https://arxiv.org/html/2609.22152#A2.F1)visualizes the sparsity pattern and the separation between I\-side support atoms and H\-side penalty atoms\. We release the per\-cell shared atom setsSa,bS\_\{a,b\}for the correlation table alongside the matrix\.
## Appendix CCross\-Axis Scoring Design
The scoring design keeps imagination\-side support signals separate from hallucination\-side penalty signals whenever a task family draws on related evidence\. In PropConj, the grounding signal forIassocI\_\{\\mathrm\{assoc\}\}stays distinct from the context\-support check forHctxH\_\{\\mathrm\{ctx\}\}\. GCW likewise separates narrative support from unsupported\-detail evidence, and CJST separates counterfactual compliance from logic violations\. For the cells reported in Table 5 of the main paper, these definitions leaveSa,b=∅S\_\{a,b\}\{=\}\\varnothing\. The atom\-partial correlations therefore come out nearly identical to the unadjusted ones\. The partial analysis consequently tests formula overlap in the results we report\.
## Appendix DDecisive Cells: Full Table
Table D1:All2323triple\-FDR\-significant cells in the main panel; column notation follows Table 5 of the main paper\.The main paper shows only the strongest decisive cells\. Table[D1](https://arxiv.org/html/2609.22152#A4.T1)lists every cell clearing BH\-FDR atq≤0\.05q\{\\leq\}0\.05in its own view and in both partial controls, grouped by view\. Each row carries the raw Spearman correlation, its atom\- and capability\-partial versions, the threeqqvalues, and the bootstrap interval\. Twenty\-three cells qualify: six in the raw view, ten in the gated view, and seven in the residual view\. Twenty\-two of the twenty\-three are negative\. The negative cells recur across views rather than appearing once\. Five distinct pairs qualify in all three views \(narrative×\\timesdetail, narrative×\\timesdrift, counterfactual×\\timeslogic, code×\\timesfact, hypothesis×\\timesboundary\), and three more qualify in two\. The one positive cell qualifies in a single view, at the largest correctedqqin the table\.
## Appendix EForest Plot of Decisive Cells
Figure E1:Forest plot of the2323decisive cells\. Raw, atom\-partial, and capability\-partial estimates with95%95\\%bootstrap CIs\.Figure[E1](https://arxiv.org/html/2609.22152#A5.F1)puts effect sizes and uncertainty next to the counts above\. The figure draws each decisive cell three times, once per estimator \(raw, atom\-partial, capability\-partial\), with95%95\\%bootstrap intervals\. Twenty\-two intervals lie entirely below zero\. One interval, analogy×\\timesdetail in the gated view, lies entirely above zero\. For almost every cell the three estimators fall within a few hundredths of each other, so the couplings do not depend on which control we apply\. The widest intervals belong to the smallest effects, not to the strongest ones\.
## Appendix FPaired Discriminative Power and Reliability
Figure F1:Model pairs distinguishable after BH\-FDR under paired and unpaired analyses of the same scores\.Figure F2:Projected reliability as the number of anchor items changes\. Panel \(a\) reports imagination and panel \(b\) hallucination; solid green lines show relative reliabilityGG, and dashed black lines show absolute reliabilityΦ\\Phi\.We first ask how much detection power the shared items buy\. On the5858models with complete coverage, we tested all1,6531\{,\}653model pairs twice over the same scores: a Wilcoxon signed\-rank test over shared items, and a Mann–Whitney test without pairing\. Both runs then took BH\-FDR correction\. The paired analysis separates5454pairs, and the unpaired analysis separates none \(Figure[F1](https://arxiv.org/html/2609.22152#A6.F1)\)\. Holding the item fixed removes prompt difficulty from the model contrast, and the removal accounts for the gain\.
How many items a comparison needs is a separate question\. For the same5858complete models, we decomposed model, item, and interaction variance, then projected relative reliabilityGGand absolute reliabilityΦ\\Phifrom1010to160160items\. At8080items, imagination reachesG=0\.567G\{=\}0\.567andΦ=0\.248\\Phi\{=\}0\.248, and hallucination reachesG=0\.765G\{=\}0\.765andΦ=0\.663\\Phi\{=\}0\.663\(Figure[F2](https://arxiv.org/html/2609.22152#A6.F2)\)\. Item difficulty carries the larger share of variance on the imagination side, and the lower absolute reliability follows\. Panel\-level relative comparisons therefore rest on firmer ground than absolute score transfer to another prompt set\.
## Appendix GSensitivity Analysis
Figure G1:Provider leave\-one\-out retention of decisive cells, by analysis view\.Figure G2:Sensitivity to decoding temperature on4343matched models\. Panel \(a\) compares the aggregate imagination–hallucination correlation, and panel \(b\) reports rank agreement between the two temperature settings on each axis\.We ran two robustness checks\. First, we deleted each of the1818providers in turn and recomputed the map\. Second, we applied the same fixed scoring rule to4343models generated at temperatures0\.850\.85and00, keeping the same7373–8080valid items per model across settings\. Decisive\-cell retention under provider deletion is1717–100%100\\%in the raw view,7070–100%100\\%in the gated view, and8686–100%100\\%in the residual view\. At temperature00, the aggregate correlation moves from−0\.332\-0\.332to−0\.722\-0\.722, with cross\-temperature rank correlations of0\.5850\.585for imagination and0\.8770\.877for hallucination \(Figures[G1](https://arxiv.org/html/2609.22152#A7.F1)and[G2](https://arxiv.org/html/2609.22152#A7.F2)\)\. The raw view reacts to removal of a large provider\. The gated and residual views stay more stable, and deterministic decoding does not remove the negative association\. The gated and residual findings therefore hold across provider composition and both decoding settings, and the raw view carries no claim on its own\.
## Appendix HSubtype Profiles by Tier
Figure H1:Subtype profiles by model tier \(frontier vs comparison\)\.The main paper reports a null tier difference on both axes, and this section shows the profiles behind the null\. Figure[H1](https://arxiv.org/html/2609.22152#A8.F1)plots the mean score of each imagination and destructive\-hallucination subtype for the2121frontier and the5858comparison models\. The two profiles overlap on both axes, and no subtype puts one tier consistently outside the other\. The tiers differ in product position and release recency, not in the subtype composition of either axis\. The figure serves orientation only\. Every inferential claim rests on the triple\-FDR map \(Figure 2 of the main paper\) and on the tests in section “Who Trades Imagination for Reliability?”\.
## Appendix ISensitivity Across Scoring Views
Some cells read differently under the raw, gated, and residual scoring views\. The released tables report all three estimates and flag every cell whose largest difference reaches\|Δρ\|≥0\.30\|\\Delta\\rho\|\{\\geq\}0\.30\. A flag alone removes nothing, because retention still requires the view test and both partial controls to pass the same false\-discovery threshold\. A large cross\-view difference instead signals that a support gate or a correction term contributes to the observed association\. Reporting all three estimates keeps a single\-view result from passing as view\-invariant\.
## Appendix JCase Study: Six Audited Outputs
Six outputs make the joint reading concrete: for each of three task families we pair one output whose invention stays inside the support boundary with one whose invention crosses it\. The examples are qualitative illustrations from the benchmark corpus and carry no statistical claim\. The three families carry different imagination subtypes; narrative×\\timesdetail and cf×\\timeslogic are decisive cells of the correlation map, while the code pair is illustrative because code×\\timesintent does not survive the triple\-FDR criterion\. Table[J1](https://arxiv.org/html/2609.22152#A10.T1)states what each prompt provides and licenses; Tables[J2](https://arxiv.org/html/2609.22152#A10.T2),[J3](https://arxiv.org/html/2609.22152#A10.T3), and[J4](https://arxiv.org/html/2609.22152#A10.T4)present the three pairs\.
Table J1:Task settings for the three case\-study families: what each prompt fixes, what it licenses, and which hallucination subtypes read the same output\.#### Grounded Fiction \(GCW\)\.
On the card*The Broken Compass*, one output turns the floor itself into the way out: Priya reads the raised cracks in the marble with her touch\-based left–right memory and uses them as an exit map\. The turn appears nowhere in the fact sheet, yet every element it uses does, andWhiteboardscored it at the top of the narrative scale with no unsupported\-detail evidence\. On*The Fog Market*, a second output reaches the goal only by asserting what the sheet never supports: Otto “recognizes the weight of the coins as those used at Sima’s spice stall,” and Sima “confirms the pouch belongs to a regular customer who always buys cumin\.” Coin weight identifies denominations, not a stall’s customers, and no fact grants Sima knowledge of the owner;Whiteboardflags both claims \(HdetH\_\{\\mathrm\{det\}\},HfactH\_\{\\mathrm\{fact\}\}\)\.
Table J2:The two GCW outputs of the case study at a glance\. The same license to invent splits into grounded invention and fabricated shortcut exactly at the support boundary\.
#### Counterfactual Extension \(CJST\)\.
Both outputs were rated at the top of the counterfactual scale; only one keeps the premise closed\. Under the premise “every cup remembers the last drink poured into it,” the first output derives consequences that never leave the premise: people stop sharing cups, baristas check a cup’s memory before refilling it, parents discard cups after allergenic drinks; each consequence carries a causal chain anchored in the premise, andWhiteboardrecords no impossible\-physics evidence\. Under the premise “footsteps leave faint glowing marks for one minute,” the second output starts to engineer the miracle: flooring products that “minimize or enhance glow visibility” and security systems with “one\-minute glow detection algorithms\.” The premise licenses the glow; nothing licenses machines that read or amplify it\. The prompt names this move explicitly, no extra magic beyond the premise, andWhiteboardscored it at the top of the logic\-violation scale \(HlogicH\_\{\\mathrm\{logic\}\}\)\.
Table J3:The two CJST outputs\. Both are maximally imaginative on the imagination axis; they differ only in whether the invention stays inside the counterfactual premise\.
#### Creative Code \(NeoCoder\)\.
The third pair shares one problem, counting 4\-connected islands in a 0/1 grid, under the same JSON\-only contract: return exactly one object with the required keys, no markdown, no text before or after it\. The first output delivers a recursive depth\-first search behind the required entry\-point signature, respects the empty\-imports list, and parses cleanly;Whiteboardrecords no hallucination evidence on any subtype\. The second opens with reasoning text before the JSON object, so the object never parses and the code inside never reaches the hidden tests\.Whiteboardstill rated its algorithmic content at the top of the code scale, but the delivery is exactly whatHintH\_\{\\mathrm\{int\}\}records: a contract miss that no amount of algorithmic creativity repairs\.
Table J4:The two NeoCoder outputs on the same count\-islands task\. The pair isolates the delivery contract: identical problem, identical license, oppositeHintH\_\{\\mathrm\{int\}\}readings\.A distance\-only creativity score cannot tell the outputs in any of these pairs apart: within each pair both leave the obvious answer, and in every pair both outputs were rated at the top of the imagination scale by the scorer\. The joint reading separates them, and for two of the three families the separation it makes on single outputs is the one the panel\-scale couplings recover across7979models: narrative×\\timesdetail atρ=−0\.66\\rho\{=\}\{\-\}0\.66and cf×\\timeslogic at−0\.61\-0\.61\(Table 5 of the main paper\)\.
## Appendix KPanel Manifest
The full panel manifest \(model identifiers, providers, families, release dates, tier assignment, and per\-item coverage\) is released with the artifact\. Provider distribution of the7979\-model panel: OpenAI \(1818\), xAI \(99\), Google \(99\), Alibaba \(99\), DeepSeek \(55\), Anthropic \(55\), MiniMax \(44\), Zhipu \(44\), Moonshot \(33\), Amazon \(33\), ByteDance \(22\), Mistral \(22\), and one model each from StepFun, Meta, Tencent, Microsoft, Upstage, and Xiaomi\. The frontier tier covers the GPT\-4\.1 and GPT\-5 families\([OpenAI 2026](https://arxiv.org/html/2609.22152#bib.bib34)\), Claude Opus 4\.6/4\.7 and Sonnet 4/4\.5/4\.6\([Anthropic 2026](https://arxiv.org/html/2609.22152#bib.bib35)\), Gemini 3\.1 Pro\([Google DeepMind 2026](https://arxiv.org/html/2609.22152#bib.bib36)\), and Grok 4\.x\([xAI 2026](https://arxiv.org/html/2609.22152#bib.bib37)\); the comparison tier adds DeepSeek\([DeepSeek 2026](https://arxiv.org/html/2609.22152#bib.bib38)\), Qwen\([Qwen Team 2026a](https://arxiv.org/html/2609.22152#bib.bib39);[Qwen Team 2026b](https://arxiv.org/html/2609.22152#bib.bib40)\), Kimi\([Moonshot AI 2026a](https://arxiv.org/html/2609.22152#bib.bib41);[Moonshot AI 2026b](https://arxiv.org/html/2609.22152#bib.bib42)\), MiMo\([Xiaomi MiMo 2026](https://arxiv.org/html/2609.22152#bib.bib43)\), Gemma\([Google DeepMind 2026](https://arxiv.org/html/2609.22152#bib.bib36)\), GLM\([GLM Team 2025](https://arxiv.org/html/2609.22152#bib.bib44);[Z\.ai 2026](https://arxiv.org/html/2609.22152#bib.bib45)\), Hunyuan\([Tencent Hy Team 2026](https://arxiv.org/html/2609.22152#bib.bib46)\), Llama\([Meta 2024](https://arxiv.org/html/2609.22152#bib.bib47)\), MiniMax\([MiniMax 2026a](https://arxiv.org/html/2609.22152#bib.bib48);[MiniMax 2026b](https://arxiv.org/html/2609.22152#bib.bib49)\), and earlier GPT, Claude, Gemini, and ByteDance Seed variants\([ByteDance Seed 2025](https://arxiv.org/html/2609.22152#bib.bib51);[ByteDance Seed 2026](https://arxiv.org/html/2609.22152#bib.bib50)\)\. Tier assignment reads model metadata only and never reads scores:2121models are frontier and5858comparison\. The panel is fully crossed by construction; the9191model\-item cells that returned no scorable output \(7979with no generation,1212failing the task’s parse contract\) are recorded as missing rather than refilled, which is why realized item overlap is0\.980\.98and not1\.001\.00, and why the complete\-case analyses \(paired power, variance components\) run on the5858models that carry a scored output for every one of the8080items\.
## Appendix LScoring Configuration Reproducibility
Every downstream analysis should run off one scoring configuration\. We compared the configuration identifier attached to each of the6,4436\{,\}443matched annotation rows and rebuilt model aggregates from their per\-task contributions\. Every matched row uses the same configuration, and the maximum reconstruction errors are1\.03×10−151\.03\{\\times\}10^\{\-15\}for imagination and4\.19×10−154\.19\{\\times\}10^\{\-15\}for hallucination\. Model identity, tier, provider, and release date are not inputs to runtime scoring, so the reconstruction does not depend on model\-specific settings\.
#### Resource Versions and Access\.
Embedding atoms use theall\-mpnet\-base\-v2encoder at revisione8c3b32\([Reimers and Gurevych 2019](https://arxiv.org/html/2609.22152#bib.bib21);[Sentence Transformers 2025](https://arxiv.org/html/2609.22152#bib.bib22)\)under the model repository’s Apache 2\.0 license\. Lexical resources are SWOW\-EN2018 R123\([De Deyne et al\. 2019](https://arxiv.org/html/2609.22152#bib.bib18)\), Word Norms 2\([Buchanan et al\. 2019](https://arxiv.org/html/2609.22152#bib.bib19)\), and WordNet 3\.0\([Miller 1995](https://arxiv.org/html/2609.22152#bib.bib20)\)\. The raw SWOW file is obtained from the official project under its research\-use terms and is not redistributed\. Word Norms 2 data and its GPL\-3\.0 reference code are public, and the WordNet 3\.0 license permits use, copying, modification, and distribution with its notices retained\. The official Arena reference is the LM Arena method and leaderboard dataset\([Chiang et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib24);[LM Arena 2026](https://arxiv.org/html/2609.22152#bib.bib52)\), using the 2026\-04\-27 snapshot\. The frozen configuration records the final\-panel Seed 1\.6 anchor as a value from the third\-party BenchLM model page\([BenchLM 2026](https://arxiv.org/html/2609.22152#bib.bib53)\), and not an official LMArena row\. The current page exposes no sourced score, so the citation documents that limitation; it does not validate the value\. The common\-answer banks are author\-built: the static bank and curated overlay are included with the artifact; mined extensions require regeneration from the documented settings and access to the listed model APIs\.
## Appendix MHeld\-Out Annotator Agreement
Figure M1:Human\-rating agreement\. Panel \(a\) compares the residual scores with the held\-out annotator and the two\-rater mean at output and model levels; the raw model\-level points provide an uncalibrated reference\. Filled dark circles mark the held\-out comparison, and light squares mark the two\-rater mean\. Error bars are95%95\\%bootstrap intervals for the held\-out comparisons\. Panel \(b\) reports inter\-annotator agreement, and panel \(c\) shows the licensed share of unsupported content by subtype\.Two annotators independently rated6,6406\{,\}640outputs on00–44imagination and unsupported\-content scales\. The ratings of annotator C selected the subtype score boundaries; the ratings of annotator D were never read by any fitting step\. Machine correlations with the held\-out annotator are0\.7100\.710and0\.6320\.632at output level \(n=6,443n\{=\}6\{,\}443\) and0\.5460\.546and0\.6340\.634at model level \(n=79n\{=\}79\)\. Against the two raters’ mean, the corresponding values are0\.7160\.716and0\.6940\.694, and0\.5600\.560and0\.6700\.670\. Reversing the split and selecting boundaries on annotator D gives model\-level correlations of0\.5110\.511and0\.6440\.644on annotator C, against0\.5460\.546and0\.6340\.634in the deployed direction\. The two boundary sets do not match cell by cell, but their model\-score rankings agree atρ=0\.977\\rho\{=\}0\.977and0\.9960\.996for imagination and hallucination\. The raw view, which the boundary and residual calibration does not touch, agrees with the held\-out annotator atρ=0\.50\\rho\{=\}0\.50and0\.570\.57at model level \(Figure[M1](https://arxiv.org/html/2609.22152#A13.F1)a\)\. Inter\-annotator correlations are0\.7120\.712for imagination and0\.7590\.759for unsupported content, with agreement within one point at97\.70%97\.70\\%and91\.07%91\.07\\%\. Licensing rates span18\.74%18\.74\\%–22\.85%22\.85\\%across subtypes and false transfer does not differ from the others \(Fisherp=\.218p\{=\}\.218\)\. The hold\-out is at the rater level\. The held\-out annotator worked from the same rubric on the same outputs, so this design controls for rater\-specific fitting, not for circularity introduced by the rubric itself\.
## Appendix NPrompt Examples
Figure N1:Human–LLM examples for UUT, PropConj, MacGyver, and CJST\. Each row places evaluated user\-prompt passages on the left and selected field values from the corresponding saved response on the right\. Bracketed ellipses mark omitted prompt text, and task IDs identify the full saved records\. Colors distinguish benchmark components and do not encode scores\.Figure N2:Human–LLM examples for HypoUseSpace, GCW, NeoCoder, and AnalogyTransfer\. The left column reproduces passages from the evaluated user prompts; the right column reports selected field values from the saved responses\. Bracketed ellipses mark omitted prompt text\. Task IDs identify the full saved records, and colors carry no score information\.The crossed panel used8080anchor prompts from eight scored components\. We sent the same8080task IDs to all7979models\. Each request contained a task\-specific system message and one user message, and asked for a single model response\. Figures[N1](https://arxiv.org/html/2609.22152#A14.F1)and[N2](https://arxiv.org/html/2609.22152#A14.F2)pair prompt passages with selected fields from the saved responses for all eight components\.
### N\.1Request Format
The system message defined the benchmark role and output format\. The user message supplied the task instance and its support rule, followed by the required JSON schema and output count\. Depending on the task, the instance could contain a tool inventory, a fact sheet, an evidence table, or a function signature\. Responses had to contain JSON without surrounding prose\. We did not add demonstrations or follow\-up turns for particular models\.
The support rule was therefore part of the input seen by the model\. Its form varied by task\. Some prompts limited proposals to physical properties or listed tools, while others licensed a single counterfactual premise, a fictional world, or an analogy\. The prompt also stated which additions fell outside that license\.
### N\.2Task\-Specific Instructions
#### UUT and PropConj\.
UUT requested eight unusual but physically implementable uses of an ordinary object and ruled out invented capabilities\. PropConj requested six real objects that satisfied every named property\. An uncommon answer was acceptable only when it remained possible and met the full conjunction\.
#### MacGyver and HypoUseSpace\.
MacGyver supplied a goal, a tool inventory, and physical or safety constraints\. Plans could combine listed tools but could not introduce new equipment\. HypoUseSpace instead supplied entities, relations, predicates, and evidence IDs\. Its answers had to cite the records used to support the proposed hypothesis\.
#### CJST and GCW\.
CJST introduced one impossible premise and requested immediate, adaptive, and second\-order consequences while keeping ordinary constraints in force\. GCW paired a story card with a fact sheet and a list of forbidden claims\. Narrative detail was open, but the stated world facts and exclusions remained binding\.
#### NeoCoder and AnalogyTransfer\.
NeoCoder fixed the entry point, required behavior, permitted imports, and prohibited techniques for an executable coding task\. AnalogyTransfer named a source system and a target domain, then specified which source relations could transfer and which source details could not be treated as literal facts about the target\.
### N\.3Source of the Examples
The cards draw from the prompt manifest and task results saved with one panel report\. The build script checks each displayed prompt passage against the full prompt and each response field value against the raw model output; it stops if either check fails\. A bracketed ellipsis marks text omitted between prompt passages\. Each task ID locates the complete prompt and response\. Colors distinguish components and do not encode scores\.
## Appendix OLimitations
#### Correlation, Not Causation\.
Triple controls and per\-item replication identify co\-variation, not directional generative mechanism\. A causal account would require interventional fine\-tuning that selectively reduces one subtype, which we leave to follow\-up work\. The same controls reduce but cannot fully eliminate formula\-driven correlation, since some atoms straddle the I/H boundary; cells whose estimate shifts across views are flagged in Appendix[I](https://arxiv.org/html/2609.22152#A9)\.
#### The Machine–Human Coupling Gap\.
The two axes track human ratings individually, but the human ratings do not independently reproduce the negative coupling between them \(−0\.14\-0\.14,95%95\\%CI\[−0\.37,\+0\.11\]\[\{\-\}0\.37,\{\+\}0\.11\], against−0\.41\-0\.41on the machine side, both reported in the main paper\)\. A single00–44rating per axis may simply lack the resolution of sixteen subtype channels, and the human interval does not exclude a moderate negative coupling, but on present evidence the coupling is a property ofWhiteboard’s scores that human raters have not confirmed at strength\.
#### Measurement Resolution\.
Items carry74\.5%74\.5\\%of the score variance and models0\.41%0\.41\\%, givingG=0\.567G\{=\}0\.567andΦ=0\.248\\Phi\{=\}0\.248at8080items and leaving only3\.3%3\.3\\%of model pairs separable\. Absolute scores are therefore not comparable across item sets, and individual model rankings should not be read off these numbers\. Five of the sixteen subtype channels also place more than half their outputs at the floor, which compresses the correlations those channels can express\.
#### Scope and Resource Dependence\.
The map is only as general as the panel; generalization beyond20262026\-vintage instruction\-tuned LLMs is unclaimed\.Whiteboardis deterministic and auditable but not resource\-free: it depends on SBERT\-class embeddings, association and feature norms, WordNet, curated common\-answer banks, closed\-world manifests, and hidden tests, so we name the scoring*deterministic and embedding\-anchored*rather than symbolic\. Proxy swaps of the embedding model and of bank coverage, run on the earlier prompt\-collection panel, retained only a minority of decisive cells, so these resources carry structural weight on the imagination side, and extension to another language requires rebuilding them rather than translating prompts\. Deterministic atoms are also blind to higher\-order qualities such as narrative tension, which a rubric\-based judge would see at the cost of auditability\.
## Appendix PStatistical Power and Design Sensitivity
Figure P1:Statistical power under the current benchmark design\. Panel \(a\) reports Monte Carlo power for a two\-sided Spearman correlation over7979models; the dashed curve usesα=\.05/63\\alpha\{=\}\.05/63as a conservative familywise reference for one6363\-cell view, and the vertical lines mark the median raw and weakest controlled coefficients among the selected cells\. Panel \(b\) reports variance\-calibrated power for an imagination\-score difference over8080shared items; the curves compare the paired Wilcoxon test atα=\.05\\alpha\{=\}\.05, a conservativeα=\.05/1653\\alpha\{=\}\.05/1653reference, and the Mann–WhitneyUUtest after discarding item pairing\. Points mark the minimum effect reaching80%80\\%power, and the vertical line marks the9090th percentile of observed model\-mean gaps\.What effect sizes can7979models and8080shared items actually resolve? Because data collection was complete, we estimated design sensitivity rather than achieved power, in10,00010\{,\}000Monte Carlo runs \(seed2026072920260729\), applying Spearman tests to Gaussian\-copula samples of7979models and paired Wilcoxon tests to variance\-calibrated Gaussian random\-effects samples for5858complete models and8080shared items\. At nominalα=\.05\\alpha\{=\}\.05,80%80\\%power was reached at\|ρ\|=\.31\|\\rho\|\{=\}\.31for one coupling and an imagination\-score gap of\.041\.041\(dz=\.33d\_\{z\}\{=\}\.33\); conservative familywise references raised the thresholds to\|ρ\|=\.45\|\\rho\|\{=\}\.45for6363cells and\.078\.078\(dz=\.62d\_\{z\}\{=\}\.62\) for1,6531\{,\}653pairs \(Figure[P1](https://arxiv.org/html/2609.22152#A16.F1)a–b\)\. The selected cells have median raw\|ρ\|=\.56\|\\rho\|\{=\}\.56, above the conservative correlation threshold, and shared\-item pairing reduced the detectable imagination gap by41%41\\%relative to the\.069\.069unpaired threshold, a gain consistent with items carrying74\.5%74\.5\\%of the score variance\. The current design therefore has adequate sensitivity for the moderate\-to\-large couplings that support the main finding, while the shared anchor set improves model contrasts, although small controlled associations and closely spaced model pairs remain below its resolution\.
## Appendix QSupplementary Related Work
The main paper focuses on work that measures imagination and hallucination together\. This section expands that discussion by reviewing the two single\-axis benchmark traditions that motivateWhiteboard’s paired design\.
### Q\.1Creativity Benchmarks for LLMs
Creativity benchmarks operationalize divergent production through several task\-specific measures\. The Divergent Association Task scores semantic distance among words generated to be mutually unrelated\([Olson et al\. 2021](https://arxiv.org/html/2609.22152#bib.bib2)\)\. Automated scoring of the Alternative Uses Test instead measures semantic distance between an alternative use and the object named in the prompt\([Beaty et al\. 2022](https://arxiv.org/html/2609.22152#bib.bib4)\)\. Forward flow measures how far each word in a free\-association chain moves from earlier words\([Gray et al\. 2019](https://arxiv.org/html/2609.22152#bib.bib3)\)\.[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.22152#bib.bib5)apply the Alternative Uses Test to GPT\-3 and compare its outputs with human responses using expert ratings of originality, usefulness, surprise, and flexibility, together with automated semantic\-distance scores\. CreativityPrism groups tasks from divergent thinking, creative writing, and logical reasoning under quality, novelty, and diversity, using automatic judges validated against human annotations\([Hou et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib9)\)\. LiveIdeaBench elicits scientific ideas from single\-keyword prompts and evaluates them with a dynamic multi\-model judging panel\([Ruan et al\. 2026](https://arxiv.org/html/2609.22152#bib.bib8)\)\. MacGyver tests object reuse under physical constraints\([Tian et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib6)\), while CS4 varies the number of story\-writing constraints to assess creativity without human ratings\([Atmakuru et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib31)\)\. NeoCoder uses denial prompting to test divergent and convergent thinking on programming problems\([Lu et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib30)\)\.[Al Rabeyah et al\. \(2025\)](https://arxiv.org/html/2609.22152#bib.bib7)test whether four LLM judges agree with an oracle set of Alternative Uses Test responses and with one another\. These benchmarks measure creativity or its evaluation, but they do not assign paired imagination and support\-boundary labels to the same generation\.Whiteboardadds that paired reading while retaining task\-specific measures of divergent production\.
### Q\.2Hallucination Benchmarks for LLMs
Hallucination benchmarks range from answer\-level truthfulness tests to fine\-grained support checks\. TruthfulQA tests whether models reproduce common human misconceptions in answers to adversarially selected questions, using generation and multiple\-choice scores\([Lin et al\. 2022](https://arxiv.org/html/2609.22152#bib.bib14)\)\. HaluEval provides generated and human\-annotated hallucination examples for question answering, dialogue, and summarization\([Li et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib13)\)\. FActScore decomposes long\-form generations into atomic facts and verifies each against a knowledge source\([Min et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib15)\)\. LongFact supplies open\-domain prompts, while SAFE decomposes responses and checks individual facts with search\([Wei et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib29)\)\. ALCE evaluates retrieval\-augmented answers using citation correctness and completeness, alongside fluency and correctness\([Gao et al\. 2023](https://arxiv.org/html/2609.22152#bib.bib28)\)\. RAGTruth annotates hallucinations at both response and word levels in retrieval\-grounded question answering and summarization\([Niu et al\. 2024](https://arxiv.org/html/2609.22152#bib.bib25)\)\. HALoGEN verifies atomic units across nine domains and distinguishes errors associated with incorrect recall, source knowledge, and fabrication\([Ravichander et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib26)\)\. HalluLens separates intrinsic and extrinsic hallucination tasks\([Bang et al\. 2025](https://arxiv.org/html/2609.22152#bib.bib27)\), while[Huang et al\. \(2025\)](https://arxiv.org/html/2609.22152#bib.bib32)organize definitions, causes, detection, and mitigation methods\. Across these benchmarks, the unit of evaluation becomes progressively finer, but the labels remain on the truthfulness or support side\. The cited benchmarks do not also assign an imagination subtype to the same output\.Whiteboardpairs that support\-boundary reading with ten hallucination subtypes and seven imagination subtypes on each generation\.
## Appendix RReliability and the Boundaries of the Instrument
We close with the resolution of the instrument itself\. Among the5858complete models, items account for74\.5%74\.5\\%of variance and models for0\.41%0\.41\\%; reliability isG=0\.567G\{=\}0\.567andΦ=0\.248\\Phi\{=\}0\.248\. Five subtypes are compressed at the floor, while provider deletion preserves at least70%70\\%of decisive cells in the gated and residual views \(Appendix[G](https://arxiv.org/html/2609.22152#A7)\)\. With only3\.3%3\.3\\%of model pairs separable \(section “The Anchor Set and Paired Discriminative Power” of the main paper\), the instrument supports panel\-level structure rather than individual rankings\. Further boundaries are detailed in Appendix[O](https://arxiv.org/html/2609.22152#A15)\.相似文章
HalluWorld:基于参考世界模型的可控幻觉基准
HalluWorld 是一个可控基准框架,通过显式的参考世界模型在网格世界、国际象棋和实际终端任务等合成环境中评估大型语言模型中的幻觉。它可以细粒度分析各种故障模式,例如感知幻觉、多步状态追踪和因果模拟,揭示出前沿模型在处理扩展思维无法解决的复杂推理时仍然存在困难。
OmniHallu:用于多模态大语言模型中跨模态理解与生成的统一幻觉检测
OmniHallu引入了一个统一的幻觉检测框架,适用于多模态大语言模型,涵盖图像、视频和音频模态的理解与生成任务,并包括一个基准测试和多智能体架构。
从架构到输出:大型语言模型中幻觉的结构根源及数据的放大作用
本文分析了大型语言模型中的幻觉,将其视为三个架构决策的结构性后果:自注意力的共现学习、最大似然估计训练目标以及自回归解码的左到右承诺。它将每种机制映射到特定的幻觉类型,并论证了数据集病态会放大但不会导致这些脆弱性。
多模态大语言模型的统一幻觉模糊测试
本文介绍了UniHall,一个基于统一分类法的细粒度幻觉基准,以及自适多模态模糊测试(SAMF),一个用于多模态LLM的自进化压力测试框架。实验表明,最先进的模型在模糊测试下性能显著下降,并揭示了有用性与幻觉之间的权衡。
注意力分散作为大型语言模型中幻觉的诊断信号
本文提出了一种基于注意力分散的无监督指标,用于检测大型语言模型中的幻觉,在使用Qwen2.5模型家族的数学推理基准测试中显示出显著的AUC提升。