Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

arXiv cs.CL Papers

Summary

Introduces Ekphrasis, a 400-task benchmark for measuring visual creative ideation in text-only LLMs, separating usefulness, expressiveness, and novelty. The paper validates it with cross-modal grounding, showing text-level visual ideation ordering survives rendering.

arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clich\'es or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clich\'es into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clich\'ed. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.
Original Article
View Cached Full Text

Cached at: 08/10/26, 08:04 AM

# Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMsgithub.com/Imhongyu/Ekphrasis
Source: [https://arxiv.org/html/2608.06967](https://arxiv.org/html/2608.06967)
Hongyu Luo, He Wang, Huihao Jing, Hong Ting Tsang, Yuxuan Liu, Wuganjing Song, Yauwai Yim, Chunyang Li, Yangqiu Song The Hong Kong University of Science and Technology, Hong Kong SAR, China \{hluoay,hwangje,hjingaa,httsangaj,yliurk,wsongan,ywyimaa,cliej\}@connect\.ust\.hk,yqsong@cse\.ust\.hk

###### Abstract

Current evaluations do not isolate whether text\-only language models can originate visual concepts before image generation\. Fluent visual prose can hide visual\-plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene\. We define Visual Creative Ideation \(VCI\) as the ability to produce textual visual plans that are useful, expressive, and population\-novel, and introduceEkphrasis, a 400\-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation\.Ekphrasisscores anonymized pairwise comparisons with dimension\-specific checklists, aggregates preferences with Bradley–Terry models, and uses Typed Idea Graphs to convert task\-specific population clichés into novelty references\. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd\. A cross\-modal grounding study further shows that text\-level VCI ordering largely survives faithful rendering and blind image\-level preference judgment, supportingEkphrasisas a measure of visual ideation beyond prose quality\.

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text\-Only LLMs††thanks:[github\.com/Imhongyu/Ekphrasis](https://github.com/Imhongyu/Ekphrasis)

Hongyu Luo, He Wang, Huihao Jing, Hong Ting Tsang, Yuxuan Liu,Wuganjing Song, Yauwai Yim, Chunyang Li, Yangqiu SongThe Hong Kong University of Science and Technology, Hong Kong SAR, China\{hluoay,hwangje,hjingaa,httsangaj,yliurk,wsongan,ywyimaa,cliej\}@connect\.ust\.hk,yqsong@cse\.ust\.hk

![Refer to caption](https://arxiv.org/html/2608.06967v1/x1.png)Figure 1:Construct and evaluation workflow ofEkphrasis\.Ekphrasisisolates Visual Creative Ideation as a text\-only, pre\-image planning construct distinct from image quality, writing quality, and general ideation\. The workflow elicits visual blueprints across four brief families, scores plans through checklist pairwise judgments with Typed Idea Graph\-grounded novelty and active Bradley–Terry aggregation, and summarizes coverage across usefulness, expressiveness, novelty, and task subtypes\.## 1Introduction

Before an image is rendered, many creative decisions have already been made\. In concept\-art brainstorming, advertising, pre\-visualization, and text\-to\-image workflows, a text\-only LLM may choose the scene, carrier, composition, and design direction that downstream artists, users, or renderers later realize\. The product is therefore neither an image nor ordinary prose, but apre\-image visual plan: language that must carry visual substance\.

This setting creates a measurement trap\. A model can write vivid visual language while recycling safe motifs, familiar composition recipes, and recurring LLM\-population clichés\. We call thisnovelty illusion: outputs read as creative even when their visual invention is thin\. The problem is hard to detect in text alone because prose can explain away what an image could not sustain\.

Consider a brief about time pressure\. A fluent plan might use an hourglass, storm clouds, and a figure running out of sand\. It fits the brief but repeats a familiar carrier\. Another plan might turn a subway map into branching deadlines, with stations collapsing into missed opportunities\. The difference is not polish but mechanism: the latter changes how time pressure becomes visible\.

Existing evaluations largely miss this construct\. Adjacent benchmarks—verbal creativity\(Guilford,[1967](https://arxiv.org/html/2608.06967#bib.bib11); Olsonet al\.,[2021](https://arxiv.org/html/2608.06967#bib.bib12); Chakrabartyet al\.,[2024](https://arxiv.org/html/2608.06967#bib.bib20); Feinet al\.,[2026](https://arxiv.org/html/2608.06967#bib.bib21); Paech,[2025](https://arxiv.org/html/2608.06967#bib.bib22)\), image\-generation and image\-editing artifacts\(Huanget al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib13); Ghoshet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib15); Wuet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib19); Zhaoet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib34); Wuet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib35); Hanet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib36)\), visual reasoning\(Antolet al\.,[2015](https://arxiv.org/html/2608.06967#bib.bib28); Hudson and Manning,[2019](https://arxiv.org/html/2608.06967#bib.bib29)\), and instruction following\(Zhouet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib30)\)—target final texts, rendered images, or constraint adherence \(Figure[1](https://arxiv.org/html/2608.06967#S0.F1)\), but do not isolate whether a text\-only plan is useful, imageable, and non\-clichéd before pixels exist\.

Can language models imagine visually, or only write as if they can?

We introduceVisual Creative Ideation\(VCI\): the ability to originate textual visual plans that areusefulfor a brief,expressiveenough to support a concrete image, andnovelrelative to task\-specific patterns in the LLM population\. This construction follows the standard view that creativity combines originality with usefulness or effectiveness\(Runco and Jaeger,[2012](https://arxiv.org/html/2608.06967#bib.bib8)\)\. The object of evaluation is an externalized plan, not mental imagery, prompt\-engineering syntax, or final image quality\. Cross\-modal validation then asks whether the text\-level signal survives when the idea is rendered and judged without the original prose\.

Our contributions are:

- •\(Construct\)We define VCI as pre\-image visual planning that separates brief satisfaction, visual specificity, and population\-novel departure from recurring LLM clichés\.
- •\(Measurement\)We releaseEkphrasis, a 400\-instance suite spanningAbstraction,Combination,Transformation, andAdaptation, with checklist\-assisted Bradley–Terry comparison\(Bradley and Terry,[1952](https://arxiv.org/html/2608.06967#bib.bib26)\)and Typed Idea Graphs for population\-anchored Novelty\.
- •\(Evidence and validation\)We evaluate 14 language models, show that VCI is multidimensional rather than a fluency leaderboard, and validate a faithful rendered subset with blind image\-level human preferences\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/x2.png)Figure 2:Evaluation\-method overview\.Ekphrasisanonymizes paired visual plans, evaluates them with dimension\-specific checklists for Usefulness, Expressiveness, and Novelty, aggregates pairwise outcomes with Bradley–Terry models, and constructs population\-anchored Novelty references through leave\-two\-models\-out Typed Idea Graphs\.
## 2Related Work

Ekphrasisoccupies the visual\-ideation cell in Table[1](https://arxiv.org/html/2608.06967#S2.T1): the output is text, but the target construct is visual concept origination before image generation\. Prior benchmarks cover adjacent cells without isolating text\-only prompt\-side visual ideation with population\-anchored novelty and cross\-modal validity evidence\.

### 2\.1Creativity Evaluation in Language Models

Language\-model creativity benchmarks mostly evaluate verbal artifacts, procedural invention, or associative distance\. Divergent\-thinking tasks such as Alternative Uses and Divergent Association probe originality\(Guilford,[1967](https://arxiv.org/html/2608.06967#bib.bib11); Olsonet al\.,[2021](https://arxiv.org/html/2608.06967#bib.bib12)\), while recent benchmarks extend to physical problem solving, code creativity, marketing ideation, and creative writing\(Tianet al\.,[2024](https://arxiv.org/html/2608.06967#bib.bib23); Luet al\.,[2024](https://arxiv.org/html/2608.06967#bib.bib31); Houet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib24); Bhatet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib25); Chakrabartyet al\.,[2024](https://arxiv.org/html/2608.06967#bib.bib20); Feinet al\.,[2026](https://arxiv.org/html/2608.06967#bib.bib21); Paech,[2025](https://arxiv.org/html/2608.06967#bib.bib22)\)\. These tasks motivate preference\-based evaluation, but a strong answer need not specify a renderable scene\.Ekphrasisinstead evaluates an externalized plan for an absent image, and defines novelty relative to recurring visual patterns in the model population for the same brief\.

### 2\.2Visual and Multimodal Evaluation

Visual benchmarks usually score artifacts or understanding\. Text\-to\-image benchmarks such as T2I\-CompBench, GenEval, and HPS v2 evaluate rendered images for compositionality, alignment, or human preference\(Huanget al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib13); Ghoshet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib15); Wuet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib19)\); related image–text compatibility and faithfulness metrics evaluate whether an image preserves its textual input\(Hesselet al\.,[2021](https://arxiv.org/html/2608.06967#bib.bib14); Huet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib17)\)\. Visual reasoning and instruction\-following benchmarks test image understanding or constraint adherence\(Antolet al\.,[2015](https://arxiv.org/html/2608.06967#bib.bib28); Hudson and Manning,[2019](https://arxiv.org/html/2608.06967#bib.bib29); Zhouet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib30)\)\. Recent multimodal creativity benchmarks study creative intelligence with image inputs or visual assets\(Fanget al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib32); Xiaet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib33)\)\. Image\-editing benchmarks such as SmartEdit, RISEBench, KRISBench, and UniREditBench add stronger reasoning coverage and, in UniREditBench, dual image/text references\(Huanget al\.,[2024](https://arxiv.org/html/2608.06967#bib.bib16); Zhaoet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib34); Wuet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib35); Hanet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib36)\), but still score edited images\.Ekphrasisevaluates the plan that precedes any image artifact\.

### 2\.3Pairwise Evaluation and LLM\-as\-Judge

Open\-ended creative outputs rarely admit a single gold answer, making absolute ratings sensitive to preference and surface fluency\. Pairwise comparison with Bradley–Terry aggregation provides local judgments and model\-level scores\(Bradley and Terry,[1952](https://arxiv.org/html/2608.06967#bib.bib26)\), including in creative\-writing and marketing\-ideation evaluation\(Feinet al\.,[2026](https://arxiv.org/html/2608.06967#bib.bib21); Bhatet al\.,[2025](https://arxiv.org/html/2608.06967#bib.bib25)\)\. Because LLM\-as\-judge systems can show position, verbosity, and fluency biases\(Zhenget al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib27)\),Ekphrasisuses construct\-derived checklists, reports the three VCI dimensions separately, anchors novelty to a leave\-out task population, and calibrates scalable judgments against human annotations\.

Table 1:Coverage matrix positioningEkphrasisrelative to adjacent benchmark families\.Ekphrasisis designed for text\-only, pre\-image visual plans, with population\-anchored novelty and a separate image\-level validation path\.Benchmark familyExamplesSubject / outputConstructValidationText\-onlyinputText planscoredImageartifact scoredVisualtargetPre\-imageideationPopulationnoveltyImage\-levelcheckVerbal creativity / writingAUT / DAT; LitBench✓––––––Text\-to\-image generationT2I\-CompBench; GenEval✓–✓✓–––Visual reasoningVQA; GQA–––✓–––Multimodal creativityCreation\-MMBench; VISIAR––✓✓–––Image editingUniREditBench; KRISBench––✓✓––✓EkphrasisText\-only visual briefs✓✓–✓✓✓✓

- •Examples are representative; the surrounding related\-work text gives the full cited set\. A checkmark denotes a primary evaluation target, except thatEkphrasisuses image\-level checking only for cross\-modal validation rather than official text\-level scoring\.

## 3Visual Creative Ideation andEkphrasis

Ekphrasisevaluates prompt\-side visual ideation before pixels are generated\. Inputs and outputs remain textual, but the score targets whether the plan fits the brief, supports a concrete image, and avoids common LLM\-population solutions\.

### 3\.1Task and Response Format

###### Definition 1\(Visual Creative Ideation\)\.

Given a visual briefbb,*Visual Creative Ideation*\(VCI\) is the ability to originate a textual visual plan that is useful for the brief, expressive enough to support a stable image, and non\-clichéd with respect to common responses to the same brief\.

AnEkphrasistask instance is a visual brief

b=\(x,τ,C\),b∈ℬ,b=\(x,\\tau,C\),\\qquad b\\in\\mathcal\{B\},wherexxis the input content,τ\\tauis the visual\-ideation operation, andCCdenotes contextual or medium constraints\. Givenbb, modelmmreturns

pm​\(b\)=\(im​\(b\),dm​\(b\)\),p\_\{m\}\(b\)=\\big\(i\_\{m\}\(b\),d\_\{m\}\(b\)\\big\),where the*Core Idea*im​\(b\)i\_\{m\}\(b\)states the carrier, metaphor or mechanism, and key state or action, and the*Visual Description*dm​\(b\)d\_\{m\}\(b\)states what would be visible in one bounded static frame\.

### 3\.2Protocol and Evaluation Dimensions

###### Definition 2\(EkphrasisTask Protocol\)\.

Given a briefb∈ℬb\\in\\mathcal\{B\}and a language modelmm, theEkphrasisprotocol requiresmmto output a two\-field textual visual plan

pm​\(b\)=\(im​\(b\),dm​\(b\)\),p\_\{m\}\(b\)=\\big\(i\_\{m\}\(b\),d\_\{m\}\(b\)\\big\),whereim​\(b\)i\_\{m\}\(b\)states the core visual idea anddm​\(b\)d\_\{m\}\(b\)provides a concrete visual description of the same idea\.

A valid response describes one static image and excludes tool names, renderer parameters, multi\-image narratives, and prompt\-engineering syntax\. The two\-field format prevents fluent scene prose from replacing an appropriate, imageable, and non\-obvious visual plan\.

###### Definition 3\(VCI Evaluation Dimensions\)\.

For a model responsepm​\(b\)p\_\{m\}\(b\),Ekphrasisevaluates VCI along

Eval​\(m,b\)=\(Um​\(b\),Em​\(b\),Nm​\(b\)\),\\mathrm\{Eval\}\(m,b\)=\\big\(U\_\{m\}\(b\),E\_\{m\}\(b\),N\_\{m\}\(b\)\\big\),whereUm​\(b\)U\_\{m\}\(b\)denotes*Usefulness*,Em​\(b\)E\_\{m\}\(b\)denotes*Expressiveness*, andNm​\(b\)N\_\{m\}\(b\)denotes*Novelty*\. Usefulness measures whether the idea satisfies the brief\. Expressiveness measures whether the description supports a concrete mental image\. Novelty measures whether the core idea departs from population\-level visual clichés for the same brief\.

Usefulness and Novelty primarily inspectim​\(b\)i\_\{m\}\(b\); Expressiveness primarily inspectsdm​\(b\)d\_\{m\}\(b\)\. Thus VCI is not image quality, verbal creativity, instruction following alone, or renderer\-specific prompting\.

### 3\.3Task Design

Ekphrasisorganizes its briefs into four controlled task families:

ℬ=ℬabs∪ℬcomb∪ℬtrans∪ℬadapt,\\mathcal\{B\}=\\mathcal\{B\}\_\{\\mathrm\{abs\}\}\\cup\\mathcal\{B\}\_\{\\mathrm\{comb\}\}\\cup\\mathcal\{B\}\_\{\\mathrm\{trans\}\}\\cup\\mathcal\{B\}\_\{\\mathrm\{adapt\}\},where the first three families specify the creative operation and Adaptation specifies a realistic context while leaving the visual strategy open\. Table[2](https://arxiv.org/html/2608.06967#S3.T2)summarizes the sampling control, balancing axis, and diagnostic failure mode for each family; exact family and subtype counts are reported in Appendix[A](https://arxiv.org/html/2608.06967#A1)\. Source pools draw on concreteness norms\(Brysbaertet al\.,[2014](https://arxiv.org/html/2608.06967#bib.bib37)\), SUBTLEX\-US frequency estimates\(Brysbaert and New,[2009](https://arxiv.org/html/2608.06967#bib.bib38)\), THINGS object concepts\(Hebartet al\.,[2019](https://arxiv.org/html/2608.06967#bib.bib39)\), and Places365 scene categories\(Zhouet al\.,[2018](https://arxiv.org/html/2608.06967#bib.bib40)\)\.

Table 2:Task\-family construction controls\.FamilyConstruction controlStress testedAbstractionCurated abstract nouns; semantic subtype, abstractness, frequency, and visualizability controlsVague metaphor; decorative symbolCombinationConcrete–concrete, abstract–concrete, and abstract–abstract concept pairs; domain balance, distance band, and repetition capJuxtaposition; trivial or incoherent fusionTransformationObject pool crossed with material, temporal, and contextual specifications; category and distance balanceSurface edit; loss of object identityAdaptationWeb\-collected real\-world visual\-design needs rewritten as applied briefs; domain, medium, audience, and deliverable balanceGeneric campaign trope; missed constraintAbstraction tests abstract\-to\-visual grounding; Combination tests whether models fuse concepts rather than merely juxtapose them; Transformation tests condition\-guided object rewriting under recognizability constraints; and Adaptation tests visual strategy selection under realistic communication or design constraints\. Representative prompts and subtype distributions are reported in Appendix[A](https://arxiv.org/html/2608.06967#A1)\.

### 3\.4Dataset Construction and Quality Control

Tasks are constructed through template\-based authoring, controlled sampling, real\-brief rewriting, and human curation\. Each prompt is screened for ambiguity, single\-image renderability, family/subtype overlap, sensitive or copyrighted content, and distributional balance\. A pilot pass then verifies that prompts elicit the required Core Idea and Visual Description fields rather than essays, tool commands, multi\-image sequences, or non\-visual explanations\. Detailed sampling protocols and prompt templates are given in Appendix[A](https://arxiv.org/html/2608.06967#A1)\.

## 4Evaluation Method

Ekphrasisscores open\-ended visual plans through checklist\-assisted pairwise comparison\. Usefulness and Expressiveness use dimension\-specific rubrics; Novelty uses task\-specific cliché references from leave\-out Typed Idea Graphs\. Pairwise preferences are aggregated with Bradley–Terry \(BT\) models\(Bradley and Terry,[1952](https://arxiv.org/html/2608.06967#bib.bib26)\)\. Figure[2](https://arxiv.org/html/2608.06967#S1.F2)summarizes this measurement engine\.

### 4\.1Checklist\-Assisted Pairwise Evaluation

For each briefbb, evaluation dimensiond∈\{U,E,N\}d\\in\\\{U,E,N\\\}, and model pair\(m,n\)\(m,n\), the judge receives anonymized outputspm​\(b\)p\_\{m\}\(b\)andpn​\(b\)p\_\{n\}\(b\)together with a dimension\-specific checklist𝒞b,d\\mathcal\{C\}\_\{b,d\}, then returns

Jd​\(b,pm​\(b\),pn​\(b\),𝒞b,d\)→yb,dm,n∈\{1,12,0\},J\_\{d\}\\\!\\left\(b,p\_\{m\}\(b\),p\_\{n\}\(b\),\\mathcal\{C\}\_\{b,d\}\\right\)\\rightarrow y^\{m,n\}\_\{b,d\}\\in\\left\\\{1,\\frac\{1\}\{2\},0\\right\\\},where11means thatmmis preferred tonn,0means thatnnis preferred tomm, and12\\frac\{1\}\{2\}denotes a tie\. Model identities are hidden, A/B order is randomized, and swapped\-order repeats estimate position bias\.

For each dimension, the BT model assigns modelmma latent skillθm,d\\theta\_\{m,d\}and models the probability thatmmis preferred tonnas

Pr⁡\(m≻n∣d\)=σ​\(θm,d−θn,d\)\.\\Pr\(m\\succ n\\mid d\)=\\sigma\(\\theta\_\{m,d\}\-\\theta\_\{n,d\}\)\.Only decisive comparisons enter the main BT likelihood; ties are kept for diagnostics\. We fit one BT model per dimension and reportθ^m,U\\hat\{\\theta\}\_\{m,U\},θ^m,E\\hat\{\\theta\}\_\{m,E\}, andθ^m,N\\hat\{\\theta\}\_\{m,N\}\. Regularization, active sampling, and confidence intervals are specified in Appendix[E](https://arxiv.org/html/2608.06967#A5)\.

### 4\.2Typed Idea Graphs for Population\-Anchored Novelty

Novelty is task\-relative: a response is novel when it avoids, recombines, or transforms visual patterns that other models repeatedly use for the same brief\. For each judged pairq=\{m,n\}q=\\\{m,n\\\}, we build a leave\-two\-models\-out reference population from all other models’ responses to that brief, preventing either candidate from defining the cliché profile used to judge it\.

Responses in the reference population are mapped into typed idea atoms \(*concept*,*motif*,*mechanism*,*purpose*, and*audience*\), connected by co\-occurrence, semantic similarity, and typed relations, and clustered into population\-common patterns\. The Novelty checklist is constructed as

ℛb−q→EDCAb−q→graphGb−q→cluster\+score𝒞b,Nq,\\mathcal\{R\}^\{\-q\}\_\{b\}\\xrightarrow\{\\;\\mathrm\{EDC\}\\;\}A^\{\-q\}\_\{b\}\\xrightarrow\{\\;\\mathrm\{graph\}\\;\}G^\{\-q\}\_\{b\}\\xrightarrow\{\\;\\mathrm\{cluster\+score\}\\;\}\\mathcal\{C\}^\{q\}\_\{b,N\},where EDC denotes Extract–Define–Canonicalize\. Each cluster is scored by prevalence, structural centrality, and semantic convergence:

S​\(c\)=α​P¯​\(c\)\+β​C¯​\(c\)\+γ​Q¯​\(c\),α\+β\+γ=1\.S\(c\)=\\alpha\\bar\{P\}\(c\)\+\\beta\\bar\{C\}\(c\)\+\\gamma\\bar\{Q\}\(c\),\\qquad\\alpha\+\\beta\+\\gamma=1\.High\-scoring clusters are verbalized as checklist anchors\. During judging, the evaluator marks whether each candidaterepeats,transforms, oravoidsthose anchors, so Novelty is measured against task\-specific visual patterns rather than generic semantic distance\. Extraction prompts, graph construction, and audit criteria are detailed in Appendix[G](https://arxiv.org/html/2608.06967#A7)\.

### 4\.3Score Reporting and Calibration

The main results use dimension\-specific BT expected\-winrate scores\. Because the three dimensions have different empirical spreads, we standardize each dimension before computing Overall VCI:

zm,d=sm,d−μdσd,z\_\{m,d\}=\\frac\{s\_\{m,d\}\-\\mu\_\{d\}\}\{\\sigma\_\{d\}\},wheresm,ds\_\{m,d\}is the BT expected\-winrate score andμd,σd\\mu\_\{d\},\\sigma\_\{d\}are the mean and standard deviation over subject models\. Overall VCI is the equal\-weight mean of the threezzscores\. For visualization only, Figure[3](https://arxiv.org/html/2608.06967#S4.F3)uses median\-centered rank percentiles to show the Usefulness–Novelty tradeoff; these coordinates are not absolute quality percentages\.

We use a two\-tier evaluation design: a scalable checklist\-assisted BT judge produces the full leaderboard, while human annotations provide targeted calibration and validation slices for agreement, reliability, and cross\-modal grounding\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/x3.png)Figure 3:Main VCI results\. Panel A ranks models by Overall VCI, computed as the mean of standardized BT expected\-winrate scores across Usefulness, Expressiveness, and Novelty\. Bars show Overall VCI and heatmap cells show dimension\-specificzzscores\. Panel B keeps the Usefulness–Novelty tradeoff display using median\-centered rank percentiles for visualization only; point size and opacity encode Overall VCI\.

## 5Experiments

We evaluate 14 text\-only LLMs on the 400\-taskEkphrasissuite\. The official leaderboard uses the first pre\-registered sample for each model–task pair; additional samples only estimate task\-specific population clichés for Typed Idea Graph construction\.

### 5\.1Experimental Setup

Usefulness, Expressiveness, and Novelty are judged separately with checklist\-assisted pairwise comparison and one BT model per dimension\. The official DeepSeek\-V4\-Flash scorer uses active BT sampling: after warm\-start coverage, it prioritizes comparisons that are informative under current uncertainty\. Order swaps, duplicate checks, auxiliary judges, and human validation slices characterize reliability rather than redefine the leaderboard\. Budgets, pilot criteria, and validation scale are reported in Appendix[D](https://arxiv.org/html/2608.06967#A4)\.

### 5\.2Main Results

Figure[3](https://arxiv.org/html/2608.06967#S4.F3)summarizes the primary evidence\. Gemini 3\.1 Pro \(\+0\.96\), GPT\-5\.4 \(\+0\.92\), Claude 4\.7 \(\+0\.73\), DeepSeek V4 \(\+0\.70\), and Kimi K2\.6 \(\+0\.68\) form a close, profile\-diverse top tier\. GPT\-5\.4 leads Usefulness \(\+1\.53\), Claude 4\.7 and Gemini 3\.1 Pro lead Expressiveness \(\+1\.20/\+1\.19\), and Kimi K2\.6 leads Novelty \(\+1\.66\)\. Grok 4\.3 shows the dissociation: high Usefulness \(\+1\.10\) can coexist with below\-average Novelty \(\-0\.17\)\.

### 5\.3Model Profiles and Task\-Family Stress Tests

Task families reveal profile\-specific strengths: Gemini 3\.1 Pro leads Abstraction \(\+1\.03\), GPT\-5\.4 leads Adaptation \(\+1\.04\), and Kimi K2\.6 leads Transformation \(\+1\.07\)\. Combination is most demanding because it requires shared visual structure rather than side\-by\-side placement, making it the clearest family\-level stress test\. Table[3](https://arxiv.org/html/2608.06967#S5.T3)reports the full model\-by\-family breakdown\.

Table 3:Model performance by task family\. Cells are family\-specific standardized Bradley–Terry scores; higher is better, and bold marks the row\-best model\.MetricGeminiGPT\-5\.4ClaudeDeepSeekKimiGrokR1GLM\-5\.1Qwen\-MaxMiniMaxQwen\-8BLlama\-8BPhi\-4Llama\-MavAbstractionabstractionU0\.850\.851\.421\.420\.750\.750\.580\.58−0\.06\-0\.060\.770\.77−0\.13\-0\.130\.590\.59−0\.00\-0\.000\.850\.85−1\.31\-1\.31−2\.13\-2\.13−1\.27\-1\.27−0\.91\-0\.91E0\.990\.990\.650\.651\.051\.050\.630\.630\.340\.340\.500\.500\.850\.850\.510\.510\.610\.61−0\.03\-0\.03−1\.57\-1\.57−1\.53\-1\.53−1\.60\-1\.60−1\.42\-1\.42N1\.261\.260\.930\.930\.540\.540\.640\.641\.251\.25−0\.27\-0\.27−0\.01\-0\.010\.100\.100\.570\.57−0\.63\-0\.63−0\.47\-0\.47−0\.17\-0\.17−1\.08\-1\.08−2\.66\-2\.66Task avg\.1\.031\.031\.001\.000\.780\.780\.620\.620\.510\.510\.340\.340\.240\.240\.400\.400\.390\.390\.070\.07−1\.12\-1\.12−1\.27\-1\.27−1\.32\-1\.32−1\.66\-1\.66CombinationcombinationU0\.670\.671\.221\.220\.010\.010\.680\.680\.110\.111\.061\.060\.610\.610\.710\.710\.370\.370\.430\.43−1\.08\-1\.08−2\.03\-2\.03−1\.67\-1\.67−1\.10\-1\.10E1\.121\.120\.920\.920\.830\.830\.690\.690\.230\.230\.590\.590\.760\.760\.650\.650\.470\.47−0\.27\-0\.27−1\.24\-1\.24−1\.59\-1\.59−1\.63\-1\.63−1\.53\-1\.53N0\.120\.120\.760\.760\.570\.570\.780\.781\.771\.770\.190\.190\.150\.150\.700\.700\.300\.30−0\.34\-0\.34−1\.10\-1\.10−0\.28\-0\.28−1\.22\-1\.22−2\.41\-2\.41Task avg\.0\.630\.630\.970\.970\.470\.470\.720\.720\.700\.700\.620\.620\.510\.510\.690\.690\.380\.38−0\.06\-0\.06−1\.14\-1\.14−1\.30\-1\.30−1\.50\-1\.50−1\.68\-1\.68TransformationtransformationU0\.540\.540\.960\.960\.870\.870\.620\.620\.580\.581\.081\.080\.670\.670\.400\.400\.010\.010\.260\.26−1\.70\-1\.70−1\.78\-1\.78−1\.62\-1\.62−0\.88\-0\.88E0\.900\.900\.760\.761\.141\.140\.720\.720\.420\.420\.520\.520\.750\.750\.540\.540\.410\.41−0\.05\-0\.05−1\.36\-1\.36−1\.56\-1\.56−1\.65\-1\.65−1\.52\-1\.52N0\.940\.94−0\.49\-0\.490\.100\.100\.470\.472\.202\.20−0\.27\-0\.270\.590\.59−0\.08\-0\.080\.600\.60−0\.02\-0\.02−0\.85\-0\.85−0\.04\-0\.04−0\.78\-0\.78−2\.38\-2\.38Task avg\.0\.800\.800\.410\.410\.700\.700\.600\.601\.071\.070\.440\.440\.670\.670\.290\.290\.340\.340\.060\.06−1\.30\-1\.30−1\.13\-1\.13−1\.35\-1\.35−1\.59\-1\.59AdaptationadaptationU0\.200\.202\.122\.120\.540\.540\.120\.120\.040\.041\.261\.260\.440\.44−0\.04\-0\.040\.270\.270\.180\.18−1\.14\-1\.14−1\.99\-1\.99−1\.03\-1\.03−0\.97\-0\.97E0\.910\.910\.360\.360\.890\.890\.730\.730\.630\.630\.800\.800\.830\.830\.480\.480\.480\.480\.010\.01−1\.38\-1\.38−1\.39\-1\.39−1\.81\-1\.81−1\.55\-1\.55N1\.271\.270\.630\.630\.250\.251\.271\.271\.511\.51−0\.09\-0\.090\.140\.140\.520\.52−0\.05\-0\.05−0\.64\-0\.64−0\.84\-0\.84−0\.48\-0\.48−1\.38\-1\.38−2\.12\-2\.12Task avg\.0\.800\.801\.041\.040\.560\.560\.710\.710\.730\.730\.660\.660\.470\.470\.320\.320\.230\.23−0\.15\-0\.15−1\.12\-1\.12−1\.29\-1\.29−1\.41\-1\.41−1\.55\-1\.55Overall avg\.0\.820\.820\.850\.850\.630\.630\.660\.660\.750\.750\.510\.510\.470\.470\.420\.420\.340\.34−0\.02\-0\.02−1\.17\-1\.17−1\.25\-1\.25−1\.39\-1\.39−1\.62\-1\.62

- •U, E, and N denote Usefulness, Expressiveness, and Novelty\. Model abbreviations: Gemini = Gemini 3\.1 Pro; Claude = Claude Opus 4\.7; DeepSeek = DeepSeek V4 Pro; R1 = DeepSeek R1; Qwen\-Max = Qwen 3\.6\-Max; MiniMax = MiniMax M2\.7; Llama\-8B = Llama 3\.1 8B Instruct; Llama\-Mav = Llama 4 Maverick\.

### 5\.4Novelty Ablation

To test whether task\-specific Typed Idea Graph \(TIG\) anchors are needed for Novelty judgment, we rerun the same human gold subset with generic anchors in place of case\-specific checklist items\. TIG anchoring improves human\-majority agreement by 20\.9 percentage points \(Table[4](https://arxiv.org/html/2608.06967#S5.T4)\)\.

Table 4:TIG anchors improve novelty agreement\.Human\-aligned agreement is measured on 235 paired Novelty comparisons\.ConditionCorrect /nnAgreementTIG anchors180 / 23576\.6%No TIG131 / 23555\.7%
Gain,\+20\.9 pp\(95% CI, \[\+14\.0, \+28\.1\] pp\); discordant correct, 64 versus 15; McNemarp<0\.001p<0\.001\.

### 5\.5Reliability Checks

Reliability diagnostics support the model ranking: JSON repair is rare, tie rates are non\-degenerate, order\-swap rank correlations remain high, and auxiliary judges recover broadly similar rankings\. Length remains a residual confound, especially for Expressiveness; dimension\-wise diagnostics are reported in Appendix[D\.2](https://arxiv.org/html/2608.06967#A4.SS2)\.

## 6Cross\-Modal Grounding Validation

We further test whether text\-level VCI scores capture visual content rather than only verbal quality\. Textual plans are rendered with a fixed text\-to\-image system, screened for faithfulness to the source idea, and then compared by humans using only the rendered images\. The resulting image\-level preferences are compared with the original text\-level VCI rankings, following prior emphasis on image–text compatibility, text\-to\-image faithfulness, and reproducible human evaluation for generated images\(Hesselet al\.,[2021](https://arxiv.org/html/2608.06967#bib.bib14); Huet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib17); Otaniet al\.,[2023](https://arxiv.org/html/2608.06967#bib.bib18)\)\.

### 6\.1Faithfulness Gate

The faithfulness gate ensures that image\-level validation is performed only on renderings that preserve the source visual plan\. A rendering passes when it preserves the central carrier, objects, relation, and context needed for image\-level comparison; hard failures include semantic drift, concept splitting, object identity loss, and style/context mismatch\. In the completed annotation pass, each of the 240 rendered images receives three independent faithfulness annotations, for 720 judgments in total\. Majority vote passes 179 of 240 renderings \(74\.6%\)\. The mean faithfulness score is 3\.86/5; pass/fail agreement is substantial \(Fleissκ=0\.637\\kappa=0\.637\), and ordinal score reliability is high for a visual screening task \(ICC=0\.711=0\.711\), using standard multi\-rater reliability measures\(Fleiss,[1971](https://arxiv.org/html/2608.06967#bib.bib9); Shrout and Fleiss,[1979](https://arxiv.org/html/2608.06967#bib.bib10)\)\.

### 6\.2Image\-Level Preference Agreement

For the faithful subset, humans compare rendered images without seeing the original text, model identity, or text\-level ranking\. The completed image\-preference subset contains 160 image pairs and 480 judgments\. Annotators reach a majority preference on 158 of 160 pairs \(98\.8%\), with 73\.8% unanimous agreement and Fleissκ=0\.690\\kappa=0\.690\. At the pair level, decisive image preferences agree with the text\-level model ordering on 142 of 147 decisive pairs \(96\.6%\); counting image ties as neutral gives 142 of 158 non\-missing majority outcomes \(89\.9%\)\. We therefore use the rendered\-image study as convergent validity evidence, not as a replacement for the text\-level benchmark or as a separate full\-scale ranking benchmark\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/x4.png)Figure 4:Agreement between text\-level VCI ordering and image\-level human preferences in the six\-model cross\-modal subset\. Faint points are individual image\-pair majority outcomes; larger points summarize the corresponding model\-pair preference rate\. The horizontal axis shows the text\-level Overall VCI gap, and the vertical axis shows the image\-level human preference for modelmm\.

## 7Discussion

Ekphrasisanswers the title question with a qualified operational yes: some text\-only outputs specify visual plans whose structure remains recoverable after rendering and blind image comparison\. This is not evidence of private mental imagery\. It shows that pre\-image plans can carry task\-relevant visual information across a text\-to\-image modality shift\. VCI is therefore not monolithic: similar aggregate scores can reflect brief following, vivid specification, or departure from population clichés, and plausible scenes can remain visually familiar\.

The central boundary exposed by the benchmark is safe plausibility\. Many outputs satisfy the brief and read well while falling back on familiar carriers, stock metaphors, or predictable compositions\. This failure is subtler than missing a constraint because ordinary text evaluation can reward it: the output appears thoughtful until it is compared with the task\-specific pattern of what other models also produce\. Combination tasks expose this most clearly when weak plans place two concepts beside each other rather than inventing a shared visual structure; Transformation tasks show a related surface edit in which the object is decorated but its visible mode of existence is unchanged\.

Population\-anchored Novelty is therefore not an ornamental creativity score\. Typed Idea Graphs make the reference set explicit: a plan is judged against recurring solutions for the same brief, with the candidate models left out of the cliché profile\. The TIG ablation suggests that this reference mainly improves human\-aligned Novelty judgments rather than acting as a generic preference boost\. This matters because novelty in VCI is relational: a plan can be semantically appropriate and visually clear while still reproducing the model population’s default answer\.

The cross\-modal study provides a second boundary\. It gives convergent validity for the text\-level VCI signal beyond verbal fluency, but only for ideas that a fixed renderer preserves well enough to compare\. Rendering failures can reflect the image model as much as the source plan; image\-level agreement is validation evidence, not a replacement image\-generation leaderboard\. More broadly,Ekphrasismeasures externalized pre\-image planning, not private mental imagery or a general theory of human creativity\.

## 8Conclusion

Ekphrasisoperationalizes Visual Creative Ideation as useful, expressive, and population\-novel visual planning before pixels exist\. Across 400 tasks and 14 language models, it shows that brief satisfaction, imageability, and escape from recurring visual clichés are related but separable: strong systems reach similar overall VCI through different profiles, and high usefulness can remain visually familiar\. The results support a scoped answer to the title question: text\-only models can specify imageable visual plans without direct image input, but this ability must be tested against task fit, recoverability, and the clichés of their own model population\. For evaluation, the practical implication is direct: creative\-looking prose should not be treated as visual creativity without separate checks for visual mechanism, imageability, and population novelty\.

## Limitations

Ekphrasisis limited by its task distribution, language and cultural assumptions, and the current model population used to define population\-level clichés\. Checklist\-assisted LLM judgment reduces annotation cost but still exhibits position and length effects, which we report rather than eliminate\. Cross\-modal grounding also depends on the chosen renderer: a failed rendering can reflect image\-generation limits rather than poor text\-level ideation\. Finally, population\-anchored Novelty is relative to the evaluated model pool and should not be read as an absolute theory of human originality\.

## Ethical Considerations

Ekphrasisscores should be used as diagnostic evidence about model behavior, not as broad claims about human creativity or artistic value\. Novelty judgments may encode cultural assumptions about what counts as a cliché, so released checklist materials and annotation guidelines should be auditable\. Human annotators should be informed about task content, data use, and disagreement handling\. Because the benchmark can rank commercial systems, reporting should avoid overclaiming beyond the tested task families and evaluation protocol\.

## References

- VQA: visual question answering\.InProceedings of the IEEE International Conference on Computer Vision,pp\. 2425–2433\.External Links:[Document](https://dx.doi.org/10.1109/ICCV.2015.279)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- N\. Bhat, K\. Browne, and P\. Bingemann \(2025\)Creativity benchmark: a benchmark for marketing creativity for LLM models\.External Links:2509\.09702,[Link](https://arxiv.org/abs/2509.09702)Cited by:[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2608.06967#S2.SS3.p1.1)\.
- R\. A\. Bradley and M\. E\. Terry \(1952\)Rank analysis of incomplete block designs: i\. the method of paired comparisons\.Biometrika39\(3/4\),pp\. 324–345\.External Links:[Document](https://dx.doi.org/10.1093/biomet/39.3-4.324)Cited by:[2nd item](https://arxiv.org/html/2608.06967#S1.I1.i2.p1.1),[§2\.3](https://arxiv.org/html/2608.06967#S2.SS3.p1.1),[§4](https://arxiv.org/html/2608.06967#S4.p1.1)\.
- M\. Brysbaert and B\. New \(2009\)Moving beyond kucera and francis: a critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for american english\.Behavior Research Methods41\(4\),pp\. 977–990\.External Links:[Document](https://dx.doi.org/10.3758/BRM.41.4.977)Cited by:[§3\.3](https://arxiv.org/html/2608.06967#S3.SS3.p1.2)\.
- M\. Brysbaert, A\. B\. Warriner, and V\. Kuperman \(2014\)Concreteness ratings for 40 thousand generally known english word lemmas\.Behavior Research Methods46\(3\),pp\. 904–911\.External Links:[Document](https://dx.doi.org/10.3758/s13428-013-0403-5)Cited by:[§3\.3](https://arxiv.org/html/2608.06967#S3.SS3.p1.2)\.
- T\. Chakrabarty, P\. Laban, D\. Agarwal, S\. Muresan, and C\. Wu \(2024\)Art or artifice? large language models and the false promise of creativity\.InProceedings of the CHI Conference on Human Factors in Computing Systems,External Links:[Document](https://dx.doi.org/10.1145/3613904.3642731),[Link](https://doi.org/10.1145/3613904.3642731)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- X\. Fang, Z\. Chen, K\. Lan, L\. Ma, S\. Ding, Y\. Liang, X\. Zhao, F\. Wen, Z\. Zhang, G\. Zhang, H\. Duan, K\. Chen, and D\. Lin \(2025\)Creation\-mmbench: assessing context\-aware creative intelligence in MLLM\.External Links:2503\.14478,[Link](https://arxiv.org/abs/2503.14478)Cited by:[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- D\. Fein, S\. Russo, V\. Xiang, K\. Jolly, R\. Rafailov, and N\. Haber \(2026\)LitBench: a benchmark and dataset for reliable evaluation of creative writing\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),Rabat, Morocco,pp\. 7740–7755\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.362),[Link](https://aclanthology.org/2026.eacl-long.362/)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2608.06967#S2.SS3.p1.1)\.
- J\. L\. Fleiss \(1971\)Measuring nominal scale agreement among many raters\.Psychological Bulletin76\(5\),pp\. 378–382\.External Links:[Document](https://dx.doi.org/10.1037/h0031619)Cited by:[§H\.6](https://arxiv.org/html/2608.06967#A8.SS6.p1.3),[§6\.1](https://arxiv.org/html/2608.06967#S6.SS1.p1.2)\.
- D\. Ghosh, H\. Hajishirzi, and L\. Schmidt \(2023\)GenEval: an object\-focused framework for evaluating text\-to\-image alignment\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/a3bf71c7c63f0c3bcb7ff67c67b1e7b1-Abstract-Datasets_and_Benchmarks.html)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- J\. P\. Guilford \(1967\)The nature of human intelligence\.McGraw\-Hill,New York\.External Links:[Link](https://search.worldcat.org/title/The-nature-of-human-intelligence/oclc/204270)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- F\. Han, Y\. Wang, C\. Li, Z\. Liang, D\. Wang, Y\. Jiao, Z\. Wei, C\. Gong, C\. Jin, J\. Chen, and J\. Wang \(2025\)UniREditBench: a unified reasoning\-based image editing benchmark\.External Links:2511\.01295,[Link](https://arxiv.org/abs/2511.01295)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- M\. N\. Hebart, A\. H\. Dickter, A\. Kidder, W\. Y\. Kwok, A\. Corriveau, C\. Van Wicklin, and C\. I\. Baker \(2019\)THINGS: a database of 1,854 object concepts and more than 26,000 naturalistic object images\.PLOS ONE14\(10\),pp\. e0223792\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0223792)Cited by:[§3\.3](https://arxiv.org/html/2608.06967#S3.SS3.p1.2)\.
- J\. Hessel, A\. Holtzman, M\. Forbes, R\. Le Bras, and Y\. Choi \(2021\)CLIPScore: a reference\-free evaluation metric for image captioning\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 7514–7528\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.595)Cited by:[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1),[§6](https://arxiv.org/html/2608.06967#S6.p1.1)\.
- Z\. J\. Hou, B\. A\. Zhang, Y\. Lu, B\. K\. Baghel, A\. Brei, X\. Lu, M\. Jiang, F\. Brahman, S\. Chaturvedi, H\. Chang, D\. Khashabi, and X\. L\. Li \(2025\)CreativityPrism: a holistic evaluation framework for large language model creativity\.External Links:2510\.20091,[Link](https://arxiv.org/abs/2510.20091)Cited by:[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- Y\. Hu, B\. Liu, J\. Kasai, Y\. Wang, M\. Ostendorf, R\. Krishna, and N\. A\. Smith \(2023\)TIFA: accurate and interpretable text\-to\-image faithfulness evaluation with question answering\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 20406–20417\.External Links:[Document](https://dx.doi.org/10.1109/ICCV51070.2023.01866)Cited by:[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1),[§6](https://arxiv.org/html/2608.06967#S6.p1.1)\.
- K\. Huang, K\. Sun, E\. Xie, Z\. Li, and X\. Liu \(2023\)T2I\-CompBench: a comprehensive benchmark for open\-world compositional text\-to\-image generation\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=weHBzTLXpH)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- Y\. Huang, L\. Xie, X\. Wang, Z\. Yuan, X\. Cun, Y\. Ge, J\. Zhou, C\. Dong, R\. Huang, R\. Zhang, and Y\. Shan \(2024\)SmartEdit: exploring complex instruction\-based image editing with multimodal large language models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 8362–8371\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2024/html/Huang_SmartEdit_Exploring_Complex_Instruction-based_Image_Editing_with_Multimodal_Large_Language_CVPR_2024_paper.html)Cited by:[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- D\. A\. Hudson and C\. D\. Manning \(2019\)GQA: a new dataset for real\-world visual reasoning and compositional question answering\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 6700–6709\.External Links:[Link](https://openaccess.thecvf.com/content_CVPR_2019/html/Hudson_GQA_A_New_Dataset_for_Real-World_Visual_Reasoning_and_Compositional_CVPR_2019_paper.html)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- Y\. Lu, D\. Wang, T\. Li, D\. Jiang, S\. Khudanpur, M\. Jiang, and D\. Khashabi \(2024\)Benchmarking language model creativity: a case study on code generation\.External Links:2407\.09007,[Link](https://arxiv.org/abs/2407.09007)Cited by:[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- J\. A\. Olson, J\. Nahas, D\. Chmoulevitch, S\. J\. Cropper, and M\. E\. Webb \(2021\)Naming unrelated words predicts creativity\.Proceedings of the National Academy of Sciences118\(25\),pp\. e2022340118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2022340118)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- M\. Otani, R\. Togashi, Y\. Sawai, R\. Ishigami, Y\. Nakashima, E\. Rahtu, J\. Heikkila, and S\. Satoh \(2023\)Toward verifiable and reproducible human evaluation for text\-to\-image generation\.External Links:2304\.01816,[Link](https://arxiv.org/abs/2304.01816)Cited by:[§6](https://arxiv.org/html/2608.06967#S6.p1.1)\.
- S\. J\. Paech \(2025\)EQ\-Bench creative writing benchmark v3\.GitHub\.Note:[https://github\.com/EQ\-bench/creative\-writing\-bench](https://github.com/EQ-bench/creative-writing-bench)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- M\. A\. Runco and G\. J\. Jaeger \(2012\)The standard definition of creativity\.Creativity Research Journal24\(1\),pp\. 92–96\.External Links:[Document](https://dx.doi.org/10.1080/10400419.2012.650092)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p6.1)\.
- P\. E\. Shrout and J\. L\. Fleiss \(1979\)Intraclass correlations: uses in assessing rater reliability\.Psychological Bulletin86\(2\),pp\. 420–428\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.86.2.420)Cited by:[§H\.6](https://arxiv.org/html/2608.06967#A8.SS6.p1.3),[§6\.1](https://arxiv.org/html/2608.06967#S6.SS1.p1.2)\.
- Y\. Tian, A\. Ravichander, L\. Qin, R\. Le Bras, R\. Marjieh, N\. Peng, Y\. Choi, T\. Griffiths, and F\. Brahman \(2024\)MacGyver: are large language models creative problem solvers?\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Mexico City, Mexico,pp\. 5303–5324\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.297),[Link](https://aclanthology.org/2024.naacl-long.297/)Cited by:[§2\.1](https://arxiv.org/html/2608.06967#S2.SS1.p1.1)\.
- X\. Wu, Y\. Hao, K\. Sun, Y\. Chen, F\. Zhu, R\. Zhao, and H\. Li \(2023\)Human preference score v2: a solid benchmark for evaluating human preferences of text\-to\-image synthesis\.External Links:2306\.09341,[Link](https://arxiv.org/abs/2306.09341)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- Y\. Wu, Z\. Li, X\. Hu, X\. Ye, X\. Zeng, G\. Yu, W\. Zhu, B\. Schiele, M\. Yang, and X\. Yang \(2025\)KRIS\-Bench: benchmarking next\-level intelligent image editing models\.External Links:2505\.16707,[Link](https://arxiv.org/abs/2505.16707)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- Z\. Xia, S\. Sarkhel, M\. Tanjim, S\. Petrangeli, I\. Dasgupta, Y\. Chen, J\. Xu, D\. Liu, S\. Mitra, and D\. N\. Metaxas \(2025\)VISIAR: empower MLLM for visual story ideation\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 18384–18402\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.945),[Link](https://aclanthology.org/2025.findings-acl.945/)Cited by:[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- X\. Zhao, P\. Zhang, K\. Tang, X\. Zhu, H\. Li, W\. Chai, Z\. Zhang, R\. Xia, G\. Zhai, J\. Yan,et al\.\(2025\)Envisioning beyond the pixels: benchmarking reasoning\-informed visual editing\.External Links:2504\.02826,[Link](https://arxiv.org/abs/2504.02826)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica \(2023\)Judging LLM\-as\-a\-judge with MT\-Bench and chatbot arena\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html)Cited by:[§2\.3](https://arxiv.org/html/2608.06967#S2.SS3.p1.1)\.
- B\. Zhou, A\. Lapedriza, A\. Khosla, A\. Oliva, and A\. Torralba \(2018\)Places: a 10 million image database for scene recognition\.IEEE Transactions on Pattern Analysis and Machine Intelligence40\(6\),pp\. 1452–1464\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2017.2723009)Cited by:[§3\.3](https://arxiv.org/html/2608.06967#S3.SS3.p1.2)\.
- J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. Hou \(2023\)Instruction\-following evaluation for large language models\.External Links:2311\.07911,[Link](https://arxiv.org/abs/2311.07911)Cited by:[§1](https://arxiv.org/html/2608.06967#S1.p4.1),[§2\.2](https://arxiv.org/html/2608.06967#S2.SS2.p1.1)\.

## Appendix AFull Task Templates

This appendix gives the operational task templates used to instantiateEkphrasis\. The main paper describes the construct and task families; here we make the elicitation format explicit enough for replication\.

Unified output format\.For every taskttand modelmm, the expected response is a pairym,t=\(cm,t,vm,t\),y\_\{m,t\}=\(c\_\{m,t\},v\_\{m,t\}\),wherecm,tc\_\{m,t\}is a one\-sentenceCore Ideaandvm,tv\_\{m,t\}is a 100–140 wordVisual Description\. The Core Idea states the visual metaphor, carrier, state, or action; it should not spend its budget on color, lighting, or atmosphere\. The Visual Description states only what is literally visible in a single static frame, binding visual attributes such as color, light, texture, and mood to named elements\.

Formally, a task instance is

xt=\(ft,bt,Ct,St\),x\_\{t\}=\(f\_\{t\},b\_\{t\},C\_\{t\},S\_\{t\}\),whereftf\_\{t\}is the task family,btb\_\{t\}is the natural\-language brief,CtC\_\{t\}is an optional set of constraints, andStS\_\{t\}is the family\-specific source structure\. The four task families differ inStS\_\{t\}and in the creative operation they elicit\.

Algorithm 1Controlled task construction forEkphrasis1:Family set

ℱ\\mathcal\{F\}, source pools

𝒫f\\mathcal\{P\}\_\{f\}, target counts

nfn\_\{f\}
2:Task set

𝒯\\mathcal\{T\}
3:

𝒯←∅\\mathcal\{T\}\\leftarrow\\emptyset
4:for all

f∈ℱf\\in\\mathcal\{F\}do

5:Construct subtype grid

𝒢f\\mathcal\{G\}\_\{f\}and cell quotas

6:Sample candidate variables

St∼𝒫fS\_\{t\}\\sim\\mathcal\{P\}\_\{f\}under each cell constraint

7:Compose

xt=\(f,bt,Ct,St\)x\_\{t\}=\(f,b\_\{t\},C\_\{t\},S\_\{t\}\)
8:Keep

xtx\_\{t\}iff it is visual, unambiguous, safe, and non\-duplicate

9:Add validated instances until family quota

nfn\_\{f\}is met

10:endfor

11:Assign stable ids and return balanced task set

𝒯\\mathcal\{T\}

### A\.1Abstraction

Abstraction tasks elicit the operation

ϕabs:a↦v,\\phi\_\{\\mathrm\{abs\}\}:a\\mapsto v,whereaais an abstract target andvvis a concrete visual plan\. The prompt asks the model to express the abstract concept through a visible object, scene, relation, or action rather than through text in the image\.

Template\.Task Type:Abstraction Brief:Design a visual concept that expresses the abstract conceptaa\. Output:Give exactly two labeled blocks: Core Idea and Visual Description\.

The evaluation checks whether the response identifies a concrete carrier for the target concept, whether the carrier maps specifically to the brief rather than to a generic emotion or theme, and whether the final image can be imagined without explanatory text\.

### A\.2Combination

Combination tasks elicit visual fusion:

ϕcomb:\(c1,c2,δ\)↦v,\\phi\_\{\\mathrm\{comb\}\}:\(c\_\{1\},c\_\{2\},\\delta\)\\mapsto v,wherec1c\_\{1\}andc2c\_\{2\}are source concepts andδ\\deltais a semantic distance band\. The intended output is a single integrated visual entity, not a collage containing two unrelated source objects\.

Template\.Task Type:Combination Brief:Fusec1c\_\{1\}andc2c\_\{2\}into one coherent visual entity\. Constraint:Both source concepts must jointly participate in structure, function, or form\.

The distance band controls how much conceptual bridging is required: near pairs test subtle synthesis, medium pairs test nontrivial blending, and far pairs test whether a model can invent a shared visual logic\.

### A\.3Transformation

Transformation tasks elicit object rewriting:

ϕtfm:\(o,r\)↦v,\\phi\_\{\\mathrm\{tfm\}\}:\(o,r\)\\mapsto v,whereoois a source object andrris a transformation condition\. We use three transformation subtypes: material, temporal, and contextual\.

Template\.Task Type:Transformation Brief:Re\-present objectoounder transformation conditionrr\. Constraint:The condition should reshape the object itself rather than merely decorate it or provide a background\.

Material transformations require the new material to alter the object’s surface and structure; temporal transformations require visible state change; contextual transformations require the environment’s rules to rewrite the object’s mode of existence\.

### A\.4Adaptation

Adaptation tasks elicit applied visual problem solving:

ϕadp:\(s,C\)↦v,\\phi\_\{\\mathrm\{adp\}\}:\(s,C\)\\mapsto v,wheressis an applied situation andCCis a set of audience, medium, domain, or communication constraints\. In the implementation, the legacy directory name isadaption; in the paper we useAdaptation\.

Template\.Task Type:Adaptation Brief:Produce a visual deliverable for applied situationss\. Constraints:Satisfy all explicit audience, medium, domain, and communication constraints inCC\.

Unlike the first three families, Adaptation leaves the creative operation less prescribed\. This makes it useful for analyzing strategy choice, but it also makes usefulness more sensitive to constraints\.

## Appendix BDataset Statistics

The released benchmark contains 400 task instances: 120 Abstraction, 90 Combination, 90 Transformation, and 100 Adaptation briefs\.

Families are elicitation operations, not topics\.A task family specifies the operation that the model must perform: abstract\-to\-visual mapping, concept fusion, object rewriting, or constrained applied design\. The same topic can appear in different families, but it tests a different form of VCI depending on the operation\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/figures/app_dataset_taxonomy.png)Figure 5:Task\-family and subtype distribution in the 400\-instanceEkphrasistask suite\.### B\.1Subtype Distribution

Table[5](https://arxiv.org/html/2608.06967#A2.T5)gives the subtype grid used to balance the benchmark\. The combination grid crosses concept type \(AA,AC,CC\) with semantic distance \(near, medium, far\), whereAAdenotes abstract–abstract,ACdenotes abstract–concrete, andCCdenotes concrete–concrete\.

Table 5:Subtype distribution used for controlled elicitation\.FamilySubtype gridCountAbstractionEmotional/psychological, perceptual/qualia, process/temporal, metaphysical, relational/structural, mathematical/logical20 eachCombinationAA\-near,AC\-near,CC\-near8, 8, 8CombinationAA\-medium,AC\-medium,CC\-medium15 eachCombinationAA\-far,AC\-far,CC\-far7 eachTransformationMaterial, temporal, contextual30 eachAdaptationPublic\-interest campaign, commercial communication, digital product, exhibition/public art, educational communication20 eachFigures[6](https://arxiv.org/html/2608.06967#A2.F6)–[8](https://arxiv.org/html/2608.06967#A2.F8)give construction diagnostics for the three families whose sampling depends on lexical or embedding\-space controls\. These diagnostics are included to document the elicitation design rather than to report model performance\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/figures/app_abstraction_source_pool.png)Figure 6:Abstraction source\-pool diagnostic\. Each panel shows the log\-scaled SUBTLEX frequency distribution for nouns in one abstraction subtype; the shaded interval marks the within\-subtype frequency belt used to avoid extremely rare or ubiquitous lexical items\.![Refer to caption](https://arxiv.org/html/2608.06967v1/figures/app_combination_distance_sampling.png)Figure 7:Combination sampling diagnostic\. Concept pairs are stratified by pair type and embedding\-distance band; numbers inside bands indicate the selected task count for each cell\.![Refer to caption](https://arxiv.org/html/2608.06967v1/figures/app_transformation_sampling.png)Figure 8:Transformation sampling diagnostic\. Material, temporal, and contextual transformations each contribute 30 tasks, with equal counts from near, medium, and far percentile buckets\.
### B\.2Prompt Examples

Table[6](https://arxiv.org/html/2608.06967#A2.T6)shows compact examples\. These examples are illustrative; full task JSON files store the same fields using stable ids, family labels, subtype labels, and task properties\.

Table 6:Representative task prompts and intended diagnostic use\.FamilyExample briefWhat the task testsAbstractionExpress “time passing” for a meditation\-app launch image while avoiding clock imagery\.Whether the model finds a concrete carrier for an abstract temporal concept\.CombinationFuse coffee culture with Chinese landscape painting and modern minimal design for a vertical brand poster\.Whether sources become one integrated visual system rather than a collage\.TransformationPresent a familiar object as if its material, temporal state, or context rewrites its form\.Whether the transformation changes object identity and structure\.AdaptationCreate an applied visual deliverable under audience, domain, and medium constraints\.Whether the model chooses a viable strategy under real design constraints\.

## Appendix CModel List and Generation Details

This section reports the model and generation protocol used by the paper\-level evaluation\. Internal scripts support smaller pilot runs, but the appendix follows the main\-paper roster and does not treat pilot\-only outputs as final experimental results\.

### C\.1Model Selection

Subject models are selected to cover proprietary frontier systems, strong open or open\-weight systems, and efficient baselines from several model families\. Models used as evaluation judges are excluded from the subject\-generation roster and are listed in the generation configuration summary\. Table[7](https://arxiv.org/html/2608.06967#A3.T7)lists the display roster used for reporting\. Provider\-specific ids are stored in configuration files and normalized to these display names for the paper\.

Table 7:Subject model roster for the official VCI leaderboard\.\#Display modelRoleNotes1GPT\-5\.4subjectFrontier OpenAI model\.2Claude Opus 4\.7subjectHigh\-end Claude model\.3Gemini 3\.1 Pro PreviewsubjectStrong Gemini\-family model\.4Grok 4\.3subjectProprietary non\-OpenAI comparison\.5DeepSeek V4 ProsubjectStrong low\-cost reasoning/text model\.6Qwen3\.6 Max PreviewsubjectAlibaba/Qwen flagship\-style model\.7GLM 5\.1subjectChinese commercial model family\.8Kimi K2\.6subjectLong\-context Chinese commercial model family\.9Llama 4 MavericksubjectStrong open\-weight Meta model\.10MiniMax M2\.7subjectCommercial model family with reasoning support\.11DeepSeek R1subjectReasoning\-oriented DeepSeek model\.12Qwen3\-8BsubjectSmall open\-weight Qwen baseline\.13Phi\-4subjectCompact model baseline\.14Llama 3\.1 8B InstructsubjectSmall open\-weight Meta baseline\.
### C\.2Decoding Parameters

The implementation uses provider\-specific configuration through OpenRouter\-compatible model entries\. The common generation defaults are shown in Table[8](https://arxiv.org/html/2608.06967#A3.T8)\. When a provider exposes reasoning controls, the configuration records whether reasoning is enabled, but the output format remains identical across models\.

Table 8:Generation and evaluation configuration summary\.ParameterDefault or policyProvider routeOpenRouter\-compatible chat APITemperature0\.7 for subject generationMaximum tokens4096 or 8192 by model familySamples per model–task3 candidate samples; sample 1 is the official primary outputOfficial leaderboard output1 pre\-registered primary response per model–taskTyped Idea Graph poolAll 3 samples per model–taskPrimary judge modelDeepSeek\-V4\-Flash routeConsistency judgesGPT\-5\.4\-mini and Claude\-Sonnet\-4\.6 on fixed sampled pairsAuxiliary samplingFixed\-count sampled pairs; not proportional to the full pairwise poolJSON repair modelDeepSeek\-V4\-Flash routeEmbedding modeltext\-embedding\-3\-large routeRandom seed42 for sampling and BT resampling
### C\.3Response Normalization

Raw model outputs are normalized before evaluation\. The extraction stage uses an LLM\-based splitter that returns strict JSON with two keys:ideaandvisual\_description\. The extractor is instructed to copy wording faithfully, avoid stylistic rewriting, and set missing fields tonull\. The normalized record stores the source model id, task id, raw\-output hash, extraction model, extraction attempt count, and timestamp\.

Normalization invariant\.Usefulness and Novelty are evaluated primarily from the extractedideafield, while Expressiveness is evaluated fromvisual\_description\. This prevents vivid scene prose from inflating task\-fit judgments and prevents brief\-fit reasoning from contaminating expressiveness scores\.

### C\.4Extracted Output Length Diagnostics

Figure[9](https://arxiv.org/html/2608.06967#A3.F9)reports a lightweight format diagnostic for the official extracted responses\. The diagnostic uses the 5,600 primary responses in the clean release \(14 models×\\times400 tasks\), after response normalization and weak\-model extraction audit\. The top row shows the length distribution of the extracted Core Idea field; the bottom row shows the corresponding Visual Description field\. Across task families, Core Ideas are concentrated around the requested one\-sentence length, while Visual Descriptions cluster around the requested 100–140 word range with model\-specific variation\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/x5.png)Figure 9:Extracted output length distributions by task type for the official clean dataset\. Each box summarizes one subject model within a task family\. The top row reports extracted Core Idea length; the bottom row reports extracted Visual Description length\. This figure is a format diagnostic rather than a performance result\.

## Appendix DExperimental Matrix and Validation Budget

Section[4](https://arxiv.org/html/2608.06967#S4)reports the conceptual structure of the evaluation; this appendix fixes the operational scale used by the final automatic evaluation\. The key design choice is that three samples are generated for each model–task pair, but only one pre\-registered primary sample enters the official leaderboard\. The remaining samples are used for population\-cliché discovery in Typed Idea Graphs, not for best\-of\-three selection\. Because the official scorer uses active BT rather than exhaustive pairwise comparison, leave\-two\-models\-out Novelty profiles are built lazily for the model pairs that are actually selected by the Novelty sampler\.

Table 9:Operational scale of the final automatic run\.LayerPurposeScaleReported evidenceSubject generationProduce comparable visual plans and a larger population pool for cliché discovery\.14 models×\\times400 tasks×\\times3 samples = 16,800 outputsOfficial sample id, output hashes, normalization logsOfficial leaderboardScore one fixed response per model–task pair\.14 models×\\times400 tasks = 5,600 primary responsesOverall and dimension\-level BT scoresTyped Idea GraphsBuild task\-specific population cliché profiles for Novelty\.400 task graphs; 42 ideas per task before leave\-out maskingCliché clusters, checklist items, stability diagnosticsLazy leave\-two\-models\-out NoveltyAvoid letting either candidate model define the clichés used to judge it\.Built on demand for Novelty comparisons selected by active BT; expected≤\\leq10,000 contextsNovelty checklist contexts and TIG audit logsPrimary active\-BT judgeRun scalable checklist\-assisted A/B comparison\.400 tasks×\\times3 dimensions×\\times50 judge calls = 60,000 callsBT estimates, confidence intervals, rank tiersOrder\-swap diagnosticsDetect A/B position effects\.Approximately 10% sampled model pairs, about 12,000 callsPosition bias and swap consistencyAuxiliary judge consistencyTest whether rankings depend on DeepSeek\-V4\-Flash\.2 auxiliary judges×\\times3 dimensions×\\times100 fixed sampled pairs = 600 callsAgreement, rank correlation, tie\-rate differencesDuplicate doublecheck diagnosticsEstimate within\-judge repeat stability on sampled pairs\.2 dimensions×\\times200 fixed sampled pairs×\\times2 repeated judgments = 800 callsDuplicate consistency, tie rates, JSON repair ratesNovelty ablationCompare TIG\-anchored Novelty with a generic Novelty checklist on the full text\-level Novelty validation slice\.240 paired Novelty cases; 235 with human\-majority labels; 240 matched TIG judgments plus 240 no\-TIG rerunsHuman alignment and checklist specificityOptional construct probesCheck synthetic failure cases such as fluent\-but\-shallow or cliché\-template responses\.Planned subset only if automated diagnostics suggest a confoundProbe\-specific score behaviorCompleted pairwise judge budgetIncludes primary, swap, auxiliary consistency, and duplicate doublecheck calls\.73,400 completed judge calls before repair overheadBudget accounting and reliability diagnosticsApproximate total chat\-call budgetAdds subject generation, atom extraction, canonicalization, and lazy L2MO profile summaries\.About 118,000 calls, excluding embeddingsBudget planning onlyTable 10:Human validation budget\. Human judgments support calibration, reliability analysis, and cross\-modal validation for the scalable evaluation protocol\.Human validation componentDesignJudgmentsText\-level validation40 tasks×\\times6 pairs×\\times3 dimensions×\\times3 annotators2,160TIG cliché checklist audit300 checklist items×\\times3 annotators900Faithfulness gate40 tasks×\\times6 models×\\times1 image×\\times3 annotators720Image\-level preference40 tasks×\\times4 pairs×\\times1 overall judgment×\\times3 annotators480Total4,260### D\.1Novelty Ablation Details

The completed ablation uses the full text\-level Novelty validation slice: 240 paired Novelty cases from the 40\-task validation package\. Of these, 235 cases have a final human\-majority label\. Both conditions use DeepSeek\-V4\-Flash as judge and preserve the original response order\. The TIG condition uses the existing case\-specific Novelty checklist items materialized from leave\-two\-models\-out Typed Idea Graphs; the no\-TIG condition reruns the same rows after replacing those task\-population anchors with generic Novelty anchors\. All no\-TIG judge calls parsed successfully\. The main readout is reported in Table[4](https://arxiv.org/html/2608.06967#S5.T4)\.

### D\.2Detailed Reliability Diagnostics

The main text reports only a compact reliability summary\. Tables[11](https://arxiv.org/html/2608.06967#A4.T11)and[12](https://arxiv.org/html/2608.06967#A4.T12)give the dimension\-wise diagnostics used to support that summary\.

Table 11:Format and order diagnostics by dimension\.DiagnosticUENInterpretationPrimary JSON repair rate0\.41%1\.01%1\.36%Low parser interventionPrimary tie rate13\.26%9\.90%10\.28%Non\-degenerate tie behaviorA\-position bias, decisive\+9\.9 pp\+11\.8 pp\+3\.9 ppDetectable, strongest for ExpressivenessSwap decisive consistency0\.8200\.8010\.888Stable enough for ranking diagnosticsPrimary–swap BT rank corr\.0\.9740\.9910\.991Leaderboard stable under order reversalTable 12:Auxiliary, duplicate, and length diagnostics by dimension\.DiagnosticUENInterpretationWinner\-longer rate0\.5910\.6680\.564Verbosity association, strongest for ExpressivenessDuplicate consistencyn/a0\.6750\.765Pair\-level repeatability is higher for Novelty than ExpressivenessAux rank corr\. \(GPT\-5\.4\-mini\)0\.5160\.7630\.727Moderate to strong auxiliary agreementAux rank corr\. \(Claude\-Sonnet\-4\.6\)0\.5820\.7930\.780Stronger auxiliary agreementAux JSON parse rate100%100%100%No auxiliary repair required
### D\.3LLM\-Only Pilot and Scaling Criteria

The full evaluation is preceded by two API\-facing pilot runs\. The goal is to test the measurement machinery, not to report final model rankings\. All API\-facing stages are run sequentially by stage but with a global concurrency cap of 100 workers\.

Table 13:LLM\-only pilot scale before the full evaluation\. The smoke stage may use a smaller pairwise budget when it is used only for format and parser checks\.Pilot stageGenerationPrimary judge callsAdditional checksSmoke test8 tasks×\\times4 models×\\times3 samples = 96 outputsSmall sampled or capped pairwise budgetFormat, extraction, L2MO profiles, JSON validityFormal active\-BT pilot20 tasks×\\times14 models×\\times3 samples = 840 outputs20 tasks×\\times3 dimensions×\\times50\-call budget = 3,00010% order swap; two auxiliary judges on fixed sampled pairsThe pilot is considered launch\-ready only if \(i\) Core Idea and Visual Description parsing exceeds 95%, \(ii\) L2MO cliché profiles are available and readable for at least 95% of active\-selected Novelty comparisons, \(iii\) judge JSON repair remains below 2%, \(iv\) tie rates remain in a reasonable range rather than collapsing to forced A/B choices, \(v\) order\-swap position effects are small, \(vi\) length and adjective\-density correlations do not dominate scores, and \(vii\) auxiliary judges produce non\-degenerate rank correlations\. Human validation and cross\-modal validation are launched only after these pilot gates pass\.

## Appendix EEvaluation Prompts

The scalable evaluator uses hidden A/B pairwise comparison\. For every comparison, the judge receives a fixed dimension definition, a dimension\-specific evaluation scope, task context when relevant, the checklist items, and Response A/Response B\. The judge must return strict JSON\.

Judge output schema\.\{ "dimension": "usefulness\|expressiveness\|novelty", "winner": "A\|B\|tie", "confidence": 0\.0, "overall\_rationale": "one or two sentences", "per\_item\_outcomes": \[ \{ "item\_index": 1, "item": "\.\.\.", "winner": "A\|B\|tie", "confidence": 0\.0, "evidence\_a": "\.\.\. or null", "evidence\_b": "\.\.\. or null", "rationale": "\.\.\." \} \] \}

### E\.1Usefulness Checklist

Usefulness measures whether the Core Idea maps specifically to the task brief\. The prompt explicitly instructs the judge not to score novelty or visual density\. Checklist generation uses three fixed slots: brief fit, operation completion, and constraint\-context fit\. A typical item has the form:

> \[Inspects: Core Idea\] Which response maps more specifically to the required brief, rather than offering a generic visual theme?

The usefulness judge receives the task brief because task fit cannot be evaluated without knowing the requested concept, sources, audience, medium, or constraints\.

### E\.2Expressiveness Checklist

Expressiveness measures whether the Visual Description gives enough specific visual information that two readers would imagine nearly the same image\. It is not a task\-fit score and does not reward novelty\. The prompt includes task context before generating checklist items: in Combination it highlights the fused entity, in Transformation the transformed object, and in Adaptation the entire intended deliverable\. Checklist generation then uses three fixed slots: visual grounding, spatial organization, and appearance specificity\. A typical item has the form:

> \[Inspects: Visual Description\] Which description shows more of the actual scene, instead of explaining what it means?

The expressiveness judge is instructed to compare only the response’s own visual language\. This separation is important because an expressive description can still be useless for the brief, and a useful concept can be visually under\-specified\.

### E\.3Novelty Checklist

Novelty measures whether a response avoids obvious, stock, or population\-common visual ideas for the same task\. The prompt uses the brief only as context for what would be obvious\. The first checklist item is replaced by a task\-specific item materialized from the Typed Idea Graph profile\. Checklist generation uses three fixed slots: cliché distance, source\-domain shift, and perspective or relation shift\. This keeps Novelty separate from brief fit and visual vividness\.

## Appendix FBradley–Terry Aggregation Details

Pairwise preferences are aggregated with a Bradley–Terry model\. Letθm\\theta\_\{m\}denote the latent skill of modelmmfor one dimension\. The core quantities are summarized below\.

Bradley–Terry aggregation summary\.

P​\(i≻j\)\\displaystyle P\(i\\succ j\)=σ​\(θi−θj\)=11\+exp⁡\[−\(θi−θj\)\],\\displaystyle=\\sigma\(\\theta\_\{i\}\-\\theta\_\{j\}\)=\\frac\{1\}\{1\+\\exp\[\-\(\\theta\_\{i\}\-\\theta\_\{j\}\)\]\},ℒ​\(θ\)\\displaystyle\\mathcal\{L\}\(\\theta\)=−∑r∈𝒟wr​log⁡σ​\(θar−θbr\)\+λ​∑mθm2,∑mθm=0,\\displaystyle=\-\\sum\_\{r\\in\\mathcal\{D\}\}w\_\{r\}\\log\\sigma\(\\theta\_\{a\_\{r\}\}\-\\theta\_\{b\_\{r\}\}\)\+\\lambda\\sum\_\{m\}\\theta\_\{m\}^\{2\},\\qquad\\sum\_\{m\}\\theta\_\{m\}=0,s​\(i,j\)\\displaystyle s\(i,j\)=\(Σi​i\+Σj​j−2​Σi​j\)​\(1\+ni​j\)−1/2,\\displaystyle=\\left\(\\Sigma\_\{ii\}\+\\Sigma\_\{jj\}\-2\\Sigma\_\{ij\}\\right\)\(1\+n\_\{ij\}\)^\{\-1/2\},W^m\\displaystyle\\widehat\{W\}\_\{m\}=Wm\+0\.5​TmWm\+Lm\+Tm\.\\displaystyle=\\frac\{W\_\{m\}\+0\.5T\_\{m\}\}\{W\_\{m\}\+L\_\{m\}\+T\_\{m\}\}\.
Herewrw\_\{r\}is judge confidence,ara\_\{r\}andbrb\_\{r\}are the preferred and non\-preferred models,s​\(i,j\)s\(i,j\)is the active\-pair score, andW^m\\widehat\{W\}\_\{m\}is the diagnostic win rate with ties as half\-wins\.

### F\.1Pairwise Sampling

For each task, the system first covers the model pool with warm\-start pairs and then allocates remaining budget to uncertain active pairs\.

Algorithm 2Checklist\-assisted pairwise BT evaluation1:Tasks

𝒯\\mathcal\{T\}, models

ℳ\\mathcal\{M\}, dimension

dd, budget

BB
2:Leaderboard for dimension

dd
3:

ℛ←∅\\mathcal\{R\}\\leftarrow\\emptyset
4:for all

t∈𝒯t\\in\\mathcal\{T\}do

5:Select warm\-start and active pairs under budget

BB
6:Randomize A/B order and query checklist judge for dimension

dd
7:Add decisive preferences to

ℛ\\mathcal\{R\}; retain ties for diagnostics

8:endfor

9:Fit BT skills

θ^\\hat\{\\theta\}from

ℛ\\mathcal\{R\}
10:Bootstrap tasks to estimate CIs and rank probabilities

11:returndimension leaderboard from

θ^\\hat\{\\theta\}

### F\.2Tie Handling

The judge may returntie\. Ties are retained for diagnostics, tie rate, raw win rate, Borda\-style summaries, and tie\-as\-half\-win diagnostic win rates\. Only decisive comparisons enter the main BT likelihood\. If a response is empty, an automatic comparison record is generated: non\-empty responses beat empty responses with confidence 1\.0, and two empty responses tie\.

### F\.3Confidence Intervals

Task\-level bootstrap intervals are computed by resampling tasks with replacement, refitting BT on the sampled comparisons, and taking percentiles ofθm\\theta\_\{m\}\. The reported appendix intervals use the 5th, 50th, and 95th percentiles unless otherwise stated\. Rank probability is estimated by drawing

θ~∼𝒩​\(θ^,Σ^\)\\tilde\{\\theta\}\\sim\\mathcal\{N\}\(\\hat\{\\theta\},\\hat\{\\Sigma\}\)and counting the frequency with which each model occupies each rank\.

### F\.4Rank Stability

Rank stability is reported using bootstrap intervals, rank probability mass, and pairwise indistinguishability tiers\. Two neighboring models are treated as indistinguishable when their BT difference is smaller than the corresponding Wald interval\.

## Appendix GTyped Idea Graphs for Population\-Anchored Novelty

Typed Idea Graphs operationalize population\-anchored novelty\. The goal is not to ask whether a response is far from a generic embedding centroid, but whether it avoids the specific ideas that many LLMs converge on for the same task\.

### G\.1Atom Schema

For each Core Idea, the extractor returns 3–7 typed atoms:

Table 14:Typed idea atom schema\.TypeMeaningconceptCentral solution or representational strategy\.motifVisual or semantic element used by the idea\.mechanismTechnique, transformation, or implementation move\.purposeIntended effect, meaning, or communicative function\.audienceTarget audience or user group, only if explicit\.Typed edges connect atoms within a response, such asmotif supports purposeormechanism realizes concept\.

### G\.2Extraction and Canonicalization Prompt

The extraction prompt takes the task brief and Core Idea as input\. It instructs the model to avoid extracting lighting, color, texture, or composition unless those properties are central to the idea\. The canonicalization prompt groups synonymous or near\-synonymous atoms within each type, requiring every input atom to appear in exactly one alias list and prohibiting invented atoms\.

For an atomuu, the canonical record stores id, type, canonical label, aliases, provenance, and optionally an embedding vector\. Provenance links the atom back to response id, model id, task id, and sample index\.

### G\.3Graph Edge Definitions

The graph contains co\-occurrence, same\-type similarity, and extracted typed\-relation edges\. The scoring quantities are:

Typed Idea Graph scoring summary\.

wcooc​\(u,v\)\\displaystyle w\_\{\\mathrm\{cooc\}\}\(u,v\)=max⁡\(0,log⁡p​\(u,v\)p​\(u\)​p​\(v\)−log⁡p​\(u,v\)\),\\displaystyle=\\max\\left\(0,\\frac\{\\log\\frac\{p\(u,v\)\}\{p\(u\)p\(v\)\}\}\{\-\\log p\(u,v\)\}\\right\),wsim​\(u,v\)\\displaystyle w\_\{\\mathrm\{sim\}\}\(u,v\)=cos⁡\(eu,ev\)⋅𝟏​\[cos⁡\(eu,ev\)≥0\.78\],\\displaystyle=\\cos\(e\_\{u\},e\_\{v\}\)\\cdot\\mathbf\{1\}\[\\cos\(e\_\{u\},e\_\{v\}\)\\geq 0\.78\],wtyped​\(u,v,r\)\\displaystyle w\_\{\\mathrm\{typed\}\}\(u,v,r\)=∑ℓ𝟏​\[\(u,v,r\)∈Eℓ\],\\displaystyle=\\sum\_\{\\ell\}\\mathbf\{1\}\[\(u,v,r\)\\in E\_\{\\ell\}\],ModelCov​\(C\)\\displaystyle\\mathrm\{ModelCov\}\(C\)=\|\{m:∃u∈C,m∈prov​\(u\)\}\|\|ℳ\|,\\displaystyle=\\frac\{\|\\\{m:\\exists u\\in C,\\;m\\in\\mathrm\{prov\}\(u\)\\\}\|\}\{\|\\mathcal\{M\}\|\},Cent​\(C\)\\displaystyle\\mathrm\{Cent\}\(C\)=∑u∈Cdegw⁡\(u\)\|C\|​maxu′⁡degw⁡\(u′\),Conv​\(C\)=2\|C\|​\(\|C\|−1\)​∑u<v∈Cmax⁡\(0,cos⁡\(eu,ev\)\),\\displaystyle=\\frac\{\\sum\_\{u\\in C\}\\deg\_\{w\}\(u\)\}\{\|C\|\\max\_\{u^\{\\prime\}\}\\deg\_\{w\}\(u^\{\\prime\}\)\},\\qquad\\mathrm\{Conv\}\(C\)=\\frac\{2\}\{\|C\|\(\|C\|\-1\)\}\\sum\_\{u<v\\in C\}\\max\(0,\\cos\(e\_\{u\},e\_\{v\}\)\),ClicheScore​\(C\)\\displaystyle\\mathrm\{ClicheScore\}\(C\)=0\.45​ModelCov​\(C\)\+0\.35​Cent​\(C\)\+0\.20​Conv​\(C\)\.\\displaystyle=0\.45\\,\\mathrm\{ModelCov\}\(C\)\+0\.35\\,\\mathrm\{Cent\}\(C\)\+0\.20\\,\\mathrm\{Conv\}\(C\)\.

### G\.4ClichéScore Computation

Clusters are formed within same\-type subgraphs using hierarchical Leiden when available and weighted connected components otherwise\. High\-scoring clusters are common across models, central in the task’s idea graph, and internally convergent; these clusters become the population\-anchored cliche candidates\.

### G\.5Leave\-Two\-Models\-Out Scoring

Anti\-leakage rule\.When comparing candidate modelsmam\_\{a\}andmbm\_\{b\}on tasktt, the population reference graph is built fromℳ∖\{ma,mb\}\\mathcal\{M\}\\setminus\\\{m\_\{a\},m\_\{b\}\\\}and uses all three samples from each remaining model\. With 14 subject models, each pairwise comparison uses 36 reference ideas\. Neither candidate model is allowed to define the clichés used to judge the pair\.

Algorithm 3Leave\-two\-models\-out Typed Idea Graph novelty1:Task

tt, model population

ℳ\\mathcal\{M\}, sample pool

YtY\_\{t\}
2:Novelty checklist for each judged model pair

3:for alljudged pair

q=\{ma,mb\}⊂ℳq=\\\{m\_\{a\},m\_\{b\}\\\}\\subset\\mathcal\{M\}do

4:

Yt−q←\{ym′,t,s:m′∉q,s∈\{1,2,3\}\}Y^\{\-q\}\_\{t\}\\leftarrow\\\{y\_\{m^\{\\prime\},t,s\}:m^\{\\prime\}\\notin q,\\;s\\in\\\{1,2,3\\\}\\\}
5:Extract and canonicalize typed atoms from

Yt−qY^\{\-q\}\_\{t\}
6:Build graph

Gt−qG^\{\-q\}\_\{t\}with co\-occurrence, similarity, and typed edges

7:Cluster

Gt−qG^\{\-q\}\_\{t\}and rank clusters by ClicheScore

8:Materialize top clusters as novelty checklist items

9:Score the primary outputs of

mam\_\{a\}and

mbm\_\{b\}with checklist\-BT

10:endfor

This procedure makes novelty population\-anchored but not self\-anchored\. It also makes the novelty target task\-specific: the reference population for a coffee\-brand poster is different from the reference population for a temporal object transformation\.

### G\.6Checklist Materialization Examples

Ranked clusters are converted into natural\-language checklist items that ask whether a response avoids, transforms, or merely repeats the common pattern\. For example, if a task’s reference graph contains a high\-scoring cluster around “plants, leaves, soft green growth” for calmness, the materialized item can ask:

> Which response avoids relying on common plant\-growth imagery unless it transforms that motif into a more specific visual mechanism?

Why not embedding distance alone?Embedding distance measures generic semantic separation, but visual clichés are task\-conditioned\. A sunset silhouette may be semantically appropriate and far from some generic centroid, yet it is clichéd if many models independently use it for the same prompt\. Typed Idea Graphs measure population convergence within the task, not just semantic distance in embedding space\.

### G\.7TIG Checklist Audit Schema

The TIG checklist audit is designed to validate the novelty checklist items themselves before treating them as reliable judge anchors\. The audit does not ask annotators a single vague question such as whether an item is “reasonable\.” Instead, each materialized Novelty checklist item is evaluated along five concrete criteria, using the item text and its underlying cluster/provenance summary\.

Table 15:TIG Novelty checklist audit schema\. Each materialized checklist item is audited as a population\-cliché anchor, not as a direct preference judgment between A and B\.Audit itemQuestionCluster validityDoes the item represent a visual pattern that recurs across multiple models, rather than a one\-off phrasing from a single response?Visual specificityIs the item a visual cliché rather than a language\-expression cliché, abstract rhetorical move, or evaluative phrase?FaithfulnessDoes the item faithfully summarize the original idea cluster without adding objects, mechanisms, or intents not supported by the cluster?Non\-leakageDoes the item avoid leaking content from either A/B candidate response in the leave\-two\-models\-out comparison?Usefulness for judgingDoes the item help a judge decide whether a candidate repeats, transforms, or avoids the population pattern?
- •All criteria use the same codes:pass,partial,fail, andnot\_enough\_evidence\.

The audit record storesaudit\_id,task\_id,comparison\_key,checklist\_item\_id,checklist\_text,cluster\_id,cluster\_summary, the five criterion codes in Table[15](https://arxiv.org/html/2608.06967#A7.T15),overall\_action,revision\_note,annotator\_id, and timestamp\.overall\_actionis one ofaccept,revise, orexclude\. We accept items that pass all criteria or contain only minor partial issues, revise items whose cluster is valid but whose wording is underspecified or not sufficiently visual, and exclude items that fail cluster validity, visual specificity, faithfulness, or non\-leakage\. This audit is independent of pairwise A/B preference annotation; pairwise annotators may flag an item and leave a note if they observe an audit failure during judging\.

## Appendix HHuman Annotation Protocol

Human annotations support calibration and reliability analysis on targeted validation slices\. The main text\-level validation subset contains 40 tasks, stratified as 10 tasks from each family\. For each task, we sample six model pairs and ask annotators to judge Usefulness, Expressiveness, and Novelty separately, with three annotators per item\. This yields 2,160 text\-level human judgments\.

### H\.1Annotator Instructions

Annotators see one task and two anonymized responses\. Model identity is hidden, order is randomized, and a tie option is available\. The interface asks annotators to judge only one dimension at a time:

- •Usefulness:Which Core Idea better satisfies the brief, constraints, audience, medium, and required source concepts?
- •Expressiveness:Which Visual Description gives a more concrete and stable mental image?
- •Novelty:Which response better avoids task\-specific population clichés while remaining meaningful for the brief?

Annotators are instructed not to reward verbosity by itself and not to infer unstated intentions\.

### H\.2Pairwise Annotation Guidelines

The annotation interface exposes one dimension at a time\. Annotators are instructed to treat the checklist as a reference cue rather than a rubric with equal\-weight items\. They must not count how many checklist items each side appears to satisfy and mechanically choose the side with more item\-level wins\. The final label is a single overall A/B/tie preference for the current dimension\.

Table 16:Human pairwise annotation guidelines by dimension\.DimensionPrimary fieldDecision ruleUsefulnessCore IdeaPrefer the idea that more specifically satisfies the brief, constraints, audience, medium, and required visual operation\. Do not reward verbal polish or visual density by itself\.ExpressivenessVisual DescriptionPrefer the description that supports a more concrete, stable, and drawable mental image\. Do not reward task fit, novelty, length, or atmospheric language by itself\.NoveltyCore Idea plus task\-specific cliché referencesPrefer the idea that better avoids or transforms population\-common visual patterns while remaining meaningful for the brief\. Do not reward randomness, obscurity, or off\-brief weirdness\.Ties are allowed only when the current dimension has no stable, explainable winner: both responses are similarly strong, similarly weak, or their advantages cancel out within the current dimension\. Annotators are explicitly told not to use tie as a substitute for uncertainty when one response has a clear dimension\-specific advantage\.

### H\.3Flagging and Notes

The interface includes a flag option for records that require later review\. Flagging is not a fourth preference label and is not a substitute for tie\. When possible, annotators still choose A, B, or tie, and use the note field to explain why the record was flagged\.

Annotators are instructed to flag missing or truncated responses, field–dimension mismatches, ambiguous briefs, checklist/task mismatch, translation problems, suspected A/B duplication or model\-identity leakage, violations of the single\-static\-image protocol, and Novelty checklist items that appear to fail the TIG audit criteria in Table[15](https://arxiv.org/html/2608.06967#A7.T15)\. They are instructed not to flag merely because a case is difficult, because A and B are close, or because they personally dislike a style\. Notes should identify the issue concisely, for exampleTIG audit: non\-leakage concernorvisual field truncated\.

### H\.4Calibration Examples

Calibration examples cover three common borderline cases: useful but clichéd responses, vivid but task\-misaligned responses, and unusual but under\-specified responses\. Annotators first inspect the checklist anchors, then compare A/B responses, and finally select A, B, or tie\. The examples are designed to make the three dimensions separable rather than to teach annotators a single global notion of quality\.

### H\.5Pair Sampling

Pairs are sampled from preliminary DeepSeek\-V4\-Flash rankings but shown to annotators without model identities or scores\. For each task, the six pairs include top\-vs\-bottom, top\-vs\-middle, middle\-vs\-middle, and dimension\-tradeoff comparisons, such as a highly expressive response against a more novel but less polished response\. This makes the validation subset include both obvious and borderline comparisons\.

### H\.6Agreement Metrics

We report raw agreement, tie rate, human\-vs\-judge rank correlation, and pairwise consistency, following standard multi\-rater reliability practice\(Fleiss,[1971](https://arxiv.org/html/2608.06967#bib.bib9); Shrout and Fleiss,[1979](https://arxiv.org/html/2608.06967#bib.bib10)\)\. Ifhi​jh\_\{ij\}is the human preference sign for model pair\(i,j\)\(i,j\)andai​ja\_\{ij\}is the automated preference sign, pairwise agreement is

Agree=1\|𝒫\|​∑\(i,j\)∈𝒫𝟏​\[hi​j=ai​j\],\\mathrm\{Agree\}=\\frac\{1\}\{\|\\mathcal\{P\}\|\}\\sum\_\{\(i,j\)\\in\\mathcal\{P\}\}\\mathbf\{1\}\[h\_\{ij\}=a\_\{ij\}\],with ties counted as agreement only when both sides tie\. Human\-subset BT rankings are compared to automated BT rankings with Spearman correlation and Kendall correlation\.

### H\.7Completed Human Annotation Reliability

Table[17](https://arxiv.org/html/2608.06967#A8.T17)summarizes the completed human annotation results used in this paper\. For the TIG checklist audit, we report only the finalaccept/revise/excludeaction label; the five finer\-grained TIG audit fields are reserved for internal revision rather than main reliability reporting\.

Table 17:Completed human annotation reliability summary\.SubsetItems / judgmentsAgreementConsensusMajority outcomeText\-level pairwise validation720 / 2,160Fleissκ=0\.596\\kappa=0\.59662\.5% unanimous; 97\.6% majorityA=312, B=276, tie=115, no majority=17TIG checklist audit action300 / 900Fleissκ=0\.762\\kappa=0\.76278\.3% unanimous; 99\.0% majorityaccept=86, revise=153, exclude=58, no majority=3Faithfulness gate240 / 720pass/failκ=0\.637\\kappa=0\.637; score ICC=0\.71177\.9% unanimous; 100\.0% majoritypass=179, fail=61Image\-level preference160 / 480Fleissκ=0\.690\\kappa=0\.69073\.8% unanimous; 98\.8% majorityA=74, B=73, tie=11, no majority=2

## Appendix ICross\-Modal Grounding Protocol

Cross\-modal grounding tests whether text\-level VCI scores reflect visual content rather than surface language quality\. The protocol has two stages: a faithfulness gate and image\-level pairwise preference\.

### I\.1Text\-to\-Image Rendering Setup

The grounding subset contains 40 tasks, stratified as 10 from each task family\. We select six representative subject models from the text\-level leaderboard: two top models, two middle models, and two bottom models\. Each textual visual plan is converted into a rendering prompt and passed to a fixed text\-to\-image system with one seed, yielding40×6×1=24040\\times 6\\times 1=240images\. The rendering prompt is constructed from the Core Idea and Visual Description, with no model identity and no evaluation score exposed\. This subset is designed as a targeted grounding check rather than a second full\-scale benchmark\.

### I\.2Faithfulness Annotation

Renderings pass the faithfulness gate only if they preserve the source idea at the level needed for image\-level preference\. The default rule is

Pass​\(I,y\)\\displaystyle\\mathrm\{Pass\}\(I,y\)=𝟏\[Faith\(I,y\)≥4/5\\displaystyle=\\mathbf\{1\}\\bigl\[\\mathrm\{Faith\}\(I,y\)\\geq 4/5∧¬HardFail\(I,y\)\]\.\\displaystyle\\qquad\\land\\neg\\mathrm\{HardFail\}\(I,y\)\\bigr\]\.Hard failures include major semantic drift, missing core object, concept splitting in fusion tasks, object identity loss in transformation tasks, and style/context mismatch in applied design tasks\.

Each image receives three independent faithfulness annotations, for a total of 720 judgments\. We report pass rates by model and by task family, the correlation between faithfulness pass rate and text\-level VCI score, and both faithful\-only and rendering\-failure\-penalized versions of the cross\-modal agreement analysis\.

Faithfulness gate\.Image\-level validation is conducted only on faithful renderings\. This prevents a poor renderer from being mistaken for poor visual ideation by the text model\.

### I\.3Image\-Level Preference Annotation

After filtering, human judges compare images alone\. They do not see the original text, model identity, or text\-level ranking\. The image\-level criteria mirror the VCI dimensions: task usefulness when task context is shown, visual expressiveness in the rendered image, and visible novelty relative to alternatives for the same task\.

For each of the 40 tasks, we sample four image pairs after the faithfulness gate and collect three annotations for one overall visual preference judgment, yielding 480 image\-level preference judgments\. The overall judgment asks which image better realizes a useful, expressive, and non\-clichéd visual plan as a whole; dimension\-specific image judgments are reserved for optional follow\-up analysis\.

### I\.4Agreement and Correlation Metrics

For each model pair\(m,n\)\(m,n\), letΔm​ntext\\Delta^\{\\mathrm\{text\}\}\_\{mn\}be the text\-level BT difference and letpm​nimgp^\{\\mathrm\{img\}\}\_\{mn\}be the image\-level preference rate for modelmmover modelnn\. The primary reported readout is pairwise sign agreement, with rank correlations reserved as exploratory diagnostics rather than headline evidence:

PairAgree\\displaystyle\\mathrm\{PairAgree\}=1\|𝒫\|∑\(m,n\)∈𝒫𝟏\[\\displaystyle=\\frac\{1\}\{\|\\mathcal\{P\}\|\}\\sum\_\{\(m,n\)\\in\\mathcal\{P\}\}\\mathbf\{1\}\\Bigl\[sign\(Δm​ntext\)=sign\(pm​nimg−0\.5\)\]\.\\displaystyle\\quad\\operatorname\{sign\}\(\\Delta^\{\\mathrm\{text\}\}\_\{mn\}\)=\\operatorname\{sign\}\(p^\{\\mathrm\{img\}\}\_\{mn\}\-5\)\\Bigr\]\.The main text reports the agreement counts rather than emphasizing aggregate rank correlations, because the six\-model subset is a targeted validation slice rather than a full image\-level leaderboard\. Cross\-modal agreement is interpreted as convergent validity evidence, not as proof that text\-level evaluation is sufficient on its own\.

## Appendix JAdaptation Strategy Taxonomy

Adaptation tasks are useful for strategy analysis because the prompt does not prescribe a single creative operation\. We therefore use a multi\-label taxonomy to describe how models respond to applied constraints\.

Figure[10](https://arxiv.org/html/2608.06967#A10.F10)summarizes the constraint profile of the released adaptation tasks\. The taskset is balanced across applied domains while retaining variation in the number of explicit requirements and in the difficulty tertile assigned during task construction\.

![Refer to caption](https://arxiv.org/html/2608.06967v1/figures/app_adaptation_constraint_distribution.png)Figure 10:Adaptation constraint diagnostic by applied domain\.### J\.1Strategy Definitions

Table 18:Multi\-label adaptation strategy taxonomy\.StrategyDefinitionConstraint literalizationDirectly turns a written constraint into a visible object, label, or layout device\.Audience reframingChanges subject, tone, or visual language to fit the intended audience\.Medium adaptationUses the affordances of the requested medium, such as poster, app screen, package, or public installation\.Metaphor substitutionReplaces an obvious applied\-design trope with a new metaphorical carrier\.Scenario embeddingPlaces the message inside a specific scene of use rather than a generic symbolic image\.Style transferApplies a recognizable art, branding, or cultural style to the solution\.Risk\-avoidant generic designProduces a safe but low\-specificity design that satisfies surface constraints without a distinctive idea\.
### J\.2Positive and Negative Examples

A positivemedium adaptationexample uses the physical or interaction constraints of the medium as part of the image concept\. A negative example merely says that the design is “poster\-like” or “suitable for an app” without making the medium visible\. A positiveaudience reframingexample changes scale, tone, imagery, or information density for the intended audience\. A negative example only names the audience in the rationale\.

### J\.3Multi\-Label Annotation Schema

Each response may receive zero or more labels\. The annotation record is

zm,t=\{s1,…,sk\},si∈𝒮,z\_\{m,t\}=\\\{s\_\{1\},\\ldots,s\_\{k\}\\\},\\quad s\_\{i\}\\in\\mathcal\{S\},where𝒮\\mathcal\{S\}is the strategy set in Table[18](https://arxiv.org/html/2608.06967#A10.T18)\. Annotators also mark whether the strategy is central or incidental\. A strategy is central if removing it would change the core visual idea; it is incidental if it appears only as surface styling\.

## Appendix KAdditional Results

This section provides appendix\-only result formats and diagnostic templates\. These materials support the main argument but are not needed to understand the benchmark\.

### K\.1Full Model Rankings

Table[19](https://arxiv.org/html/2608.06967#A11.T19)gives the final official 14\-model leaderboard\. Overall VCI is the mean of the three standardized dimension scores\. Expected winrates are BT\-derived pairwise win probabilities against the model population for each dimension\.

Table 19:Full model ranking with standardized BT scores and expected winrates\.RankModelOverallUzzEzzNzzU winE winN win1Gemini\-3\.1 Pro0\.960\.960\.610\.611\.191\.191\.081\.080\.610\.610\.800\.800\.660\.662GPT\-5\.40\.920\.921\.531\.530\.650\.650\.600\.600\.780\.780\.660\.660\.590\.593Claude\-4\.70\.730\.730\.590\.591\.201\.200\.390\.390\.610\.610\.800\.800\.560\.564DeepSeek\-V40\.700\.700\.490\.490\.700\.700\.900\.900\.590\.590\.680\.680\.640\.645Kimi\-K2\.60\.680\.680\.130\.130\.260\.261\.661\.660\.520\.520\.570\.570\.750\.756Grok\-4\.30\.490\.491\.101\.100\.550\.55−0\.17\-0\.170\.700\.700\.640\.640\.470\.477DeepSeek\-R10\.490\.490\.380\.380\.880\.880\.200\.200\.570\.570\.720\.720\.530\.538GLM\-5\.10\.390\.390\.430\.430\.440\.440\.310\.310\.580\.580\.610\.610\.550\.559Qwen3\.6\-Max0\.290\.290\.120\.120\.370\.370\.370\.370\.520\.520\.590\.590\.560\.5610MiniMax\-M2\.7−0\.15\-0\.150\.450\.45−0\.37\-0\.37−0\.52\-0\.520\.580\.580\.410\.410\.420\.4211Qwen3\-8B−1\.22\-1\.22−1\.37\-1\.37−1\.37\-1\.37−0\.92\-0\.920\.250\.250\.150\.150\.360\.3612Llama\-3\.1\-8B−1\.23\-1\.23−1\.95\-1\.95−1\.46\-1\.46−0\.28\-0\.280\.140\.140\.130\.130\.460\.4613Phi\-4−1\.42\-1\.42−1\.46\-1\.46−1\.57\-1\.57−1\.24\-1\.240\.230\.230\.110\.110\.310\.3114Llama\-4\-Mav−1\.64\-1\.64−1\.06\-1\.06−1\.46\-1\.46−2\.39\-2\.390\.310\.310\.130\.130\.140\.14
### K\.2Full Task\-Family Results

Task\-family residuals are reported rather than only raw means\. For dimensionddand familyff, the residual is

Rf,d=s¯f,d−s¯⋅,d\.R\_\{f,d\}=\\bar\{s\}\_\{f,d\}\-\\bar\{s\}\_\{\\cdot,d\}\.Positive values indicate that a family is above the dimension average; negative values indicate that it is below the dimension average\. The main analysis shows that Combination is below average across dimensions, Transformation is relatively strong for Usefulness and Novelty, and Abstraction is especially strong for Expressiveness\.

### K\.3Additional Failure Cases

We use the following diagnostic template for failure cases:

Failure\(y\)=\(\\displaystyle\\mathrm\{Failure\}\(y\)=\(family,dimension,symptom,\\displaystyle\\mathrm\{family\},\\mathrm\{dimension\},\\mathrm\{symptom\},likelycause,evidence\)\.\\displaystyle\\mathrm\{likely\\ cause\},\\mathrm\{evidence\}\)\.Common symptoms include:

- •Cliché\-heavy abstraction:the response uses a common symbol such as a sunset, mirror, plant, or broken object without a task\-specific transformation\.
- •Juxtaposition instead of fusion:combination tasks place two sources next to each other rather than inventing a shared structure\.
- •Surface\-level transformation:the condition decorates an object but does not change its form, material behavior, or context\.
- •Verbose but unstable visualization:the response contains many adjectives, but readers would not converge on the same image\.
- •Constraint omission:adaptation tasks ignore medium, audience, or domain constraints while producing a plausible generic design\.

Appendix use\.These additional results are diagnostic rather than argumentative\. The main paper should remain understandable without reading this appendix; the appendix supplies implementation details, replication targets, and failure\-analysis scaffolding\.

Similar Articles

Can a Language Model Paint?

Hacker News Top

The author explores whether language models can create art through an iterative painting process rather than one-shot generation, building an app that uses a vision-language model to apply strokes one at a time. The experiment highlights the fragility of LLM-generated artefacts and reflects on artistic sincerity.

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

arXiv cs.CL

This paper introduces ArtECulture, a benchmark for culture-conditioned visual emotion understanding in multimodal large language models, covering English, Chinese, and Arabic cultures with balanced Western and non-Western artwork. Evaluations reveal the task remains challenging, and the authors propose a retrieval-augmented framework to inject cultural knowledge into MLLMs.