Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture

arXiv cs.CL Papers

Summary

This paper explores calibrated ambiguity as a generative resource in human communication versus multimodal language models, using the Dixit game to show that AI exhibits ambiguity collapse and lacks cultural references compared to humans.

arXiv:2609.12575v1 Announce Type: new Abstract: Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable. We operationalise this notion of calibrated ambiguity with a task drawn from the parlour game Dixit. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse (i.e., their outputs are over-specified, leaving no room for multiple legitimate interpretations). Unlike human clues, AI-generated clues also exhibit cultural flattening; they almost never make reference to culturally-situated knowledge, even when prompted to use allusion and figurative language.
Original Article
View Cached Full Text

Cached at: 09/14/26, 08:37 AM

# Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture
Source: [https://arxiv.org/html/2609.12575](https://arxiv.org/html/2609.12575)
DOI:[XXXXXXX\.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)Conference:CHI Conference on Human Factors in Computing Systems; TBD; TBDISBN:978\-1\-4503\-XXXX\-X/27/XXCCS:Human\-centered computing Human computer interaction \(HCI\)Cody Kommers[https://orcid.org/https://orcid.org/0009-0007-8985-0085](https://orcid.org/https://orcid.org/0009-0007-8985-0085)Note:Both authors contributed equally to this research\.email:[ckommers@turing\.ac\.uk](mailto:[email protected])Affiliation:The Alan Turing Institute,London,United KingdomMingrui Ye[https://orcid.org/https://orcid.org/0009-0002-8338-8778](https://orcid.org/https://orcid.org/0009-0002-8338-8778)email:[mingrui\.ye@kcl\.ac\.uk](mailto:[email protected])Affiliation:King’s College London,London,United Kingdom,Evelyn Gius[https://orcid.org/https://orcid.org/0000-0001-8888-8419](https://orcid.org/https://orcid.org/0000-0001-8888-8419)Affiliation:Technical University of Darmstadt,Darmstadt,Germany,Daniela Mihai[https://orcid.org/https://orcid.org/0000-0003-3368-9062](https://orcid.org/https://orcid.org/0000-0003-3368-9062)Affiliation:University of Southampton,Southampton,United Kingdom,Hoyt Long[https://orcid.org/https://orcid.org/0000-0002-8562-5426](https://orcid.org/https://orcid.org/0000-0002-8562-5426)Affiliation:University of Chicago,Chicago,United States of America,Zheng Yuan[https://orcid.org/https://orcid.org/0000-0003-2406-1708](https://orcid.org/https://orcid.org/0000-0003-2406-1708)Affiliation:University of Sheffield,Sheffield,United KingdomandDrew Hemment[https://orcid.org/https://orcid.org/0000-0002-0068-5500](https://orcid.org/https://orcid.org/0000-0002-0068-5500)Affiliation:The Alan Turing Institute,London,United KingdomAffiliation:University of Edinburgh,Edinburgh,United Kingdom

2027

###### Abstract\.

Ambiguity is often treated as a bug for AI systems to resolve—but in human communication and culture, ambiguity can also be a generative resource\. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable\. We operationalise this notion ofcalibrated ambiguitywith a task drawn from the parlour gameDixit\. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse \(i\.e\., their outputs are over\-specified, leaving no room for multiple legitimate interpretations\)\. Unlike human clues, AI\-generated clues also exhibit cultural flattening; they almost never make reference to culturally\-situated knowledge, even when prompted to use allusion and figurative language\.

###### Keywords:

ambiguity, culture, calibration, homogenization, collapse, cultural flattening, multimodal language models, interpretive technologies

![Five Dixit cards labelled A to E, with card B, a cloaked figure in the snow, outlined in red as the target. Below, three clues for this board with rubric tags: the human clue Let it go, tagged calibrated and cultural reference; Gemma 3's clue Lost in wintry contemplation, tagged calibrated and no cultural reference; GLM 4.6V's clue A woman in a long blue cloak, tagged overspecified and no cultural reference.](https://arxiv.org/html/2609.12575v1/fig1_teaser_b_three.png)Figure 1\.We use the parlour game Dixit as a paradigm for measuring calibrated ambiguity\. The task requires a “storyteller” to generate a “clue” that ambiguously picks out a target card from a set of distractor—as judged by being correctly selected by some but not all of the other players\. In this example, a human player constructs their clue around an allusion to a song from the movieFrozen—a widely\-known cultural reference\. Our results show how model\-generated clues differ systematically from those given by humans\. In this example, Gemma 3 offers a clue that is calibrated for ambiguity, but describes the image without reference to culturally\-situated knowledge; while GLM\-4\.6V simply provides an unambiguously literal clue\.Five Dixit cards labelled A to E, with card B, a cloaked figure in the snow, outlined in red as the target\. Below, three clues for this board with rubric tags: the human clue Let it go, tagged calibrated and cultural reference; Gemma 3's clue Lost in wintry contemplation, tagged calibrated and no cultural reference; GLM 4\.6V's clue A woman in a long blue cloak, tagged overspecified and no cultural reference\.## 1\.Introduction

In AI research, ambiguity is often treated as a quantity to be minimised\. A long tradition of work on word sense disambiguation\([Navigli, 2009](https://arxiv.org/html/2609.12575#bib.bib22);[Yuan and Strohmaier, 2021](https://arxiv.org/html/2609.12575#bib.bib77)\)and, more recently, benchmarks that test whether language models can resolve ambiguous questions and sentences\([Min et al\., 2020](https://arxiv.org/html/2609.12575#bib.bib5);[Liu et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib6)\)share the premise that a system succeeds when ambiguity is removed\. However, ambiguity is not just defined by semantic vagueness; it can be a vital aspect of human communication and culture\([Empson, 1930](https://arxiv.org/html/2609.12575#bib.bib23);[Piantadosi et al\., 2012](https://arxiv.org/html/2609.12575#bib.bib24);[Sennet, 2023](https://arxiv.org/html/2609.12575#bib.bib36);[Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74)\)\.

For example, ambiguity plays a role in narrative, where explicit exposition of themes or insights tends to trivialise a story \(e\.g\., “telling” rather than “showing”\([Booth, 1961](https://arxiv.org/html/2609.12575#bib.bib25)\)\), and where the gaps a text leaves open are precisely what a reader must fill in to make it meaningful\([Iser, 1978](https://arxiv.org/html/2609.12575#bib.bib26)\)\. It plays a role in interpersonal conversation and pragmatic implicature, in which people gesture towards meaning rather than state it outright to avoid causing undue offence or tension\([Grice, 1975](https://arxiv.org/html/2609.12575#bib.bib27);[Brown and Levinson, 1987](https://arxiv.org/html/2609.12575#bib.bib70);[Pinker et al\., 2008](https://arxiv.org/html/2609.12575#bib.bib68)\)\. Careful management of ambiguity can play a role in many deliberative processes: groups frequently reach agreement on what to do while leaving the underlying principles unresolved\([Sunstein, 1995](https://arxiv.org/html/2609.12575#bib.bib71)\)\. When juries are faced with incomplete information, their collective job is to reach a conclusion even when ambiguity remains\([Hastie et al\., 1983](https://arxiv.org/html/2609.12575#bib.bib72)\)\.

Calibrating ambiguity is crucial to designing and evaluating generative AI as a cultural technology\([Farrell et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib73);[Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74)\), as well as an important consideration for Interpretive Technologies: AI systems designed to navigate cultural complexity\([Hemment et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib59)\)\. AI systems are increasingly deployed for tasks that do not admit of straightforward ground\-truth answers\([Leibo et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib75)\)\. Evaluating the outputs of these systems requires not just binary judgments of correct and incorrect—but interpretive acts of the kind traditionally performed by humanities scholars\([Klein et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib66);[Underwood, 2025](https://arxiv.org/html/2609.12575#bib.bib34);[Hemment et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib59);[Beguš, 2025](https://arxiv.org/html/2609.12575#bib.bib41)\)\. These interpretations account for ambiguity directly, specifically by accounting for a multiplicity of conflicting but legitimately held perspectives, rather than prioritising a canonical ground\-truth answer\([Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74)\)\. Accordingly, a significant risk of generative AI systems is ambiguity collapse\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\): whereas most real\-world situations or outcomes can be described by multiple legitimate interpretations, AI interfaces often offer a single, authoritatively\-presented judgment\([Metzger et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib35)\)\.

The challenge is that, by definition, something that is ambiguous resists a unitary ground\-truth evaluation\([Sennet, 2023](https://arxiv.org/html/2609.12575#bib.bib36)\)\. In its most general form, ambiguity simply means that a word, statement, image, or other cultural artefact can lend itself to multiple interpretations or meanings\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)\. But there can be many shades of what this looks like\. For example, Clint Eastwood’s 1992 filmUnforgivenis characterised by moral ambiguity: each of the characters is trying to do the right thing, but ends up transgressing nonetheless\. Who is the archetypal “good guy” the audience is supposed to root for? By contrast, Roy Lichtenstein’s use of styles from graphic novels is also ambiguous\. Are these works celebrating comic book art, or mocking it? There is no unified way to operationalise ambiguity comprehensively across even just these two cases—except to say that the artefacts in question admit of multiple interpretations\.

The calibration of ambiguity in AI systems is therefore fundamentally a design problem\([Gaver et al\., 2003](https://arxiv.org/html/2609.12575#bib.bib65);[Di Lodovico et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib44)\)\. We want AI systems that do more than just minimise ambiguity under any circumstances—but when, where, and how can ambiguity be constructively or usefully employed? The answer is context dependent: it depends on the people involved, their background, the situation they are in, and what they are trying to do\([Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74)\)\. While there is no one\-size\-fits\-all level of ambiguity that works in all cases for all people, appropriate management of ambiguity is especially important in settings where tasks are culturally\-situated\([Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74);[Heuser, 2025](https://arxiv.org/html/2609.12575#bib.bib29);[Veselovsky et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib33);[de Rooij and Biskjaer, 2026](https://arxiv.org/html/2609.12575#bib.bib30);[Montesinos and Løvlie, 2026](https://arxiv.org/html/2609.12575#bib.bib62);[Yadav et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib60);[Liu et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib54);[Bergman et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib48);[Sorensen et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib46);[Ge et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib45);[Roland et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib40)\)\. For example, AI systems are increasingly used to shape or produce stories and other other cultural artefacts\([Gupta et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib64)\)\. Such narratives provide a crucial structure for how people make sense of themselves and the world around them\([Kommers and DeDeo, 2025](https://arxiv.org/html/2609.12575#bib.bib63);[Bruner, 1990](https://arxiv.org/html/2609.12575#bib.bib61)\)\. At scale, how AI shapes narratives and cultural artefacts can therefore have a significant impact on identity and meaning\-making\([Kommers and Holtzman, 2026](https://arxiv.org/html/2609.12575#bib.bib69);[Kommers et al\., 2026b](https://arxiv.org/html/2609.12575#bib.bib58);[Zhi et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib39)\)\. But what it looks like for a cultural artefact \(such as a story\) to be appropriate within a given cultural frame depends not just on checking specific demographic boxes, but navigating and managing tradeoffs among intrinsically ambiguous concepts\([Zhou et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib28)\)\. Likewise, determining what kinds of stories count as meaningful to a given person or community depends not just on raw engagement metrics, but by accounting for the ambiguities inherent in contextually\-situated perspectives\([Kommers et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib67);[Lowe et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib57)\)\. These considerations are notoriously difficult to quantify\([Leibo et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib75)\)\.

In this paper, we present a paradigm for measuring the calibration of ambiguity in multimodal AI systems\. Whereas “disambiguation” has a clear operational definition\([Navigli, 2009](https://arxiv.org/html/2609.12575#bib.bib22);[Yuan and Strohmaier, 2021](https://arxiv.org/html/2609.12575#bib.bib77);[Liu et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib6);[Tsai et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib43)\), an evaluation framework for calibrated ambiguity is still needed\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)\. We develop a task aimed at codifying the calibrated ambiguity and provide initial answers to the following questions: How well can generative AI systems calibrate ambiguity? What strategies do models rely on when managing ambiguity? How do these strategies differ from those of humans? And how might we calibrate ambiguity for a particular context or use\-case?

### 1\.1\.Summary of Contributions

We present a novel methodology for measuring calibrated ambiguity in multimodal language models, based on the popular parlour game*Dixit*\(Section[3](https://arxiv.org/html/2609.12575#S3)\)\. We develop a novel annotation rubric to distinguish between different dimensions of ambiguity relevant to the task \(Section[4](https://arxiv.org/html/2609.12575#S4)\)\. We find that humans and models tend to calibrate ambiguity in fundamentally different ways—with humans more likely to make use of culturally\-situated knowledge \(Section[5](https://arxiv.org/html/2609.12575#S5)\)\. In a model comparison, we show that there are significant differences between models in their default behaviour of how ambiguity is managed and show that prompting can alleviate miscalibration to a degree \(specifically by emphasising a “social cognition” framing of the task\), but cannot fully prevent ambiguity collapse \(Section[6](https://arxiv.org/html/2609.12575#S6)\)\.

## 2\.Related Work

Previous work by Gur\-Arieh, Wang, and Fazelpour\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)establishes a taxonomy of epistemic risks posed by ambiguity collapse in large language models \(LLMs\)111We use ’LLM’ rather than ’MLM’ or ’VLM’ in the text to reflect the common usage in this discussion, contextualising ambiguity collapse as a risk of generative AI systems broadly—though it should be remembered that the models we use are not only large, but multimodal\.; we explicitly aim to build on their work\. Gur\-Arieh et al\.\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)distinguish between three categories of risk associated with ambiguity collapse—process, output, and ecosystem\. In this paper, we focus specifically on output collapse: the degree to which systems produce artefacts that support a plurality of legitimate interpretations\. The concern about ambiguity collapse articulated by Gur\-Arieh et al\.\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)reflects a larger concern about the homogenising force exerted by LLMs in collapsing or flattening culturally\-situated outputs\([Heuser, 2025](https://arxiv.org/html/2609.12575#bib.bib29);[Xie and Xie, 2026](https://arxiv.org/html/2609.12575#bib.bib56);[Xiao et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib50);[Wilkens, 2026](https://arxiv.org/html/2609.12575#bib.bib38);[Jain et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib32)\)\.

### 2\.1\.Measuring ambiguity

While there is no established general measure for evaluating a computational system’s capacity to calibrate ambiguity, previous work attempts to evaluate whether systems recognise or resolve multiple \(reasonable\) interpretations of an input\([Min et al\., 2020](https://arxiv.org/html/2609.12575#bib.bib5);[Liu et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib6);[Liu et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib6);[Nam et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib8)\)\. For example AmbigQA, for instance, evaluates whether question\-answering systems recover answers corresponding to different interpretations of ambiguous questions\([Min et al\., 2020](https://arxiv.org/html/2609.12575#bib.bib5)\), while the benchmarkAmbiEnttests whether language models can represent alternative readings of ambiguous linguistic inputs\([Liu et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib6)\)\. Wildenburg et al\.\([Wildenburg et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib7)\)similarly provide a dataset for testing whether language models recognise semantic underspecification; Nam et al\.\([Nam et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib8)\)extend ambiguity evaluation to multimodal models by examining whether visual context enables disambiguation; and Karim et al\.\([Karim et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib37)\)show how models can be calibrated for uncertainty\. Previous work has also investigated how artists have utilised models as sources of ambiguity in their works\([Tsai et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib43)\)\.

However, few approaches position ambiguity as an explicitly desirable element of model output\. The closest line of research is perhaps “perspectivist” annotation, premised on the observation that in many datasets disagreement is desirable or expected rather than something to be eliminated\([Basile et al\., 2021](https://arxiv.org/html/2609.12575#bib.bib55)\)\. For example, recent work has taken this perspectivist approach to label disagreement in irony\([Frenda et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib52)\)or offensive language\([Kim et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib51)\)\. However, ambiguity is only one potential cause of disagreement\([Sandri et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib47)\); others include bias and noise\([Uma et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib53)\), or misinterpretation and deficient definitions\([Gius and Jacke, 2017](https://arxiv.org/html/2609.12575#bib.bib42)\)\. Other annotation work use variability as a proxy for ambiguity; for example, by using human annotation patterns to capture the appropriate amount of disagreement across interpretations\([Dumitrache et al\., 2019](https://arxiv.org/html/2609.12575#bib.bib9);[Nie et al\., 2020](https://arxiv.org/html/2609.12575#bib.bib10);[Marchal et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib49)\)\. In this vein, Pavlick and Kwiatkowski\([Pavlick and Kwiatkowski, 2019](https://arxiv.org/html/2609.12575#bib.bib11)\)examine the shape of human judgments by testing for multiple modes within a distribution\. Other work uses graded human ratings of ambiguity or indirectness to assess models’ ability to integrate visual context for intent disambiguation\([Nam et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib8)\)or uses semantic entropy to quantify uncertainty across semantically distinct model outputs\([Farquhar et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib12)\)\. Together, this work gestures towards the way desirable variability, disagreement, or multiplicity has been evaluated in generative AI systems—specifically by using a human sample for reference in defining what form that variability ought to take\.

### 2\.2\.Dixit\-based tasks

In an effort to assess model performance on tasks requiring complex reasoning or which do not have a single right answer, researchers have increasingly turned to strategic games as a useful testbed\([Lin et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib76)\)\. One game that has attracted significant attention in this respect is the card game Dixit\. Kunda and Rabkina\([Kunda and Rabkina, 2020](https://arxiv.org/html/2609.12575#bib.bib13)\)identified this task of “creative captioning” as a novel challenge for AI requiring the integration of visual perception, natural language understanding, and social reasoning\.

Existing computational work on Dixit has primarily treated the game as a paradigm for vision–language association\. Iwata et al\.\([Iwata et al\., 2020](https://arxiv.org/html/2609.12575#bib.bib14)\)developed an AI player to model the image–word associations involved in Dixit and found that the AI performed equally or worse than human players in clue generation, and worse in card selection and voting\. Vatsakis et al\.\([Vatsakis et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib15)\)focused on the voting task, using human data, natural language processing methods, and Internet search to predict the storyteller’s card from a deliberately vague hint\. Wei\([Wei, 2023](https://arxiv.org/html/2609.12575#bib.bib16)\)subsequently approached card–hint matching with a neural network designed to learn visual concepts from natural language, while Chang\([Chang, 2024](https://arxiv.org/html/2609.12575#bib.bib17)\)used the same neural network for hint selection\. Finally, Tanzawa et al\.\([Tanzawa et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib21)\)operationalise and quantitatively measure card\-hint relevance, which they explicitly connect to the problem of producing appropriately ambiguous hints\. They evaluate the relevance of hints for individual cards, comparing three human hint creators with one AI system, and find the lowest median relevance score and the greatest variability for the latter\.

These studies of Dixit have ranged widely in their scale and size\. On one end, the controlled human evaluation studies are very small\-scale—one 15\-turn game\([Iwata et al\., 2020](https://arxiv.org/html/2609.12575#bib.bib14)\)and the analysis of only 20 cards\([Tanzawa et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib21)\)\. At the other extreme, evaluations in Vatsakis et al\.\([Vatsakis et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib15)\), Wei\([Wei, 2023](https://arxiv.org/html/2609.12575#bib.bib16)\), and Chang\([Chang, 2024](https://arxiv.org/html/2609.12575#bib.bib17)\)rely on an archival dataset of approximately 116,000 human\-played rounds of the game rather than newly collected experimental judgments\.

## 3\.Method

### 3\.1\.Dixit Paradigm

Dixit is a parlour game designed around calibrated ambiguity\. It is a multiplayer game featuring a deck of Dixit\-specific cards, each with a surreal image of an object, person, or situation that could admit of a range of interpretations\. Each round, one player acts as the “storyteller”: they pick a card from their hand and offer a “clue” related to the image\. Based on the clue, the other players pick an image from their own hands which they think could be described by the clue; all cards are then shuffled together, and the non\-storyteller players vote for the card they believe was the storyteller’s\. Crucially, the storyteller wins the round only if*some but not all*of the other players choose the image they put into the middle\. The optimal clue is therefore one based on calibrated ambiguity: not so vague as to be uninterpretable, not so specific as to obviously map to the target image\.

### 3\.2\.Existing dataset of human\-generated clues

The dataset of human responses used in this study is based on the Dixit dataset introduced by Vatsakis et al\.\([Vatsakis et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib15)\)and made publicly available by the authors\.222[https://www\.spronck\.net/datasets/Dixit\_AI\_data\.zip](https://www.spronck.net/datasets/Dixit_AI_data.zip)The dataset was constructed from games played on the online platformboiteajeux\.netbetween July 2012 and September 2021 and is restricted to games using the 84 cards from the Dixit base game\. As described by Vatsakis et al\.\([Vatsakis et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib15)\), the original data were manually filtered to improve the quality and consistency of the textual information\. Rounds in which the hint consisted solely of a web address were removed, as were rounds containing non\-English hints\. Exceptions were retained for commonly recognized foreign\-language expressions \(e\.g\.,carpe diem\), titles of well\-known works \(e\.g\.,Le Petit Prince\), and names referring to franchises, people, or landmarks\. Web links appearing in storyteller explanations were also removed\. Following this preprocessing, the resulting dataset contained 116,226 rounds with the storyteller’s hint, the cards submitted during that round, the identity of the storyteller’s card, and the votes received by each card\. We use the dataset as released and did not alter its content; clue texts are reproduced verbatim, up to whitespace normalisation\.

##### Integrity checks and study sample

Before adopting the dataset we verified that the release matches its published description: it contains 116,226 rounds, and a storyteller’s post\-round explanation is present for 26% of them\. Because our boards require the five\-player format \(one target, four decoys, four guessers\), we restricted the data to the 19,407 five\-player rounds and checked the structural consistency of each: a round was retained only if its board comprised exactly five distinct cards from the base\-game deck, exactly one card was marked as the storyteller’s and agreed with the dataset’s target field, and the vote record was complete, with votes summing to the four guessers\. Only six rounds failed these checks, and every retained clue is English\-language text, consistent with the authors’ preprocessing\. We selected a subset of 350 rounds sampled with fixed random seeds, excluding duplicate clue texts; this was used in as the set of boards \(Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px1)\) for use in the various analyses detail in this paper\.

### 3\.3\.LLM\-generated Dixit Clues and Judgments

We sought to compare human\-generated clues with those generated by models\. As storytellers, vision–language models produced clues for the same cards that human storytellers had played, so that human\- and model\-generated clues can be compared on identical boards \(Table[1](https://arxiv.org/html/2609.12575#S3.T1)\)\. As judges, LLMs applied the ambiguity coding rubric of Section[4](https://arxiv.org/html/2609.12575#S4)to clues, following the protocol applied by human coders\.

##### Boards

All clues were generated for the same set of 350 five\-player rounds drawn from the human dataset \(see Section[3\.2](https://arxiv.org/html/2609.12575#S3.SS2)for selection criteria and integrity checks\)\. Each round supplies the human storyteller’s target card, the four decoy cards that the other players actually submitted, the human storyteller’s own clue, and the votes it received\. The set comprises 200 rounds sampled at random with a fixed seed \(the*baseline*boards\) and the 150 rounds whose human clues were hand\-coded in Section[4](https://arxiv.org/html/2609.12575#S4), so that every coded human clue has an LLM counterpart on the same board\.

##### Model selection

We used six open\-weight vision–language models as storytellers \(Table[1](https://arxiv.org/html/2609.12575#S3.T1)\)\. Three considerations guided the selection\. First, every model is open\-weight and was served locally in bfloat16 on a single 80 GB GPU, so that generation ran under our own system prompt and sampling settings, with no API layer in between\. Most models were roughly 30B total parameters, which is approximately the largest tier that fits one accelerator without quantization\. Smaller models were not of interest \(with the exception of GLM 4\.6V Flash \[9B\], to have one example of a smaller model\) because we wanted to focus systems that performed the task well, and larger ones could not be run under the same conditions\. Second, the six models come from five developers and span both dense and mixture\-of\-experts \(MoE\) architectures\. An MoE model routes each token through a few of many expert sub\-networks\([Fedus et al\., 2022](https://arxiv.org/html/2609.12575#bib.bib83)\), and analyses of open MoE models find their experts more specialized and less polysemantic than dense feed\-forward layers\([Lo et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib82)\)\. Whether such internal differences surface in generation behaviour, for instance in how specific or how varied a clue is, has not been studied; the contrast lets us ask whether architecture predicts calibration \(Section[6\.2\.1](https://arxiv.org/html/2609.12575#S6.SS2.SSS1)\)\. Third, within these constraints the lineup is not size\-matched: total parameters range from 9B to 35B and active parameters from 3B to 27B\. GLM\-4\.6V\-Flash \(9B\) is the only member of its family small enough to serve locally \(the full GLM\-4\.6V has 106B parameters\), and Kimi\-VL is released at 16B only\. We kept both because family diversity mattered more to us than exact size matching, and because size can then be examined across the lineup rather than held constant \(Section[6\.2\.1](https://arxiv.org/html/2609.12575#S6.SS2.SSS1)finds that neither total nor active parameter count orders the models\)\.

Table 1\.Storyteller models \(“active”: parameters used per token in MoE models\)\.
##### Prompts

Each model was run under two prompts\. The*minimal*prompt stated only the rules of the game and its scoring goal—that some but not all of the other players should identify the target—together with a length bound of three to six words and the instruction to respond with the clue alone\. The*social cognition*prompt added additional emphasis: it asked the model to consider how the other players will interpret the clue, to prefer figurative or abstract language that could apply to several cards, and to aim for a clue that exactly two of the four other players would resolve\.

The contrast allowed us to ask whether explicitly prompting a model to reason about other minds changes the kind of ambiguity it produces\([Strachan et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib18)\)\. Because the social cognition prompt bundled several instructions—audience modelling, a preference for figurative or abstract language, and a numerical target—the comparison tested the prompt variations as a whole rather than any one component\. Some of the clues coded by the human panel \(Section[5\.1](https://arxiv.org/html/2609.12575#S5.SS1)\) were generated under an earlier wording of the social cognition prompt that differed only in its closing instruction; the panel rated the two wordings indistinguishably, and they are not distinguished below\. Additional details of prompting protocols are available in Appendix[A](https://arxiv.org/html/2609.12575#A1)\.

##### Clue generation

Clues were sampled at temperature0\.90\.9with a fixed random seed, with extended “thinking” disabled for models that support it, one clue per round for each model–prompt combination\. The clue was extracted from the response by stripping markup and surrounding quotation marks\. Clues were never truncated: if a generation exceeded the length bound, the model was resampled up to five times and the shortest complete attempt was kept, so that clue length is a property of the model rather than of post\-processing\. \(Human clues were likewise unbounded in length; GLM\-4\.6V\-Flash exceeds seven words in17%17\\%of rounds, against10%10\\%of human clues\.\) This yields350×2350\\times 2prompts=700=700clues per model and4,2004\{,\}200LLM clues in all\.

To match the informational position of a human storyteller at the moment of clue\-giving, a model saw*only*its own target card—never the other players’ cards—and was told that four other players would each contribute a decoy before all five cards were shuffled on the table\. \(Note: human storytellers chose which card to play, often with a clue already in mind, whereas the models produced a clue for a card it did not choose\)

##### Judgments: LLMs as rubric coders

To compare producers at this scale we used LLMs to employ the coding rubric of Section[4](https://arxiv.org/html/2609.12575#S4), applied by LLM coders under the same protocol as the human annotators\([Dunivin, 2025](https://arxiv.org/html/2609.12575#bib.bib19);[Xiao et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib78)\)\. Each item was presented exactly as a human coder saw it: the clue and the five candidate card images \(labelled A–E, in the order the human coders saw them\), with the rubric as the system prompt\. The coder first selected the image it believed to be the target and then rated the clue on all four dimensions\. Each item was a single independent API call at temperature00, and the coder never saw the answer key, whether the clue was human\- or LLM\-generated, or the study hypotheses\. We used two coders from different model generations, Gemini 3\.1 Flash\-Lite and Gemini 3\.5 Flash\-Lite\. Every clue—the4,2004\{,\}200LLM clues and the 350 human clues on the same boards—was rated independently by both coders, giving9,1009\{,\}100ratings in total; unless stated otherwise, the analyses below use the mean of the two coders per clue\.

Before scaling up, both coders were checked against the 150 clues that the five\-annotator human panel coded \(78 human and 72 LLM clues, Section[4](https://arxiv.org/html/2609.12575#S4)\)\. Per\-clue coder ratings correlate with the human panel mean atr=0\.60r=0\.60–0\.790\.79across the four dimensions \(Calibration0\.620\.62–0\.680\.68, Situatedness0\.770\.77–0\.790\.79, Literalness0\.690\.69–0\.700\.70, Figurativeness0\.600\.60–0\.630\.63\), which matches the agreement of an individual human annotator with the rest of the panel \(leave\-one\-outr=0\.58r=0\.58–0\.800\.80on the same dimensions\)\. Exact agreement with the rounded panel mean ranges from0\.550\.55\(Figurativeness\) to0\.800\.80\(Situatedness\)\. LLM coders thus approximate the human panel about as well as individual human annotators do, on the subset the panel coded; on that basis we use their ratings for the model comparison of Section[6\.2](https://arxiv.org/html/2609.12575#S6.SS2)\(full comparison in Section[6\.1](https://arxiv.org/html/2609.12575#S6.SS1)\)\.

An earlier version of the pipeline scored clues behaviourally, by having a panel of small vision–language models vote for the target card as a Dixit audience would\. We retain it as a robustness check \(Appendix[B](https://arxiv.org/html/2609.12575#A2)\); it reproduced the population\-level human outcome but tracked individual clues only weakly; accordingly, we used the coding rubric as our primary instrument of assessment\.

## 4\.Ambiguity Coding Rubric

We developed a novel coding rubric to account for the different ways of generating and managing ambiguity employed by humans and models when producing Dixit clues\([Chinh et al\., 2019](https://arxiv.org/html/2609.12575#bib.bib20)\)\. The coding scheme feature four dimensions, each scored on a 3\-point ordinal scale, which are summarised in Table[2](https://arxiv.org/html/2609.12575#S4.T2)\.*Calibration*asks whether the clue is overly specific, too vague, or just right\.*Situatedness*asks what degree of cultural knowledge is needed to interpret the clue\.*Literalness*and*Figurativeness*ask, respectively, how literally and how figuratively the clue interprets the image; they are coded independently, since a single clue can be high on both \(see Section[4\.2](https://arxiv.org/html/2609.12575#S4.SS2)\)\.

These dimensions were chosen to correspond directly to our hypotheses or considerations of interests, rather than to provide an exhaustive taxonomy of the different possible textures of ambiguity\. Specifically, we wanted to investigate the following considerations:

1. \(1\)A direct test of ambiguity collapse\. This was imperfectly captured by raw accuracy scores \(i\.e\., do some but not all judges guess the correct card?\)\. In pilot studies, we assessed “collapse” as it is done in the original game\. The problem is that a lot of false positives occurred when the clue was completely ambiguous: guessing results in at least one judge getting the correct answer by chance\. This is probably not a major issue when humans play the game in a social context\. But it was insufficiently sensitive to detect when models simply failed to do the task\. We therefore defined the calibration dimension as a more targeted instrument for assessing collapse\.
2. \(2\)A direct test of cultural flattening\. In our informal initial surveying of human versus model clues, it was apparent that humans were making use of a cultural references in a way models were not\. This is consistent with accounts of cultural collapse and homogenisation of model outputs\([Heuser, 2025](https://arxiv.org/html/2609.12575#bib.bib29);[Jain et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib32)\)\. We developed the situatedness dimension, using this term from the humanities\([Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74);[Haraway, 1988](https://arxiv.org/html/2609.12575#bib.bib31)\)to capture the degree to which specialised cultural knowledge was needed to interpret the clue\.
3. \(3\)The degree to which a clue was constructed by reference to concrete visual features in the picture—or using more abstract, figurative devices\. Informally, it seemed that model\-generated clues were more likely to pick out concrete visual features from an image\. We explored several iterations on how best to capture this consideration \(see Section[4\.2](https://arxiv.org/html/2609.12575#S4.SS2)\)\.

Table 2\.The ambiguity coding rubric\. Each clue is rated on four dimensions, each on a 3\-point ordinal scale\.### 4\.1\.Procedure

Five of the authors of this paper acted as expert annotators\. Before conducting the full set of annotations, a session was held with a majority of the annotators to review specific examples within the data set\. The first 15 clues were discussed in depth, as well as several instructive examples that could be considered edge cases\. The session was recorded and shared with annotators who were unable to attend\.

Annotators viewed each clue in the context of a full 5\-image set \(target \+ four distractors\)\. They were blind to the answer key \(i\.e\., which was the correct target image\) as well as to whether the clue was human\- or model\-generated\. Annotators first selected which image they believed to be the target, then rated each clue on all four dimensions\.

A fifth rating was included, offering notes on any externally verifiable referents for each clue\. For example, the clue “Colors of the wind…” might be interpreted as being about literal, feature\-based ambiguity—unless the interpreter knows that this is a reference to a song from the moviePocahontas\. In this case, the clue does not refer to wind or colour, but to the depiction of a woman who \(allegedly\) resembles this film character\. A single rater used a combination of Gemini and Google search to verify whether each clue had an external referent\. Some were ambiguous, but many made clear reference to some specific cultural artefact\. We decided it was better to try to control for this knowledge by noting these references, rather than to leave it up to the prior cultural knowledge of the raters\.

### 4\.2\.Rubric iterations

We went through several iterations of the rubric\([Chinh et al\., 2019](https://arxiv.org/html/2609.12575#bib.bib20)\)\. In an initial version, we featured a category called “mechanism,” with four nominal labels distinguishing between whether the clue’s ambiguity was primarily based in a feature, a concept, figurative language, or an indexical reference\. However, many clues had multiple aspects of these\. In a subsequent version, we developed one dimension meant to capture whether the ambiguous language was primarily literal \(e\.g\., a vague description of visual features\) versus figurative \(e\.g\., an allusion to a cultural reference\)\. However, many clues featured aspects of both\. For example, consider the clueDavid and Goliathintended to pick out a card with figures of two difference sizes\. This is clearly figurative \(rated 3/3\)\. But the figurative language is used in reference to concrete visual features in the image; thus it also has a non\-negligible degree of literalness as well \(rated 2/2\)\. It may have been possible to define an exact trade\-off between the two, but it proved cognitively effortful for our raters to determine whether it a clue was primarily literal versus figurative\. Accordingly, we settled on a final version distinguishing literalness and figurativeness as separate dimensions\.

### 4\.3\.Reliability

We calculated inter\-rater reliability using Krippendorff’s alpha \(interval metric,n=150n=150clues, 5 raters\) for each coding dimension: Calibration \(α=0\.866\\alpha=0\.866\), Situatedness \(α=0\.927\\alpha=0\.927\), Literalness \(α=0\.884\\alpha=0\.884\), and Figurativeness \(α=0\.859\\alpha=0\.859\)\. All four dimensions exceeded conventional thresholds for good reliability \(α≥0\.80\\alpha\\geq 0\.80\)\.

## 5\.Results

Our findings demonstrate key differences in how humans and LLMs calibrate ambiguity in the Dixit task \(Fig\.[2](https://arxiv.org/html/2609.12575#acmlabel2)\)\. We investigated the degree to which models show ambiguity collapse—as well as how readily they draw on culturally situated knowledge in constructing their clues\. We then analysed whether formal linguistic features are sufficient to explain these differences\. In Section[6](https://arxiv.org/html/2609.12575#S6), we looked at the conditions under which various models differ in how they calibrate ambiguity\.

Figure 2\.Human\-coded rubric ratings on each of the four dimensions \(Table[2](https://arxiv.org/html/2609.12575#S4.T2)\) for human\- vs\. model\-generated clues\. Shaded bars reflect the share of clues at each rating level; points indicate group mean with 95% confidence interval\. \(a\) Ambiguity collapse is demonstrated by models’ calibration erring on the side of Over\-specified\. \(b\) Cultural flattening is shown by the tendency of models to generate clues that rely on universally situated knowledge \(i\.e\., not requiring cultural knowledge for interpretation\. \(c,d\) Models tend to be more literal and less figurative than humans by default\.Four panels, one per rubric dimension, each comparing human and LLM clues as stacked blocks of rating shares with a mean and confidence interval\. LLM clues are rated more overspecified, far less culturally situated, more literal and less figurative than human clues\.### 5\.1\.Humans and LLMs use different strategies to calibrate ambiguity—with varying levels of success

For comparison of human\- vs model\-generated clues, a panel of five expert human annotators used our coding rubric to analyse a set of 150 clues mixed and blinded \(78 human, 72 model\), drawn from 150 rounds that form part of our 350 seed boards \(Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px1)\)\. To models from our set of six \(see Table[1](https://arxiv.org/html/2609.12575#S3.T1)\) were used to generate clues for this comparison \(35 Gemma 3 \[27B\], 37 Qwen 3\.5 \[35B\]\)\. The model clues were generated under the*social cognition*prompt \(Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3)\); the*minimal*prompt of the expanded comparison is not represented in this set \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\)\.

#### 5\.1\.1\.LLMs exhibit ambiguity collapse

We sought explicit evidence of ambiguity output collapse in language models by coding human\- vs model\-generated clues according to “calibration” \(as described in Section[4](https://arxiv.org/html/2609.12575#S4)\)\. Ambiguity output collapse was operationalized in this scheme by a rating of 3 on the Calibration dimension \(where 2 is calibrated and 1 is overly ambiguous\)\.

In ratings from our five expert human annotators, LLM clues were rated as significantly more likely to be over\-specified than human clues \(mean calibration2\.392\.39vs\.1\.931\.93on the 1–3 scale; Mann–WhitneyUU,p<0\.001p<0\.001; Fig\.[2](https://arxiv.org/html/2609.12575#acmlabel2)a\)\. Humans tended to be calibrated for ambiguity\. But on average, LLMs skewed toward being overly specific in their clues, thus exhibiting ambiguity output collapse\.

#### 5\.1\.2\.LLMs tend to flatten cultural context

We investigated the degree to which models rely on culturally\-situated knowledge when constructing their clues—such as allusions to music or movies, widely circulated idioms or stories, or idiosyncratic interpersonal knowledge \(e\.g\., inside jokes or common acquaintances\)\. As described in Section[4](https://arxiv.org/html/2609.12575#S4), our “situatedness” dimension encoded three levels: universal \(no culturally situated knowledge\), mainstream \(broadly shared Anglophone knowledge\), or niche \(specialized cultural knowledge, e\.g\., a line from a particular film\)\.

In ratings from our five expert human annotators, LLM clues were rated as far less culturally situated than human clues \(mean situatedness1\.151\.15vs\.1\.821\.82;p<0\.001p<0\.001; Fig\.[2](https://arxiv.org/html/2609.12575#acmlabel2)b\)\. Human clues drew on a range of culturally specific references—which was almost never true of LLM clues\.

#### 5\.1\.3\.LLMs tend to be more literal and less figurative than humans

We also looked at the degree to which the clue offered a literal description of the image or relied on figurative language\. While these dimensions are strongly anti\-correlated \(Spearman’sρ=−0\.82\\rho=\-0\.82across the150150human\-coded clues,p<0\.001p<0\.001\), it is possible for a clue\-target mapping to rate high \(or low\) on both\.

In ratings from our five expert human annotators, LLM clues were rated as significantly more likely to offer literal clues than humans \(mean literalness2\.212\.21vs\.1\.551\.55; Mann–WhitneyUU,p<0\.001p<0\.001\) and significantly less likely to rely on figurative language \(mean figurativeness1\.941\.94vs\.2\.372\.37;p<0\.001p<0\.001; Fig\.[2](https://arxiv.org/html/2609.12575#acmlabel2)c–d\)\. Some of the highly literal clues contributed to ambiguity collapse \(i\.e\., increased incidence of over\-specified clues\), though much of this miscalibration reflected the models’ failed attempts to generate ambiguity by providing vague descriptions of visual features rather than exact labels \(e\.g\., using terms like “wispy” or “fluffy” for a target image featuring a cloud\)\.

Figure 3\.Model\-generated clues from Fig\.[2](https://arxiv.org/html/2609.12575#acmlabel2)separated by model\.The dashed line reflects the the human mean\. \(a\) Miscalibration is driven primarily by Qwen 3\.5 \(35B\), rather than Gemma 2 \(27B\)\. \(b\) Both avoid cultural references\. \(c,d\) The preference for literal rather than figurative language is more pronounced in Qwen\.Four panels, one per rubric dimension, showing the LLM clues of the previous figure split into Gemma and Qwen, with the human mean as a dashed line\. Both models sit at the floor of situatedness; Qwen is more overspecified and more literal than Gemma, whose calibration matches the human mean\.
#### 5\.1\.4\.Ambiguity calibration depends crucially on the model

We used two different models to generate LLM clues for comparison with humans: Gemma 3 \(27B; dense\) and Qwen 3\.5 \(35B; mixture of experts; see Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3)\)\. Separating the outputs of these two models, we saw significant differences in the clues they generated \(Fig\.[3](https://arxiv.org/html/2609.12575#acmlabel3)\)\. Clues generated by both models were rated near the floor of the situatedness scale and did not differ significantly from one another \(Gemma1\.181\.18, Qwen1\.121\.12;p=0\.06p=0\.06; Fig\.[3](https://arxiv.org/html/2609.12575#acmlabel3)b\)\. However, they differed on the literalness and figurativeness dimensions: Qwen clues were more literal than Gemma clues \(mean literalness2\.552\.55vs\.1\.861\.86;p<0\.001p<0\.001\), while Gemma was more likely to draw on figurative language \(2\.252\.25vs\.1\.641\.64;p<0\.001p<0\.001; Fig\.[3](https://arxiv.org/html/2609.12575#acmlabel3)c–d\)\.

This gives an important qualification to our calibration claim: the effect of LLM clues as significantly over\-specified compared to human clues \(Section[5\.1\.1](https://arxiv.org/html/2609.12575#S5.SS1.SSS1)\) is driven by Qwen \(mean calibration2\.752\.75vs\.1\.931\.93for the human clues;p<0\.001p<0\.001\), whereas Gemma’s calibration is statistically indistinguishable from the human clues \(2\.012\.01vs\.1\.931\.93;p=0\.36p=0\.36; Fig\.[3](https://arxiv.org/html/2609.12575#acmlabel3)a\)\. Crucially, both models exhibit a similar degree of cultural flattening \(Section[5\.1\.2](https://arxiv.org/html/2609.12575#S5.SS1.SSS2); situatedness1\.181\.18and1\.121\.12vs\.1\.821\.82for the human clues; bothp<0\.001p<0\.001\)\. In Section[6](https://arxiv.org/html/2609.12575#S6), we investigate whether this difference in calibration is an artefact of the models we initially selected or a robust difference in model behaviour\.

### 5\.2\.Can formal linguistic features capture the differences without the rubric?

One potential concern about the findings presented in Section[5\.1](https://arxiv.org/html/2609.12575#S5.SS1)is that the differences between human\- and LLM\-generated clues can be explained without reference to a complex, bespoke annotation rubric\. Perhaps it would be enough to look at formal linguistic features to distinguish between these different strategies for calibrating ambiguity\. Accordingly, we analysed whether a canonical set of formal linguistic features could predict whether a given clue was generated by a human or a model\.

#### 5\.2\.1\.Individually, formal features separate human and LLM clues only weakly

We computed twelve surface features for every clue—length \(word count, mean word length\), lexical statistics \(word frequency on the Zipf scale\([van Heuven et al\., 2014](https://arxiv.org/html/2609.12575#bib.bib2)\), computed with thewordfreqpackage\([Speer, 2022](https://arxiv.org/html/2609.12575#bib.bib3)\), and concreteness\([Brysbaert et al\., 2014](https://arxiv.org/html/2609.12575#bib.bib1)\)\), part\-of\-speech composition \(noun, adjective, determiner, and function\-word ratios, presence of a finite verb\), syntactic type, and orthography \(title\-casing, final punctuation\)—for the 4,200 LLM\-generated clues of the expanded clue set \(six models×\\timestwo prompts×\\times350 rounds; Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px1)\) and for a random sample of 2,000 human clues drawn from the 62,688 English clues in the full corpus\.

For each single feature we asked how well it alone would tell an LLM clue from a human clue\. The area under the ROC curve \(AUC\) summarised this as the probability that a randomly chosen LLM clue scored higher on the feature than a randomly chosen human clue, with0\.500\.50as chance and1\.01\.0as perfect separation\. With samples this large, small average differences are statistically reliable, so the confidence intervals of nearly every feature exclude0\.500\.50\(LLM clues are, on average, longer and more concrete\)\. But statistical reliability does not translate directly to discriminative power: a difference in means can be significant while the two distributions have significant overlap\. For example,93%93\\%of LLM clues had more than four words, while this was true of46%46\\%of human clues \(Fig\.[4](https://arxiv.org/html/2609.12575#acmlabel4)a, inset\)\. Using word count to distinguish these clues would mark the majority of LLM clues correctly, but also misidentify almost half of human clues\. Therefore, even the most discriminative formal feature on its own was not enough to distinguish human\- and model\-generated clues\.

Word count was the most discriminative feature \(AUC=0\.74=0\.74, 95% CI\[0\.72,0\.75\]\[0\.72,0\.75\]\), followed by concreteness \(0\.690\.69\)\. The degree to which word count distinguishes between human and model clues is at least partially an artefact of our prompts \(Appendix[A](https://arxiv.org/html/2609.12575#A1)\)\. We asked models to generate clues in a length of 3\-6 words\. Humans featured more variation in their clues—some were a single word, while others were lengthy phrases\. Concreteness maps well onto our literalness/figurative dimensions\. This supports their inclusion as key dimensions in the rubric\.

#### 5\.2\.2\.LLM clues tend to be more homogeneous than human clues

Taken as an aggregate, however, superficial features could identify at least one important systematic difference between human\- and model\-generated clues: homogeneity\. In a one\-to\-one comparison of individual clues, the signal of any single formal feature was weak, but model\-generated clues tended to look more like one another than human clues, which exhibit a much higher degree of variation\. In short, LLM\-generated clues tended to cluster more tightly than human\-generated ones \(Fig\.[4](https://arxiv.org/html/2609.12575#acmlabel4)b\)\.

![Left: a dot plot of the ROC AUC of twelve surface features for telling LLM from human clues, all between the chance line at 0.5 and a dashed line at 0.92 for the combined twelve-feature score; word count is highest at 0.74, and an inset shows the overlapping word-count histograms. Right: a scatter plot of clues in two principal components, with a large dashed human ellipse and six smaller model ellipses lying inside it.](https://arxiv.org/html/2609.12575v1/fig3_surface.png)Figure 4\.Formal linguistic features\. \(a\) ROC AUC of each feature alone \(95% CIs\); dashed line: all twelve combined \(Mahalanobis score, held out\); inset: word\-count distributions\. \(b\) The same clues in the first two principal components; ellipses: 2 SD\.Left: a dot plot of the ROC AUC of twelve surface features for telling LLM from human clues, all between the chance line at 0\.5 and a dashed line at 0\.92 for the combined twelve\-feature score; word count is highest at 0\.74, and an inset shows the overlapping word\-count histograms\. Right: a scatter plot of clues in two principal components, with a large dashed human ellipse and six smaller model ellipses lying inside it\.Specifically, a multivariate score \(the Mahalanobis distance to the LLM centroid in the z\-scored twelve\-feature space\) separated held\-out LLM clues from human clues with AUC=0\.92=0\.92\(the dashed line in Fig\.[4](https://arxiv.org/html/2609.12575#acmlabel4)a\), far above any single feature\. In the twelve\-feature space, LLM clues occupied a tighter subregion inside the human cloud \(Fig\.[4](https://arxiv.org/html/2609.12575#acmlabel4)b\):99\.8%99\.8\\%of LLM clues fall inside the human 95% ellipsoid, but only23%23\\%of human clues fall inside the LLM one\. The mean pairwise distance between LLM clues was3\.23\.2versus4\.74\.7for human clues, roughly1\.5×1\.5\\timestighter\. This aggregate figure potentially understates the effect: each model was narrower than the human cloud \(2\.22\.2–3\.43\.4;1\.41\.4–2\.1×2\.1\\timestighter than humans\)\. In the 2D projection \(Fig\.[4](https://arxiv.org/html/2609.12575#acmlabel4)b\), the six models visually sit in different parts of the human cloud\. The largest distances between model centroids \(3\.23\.2between Gemma and GLM;0\.70\.7–3\.23\.2across all pairs\) were twice the distance between the pooled LLM and human centroids \(1\.61\.6\)\. Thus, there was no single “LLM style”, only a set of narrow, model\-specific styles\. Each of these styles could plausibly be generated by a human, but reliance on any single one dramatically underestimates the space of clues humans actually generate\.

## 6\.Model Comparison

In Section[5\.1\.4](https://arxiv.org/html/2609.12575#S5.SS1.SSS4), our initial analysis suggested that different LLMs use different strategies for ambiguity calibration\. This was supported by the analysis in Section[5\.2\.2](https://arxiv.org/html/2609.12575#S5.SS2.SSS2), which showed that different LLMs produce clues characterised by separate clusters of formal features\. To what degree are these just artefacts of the two models we chose to compare against humans? To test this, we ran a larger comparison of six models \(Table[1](https://arxiv.org/html/2609.12575#S3.T1)\), featuring a range of parameter sizes \(9B–35B\) and two different architectures \(dense vs mixture of experts\)\.

Figure 5\.Models differ in their degree of ambiguity collapse, but all of them show evidence of cultural flattening\. Each bank reflects analysis of 350 clues across six models under the minimal and the social cognition prompt, ordered by collapse under the minimal prompt\. \(a\) Share of clues both coders rated*Overspecified*\. \(b\) Share rated*Mainstream*or*Niche*on situatedness; note the 50% scale\. Dashed lines: the human clues on the same rounds\.Two bar charts across thirteen groups: human clues, six models under the minimal prompt and the same six under the social cognition prompt\. Top: share of clues rated overspecified, 2 percent for humans, 28 to 93 percent for models under the minimal prompt, 2 to 30 percent under the social cognition prompt\. Bottom, on a 50 percent scale: share of clues rated culturally situated, 43 percent for humans and 1 to 10 percent for every model under either prompt\.### 6\.1\.Scaling annotation with LLM coders

To run this model comparison, we first needed to address the bottleneck of reliance on our human expert annotators\. While there are recognized limitations to using LLMs for qualitative coding\([Xiao et al\., 2023](https://arxiv.org/html/2609.12575#bib.bib78);[Tai et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib79);[Ziems et al\., 2024](https://arxiv.org/html/2609.12575#bib.bib80)\), a larger scale model comparison would not be feasible if it were limited by capacity of our human experts\.

To scale up our model comparison, we analysed how closely LLM coders could approximate the human coders when applying the rubric under the human protocol \(Section[4](https://arxiv.org/html/2609.12575#S4)\)\. On the 150 clues coded by the five human annotators, each LLM coder’s rating correlated with the human panel mean atr=0\.60r=0\.60–0\.790\.79depending on the dimension \(Table[3](https://arxiv.org/html/2609.12575#S6.T3)\)—close to, and on three of the four dimensions higher than, the agreement of an individual human annotator with the rest of the panel \(leave\-one\-outr=0\.58r=0\.58–0\.800\.80\)\. Exact agreement with the rounded panel mean was0\.550\.55–0\.800\.80, and the LLM coders identified the intended target card at least as often as the human annotators did \(0\.620\.62–0\.650\.65vs\.0\.570\.57\)\. Agreement was highest for situatedness, which was also the dimension on which the human annotators agree most with one another \(α=0\.90\\alpha=0\.90\), and lowest for figurativeness and calibration, the two dimensions with the lowest human reliability \(α≈0\.80\\alpha\\approx 0\.80; Section[4](https://arxiv.org/html/2609.12575#S4)\)\.

Table 3\.LLM coders vs\. the human panel \(n=150n=150clues\)\. LOOrr: one annotator vs\. the mean of the other four\.rr, exact: correlation and exact agreement with the panel mean\. Target: share of clues whose target card was identified\.On this subset the LLM coders also reached the same substantive conclusions\. Re\-running the comparisons of Section[5\.1](https://arxiv.org/html/2609.12575#S5.SS1)with each LLM coder in place of the human panel reproduced the calibration result exactly—Gemma was indistinguishable from humans and Qwen was significantly overspecified \(both coders: Human vs\. Gemma n\.s\., Human vs\. Qwenp<0\.001p<0\.001\)—as well as the Human<<Gemma<<Qwen gradient in literalness and the collapse of situatedness for both models \(allp<0\.01p<0\.01\)\.

### 6\.2\.Model Comparison Results

After establishing broad agreement with the human annotators, we applied the two LLM coders to the full expanded clue set: six models×\\timestwo prompts×\\times350 rounds=4,200=4\{,\}200LLM clues, plus the 350 human clues on the same boards, each rated independently by both coders \(9,1009\{,\}100ratings\)\. Ratings were averaged over the two coders per clue\. Each of the twelve conditions \(6 models×\\times2 prompts\) were compared with the human clues on the same 350 rounds \(Fig\.[5](https://arxiv.org/html/2609.12575#acmlabel5); Table[4](https://arxiv.org/html/2609.12575#S6.T4)\)\. This increased the number of clues being analysed, from3535–7878clues per condition to350350, and from two models to six\. As described in Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3), we tested two different prompting conditions: minimal and social cognition\.

#### 6\.2\.1\.All six models collapse, regardless of architecture or size

Table 4\.Models differ in their degree of ambiguity collapse, but all of them show evidence of cultural flattening\. Each bank reflects analysis of 350 clues across six models under the minimal and the social cognition prompt, ordered by collapse under the minimal prompt\. \(a\) Share of clues both coders rated Overspecified\. \(b\) Share rated Mainstream or Niche on situatedness; note the 50% scale\. Dashed lines: the human clues on the same rounds\.All six models exhibited ambiguity collapse under the minimal prompt\. Each model produced clues that were rated as substantially more specified than the human clues on the same boards \(allp<10−40p<10^\{\-40\}; Cliff’sδ=0\.54\\delta=0\.54–0\.970\.97; Fig\.[5](https://arxiv.org/html/2609.12575#acmlabel5)a\)\. The model\-specific collapse we observed in Section[5\.1](https://arxiv.org/html/2609.12575#S5.SS1)is therefore not just an effect of Qwen\.

However, in some models the effects were more pronounced than others \(Fig\.[5](https://arxiv.org/html/2609.12575#acmlabel5)\)\. GLM\-4\.6V\-Flash was rated as over\-specified on93%93\\%of rounds \(2%2\\%for the human clues\)\. At the other end, Gemma\-3\-27B was labeled as over\-specified on28%28\\%of rounds and Kimi\-VL on41%41\\%\. This accorded with the coders’ target\-card choices: when shown a minimal\-prompt model clue, they identified the target on8383–100%100\\%of boards, against34%34\\%for the human clues on the same boards \(Table[4](https://arxiv.org/html/2609.12575#S6.T4), target id\.\)\. Clue length did not explain the difference: the collapsed banks range from3\.53\.5to6\.46\.4words per clue around the human mean of4\.04\.0\. \(Note: Section[5\.1\.4](https://arxiv.org/html/2609.12575#S5.SS1.SSS4)found no miscalibration for Gemma, but the human\-coded set contained no minimal\-prompt clues; under the social cognition prompt Gemma’s gap remains negligible here as well, Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\.\)

We designed this comparison to contrast dense and mixture\-of\-experts models within a broadly similar scale \(9B–35B total parameters; Table[1](https://arxiv.org/html/2609.12575#S3.T1)\)\. Within this sample, calibration did not align with architecture\. The two Qwen models—Qwen3\.5\-35B\-A3B, a mixture\-of\-experts model with 3B active parameters, and Qwen3\.8\-27B, a dense model of a different size and generation—are almost indistinguishable on every dimension under both prompts \(calibration2\.872\.87vs\.2\.862\.86under the minimal prompt and2\.422\.42vs\.2\.432\.43under social cognition; literalness2\.872\.87vs\.2\.872\.87and2\.302\.30vs\.2\.372\.37; figurativeness1\.191\.19vs\.1\.181\.18and2\.102\.10vs\.1\.971\.97\)\.

Meanwhile the three dense models spanned a wider range, from the best\-calibrated \(Gemma\) to the worst \(GLM\); the same was true of the three mixture\-of\-experts models\. Neither parameter count nor active parameter count predicted the models’ ranking\.

#### 6\.2\.2\.Prompting mitigates collapse but does not remove it

The social cognition prompt substantially changed the models’ calibration\. A two\-way analysis of variance on the LLM clues yielded significant effects of prompt \(F⁡\(1,4188\)=1590F\(1,4188\)=1590\), model \(F⁡\(5,4188\)=145F\(5,4188\)=145\) and their interaction \(F⁡\(5,4188\)=37\.9F\(5,4188\)=37\.9; allp<0\.001p<0\.001\) on calibration\. The social cognition prompt influenced models to look more like the human distribution—though the size of the difference depended on the model\. GLM shifted by0\.760\.76points on the three\-point scale \(from2\.962\.96to2\.212\.21\), the two Qwen models by0\.430\.43–0\.450\.45, Gemma by0\.350\.35, ERNIE and Kimi by0\.320\.32\. In short: the social cognition prompt was most helpful for models that exhibit the sharpest initial degree of collapse\.

Table 5\.Example clues from all 6 model and 2 prompt permutations compared with humans for two boards, shown with their five cards \(target in red\)\. Cal\.: coder\-averaged calibration rating \(1 Underspecified, 2 Calibrated, 3 Overspecified\)\. min/soc: minimal / social cognition prompt\.Board 1\(3 of 4 players found the target\)Board 2\(1 of 4\)![[Uncaptioned image]](https://arxiv.org/html/2609.12575v1/board_g66818_r1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.12575v1/board_g28044_r11.png)ProducerClueCal\.ClueCal\.HumanLet it go2\.0Christina Rossetti2\.0GLM\-4\.6V\-FlashminA woman in a long blue cloak3\.0A white figure in a stormy field3\.0Qwen3\.5\-35BminGirl in green coat in snow3\.0Snowman built by a child2\.5Qwen3\.8\-27BminLong hair blowing in wind3\.0Walking through scattered white paper3\.0ERNIE\-4\.5\-VLminA figure in a long cloak3\.0A ghost made of paper3\.0Kimi\-VL\-A3BminA woman with a head injury2\.0I left my shirt in your car2\.0Gemma\-3\-27BminLost in wintry contemplation2\.0Ephemeral, fleeting, almost gone2\.0GLM\-4\.6V\-FlashsocThe lone wanderer in the vast, silent expanse2\.0The ghost of a forgotten dream2\.0Qwen3\.5\-35BsocCold winds of change are coming2\.0Walking with a heart made of paper2\.0Qwen3\.8\-27BsocWalking through frozen tears2\.0Walking on fragile thoughts2\.0ERNIE\-4\.5\-VLsocA lonely wanderer in a long coat3\.0A phantom in a tear\-stained dream2\.0Kimi\-VL\-A3Bsoca beacon of hope amidst despair2\.0The heart of winter2\.0Gemma\-3\-27BsocLost in quiet contemplation2\.0Fleeting, fragile, and hopeful2\.0Only Gemma reached a level of ambiguity calibration consistent with human\-generated clues\. Its social cognition clues were rated*Calibrated*on98%98\\%of rounds \(97%97\\%for the human clues\) and received exactly the same rating as the human clue on76%76\\%of boards; the residual difference was statistically detectable atn=350n=350\(p<0\.001p<0\.001\) but negligible in size \(δ=0\.10\\delta=0\.10; with the Gemini 3\.5 coder alone,Δ=0\.02\\Delta=0\.02,p=0\.20p=0\.20\)\. For every other model the gap remained substantial \(δ=0\.25\\delta=0\.25–0\.530\.53; more specified than the human clue on3131–57%57\\%of boards\)\.

The social cognition prompt asked the model to prefer figurative, rather than literal, language\. For the most part, models were able to oblige\. Under that prompt the figurativeness of Gemma’s and GLM’s clues \(2\.552\.55and2\.612\.61\) actually exceeded that of the human clues \(2\.422\.42;p<0\.01p<0\.01\), and Kimi’s and ERNIE’s match it \(2\.322\.32–2\.332\.33, n\.s\.; Table[4](https://arxiv.org/html/2609.12575#S6.T4)\)\. Yet GLM, Kimi and ERNIE remain significantly overspecified: GLM’s social\-cognition clues, the most figurative bank in the study, were still rated*Overspecified*on12%12\\%of rounds and more specified than the human clue on31%31\\%of boards\. Therefore, prompting is an important lever for modulating the degree to which they “describe the picture,” as we claim in our title—though it is crucial note that figurative language did not translate directly to better calibration of ambiguity\.

#### 6\.2\.3\.Cultural flattening persists across models and prompts

Our initial findings on cultural flattening in model\-based clues \(Section[5\.1\.2](https://arxiv.org/html/2609.12575#S5.SS1.SSS2)\) were supported by this analysis\. Models uniformly relied on the “universal” level of situatedness, requiring no situated cultural knowledge to interpret their target \(Fig\.[5](https://arxiv.org/html/2609.12575#acmlabel5)b\)\. Human clues on these boards were rated “mainstream” or “niche”43%43\\%of the time \(11%11\\%niche\)\. The LLM clues were rated above*Universal*on3\.9%3\.9\\%of rounds\. Every one of the twelve conditions \(6 models×\\times2 prompts\) showed near\-floor effects on the scale \(1\.011\.01–1\.071\.07;δ=−0\.34\\delta=\-0\.34to−0\.42\-0\.42against the human clues, allp<0\.001p<0\.001\)\. Unlike calibration, situatedness did not respond to the prompt \(prompt effectF⁡\(1,4188\)=3\.3F\(1,4188\)=3\.3,p=0\.07p=0\.07; interactionp=0\.24p=0\.24\); the only model effect is Kimi\-VL’s marginally higher floor\. Collapse and cultural flattening therefore have different profiles in our sample\. Collapse varies by model and is partly reversed by the social cognition prompt, whereas cultural flattening is present across models and is not modulated by prompt\.

## 7\.Discussion

We present a novel paradigm for operationalising calibrated ambiguity in multimodal generative AI systems\. Across several analyses, we compare clues from an online data set of human\-generated Dixit clues with clues generated by vision–language models playing the same role on the same boards\. We developed a coding rubric to distinguish key differences in how ambiguity is calibrated and managed by humans and models\. Our main finding is that models, as compared with humans, tend to be less well calibrated and almost never rely on the kind of cultural references found in human\-constructed clues\. In follow\-up analyses, we show that individual surface\-level features distinguish human vs model clues only weakly, and that there is significant variation in how different models calibrate ambiguity\.

### 7\.1\.Do LLMs exhibit ambiguity output collapse?

Yes\. In our paradigm, models tend to be less calibrated for ambiguity than humans \(Section[5\.1](https://arxiv.org/html/2609.12575#S5.SS1)\)\. This “output collapse” identified in our findings is only a subset of possible forms of ambiguity collapse articulated by Gur\-Arieh et al\.\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)—but an important one that has been hypothesised, yet not demonstrated, in previous work\.

There are caveats to this claim\. The prompt makes a difference \(Section[6\.2\.1](https://arxiv.org/html/2609.12575#S6.SS2.SSS1)and Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\), though it does not mitigate ambiguity collapse entirely\. A range of model sizes \(9B–35B\) and model architectures \(dense, mixture\-of\-experts\) are susceptible to this form of ambiguity collapse\. However, different models do not collapse to the same degree\. For example, under a “social cognition” prompt, Gemma 3 \(27B\) exhibited human\-level ambiguity calibration \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\)\. But overall, we found robust evidence for ambiguity collapse across a range of circumstances\.

### 7\.2\.Do LLMs avoid cultural references when calibrating ambiguity?

To a significant degree, yes\. In our paradigm, models avoided culturally situated knowledge \(Section[5\.1\.2](https://arxiv.org/html/2609.12575#S5.SS1.SSS2)\)\. This effect held even under the social cognition prompt, which invites allusion and figurative language \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\)\. This seemed to be one of the defining differences between how humans and LLMs managed ambiguity in our task \(Section[5\.1](https://arxiv.org/html/2609.12575#S5.SS1)\)\. As a default behaviour, models tended to construct clues without cultural references\.

This is a significant capability gap in model outputs, consistent with the output\-based risks described by Gur\-Arieh\([Gur\-Arieh et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib4)\)and the hermeneutic challenges offered by Kommers et al\.\([Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74)\)\. The evidence from our results suggests that culturally\-situated knowledge plays a key role in how humans generate, manage, and calibrate ambiguity\. If models avoid it by default, they are missing out an important communicative resource\. However, it remains an open question how best to elicit cultural information from models, for example by prompting them with information about the audience for whom they were generating clues or offering implicit clues about context without directly stating it \(e\.g\., using British spellings\)\([Veselovsky et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib33)\)\. Nonetheless, humans naturally reference popular movies or stories like Pocahontas and Frozen; they riff on cultural idioms and stock phrases; and they play on associations with single words or concepts \(like Australia or Metallica\)\. These rely on cultural knowledge widely shared in the Anglophone world—and depending on the clue\-target pairing in question, little expertise is required for interpretation beyond an awareness that such cultural referent exists\. This is consistent with emerging evidence that LLMs struggle to “fill in the blanks” with details about hypothetical personas\([Wang et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib81)\)\.

### 7\.3\.Do models tend to calibrate ambiguity in the same way?

No, not necessarily\. The dynamics of ambiguity calibration seem to be model\-specific \(Section[5\.1\.4](https://arxiv.org/html/2609.12575#S5.SS1.SSS4)\)\. For example, models differed in their proclivity for literal vs figurative language \(Section[5\.1\.3](https://arxiv.org/html/2609.12575#S5.SS1.SSS3)\)\. We found that models tended to produce homogeneous clues, but with each model using an idiosyncratic strategy that differed from other models \(Fig\.[4](https://arxiv.org/html/2609.12575#acmlabel4); Section[5\.2](https://arxiv.org/html/2609.12575#S5.SS2)\)\. Models also responded differently to different prompts \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\)\. The one effect that was entirely robust across conditions and models was an avoidance of clues based on culturally\-situated knowledge \(Section[6\.2\.3](https://arxiv.org/html/2609.12575#S6.SS2.SSS3)\); models were consistently and uniformly averse to making cultural references in their clues\.

### 7\.4\.Can ambiguity collapse be fixed with prompting?

A little bit, sometimes, and potentially\. We provided models with a “minimal” prompt \(just the rules of the game\), and this resulted in unanimous ambiguity collapse \(Section[6\.2\.1](https://arxiv.org/html/2609.12575#S6.SS2.SSS1)\)\. The least calibrated models performed better when we adjusted the prompt to emphasise social cognition \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\), but this did not by any means obviate the effect of collapse\. Again, there were model\-based differences\. Of the models we looked at, Gemma performed best—and in some cases at human\-level performance \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\)\.

In short, the evidence is mixed\. It seems that prompting can help some models alleviate miscalibration to some degree\. Targeted prompts can encourage models to use specific strategies \(e\.g\., more figurative language\)—but this does not automatically make them “good” at using these strategies in the sense that they result in more calibrated clues \(Section[6\.2\.2](https://arxiv.org/html/2609.12575#S6.SS2.SSS2)\)\. Thus, we stand by our primary claim that when calibrating ambiguity humans tend to reach for cultural references, while models tend to describe the picture\.

## 8\.Limitations and Alternative Perspectives

### 8\.1\.I agree the ambiguity matters\. But does your task actually measure the kind of ambiguity we care about?

The strength of this task is that it provides a measurable, operationalised definition of calibrated ambiguity, where none existed before\. The cost of that precision is that it only gets at a limited range of the possible textures of ambiguity\. We operationalise ambiguity by the presence of multiple legitimate interpretations\. This is kind of like operationalising a joke by whether it makes someone laugh\. It is relevant—but it is also far from the whole story\. The three non\-calibration dimensions of our rubric \(situatedness, literalness, and figurativeness\) are meant to be broadly applicable to different types of ambiguity, but the mapping is imperfect\. For example, in the introduction we give the example of moral ambiguity in the filmUnforgiven\. Certainly the dimension of culturally\-situated knowledge would be relevant in analysing this form of ambiguity \(what counts as morally good or bad behaviour in a particular frame of reference?\)\. But it is less clear whether the tension between literal and figurative language or depictions is especially important\. Our position is that the paradigm we present in this paper is a useful baseline\. It captures some \(not all\) important aspects of ambiguity that apply in a range of situates\. But ultimately the true “calibration” of ambiguity is tied intimately to the context in which a system is deployed\. In future work, variations on this paradigm could be used to investigate more specific kinds of ambiguity \(e\.g\., moral ambiguity in narratives produced by AI based on a set of Dixit\-like target images\) or variations could be adapted for specific use cases with a specific target population in mind \(e\.g\., calibrating the ambiguity in labelling inappropriate images for content moderation for a given platform or population\)\.

### 8\.2\.You developed a paradigm for measuring calibrated ambiguity\. But once we have identified miscalibration, how do we fix it?

Our findings suggest that prompting makes a difference, especially for overall calibration and collapse\. Of course it does, and the number of permutations for exploring how this might work are unbounded\. Again, the key will be in contextualisation\. Our paradigm offers a general framework for measuring calibrated ambiguity—but the kind of ambiguity that is desirable will depend on a specific situation or use\-case\. Future work should explore how prompting methods can achieve specific kinds of ambiguity—for specific models, operating in specific contexts\. Our findings about the effects of prompting on cultural flattening are less clear\. The relevant question is not whether models can be coaxed into drawing on cultural references, but about the processes for doing so that yield the most compelling results\. There is a tension in existing literature between evidence for widespread homogenisation and cultural collapse\([Jain et al\., 2025](https://arxiv.org/html/2609.12575#bib.bib32);[Heuser, 2025](https://arxiv.org/html/2609.12575#bib.bib29)\)versus evidence that local cultural knowledge is controllable, if you know how to elicit it\([Veselovsky et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib33)\)\. Do models produce homogeneous outputs because we are just not prompting them correctly? It is a difficult question to answer, but the evidence from Veselovsky et al\.\([Veselovsky et al\., 2026](https://arxiv.org/html/2609.12575#bib.bib33)\)suggests that the key is not just asking behaviour representing a particular cultural milieu—but actively providing cues of participation in it\. This is consistent with the view of LLMs as “context machines”\([Kommers et al\., 2026a](https://arxiv.org/html/2609.12575#bib.bib74)\)which are fundamentally designed to adapt to the contextual cues with which they are provided\. Future work can explore the transition from measuring calibratedambiguity to best practices for calibratingit\.

### 8\.3\.I suspect these effects would be obviated in bigger models\. Do you really think Claude Fable 5 would exhibit ambiguity collapse?

This is an important direction for future research\. Our primary aim in this paper was to provide a proof of concept for a novel methodology of measuring ambiguity calibration in LLM outputs\. With this foundation, there are a lot of further questions to pose and investigate\. How ambiguity calibration scales with model size is definitely one of them\. In this version, we chose to use models that were large enough to do the task \(e\.g\., Gemma 3 \[27B\] performed near human level in some cases\) but lightweight enough to deploy in a range of situations \(see Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px2)for further rationale\)\. An additional part of the consideration was finding the right match for variation in performance for a simple baseline version of the task\. For example, the task can be made arbitrarily more difficult by increasing the number of distractors\. Given that Gemma 3 \(27B\) performs well under the conditions we studied, it is probably the case that bigger, more sophisticated models could do the task as currently designed\. The task can be adapted for larger models—for example, by increasing the number of distractors, altering the kind of cards that are used, or providing background information about the shared cultural frames relevant for the other players in the game\. Future work can explore how variations of this paradigm can elucidate the strategies more sophisticated models use to calibrate ambiguity—and provide a concrete measure of when they collapse\.

## Ethics and Privacy Statement

All data for human\-generated clues in this study were publicly available\. All expert coders were authors on the paper, not participants in a study\. Dixit images are subject to copyright \(c\) Libellud / Marie Cardouat\. The work presented in our system presents minimal ethical or privacy risk\. The Dixit images do not feature harmful, triggering, or inappropriate content\. However, an adaptive version of this study could—specifically one that looks at ambiguity in a more ethically fraught context, such as content moderation\.

## References

- Basileet al\.\(2021\)V\. Basile, M\. Fell, T\. Fornaciari, D\. Hovy, S\. Paun, B\. Plank, M\. Poesio, and A\. UmaWe need to consider disagreement in evaluation\.InProceedings of the 1st workshop on benchmarking: past, present and future,pp\. 15–21\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Beguš \(2025\)N\. BegušArtificial humanities: a fictional perspective on language in ai\.University of Michigan Press\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1)\.
- Bergmanet al\.\(2023\)A\. S\. Bergman, L\. A\. Hendricks, M\. Rauh, B\. Wu, W\. Agnew, M\. Kunesch, I\. Duan, I\. Gabriel, and W\. IsaacRepresentation in ai evaluations\.InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency,pp\. 519–533\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Booth \(1961\)W\. C\. BoothThe rhetoric of fiction\.University of Chicago Press,Chicago\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Brown and Levinson \(1987\)P\. Brown and S\. C\. LevinsonPoliteness: some universals in language usage\.Cambridge University Press,Cambridge\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Bruner \(1990\)J\. S\. BrunerActs of meaning: four lectures on mind and culture\.Vol\.3,Harvard university press\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Brysbaertet al\.\(2014\)M\. Brysbaert, A\. B\. Warriner, and V\. KupermanConcreteness ratings for 40 thousand generally known English word lemmas\.Behavior Research Methods46\(3\),pp\. 904–911\.External Links:[Document](https://dx.doi.org/10.3758/s13428-013-0403-5)Cited by:[§5\.2\.1](https://arxiv.org/html/2609.12575#S5.SS2.SSS1.p1.1)\.
- Chang \(2024\)S\. ChangDixit AI: an OpenAI CLIP\-based hint selection player\.Academic Journal of Computing & Information Science7\(11\),pp\. 102–108\.External Links:[Document](https://dx.doi.org/10.25236/AJCIS.2024.071114)Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p3.1)\.
- Chinhet al\.\(2019\)B\. Chinh, H\. Zade, A\. Ganji, and C\. AragonWays of qualitative coding: a case study of four strategies for resolving disagreements\.InExtended abstracts of the 2019 CHI conference on human factors in computing systems,pp\. 1–6\.Cited by:[§4\.2](https://arxiv.org/html/2609.12575#S4.SS2.p1.1),[§4](https://arxiv.org/html/2609.12575#S4.p1.1)\.
- de Rooij and Biskjaer \(2026\)A\. de Rooij and M\. M\. BiskjaerDoes generative ai make us think alike? a systematic review and meta\-analysis of homogenisation effects in human–ai co\-creation\.Behaviour & Information Technology,pp\. 1–14\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Di Lodovicoet al\.\(2025\)C\. Di Lodovico, S\. Houben, and S\. ColomboHow to design with ambiguity: insights from self\-tracking wearables\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,pp\. 1–18\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Dumitracheet al\.\(2019\)A\. Dumitrache, L\. Aroyo, and C\. WeltyA crowdsourced frame disambiguation corpus with ambiguity\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT\),pp\. 2164–2170\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1224)Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Dunivin \(2025\)Z\. O\. DunivinScaling hermeneutics: a guide to qualitative coding with llms for reflexive content analysis\.EPJ Data Science14\(1\),pp\. 28\.Cited by:[§3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px5.p1.1)\.
- Empson \(1930\)W\. EmpsonSeven types of ambiguity\.Chatto & Windus,London\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1)\.
- Farquharet al\.\(2024\)S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. GalDetecting hallucinations in large language models using semantic entropy\.Nature630,pp\. 625–630\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07421-0)Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Farrellet al\.\(2025\)H\. Farrell, A\. Gopnik, C\. Shalizi, and J\. EvansLarge AI models are cultural and social technologies\.Science387\(6739\),pp\. 1153–1156\.External Links:[Document](https://dx.doi.org/10.1126/science.adt9819)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1)\.
- Feduset al\.\(2022\)W\. Fedus, B\. Zoph, and N\. ShazeerSwitch transformers: scaling to trillion parameter models with simple and efficient sparsity\.Journal of Machine Learning Research23\(120\),pp\. 1–39\.Cited by:[§3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px2.p1.1)\.
- Frendaet al\.\(2023\)S\. Frenda, A\. Pedrani, V\. Basile, S\. M\. Lo, A\. T\. Cignarella, R\. Panizzon, C\. Sánchez\-Marco, B\. Scarlini, V\. Patti, C\. Bosco,et al\.Epic: multi\-perspective annotation of a corpus of irony\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13844–13857\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Gaveret al\.\(2003\)W\. W\. Gaver, J\. Beaver, and S\. BenfordAmbiguity as a resource for design\.InProceedings of the SIGCHI conference on Human factors in computing systems,pp\. 233–240\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Geet al\.\(2024\)X\. Ge, C\. Xu, D\. Misaki, H\. R\. Markus, and J\. L\. TsaiHow culture shapes what people want from ai\.InProceedings of the 2024 CHI conference on human factors in computing systems,pp\. 1–15\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Gius and Jacke \(2017\)E\. Gius and J\. JackeThe hermeneutic profit of annotation: on preventing and fostering disagreement in literary analysis\.International Journal of Humanities and Arts Computing11\(2\),pp\. 233–254\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Grice \(1975\)H\. P\. GriceLogic and conversation\.InSyntax and Semantics, Vol\. 3: Speech Acts,P\. Cole and J\. L\. Morgan \(Eds\.\),pp\. 41–58\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Guptaet al\.\(2026\)N\. Gupta, M\. Antoniak, and M\. WalshAI fiction in the wild\.arXiv preprint arXiv:2606\.22748\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Gur\-Ariehet al\.\(2026\)S\. Gur\-Arieh, A\. Wang, and S\. FazelpourAmbiguity collapse by LLMs: a taxonomy of epistemic risks\.Note:arXiv:2603\.05801External Links:2603\.05801Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1),[§1](https://arxiv.org/html/2609.12575#S1.p4.1),[§1](https://arxiv.org/html/2609.12575#S1.p6.1),[§2](https://arxiv.org/html/2609.12575#S2.p1.1),[§7\.1](https://arxiv.org/html/2609.12575#S7.SS1.p1.1),[§7\.2](https://arxiv.org/html/2609.12575#S7.SS2.p2.1)\.
- Haraway \(1988\)D\. HarawaySituated knowledges: the science question in feminism and the privilege of partial perspective\.Feminist studies14\(3\),pp\. 575–599\.Cited by:[item 2](https://arxiv.org/html/2609.12575#S4.I1.i2.p1.1)\.
- Hastieet al\.\(1983\)R\. Hastie, S\. D\. Penrod, and N\. PenningtonInside the jury\.Harvard University Press,Cambridge, MA\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Hemmentet al\.\(2025\)D\. Hemment, C\. Kommers, R\. Ahnert, M\. Antoniak, G\. Arbix, V\. Belle, S\. Benford, A\. Brintrup, N\. Bryan\-Kinns, M\. Bunz,et al\.Doing ai differently: rethinking the foundations of ai via the humanities\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1)\.
- Heuser \(2025\)R\. HeuserCultural collapse: toward a generative formalism for ai cultural production\.Anthology of Computers and the Humanities3,pp\. 575–588\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1),[§2](https://arxiv.org/html/2609.12575#S2.p1.1),[item 2](https://arxiv.org/html/2609.12575#S4.I1.i2.p1.1),[§8\.2](https://arxiv.org/html/2609.12575#S8.SS2.p1.1)\.
- Iser \(1978\)W\. IserThe act of reading: a theory of aesthetic response\.Johns Hopkins University Press,Baltimore\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Iwataet al\.\(2020\)K\. Iwata, R\. Suzuki, and T\. AritaAn approach to human cognition based on AI player creation for a picture\-guessing card game Dixit\.InProceedings of the 34th Annual Conference of the Japanese Society for Artificial Intelligence \(JSAI 2020\),pp\. 4C2GS1302\.External Links:[Document](https://dx.doi.org/10.11517/pjsai.JSAI2020.0%5F4C2GS1302)Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p3.1)\.
- Jainet al\.\(2025\)S\. Jain, J\. Lanchantin, M\. Nickel, C\. Ross, K\. Ullrich, A\. Wilson, and J\. Watson\-DanielsTask\-dependent evaluation of llm output homogenization: a taxonomy\-guided framework\.arXiv preprint arXiv:2509\.21267\.Cited by:[§2](https://arxiv.org/html/2609.12575#S2.p1.1),[item 2](https://arxiv.org/html/2609.12575#S4.I1.i2.p1.1),[§8\.2](https://arxiv.org/html/2609.12575#S8.SS2.p1.1)\.
- Karimet al\.\(2025\)A\. Karim, Q\. Wang, and Z\. YuanBeyond the score: uncertainty\-calibrated llms for automated essay assessment\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 19642–19647\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p1.1)\.
- Kimet al\.\(2025\)D\. Kim, H\. Ahn, Y\. Kim, and Y\. HanAnalyzing offensive language dataset insights from training dynamics and human agreement level\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 9780–9792\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Kleinet al\.\(2025\)L\. Klein, M\. Martin, A\. Brock, M\. Antoniak, M\. Walsh, J\. M\. Johnson, L\. Tilton, and D\. MimnoProvocations from the humanities for generative ai research\.arXiv preprint arXiv:2502\.19190\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1)\.
- Kommerset al\.\(2026a\)C\. Kommers, R\. Ahnert, M\. Antoniak, E\. Benetos, S\. Benford, M\. Bunz, B\. Caramiaux, S\. Concannon, M\. Disley, J\. Dobson, Y\. Du, E\. Duéñez\-Guzmán, K\. Francksen, E\. Gius, J\. W\. Y\. Gray, R\. Heuser, S\. Immel, R\. J\. So, S\. Leigh, D\. Livingston, H\. Long, M\. Martin, G\. Meyer, D\. Mihai, A\. Noel\-Hirst, K\. Ostherr, D\. Parker, Y\. Qin, J\. Ratcliff, E\. Robinson, K\. Rodriguez, A\. Sobey, T\. Underwood, A\. Vashistha, M\. Wilkens, Y\. Wu, Y\. Zheng, and D\. HemmentComputational hermeneutics: evaluating generative AI as a cultural technology\.Frontiers in Artificial Intelligence9,pp\. 1753041\.External Links:[Document](https://dx.doi.org/10.3389/frai.2026.1753041)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1),[§1](https://arxiv.org/html/2609.12575#S1.p3.1),[§1](https://arxiv.org/html/2609.12575#S1.p5.1),[item 2](https://arxiv.org/html/2609.12575#S4.I1.i2.p1.1),[§7\.2](https://arxiv.org/html/2609.12575#S7.SS2.p2.1),[§8\.2](https://arxiv.org/html/2609.12575#S8.SS2.p1.1)\.
- Kommers and DeDeo \(2025\)C\. Kommers and S\. DeDeoSense\-making, cultural scripts, and the inferential basis of meaningful experience\.InProceedings of the annual meeting of the cognitive science society,Vol\.47\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Kommerset al\.\(2026b\)C\. Kommers, E\. Duede, J\. Gordon, A\. Holtzman, T\. McNulty, S\. Stewart, L\. Thomas, R\. Jean So, and H\. LongWhy slop matters\.ACM AI Letters1\(1\),pp\. 1–6\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Kommerset al\.\(2025\)C\. Kommers, D\. Hemment, M\. Antoniak, J\. Z\. Leibo, H\. Long, E\. Robinson, and A\. SobeyMeaning is not a metric: using llms to make cultural context legible at scale\.arXiv preprint arXiv:2505\.23785\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Kommers and Holtzman \(2026\)C\. Kommers and A\. HoltzmanAI as entertainment\.arXiv preprint arXiv:2601\.08768\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Kunda and Rabkina \(2020\)M\. Kunda and I\. RabkinaCreative captioning: an AI grand challenge based on the Dixit board game\.Note:arXiv:2010\.00048External Links:2010\.00048Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p1.1)\.
- Leiboet al\.\(2024\)J\. Z\. Leibo, A\. S\. Vezhnevets, M\. Diaz, J\. P\. Agapiou, W\. A\. Cunningham, P\. Sunehag, J\. Haas, R\. Koster, E\. A\. Duéñez\-Guzmán, W\. S\. Isaac, G\. Piliouras, S\. M\. Bileschi, I\. Rahwan, and S\. OsinderoA theory of appropriateness with applications to generative artificial intelligence\.Note:arXiv:2412\.19010External Links:2412\.19010Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1),[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Linet al\.\(2025\)W\. Lin, J\. Roberts, Y\. Yang, S\. Albanie, Z\. Lu, and K\. HanGAMEBoT: transparent assessment of LLM reasoning in games\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 7656–7682\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.378)Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p1.1)\.
- Liuet al\.\(2023\)A\. Liu, Z\. Wu, J\. Michael, A\. Suhr, P\. West, A\. Koller, S\. Swayamdipta, N\. A\. Smith, and Y\. ChoiWe’re afraid language models aren’t modeling ambiguity\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 790–807\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.51)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1),[§1](https://arxiv.org/html/2609.12575#S1.p6.1),[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p1.1)\.
- Liuet al\.\(2025\)C\. C\. Liu, I\. Gurevych, and A\. KorhonenCulturally aware and adapted nlp: a taxonomy and a survey of the state of the art\.Transactions of the Association for Computational Linguistics13,pp\. 652–689\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Loet al\.\(2025\)K\. M\. Lo, Z\. Huang, Z\. Qiu, Z\. Wang, and J\. FuA closer look into mixture\-of\-experts in large language models\.InFindings of the Association for Computational Linguistics: NAACL 2025,Albuquerque, New Mexico,pp\. 4427–4447\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.251)Cited by:[§3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px2.p1.1)\.
- Loweet al\.\(2025\)R\. Lowe, J\. Edelman, T\. Zhi\-Xuan, O\. Klingefjord, E\. Hain, V\. Wang, A\. Sarkar, M\. A\. Bakker, F\. Barez, M\. Franklin,et al\.Full\-stack alignment: co\-aligning ai and institutions with thicker models of value\.In2nd workshop on models of human feedback for AI alignment,Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Marchalet al\.\(2022\)M\. Marchal, M\. Scholman, F\. Yung, and V\. DembergEstablishing annotation quality in multi\-label annotations\.InProceedings of the 29th international conference on computational linguistics,pp\. 3659–3668\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Metzgeret al\.\(2024\)L\. Metzger, L\. Miller, M\. Baumann, and J\. KrausEmpowering calibrated \(dis\-\) trust in conversational agents: a user study on the persuasive power of limitation disclaimers vs\. authoritative style\.InProceedings of the 2024 CHI conference on human factors in computing systems,pp\. 1–19\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1)\.
- Minet al\.\(2020\)S\. Min, J\. Michael, H\. Hajishirzi, and L\. ZettlemoyerAmbigQA: answering ambiguous open\-domain questions\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 5783–5797\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.466)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p1.1)\.
- Montesinos and Løvlie \(2026\)L\. Montesinos and A\. S\. LøvlieMachine learning as design material for music\-making\.InProceedings of the 2026 Designing Interactive Systems Conference,pp\. 4768–4784\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Namet al\.\(2025\)H\. Nam, J\. Ahn, K\. Ka, J\. Chung, and Y\. YuVAGUE: visual contexts clarify ambiguous expressions\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),Note:arXiv:2411\.14137Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Navigli \(2009\)R\. NavigliWord sense disambiguation: a survey\.ACM Computing Surveys41\(2\),pp\. 10:1–10:69\.External Links:[Document](https://dx.doi.org/10.1145/1459352.1459355)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1),[§1](https://arxiv.org/html/2609.12575#S1.p6.1)\.
- Nieet al\.\(2020\)Y\. Nie, X\. Zhou, and M\. BansalWhat can we learn from collective human opinions on natural language inference data?\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 9131–9143\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.734)Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Pavlick and Kwiatkowski \(2019\)E\. Pavlick and T\. KwiatkowskiInherent disagreements in human textual inferences\.Transactions of the Association for Computational Linguistics7,pp\. 677–694\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00293)Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Piantadosiet al\.\(2012\)S\. T\. Piantadosi, H\. Tily, and E\. GibsonThe communicative function of ambiguity in language\.Cognition122\(3\),pp\. 280–291\.External Links:[Document](https://dx.doi.org/10.1016/j.cognition.2011.10.004)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1)\.
- Pinkeret al\.\(2008\)S\. Pinker, M\. A\. Nowak, and J\. J\. LeeThe logic of indirect speech\.Proceedings of the National Academy of Sciences105\(3\),pp\. 833–838\.External Links:[Document](https://dx.doi.org/10.1073/pnas.0707192105)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Rolandet al\.\(2026\)E\. Roland, R\. J\. So, and H\. LongThe social ai author: modeling creativity and distinction in simulated cultural fields\.AI & SOCIETY41\(5\),pp\. 4333–4347\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Sandriet al\.\(2023\)M\. Sandri, E\. Leonardelli, S\. Tonelli, and E\. JežekWhy don’t you do it right? analysing annotators’ disagreement in subjective tasks\.InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics,pp\. 2428–2441\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Sennet \(2023\)A\. SennetAmbiguity\.InThe Stanford Encyclopedia of Philosophy,E\. N\. Zalta and U\. Nodelman \(Eds\.\),Note:[https://plato\.stanford\.edu/archives/sum2023/entries/ambiguity/](https://plato.stanford.edu/archives/sum2023/entries/ambiguity/)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1),[§1](https://arxiv.org/html/2609.12575#S1.p4.1)\.
- Sorensenet al\.\(2024\)T\. Sorensen, J\. Moore, J\. Fisher, M\. Gordon, N\. Mireshghallah, C\. M\. Rytting, A\. Ye, L\. Jiang, X\. Lu, N\. Dziri,et al\.A roadmap to pluralistic alignment\.arXiv preprint arXiv:2402\.05070\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Speer \(2022\)R\. SpeerRspeer/wordfreq: v3\.0\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.7199437),[Link](https://doi.org/10.5281/zenodo.7199437)Cited by:[§5\.2\.1](https://arxiv.org/html/2609.12575#S5.SS2.SSS1.p1.1)\.
- Strachanet al\.\(2024\)J\. W\. Strachan, D\. Albergo, G\. Borghini, O\. Pansardi, E\. Scaliti, S\. Gupta, K\. Saxena, A\. Rufo, S\. Panzeri, G\. Manzi,et al\.Testing theory of mind in large language models and humans\.Nature human behaviour8\(7\),pp\. 1285–1295\.Cited by:[§3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px3.p2.1)\.
- Sunstein \(1995\)C\. R\. SunsteinIncompletely theorized agreements\.Harvard Law Review108\(7\),pp\. 1733–1772\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p2.1)\.
- Taiet al\.\(2024\)R\. H\. Tai, L\. R\. Bentley, X\. Xia, J\. M\. Sitt, S\. C\. Fankhauser, A\. M\. Chicas\-Mosier, and B\. G\. MonteithAn examination of the use of large language models to aid analysis of textual data\.International Journal of Qualitative Methods23\.External Links:[Document](https://dx.doi.org/10.1177/16094069241231168)Cited by:[§6\.1](https://arxiv.org/html/2609.12575#S6.SS1.p1.1)\.
- Tanzawaet al\.\(2023\)Y\. Tanzawa, K\. Tsubokura, R\. Ohashi, H\. Sakurai, and K\. KobayashiEvaluation focused on the relevance of hints for implementing Dixit game AI\.InProceedings of the 37th Annual Conference of the Japanese Society for Artificial Intelligence \(JSAI 2023\),pp\. 2F5GS505\.External Links:[Document](https://dx.doi.org/10.11517/pjsai.JSAI2023.0%5F2F5GS505)Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p3.1)\.
- Tsaiet al\.\(2026\)C\. Tsai, N\. Tacconi, A\. D\. Wilson, and P\. AbtahiUncertain pointer: situated feedforward visualizations for ambiguity\-aware ar target selection\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,pp\. 1–31\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p6.1),[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p1.1)\.
- Umaet al\.\(2022\)A\. Uma, D\. Almanea, and M\. PoesioScaling and disagreements: bias, noise, and ambiguity\.Frontiers in Artificial Intelligence5,pp\. 818451\.Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p2.1)\.
- Underwood \(2025\)T\. UnderwoodThe impact of language models on the humanities and vice versa\.Nature Computational Science5\(9\),pp\. 695–697\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p3.1)\.
- van Heuvenet al\.\(2014\)W\. J\. B\. van Heuven, P\. Mandera, E\. Keuleers, and M\. BrysbaertSUBTLEX\-UK: a new and improved word frequency database for British English\.Quarterly Journal of Experimental Psychology67\(6\),pp\. 1176–1190\.External Links:[Document](https://dx.doi.org/10.1080/17470218.2013.850521)Cited by:[§5\.2\.1](https://arxiv.org/html/2609.12575#S5.SS2.SSS1.p1.1)\.
- Vatsakiset al\.\(2022\)D\. Vatsakis, P\. Mavromoustakos\-Blom, and P\. SpronckAn internet\-assisted Dixit\-playing AI\.InProceedings of the 17th International Conference on the Foundations of Digital Games \(FDG\),pp\. 1–7\.External Links:[Document](https://dx.doi.org/10.1145/3555858.3555863)Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p3.1),[§3\.2](https://arxiv.org/html/2609.12575#S3.SS2.p1.1)\.
- Veselovskyet al\.\(2026\)V\. Veselovsky, B\. Argın, B\. Stroebl, C\. Wendler, R\. West, J\. Evans, T\. L\. Griffiths, and A\. NarayananLocalized cultural knowledge is conserved and controllable in large language models\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 43152–43178\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1),[§7\.2](https://arxiv.org/html/2609.12575#S7.SS2.p2.1),[§8\.2](https://arxiv.org/html/2609.12575#S8.SS2.p1.1)\.
- Wanget al\.\(2025\)A\. Wang, J\. Morgenstern, and J\. P\. DickersonLarge language models that replace human participants can harmfully misportray and flatten identity groups\.Nature Machine Intelligence7\(3\),pp\. 400–411\.External Links:[Document](https://dx.doi.org/10.1038/s42256-025-00986-z)Cited by:[§7\.2](https://arxiv.org/html/2609.12575#S7.SS2.p2.1)\.
- Wei \(2023\)R\. WeiDixit player with Open CLIP\.Journal of Data Analysis and Information Processing11\(4\),pp\. 536–547\.External Links:[Document](https://dx.doi.org/10.4236/jdaip.2023.114027)Cited by:[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2609.12575#S2.SS2.p3.1)\.
- Wildenburget al\.\(2024\)F\. Wildenburg, M\. Hanna, and S\. PezzelleDo pre\-trained language models detect and understand semantic underspecification? Ask the DUST\!\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 9598–9613\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.572)Cited by:[§2\.1](https://arxiv.org/html/2609.12575#S2.SS1.p1.1)\.
- Wilkens \(2026\)M\. WilkensAI as a tool for simulation\-based experiments in literary studies\.arXiv preprint arXiv:2606\.02293\.Cited by:[§2](https://arxiv.org/html/2609.12575#S2.p1.1)\.
- Xiaoet al\.\(2025\)J\. Xiao, Z\. Li, X\. Xie, E\. Getzen, C\. Fang, Q\. Long, and W\. J\. SuOn the algorithmic bias of aligning large language models with rlhf: preference collapse and matching regularization\.Journal of the American Statistical Association120\(552\),pp\. 2154–2164\.Cited by:[§2](https://arxiv.org/html/2609.12575#S2.p1.1)\.
- Xiaoet al\.\(2023\)Z\. Xiao, X\. Yuan, Q\. V\. Liao, R\. Abdelghani, and P\. OudeyerSupporting qualitative analysis with large language models: combining codebook with GPT\-3 for deductive coding\.InCompanion Proceedings of the 28th International Conference on Intelligent User Interfaces \(IUI ’23 Companion\),pp\. 75–78\.External Links:[Document](https://dx.doi.org/10.1145/3581754.3584136)Cited by:[§3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px5.p1.1),[§6\.1](https://arxiv.org/html/2609.12575#S6.SS1.p1.1)\.
- Xie and Xie \(2026\)Y\. Xie and Y\. XieWhen artificial intelligence makes everything similar: the risks of content homogenization\.Chinese Journal of Sociology12\(2\),pp\. 157–172\.Cited by:[§2](https://arxiv.org/html/2609.12575#S2.p1.1)\.
- Yadavet al\.\(2025\)S\. Yadav, L\. Tilton, M\. Antoniak, T\. Arnold, J\. Li, S\. M\. Pawar, A\. Karamolegkou, S\. Frank, Z\. An, N\. Rostamzadeh,et al\.Evaluation of cultural competence of vision\-language models\.arXiv preprint arXiv:2505\.22793\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Yuan and Strohmaier \(2021\)Z\. Yuan and D\. StrohmaierCambridge at SemEval\-2021 task 2: neural WiC\-model with data augmentation and exploration of representation\.InProceedings of the 15th International Workshop on Semantic Evaluation \(SemEval\-2021\),A\. Palmer, N\. Schneider, N\. Schluter, G\. Emerson, A\. Herbelot, and X\. Zhu \(Eds\.\),Online,pp\. 730–737\.External Links:[Link](https://aclanthology.org/2021.semeval-1.96/),[Document](https://dx.doi.org/10.18653/v1/2021.semeval-1.96)Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p1.1),[§1](https://arxiv.org/html/2609.12575#S1.p6.1)\.
- Zhiet al\.\(2026\)J\. Zhi, H\. Long, R\. J\. So, and M\. LeeWhat does ai do for cultural interpretation? a randomized experiment on close reading poems with exposure to ai interpretation\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,pp\. 1–18\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Zhouet al\.\(2025\)N\. Zhou, D\. Bamman, and I\. L\. BleamanCulture is not trivia: sociocultural theory for cultural nlp\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 25869–25886\.Cited by:[§1](https://arxiv.org/html/2609.12575#S1.p5.1)\.
- Ziemset al\.\(2024\)C\. Ziems, W\. Held, O\. Shaikh, J\. Chen, Z\. Zhang, and D\. YangCan large language models transform computational social science?\.Computational Linguistics50\(1\),pp\. 237–291\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00502)Cited by:[§6\.1](https://arxiv.org/html/2609.12575#S6.SS1.p1.1)\.

## Appendix AStoryteller prompts

Both prompts were given as the system message together with the target card image\. Text in the two prompts is identical up to the point marked “…”; the social cognition prompt then adds the coaching shown\.

##### Minimal

> You are playing a game with four other players\. In the game, you secretly hold a TARGET card\. Your job is to be the storyteller: you must generate a clue that hints at your TARGET\. The other four players each hold a hand of cards and, after hearing your clue, will each submit one of their own cards as a decoy that might fit your clue\. All five cards \(your TARGET plus their four decoys\) will then be shuffled on the table, and each of them will vote for which one is the TARGET\. You want some of the other players to guess your TARGET correctly, but not all of them\. Your clue should be three to six words in length\. Respond with only the clue\.

##### Social cognition

> You are playing a game with four other players\. In the game, you secretly hold a TARGET card\. Your job is to be the storyteller: you must generate a clue that hints at your TARGET\. However, you don’t want the answer to be obvious\. The other four players … You want some of the other players to guess your TARGET correctly, but not all of them\. When generating the clue, think about things that could plausibly apply to many cards in a way that might mislead some players, while still favoring the TARGET\. An ideal clue should be: - •Three to six words in length - •Based on emotions, metaphors, allusions, and other forms of figurative or abstract language that could be applied to multiple cards - •Ambiguous enough that not all players will agree on the same card, but more likely to refer to your TARGET than to an unrelated card Think about how the other players will perceive your clue: Respond with only a clue that two of the four other players would correctly guess applies to the TARGET rather than to a decoy\.

## Appendix BBehavioural robustness check: a vision–language judge panel

Before adopting the rubric as the primary instrument, we scored clues behaviourally with a panel of judges that voted for the target card as a Dixit audience would\. This check used an earlier configuration of the pipeline: 200 four\-player rounds from the human dataset, Gemma\-3\-27B and Qwen3\.5\-35B\-A3B as storytellers under the two wordings of the social cognition prompt \(200×2×2=800200\\times 2\\times 2=800clues\), with clues clipped to a maximum of seven words\.

##### Judge panel

The panel comprised four open\-weight vision–language judges \(Qwen3\.5\-9B, Qwen3\-VL\-8B, Gemma\-3\-12B, and InternVL3\-8B\), each served independently\. For every clue, each judge saw the four table cards from the original round—the target plus the three decoys the human players had actually contributed—in a shuffled order together with the clue, and voted for the single card it believed the clue referred to \(greedy decoding, forced single\-index reply\)\. Votes that could not be parsed were discarded; the full four\-judge panel returned valid votes in 76% of the 1,000 judged rounds, and all metrics are computed over the valid votes of each round\. Valid\-vote counts are essentially identical across producer conditions \(means of 3\.60–3\.69 judges per round\), so discarded votes do not favour any producer\. Following Dixit’s scoring rule, a clue counts as*in the target range*when some but not all valid judges select the target card; clues that every judge resolves are overspecified \(“too obvious”\), and clues that no judge resolves are underspecified \(“too obscure”\)\. The same judge panel scored human\- and LLM\-generated clues on the same boards, so any difference between producers cannot be attributed to the judging procedure\. The four judges behave as a heterogeneous audience rather than four copies of one opinion: mean pairwise agreement on the voted card is0\.450\.45\(chance0\.250\.25\), and the panel is unanimous in only14%14\\%of full\-panel rounds\.

##### Validation against human audiences

The 200 human\-clue rounds come with the real audience’s outcome, which lets us ask how well the judge panel stands in for human guessers\. At the population level the two audiences agree closely: the panel places60\.5%60\.5\\%of human clues in the target range, against60\.0%60\.0\\%for the original human audiences on the same rounds\. At the level of individual rounds the correspondence is weaker: per\-round hit rates correlate positively but modestly \(Spearmanρ=\.29\\rho=\.29,p<\.001p<\.001\), and binary in\-range agreement between the two audiences is close to chance\. Some attenuation is inevitable when comparing single samples of two small audiences \(three human guessers vs\. four model judges\), but it does not account for all of the gap: a binomial simulation in which both audiences respond to the same per\-clue transparency predicts substantially higher correspondence \(agreement≈\.61\\approx\.61,ρ≈\.53\\rho\\approx\.53\) than we observe\. The panel is therefore not a stand\-in for how a particular human audience reads a particular clue, and we use it only for what its aggregate behaviour is shown to match: population\-level comparisons of clue producers on identical boards\.

### B\.1\.Model Comparison Robustness

Three checks indicate that these results do not depend on incidental features of the pipeline\. First, the two coders agree with each other on the4,5504\{,\}550double\-coded clues atr=\.66r=\.66\(Calibration\),\.73\.73\(Situatedness\),\.75\.75\(Literalness\) and\.84\.84\(Figurativeness\), with exact agreement of\.76\.76–\.96\.96; every calibration gap reported above is significant in the same direction with either coder alone, the sole exception being the Gemma social\-cognition bank, which one coder rates as equivalent to the human clues\. Second, the 350 boards comprise two subsets selected in different ways \(Section[3\.3](https://arxiv.org/html/2609.12575#S3.SS3.SSS0.Px1)\); the mean calibration of every bank differs by at most0\.100\.10between them\. Third, length does not explain calibration\. Within the LLM clues the correlation between word count and calibration rating isρ=\.11\\rho=\.11; Kimi\-VL’s minimal\-prompt clues are the shortest of any bank \(3\.53\.5words\) yet substantially overspecified \(δ=\.61\\delta=\.61\), and GLM’s two banks have identical length \(6\.46\.4words\) butδ=\.97\\delta=\.97and\.27\.27\. Length, in other words, does not account for the calibration gap\.

Similar Articles

Evaluating Chinese Ambiguity Understanding in Large Language Models

arXiv cs.CL

This paper introduces CHA-Gen, a Chinese ambiguity dataset grounded in Potential Ambiguity theory, and evaluates several LLMs on ambiguity detection, finding that models struggle but benefit from chain-of-thought prompting, and that instruction tuning induces overconfidence.

Language models suffer from a curse of ambiguity

arXiv cs.CL

This paper identifies a curse of ambiguity in language models, where more ambiguous next-token distributions are harder to learn, tracing this to architectural and learning roots and validating on synthetic and real data.