Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

arXiv cs.CL Papers

Summary

This paper evaluates LLMs' ability to recognize unspoken beliefs (implicatures) and their updates through implicature cancellation, introducing the expert-annotated ImplicatureX dataset. Results show LLMs lag behind humans, especially in natural scenarios.

arXiv:2607.25094v1 Announce Type: new Abstract: Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, [DatasetName], crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.
Original Article
View Cached Full Text

Cached at: 07/29/26, 09:54 AM

# Evaluating Communicative Belief Updates in Large Language Modelsvia Implicature Recognition and Cancellation
Source: [https://arxiv.org/html/2607.25094](https://arxiv.org/html/2607.25094)
Cesare Spinoso\-Di Piano1Verna Dankers1​\{\}^\{1\\,\\text\{\\faIcon\{coffee\}\}\}Marius Mosbach1​\{\}^\{1\\,\\text\{\\faIcon\{coffee\}\}\}Jackie Chi Kit Cheung1,2 1Mila \- Quebec AI Institute & McGill University,2Canada CIFAR AI Chair \{cesare\.spinoso, cheungja\}@mila\.quebec

###### Abstract

Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models \(LLMs\) and their users\. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through*implicatures*and to understand their updates through*implicature cancellation*: the pragmatic phenomenon whereby an utterance’s implied meaning is*weakened*or*negated*\. We create the first expert\-annotated implicature cancellation dataset,ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations\. We find that LLM belief update understanding lags behind that of humans, especially in more naturally\-occurring scenarios\. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form\. Overall, our study suggests that current LLMs have not yet reached human\-level understanding of unspoken beliefs and belief updates\.111Code and data available at[https://github\.com/cesare\-spinoso/ImplicatureX](https://github.com/cesare-spinoso/ImplicatureX)\.

Evaluating Communicative Belief Updates in Large Language Modelsvia Implicature Recognition and Cancellation

Cesare Spinoso\-Di Piano1Verna Dankers1​\{\}^\{1\\,\\text\{\\faIcon\{coffee\}\}\}Marius Mosbach1​\{\}^\{1\\,\\text\{\\faIcon\{coffee\}\}\}Jackie Chi Kit Cheung1,21Mila \- Quebec AI Institute & McGill University,2Canada CIFAR AI Chair\{cesare\.spinoso, cheungja\}@mila\.quebec

\\faIcon\{coffee\}\\faIcon\{coffee\}footnotetext:Equal contribution\.## 1Introduction

Natural language communication consists of threads of beliefs negotiated between interlocutors in a conversational common ground\(Lewis,[1979](https://arxiv.org/html/2607.25094#bib.bib7); Heim,[1982](https://arxiv.org/html/2607.25094#bib.bib9); Clark and Brennan,[1991](https://arxiv.org/html/2607.25094#bib.bib8); Stalnaker,[1998](https://arxiv.org/html/2607.25094#bib.bib12); Kamp,[1981](https://arxiv.org/html/2607.25094#bib.bib10)\)\. These beliefs are often introduced and updated implicitly through*implicatures*and*implicature cancellations*, whereby a previously implicated belief is*negated*or*weakened*\(Grice,[1975](https://arxiv.org/html/2607.25094#bib.bib14); Sperber and Wilson,[1986](https://arxiv.org/html/2607.25094#bib.bib5)\)\. For instance, when Bo asks their friend Aya whether they can help with something, the response “I think Cai was looking for someone to buy drinks\.” triggers an implicature that Cai needs Bo to buy drinks \(Figure[1](https://arxiv.org/html/2607.25094#S1.F1)\)\. However, Aya can cancel this implicature by saying “Though they must have gotten around to it by now\.”, at which point Bo should update their beliefs and act accordingly, e\.g\., by asking Cai to confirm\. Interlocutors’ beliefs can thus change from one utterance to the next\. Understanding these beliefs and how they are updated is thus of fundamental importance for successful interactions between large language models \(LLMs\) and human system users\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x1.png)Figure 1:An example of belief negotiation: A beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}in the common groundG\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}is updated as an implicature is triggered byu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}and then cancelled byu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\.In this work, we evaluate LLMs’ ability to understand such dynamically changing beliefs in language\. We focus on implicature recognition and cancellation, two of the most widely studied and fundamental phenomena involving belief updates in linguistics and cognitive science\(Grice,[1975](https://arxiv.org/html/2607.25094#bib.bib14); Sperber and Wilson,[1986](https://arxiv.org/html/2607.25094#bib.bib5)\)\. While previous studies have evaluated the ability of LLMs to perform belief updates of world knowledge and logical reasoning\(Rudingeret al\.,[2020](https://arxiv.org/html/2607.25094#bib.bib38); Hwanget al\.,[2021](https://arxiv.org/html/2607.25094#bib.bib39)\), similar evaluations in the communicative space remain underexplored\. Our work fills this gap by evaluating LLMs’ ability to identify beliefs triggered by implicatures and to update those implicated beliefs as a result of implicature cancellation\.

We create the first implicature cancellation dataset,ImplicatureX, consisting of271271expert\-annotated implicatures and corresponding cancellations\. Beyond synthetic two\-turn conversational implicatures, it contains naturally occurring scalar implicatures, discourse implicatures and multi\-turn dialogue conversational implicatures\. In addition, we conduct and include a crowdsourcing annotation of ourImplicatureXitems to measure implicature and cancellation recognition accuracies\.

Our results show that LLMs’ pragmatic reasoning falls short of a human understanding of unspoken beliefs conveyed by implicatures\. While the strongest LLMs we tested \(e\.g\., GPT\-5\.4 Thinking\) match human performance on implicature recognition of scalar and discourse implicatures, LLMs struggle to perform at a level above random chance on naturally\-occurring conversational implicatures\. Moreover, a control experiment reveals that, in many cases, LLMs successfully recognize implicatures without observing the context or the utterance\. We, therefore, question the extent to which their success is due to legitimate pragmatic reasoning\.

In a similar vein to our implicature recognition findings, we find that belief updates following implicature cancellation—which are readily made by humans—are especially challenging for LLMs in naturally\-occurring and realistic scenarios\. Additional control experiments reveal that the type of belief update—cancelling, leaving unchanged, or strengthening—and the way in which the update is triggered—explicitly or implicitly—affect the extent to which LLMs are able to revise their existing beliefs\. For instance, our results demonstrate that even when presented with explicit belief negations, strong LLMs \(e\.g\., Qwen 3 32B Thinking\) cannot match human accuracy in cancellation recognition\.

To summarize, we study the extent to which LLMs are able to understand unspoken beliefs via implicature recognition and belief updates via implicature cancellation, two fundamental properties of human communication\. We create the first implicature cancellation dataset,ImplicatureX, verified by linguistic experts and annotated with crowdsourced judgements\. Our evaluation reveals that LLMs fall short of both a human understanding of unspoken beliefs and of belief updates\. More broadly, our work suggests that LLMs may still not be able to understand the nuances of belief productions and negotiations, especially when they are unspoken and naturally\-occurring\.

## 2Related Work

#### Philosophy of language

Communication has long been thought to be an exercise of*belief negotiation*\(Grice,[1957](https://arxiv.org/html/2607.25094#bib.bib6); Lewis,[1979](https://arxiv.org/html/2607.25094#bib.bib7)\)\. In this vein, several seminal studies have posited that we communicate by making updates to an ever\-evolving tacitly agreed upon set of propositions, i\.e\., a*common ground*\(Heim,[1982](https://arxiv.org/html/2607.25094#bib.bib9); Clark and Brennan,[1991](https://arxiv.org/html/2607.25094#bib.bib8); Stalnaker,[1998](https://arxiv.org/html/2607.25094#bib.bib12); Kamp,[1981](https://arxiv.org/html/2607.25094#bib.bib10)\)\. As such, implicatures—beliefs about a possible intended meaning suggested by a speaker—and implicature negotiations—the cancellation and strengthening of these beliefs—are a core mechanism by which we add and make updates to the common ground\(Grice,[1975](https://arxiv.org/html/2607.25094#bib.bib14)\)\. In this work, we offer an experimental account of belief updates through implicature cancellations which, unlike implicature strengthenings\(Benotti,[2010](https://arxiv.org/html/2607.25094#bib.bib13)\), remain an understudied area of belief negotiation\.

#### Experimental pragmatics

Belief negotiation in human communication has received widespread attention in experimental pragmatics\. Initially studied in the context of referring expressions\(Clark and Wilkes\-Gibbs,[1986](https://arxiv.org/html/2607.25094#bib.bib15); Isaacs and Clark,[1987](https://arxiv.org/html/2607.25094#bib.bib16); Selten and Warglien,[2007](https://arxiv.org/html/2607.25094#bib.bib17); Deemteret al\.,[2012](https://arxiv.org/html/2607.25094#bib.bib41)\), negotiations of common ground beliefs have continued to be of central relevance to other pragmatic phenomena including non\-verbal communication\(Veinottet al\.,[1999](https://arxiv.org/html/2607.25094#bib.bib20); Clark and Krych,[2004](https://arxiv.org/html/2607.25094#bib.bib19)\), discourse relations\(Fetzer,[2018](https://arxiv.org/html/2607.25094#bib.bib18)\)and scalar implicatures\(Noveck,[2001](https://arxiv.org/html/2607.25094#bib.bib21)\)\. In particular, studies have shown that the inferences made from scalar implicatures are more easily and more readily accepted into the conversational common ground based on the context in which they appear and the prior beliefs held by interlocutors\(Brehenyet al\.,[2006](https://arxiv.org/html/2607.25094#bib.bib22); Grodneret al\.,[2010](https://arxiv.org/html/2607.25094#bib.bib26); Degen,[2015](https://arxiv.org/html/2607.25094#bib.bib23); Yanget al\.,[2018](https://arxiv.org/html/2607.25094#bib.bib24); Huang and Snedeker,[2018](https://arxiv.org/html/2607.25094#bib.bib25)\)\. Our study provides a continued experimental investigation of communicative belief updates through the phenomenon of implicature cancellation\.

#### Communicative beliefs in NLP

Notions surrounding communicative beliefs have long served as inspiration for language generation systems, including dialogue systems\(Allen and Perrault,[1980](https://arxiv.org/html/2607.25094#bib.bib28); Grosz and Sidner,[1986](https://arxiv.org/html/2607.25094#bib.bib27); Dale and Reiter,[1995](https://arxiv.org/html/2607.25094#bib.bib29)\), image captioning\(Andreas and Klein,[2016](https://arxiv.org/html/2607.25094#bib.bib30)\), and conversational agents\(Körneret al\.,[2025](https://arxiv.org/html/2607.25094#bib.bib31)\)\. Furthermore, implicatures and unspoken beliefs have been used extensively to evaluate the communicative competence of LLMs\(Jereticet al\.,[2020](https://arxiv.org/html/2607.25094#bib.bib32); Kabbara and Cheung,[2022](https://arxiv.org/html/2607.25094#bib.bib52); Ruiset al\.,[2023](https://arxiv.org/html/2607.25094#bib.bib34); Cho and mook Kim,[2024](https://arxiv.org/html/2607.25094#bib.bib33); Yueet al\.,[2024](https://arxiv.org/html/2607.25094#bib.bib35); Cong,[2024](https://arxiv.org/html/2607.25094#bib.bib36)\)\. Our contribution differs in that we leverage the cancellability of implicatures as a tractable space to evaluate the ability of LLMs to*update*unspoken beliefs\. Finally, while processes similar to implicature cancellation, such as abductive and defeasible reasoning, have been studied extensively\(Lascarides and Asher,[1991](https://arxiv.org/html/2607.25094#bib.bib37); Rudingeret al\.,[2020](https://arxiv.org/html/2607.25094#bib.bib38); Hwanget al\.,[2021](https://arxiv.org/html/2607.25094#bib.bib39)\), they differ from our focus of communicative belief updates\.

## 3Operationalizing Implicature and Implicature Cancellation

We view any communicative exchange \(spoken, written, signed, etc\.\) as consisting of a sequence of*utterances*u∈𝒰\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\in\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathcal\{U\}\}, i\.e\., any unit of language produced by a participant of the exchange\. Further, let𝒰∗\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathcal\{U\}^\{\*\}\}denote the set of all possible utterance sequences and𝐮=⟨u1,…,un⟩∈𝒰∗\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}=\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\_\{1\},\\ldots,\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\_\{n\}\\rangle\\in\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathcal\{U\}^\{\*\}\}a communicative history\. Each exchange takes place in a*context*c∈𝒞\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}\\in\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}\\mathcal\{C\}\}, which captures situational and communicative information relevant to interpreting the utterances\.222c\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}is assumed to be fixed for the duration of the exchange\.

Participants share a set of*beliefs*ℬ\\mathcal\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}B\}\}, i\.e\., propositional statements such as “Cai needs help buying drinks\.”, which represent what they mutually take to be true\. Formally, the*common ground*is a functionG:ℬ×𝒞×𝒰∗→\[0,1\]\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}:\\mathcal\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}B\}\}\\times\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}\\mathcal\{C\}\}\\times\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathcal\{U\}^\{\*\}\}\\rightarrow\[0,1\]that, givenc\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}and𝐮\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}, assigns to eachb∈ℬ\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\in\\mathcal\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}B\}\}a probability reflecting how strongly that belief is mutually held by the participants\.333Prior to any exchange, the common ground is given byG​\(b∣∅,∅\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\emptyset,\\emptyset\), which may assign uniform probability to allb∈ℬ\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\in\\mathcal\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}B\}\}, or reflect presupposed propositions based on the participants’ shared history and prior beliefs\(Anderson,[2018](https://arxiv.org/html/2607.25094#bib.bib43)\)\.

A beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}becomes part of the common ground via an*implicature*when utteranceu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}, produced in contextc\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}, makesb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}more probable than its negation according to the common ground, i\.e\.,G​\(b∣c,⟨u⟩\)\>G​\(¬b∣c,⟨u⟩\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\>\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\. A beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}is subsequently*cancelled*from the common ground via an*implicature cancellation*when a speaker extends the communicative history from⟨…,u⟩\\langle\\ldots,\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangleto⟨…,u,u×⟩\\langle\\ldots,\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\ranglewith a cancelling utteranceu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}, which weakens or negates the implicature triggered byu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}, i\.e\.,G​\(b∣c,⟨u,u×⟩\)<G​\(b∣c,⟨u⟩\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)<\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\. Thus, we define a*belief update*in the context of implicature cancellation asG​\(b∣c,⟨u,u×⟩\)<G​\(b∣c,⟨u⟩\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)<\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)when provided withu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}for an utteranceu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}for whichG​\(b∣c,⟨u⟩\)\>G​\(¬b∣c,⟨u⟩\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\>\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)holds\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x2.png)Figure 2:Composition of theImplicatureXdataset with sizes, sources and illustrative examples\. Note that the sizes reported in this figure reflect the dataset post cleaning \([Section˜4\.2](https://arxiv.org/html/2607.25094#S4.SS2)\)\.
## 4TheImplicatureXDataset

Here, we describe the composition of theImplicatureXimplicature cancellation dataset along with its expert and crowdsourcing annotation\.

### 4\.1Dataset Composition

ImplicatureXis an expert\-annotated dataset of271271items which include1Scalar Implicatures,2Discourse Implicaturesand3Conversational Implicatures\.[Figure˜2](https://arxiv.org/html/2607.25094#S3.F2)provides example stimuli per implicature type\. Below, we describe the dataset’s composition\.

1

Scalar Implicatures\.Our items are based on the well\-known pragmatic phenomenon where, given a pair of lexical items⟨\\langlew1w\_\{1\},w2w\_\{2\}⟩\\ranglewithw2w\_\{2\}semantically entailingw1w\_\{1\}, the use ofw1w\_\{1\}in an utterance indicates the negation ofw2w\_\{2\}\. For instance, the triggering utteranceu=u=“*Some*places nail them for tax\.”increases the probability of the beliefb=b=“*Not all*places nail them for tax\.”belonging to the common ground\. This belief can be revised by the speaker using a cancelling utteranceu×=u^\{\\times\}=“I think they*all*do that now\.”\. We sample 50 naturally\-occurring⟨\\langle*some*,*all*⟩\\ranglescalar implicatures from the Switchboard Corpus\(Godfreyet al\.,[1992](https://arxiv.org/html/2607.25094#bib.bib42)\)as originally collected byDegen \([2015](https://arxiv.org/html/2607.25094#bib.bib23)\)and manually add an implicature cancellation to each sampled item\.

2

Discourse Implicatures\.We posit that discourse relations, especially those which involve implicit causal relations, may be recast as*discourse*implicatures—i\.e\., implicatures found in traditional forms of discourse\. In particular, we identify and extract discourse implicatures from Wall Street Journal \(WSJ\) articles, using the corresponding implicit discourse relations annotated in the Penn Discourse Tree Bank \(PDTB\) corpus\(Prasadet al\.,[2019](https://arxiv.org/html/2607.25094#bib.bib46)\)\. Using the implicit PDTB causal discourse relations, we are able to identify excerpts such asc=c=“The Fuji apple has been extensively researched\.”andu=u=“Strains have been developed to age as gracefully as the Granny apples\.”which implicate the beliefb=b=“Fuji apples last as long as Granny apples\.”This belief is negated—or at least weakened—with the cancelling utteranceu×=u^\{\\times\}=“Though flaws in experimental controls now cast doubt on the apple’s longevity\.”We identify and annotate 31 WSJ article excerpts using implicit causal relations from the PDTB corpus\.

3

Conversational Implicatures\.Our conversational implicatures consist of implicatures which are triggered by utterances produced in a transcribed exchange between two conversational participants\. These 197 conversational implicatures are further subdivided into two categories which differ principally by their source:ASynthetic conversational implicatures andBNaturally\-occurring conversational implicatures\.

A

The synthetic conversational implicatures consist of an exchange between two participants,S1S\_\{1\}andS2S\_\{2\}\. The exchange begins typically with a question \(c=c=“Do you need help with anything?”\) which is followed by a response by the other participant \(u=u=“I thinkCCwas looking for someone to doXX\.”\), introducing an implicated belief into the common ground \(b=b=“You can help by doingXXforCC\.”\)\. This belief is updated via a cancelling utterance \(u×=u^\{\\times\}=“Though they must be done by now\.”\)\. We collect 146 such two\-turn conversational implicatures from several existing implicature understanding datasets\(George and Mamidi,[2020](https://arxiv.org/html/2607.25094#bib.bib47); Louiset al\.,[2020](https://arxiv.org/html/2607.25094#bib.bib48); Wilson and Bishop,[2021](https://arxiv.org/html/2607.25094#bib.bib49); Huet al\.,[2023](https://arxiv.org/html/2607.25094#bib.bib50)\)and manually add implicature cancellations\.

B

Our scenario\-based conversational implicatures also consist of an exchange betweenS1S\_\{1\}andS2S\_\{2\}\. In this case, the exchange is a naturally\-occurring discussion between two Switchboard Corpus participants\(Godfreyet al\.,[1992](https://arxiv.org/html/2607.25094#bib.bib42)\)\. We collect 51 naturally\-occurring conversational implicatures\. The cancelling utterances are produced by one of the participants within the exchange\. For instance,u=u=“We’re close to the golf course\.”\(b=b=“They go golfing near their home”\) is cancelled byu×=u^\{\\times\}=“Unfortunately the house is taking up time\.”

### 4\.2Expert Annotation

#### Annotation details

To validate the plausibility of the implicatures and the correctness of their cancellations inImplicatureX, we ran an expert annotation of our implicature cancellation items\. We elicit expert judgements for the implicatures’ plausibility as well as the cancellations’ correctness\. In addition, we elicit alternatives for implausible implicatures or incorrect cancellations \(where possible\)\.

The stimuli were subdivided into six batches of 53 items each of which included seven attention checks, i\.e\., items with either a clearly implausible implicature or incorrect cancellation\. The seven attention checks were different for each batch and their frequency per batch reflected the makeup of the dataset: one scalar implicature, one discourse implicature, four synthetic conversational implicatures and one naturally\-occurring conversational implicature\.

We hired two professionally trained linguists for two rounds of expert annotation\. In the first annotation round, both expert annotators were given the same batch in order to compute inter\-annotator agreement of implicature plausibility and cancellation correctness\. In the second round of annotations, one of the expert annotators was tasked to annotate the five remaining batches\.

Additional details regarding the expert annotation are presented in[Appendix˜A](https://arxiv.org/html/2607.25094#A1)\.

#### Annotation results

For the first round of annotation, we report the raw agreement, Cohen’s kappa \(κ\\kappa\) and the prevalence\-adjusted bias\-adjusted kappa \(PABAK\) for implicature plausibility and cancellation correctness444To compute this second agreement, we exclude items where at least one of the annotators marked the implicature as implausible\.in Table[1](https://arxiv.org/html/2607.25094#S4.T1)\. The raw agreement and PABAK are relatively high\. The relatively low Cohen’sκ\\kappais explained by the imbalance in class distributions for both the implicature plausibility and cancellation correctness annotation\(Byrtet al\.,[1993](https://arxiv.org/html/2607.25094#bib.bib54)\)\(See confusion matrices in Tables[3](https://arxiv.org/html/2607.25094#A1.T3)and[4](https://arxiv.org/html/2607.25094#A1.T4)in Appendix[A](https://arxiv.org/html/2607.25094#A1)\)\.

RawAgreementCohen’sκ\\kappaPABAKImplicaturePlausibility0\.850\.340\.70CancellationCorrectness0\.860\.320\.72

Table 1:Inter\-annotator agreement scores per annotation task\.The second round of annotation generated 18 implicature replacements and 27 cancellation replacements\. In seven cases, the items were flagged but the annotator was not able to provide alternatives \(e\.g\., the cancellations for the naturally\-occurring conversational implicatures\)\. After dropping and replacing the flagged items, this left us with a total of 271 expert\-annotated implicature cancellation items which are distributed as follows:1:4646items,2:3131items,3A:144144items,3B:5050items\.

### 4\.3Crowdsourcing Annotation

While our expert annotation served to validate the plausibility and correctness ofImplicatureX, we perform a crowdsourcing annotation to provide a realistic topline comparison to LLM understanding of implicature and implicature cancellation\.

#### Annotation details

To obtain aggregated human judgments ofImplicatureXbelief updates, we run a crowdsourcing annotation of the271271expert annotated items\. To do so, we elicit likelihood judgements for each implicature item given its corresponding context, triggering utterance, and cancelling utterance\. In particular, we create1616stimuli batches of3030items and two stimuli batches of3232items\. For each batch, half of the items contain the cancelling utterance and half do not\. We shuffle batch items in random order and ensure that for every batch no overlap exists between the items with and without the cancelling utterance\. In addition, we include and reuse1010attention checks for every batch, five whose likelihood rating should be high and five whose likelihood rating should be low\.

To implement our likelihood judgement elicitation, we generalize the approach from experimental pragmatics for studying the strength of scalar implicatures to our entire set of stimuli\. In particular, we generalizeDegen \([2015](https://arxiv.org/html/2607.25094#bib.bib23)\)’s approach and ask participants to rate the likelihood of the implicature on a seven point Likert scale with endpoints labeled as “absolutely impossible” and “absolutely certain” and individual points labeled as 1, 2, …, 7\.

We recruit9090participants from the Prolific platform based in Canada, the U\.S\. and the U\.K\. whose self\-reported native language is English\. In addition, we select participants with an undergraduate degree and who have taken part in at least 100 studies with a hit rate greater than or equal to99/10099/100\. Each batch is annotated by 5 participants\. We exclude a participant’s responses if they fail 3 or more attention checks\.

Additional details regarding the crowdsourcing annotation are presented in[Appendix˜B](https://arxiv.org/html/2607.25094#A2)\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x3.png)Figure 3:Histogram of average likelihood judgements z\-scores fromImplicatureXwith and without the cancelling utterance\.
#### Annotation results

Since human annotators are known to interpret Likert scales differently\(Cliff,[1993](https://arxiv.org/html/2607.25094#bib.bib53)\), we apply per\-participant z\-scoring to the Likert scale likelihood judgements\. We present z\-scored likelihood judgements averaged per item in[Figure˜3](https://arxiv.org/html/2607.25094#S4.F3)\.[Figure˜3](https://arxiv.org/html/2607.25094#S4.F3)shows a clear belief update given cancelling utterances: items without a cancelling utterance tend to have a positive z\-score while items with a cancelling utterance tend to have a negative z\-score\. Additional results related to our crowdsourcing annotation can be found in[Appendix˜B](https://arxiv.org/html/2607.25094#A2)\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x4.png)Figure 4:Implicature recognition accuracies of the largest and most modern models for each class of models tested\. Original denotes the accuracy using thetb\(c,𝐮\)\)\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)prompt and prior denotes the accuracy using thetb\(∅,∅\)\)\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\\emptyset,\\emptyset\)\)prompt\.

## 5Task Definitions and Setup

We leverage theImplicatureXdataset to evaluate LLM understanding of implicature, implicature cancellation, and of belief updates induced by these two pragmatic phenomena\. To do so, we use a multiple\-choice\-question prompt template to estimate specific values of the common ground functionG:ℬ×𝒞×𝒰∗→\[0,1\]\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}:\\mathcal\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}B\}\}\\times\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}\\mathcal\{C\}\}\\times\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathcal\{U\}^\{\*\}\}\\rightarrow\[0,1\]\. Given a beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}and a conversational history𝐮∈𝒰∗\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\\in\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathcal\{U\}^\{\*\}\}, we operationalizeG​\(b∣c,𝐮\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)via an LLMℳ\\mathcal\{M\}’s next\-token probability distribution over its token space𝒱\\mathcal\{V\}, such that:G​\(b∣c,𝐮\)≈Pℳ​\(b∣tb​\(c,𝐮\)\)\.\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\\approx P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)\.Here,tb​\(c,𝐮\)\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)is a prompt containing the verbalization ofb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}, the contextc\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}, the conversational history𝐮\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}, and a question about the truthfulness ofb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}in a multiple\-choice format using “True” and “False” as options \(Example prompt in[Figures˜17](https://arxiv.org/html/2607.25094#A4.F17)and[18](https://arxiv.org/html/2607.25094#A4.F18)in[Appendix˜D](https://arxiv.org/html/2607.25094#A4)\)\. To reduce positional bias, we followShiet al\.\([2025](https://arxiv.org/html/2607.25094#bib.bib51)\)and shuffle the order of the option choices, creating two copies for each prompt\.

#### Implicature recognition

We evaluate an LLMℳ\\mathcal\{M\}’simplicature recognition accuracyby computing the proportion of items for whichℳ\\mathcal\{M\}favors the implicated belief over its negation, i\.e\.,

Pℳ​\(b∣tb​\(c,𝐮\)\)\\displaystyle P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)\>Pℳ​\(¬b∣tb​\(c,𝐮\)\)\.\\displaystyle\>P\_\{\\mathcal\{M\}\}\(\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)\.

#### Cancellation recognition

We computeℳ\\mathcal\{M\}’scancellation recognition accuracyas the proportion of items for whichℳ\\mathcal\{M\}decreases its probability ofb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}upon observingu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}, i\.e\.,

Pℳ​\(b∣tb​\(c,⟨u⟩\)\)\\displaystyle P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)\>Pℳ​\(b∣tb​\(c,⟨u,u×⟩\)\)\.\\displaystyle\>P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)\)\.

#### Belief update

Lastly, we define a model’sbelief update accuracyas the proportion of items for which both

Pℳ​\(b∣tb​\(c,⟨u⟩\)\)\\displaystyle P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)\>Pℳ​\(¬b∣tb​\(c,⟨u⟩\)\)and\\displaystyle\>P\_\{\\mathcal\{M\}\}\(\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)\\quad\\text\{and\}Pℳ​\(b∣tb​\(c,⟨u⟩\)\)\\displaystyle P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)\>Pℳ​\(b∣tb​\(c,⟨u,u×⟩\)\)\\displaystyle\>P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)\)hold, i\.e\.,ℳ\\mathcal\{M\}both recognizes the implicature triggered byu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}and its subsequent cancellation byu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\.

We take these accuracy definitions to be necessary \(but not necessarily sufficient\) conditions of an understanding of implicature, implicature cancellation, and of the belief updates these phenomena, taken together, entail\.

#### Models

We conduct our experiments on both open and closed instruction\-tuned LLMs\. The open\-weight LLMs we evaluate include Gemma 3\(4B, 12B, 27B; Gemma\-Teamet al\.,[2025](https://arxiv.org/html/2607.25094#bib.bib4)\), Llama 3\.1 \(8B, 70B\), Llama 3\.2 \(3B\), Llama 3\.3\(70B; Grattafioriet al\.,[2024](https://arxiv.org/html/2607.25094#bib.bib3)\), Qwen 2\.5\(3B, 7B, 14B, 32B, 72B; Qwen\-Teamet al\.,[2025](https://arxiv.org/html/2607.25094#bib.bib2)\)and Qwen 3\(0\.6B, 1\.7B, 4B, 8B, 14B, 32B; Yanget al\.,[2025](https://arxiv.org/html/2607.25094#bib.bib1)\)\. For Qwen 3, we also experiment with the model’s reasoning functionality\. The closed\-source LLMs we evaluate include GPT\-5\.2 and GPT\-5\.4 with and without their reasoning functionality\.555All open\-weight models are downloaded from HuggingFace and all closed\-source models are accessed via OpenAI’s API\. See[Table5](https://arxiv.org/html/2607.25094#A3.T5)for full model names\.

#### ComputingPℳ​\(b∣tb​\(c,𝐮\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)

For open\-weight LLMs,Pℳ​\(b∣tb​\(c,𝐮\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)is computed by extracting the logits corresponding to each option and renormalizing them, averaging over both order shuffles\. For closed\-source models,Pℳ​\(b∣tb​\(c,𝐮\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}\)\)is computed by sampling and using the resulting relative frequencies, sampling five times for each order shuffle\. In our experimental setup,𝐮\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}is instantiated as a specific utterance sequence, e\.g\.,⟨u⟩\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangleto estimateG​\(b∣c,⟨u⟩\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\), or⟨u,u×⟩\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangleto estimateG​\(b∣c,⟨u,u×⟩\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\), allowing us to evaluate how each additional utterance affects the common ground probability assigned tob\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\.

## 6Implicature Recognition Experiment

Here, we present the experiments and results for the implicature recognition task\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x5.png)Figure 5:Cancellation recognition accuracies of the largest and most modern models for each class of models tested\.![Refer to caption](https://arxiv.org/html/2607.25094v1/x6.png)Figure 6:Belief update accuracies of the largest and most modern models for each class of models tested\.### 6\.1Human Topline and Prior Common Ground Control

We use our crowdsourcing annotation results from Section[4\.3](https://arxiv.org/html/2607.25094#S4.SS3)to compute an implicature recognitionhuman accuracy topline\. To do so, we compute the proportion of items for which the z\-scored Likert ratings averaged across participants are above0\.

We run an additional experiment estimating theprior common groundG​\(b∣∅,∅\)\{\\color\[rgb\]\{0,0,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0,0\}G\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\emptyset,\\emptyset\)for every item, i\.e\.,Pℳ​\(b∣tb​\(∅,∅\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\\emptyset,\\emptyset\)\), by removingc\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}and𝐮\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}\\mathbf\{u\}\}from the prompt and rewording the question to ask the model about its prior knowledge\. We run this experiment to isolate the effect of the LLM’s prior belief on implicature recognition\. For instance, in the case of scalar implicatures, this prior control may reveal to what extent the “not all” pragmatic interpretation of “some” has been “memorized”\.

### 6\.2Results

The results for the implicature recognition task experiment as well as the prior belief control experiment are shown in Figure[4](https://arxiv.org/html/2607.25094#S4.F4)\. Overall, we see that models are able to recognize scalar implicatures and discourse implicatures similarly to humans with Gemma 3 27B matching human scalar implicature recognition accuracy of1\.01\.0and Llama 3\.3 70B outperforming the human discourse implicature recognition accuracy of0\.900\.90by0\.040\.04\. On the other hand, models tend to perform worse on conversational implicatures\. Compared to the human accuracies of0\.910\.91and0\.780\.78on synthetic and naturally\-occurring implicatures, the best\-performing models achieve implicature recognition accuracies of0\.810\.81\(GPT\-5\.4 Thinking\) and0\.600\.60\(Llama 3\.3 70B\), respectively\. Furthermore, all of the models except for Llama 3\.3 70B perform worse than random at identifying naturally\-occurring implicatures, demonstrating that understanding unspoken beliefs remains a challenging task for LLMs\. Full results are shown in Table[6](https://arxiv.org/html/2607.25094#A6.T6)\(Appendix[F](https://arxiv.org/html/2607.25094#A6)\)\.

#### LLMs have strong priors for certain types of implicatures

In addition, our prior belief experiment reveals that models have a moderately strong prior for synthetic conversational implicatures and an extremely strong prior for⟨\\langle*some*,*all*⟩\\ranglescalar implicatures\. For instance, both Llama 3\.3 70B and Qwen 2\.5 72B are able to achieve perfect implicature recognition accuracy on scalar implicature items*without*having access to their corresponding utterances\. This result suggests that the performance of models on implicature recognition may not just stem from an understanding of implicated beliefs, but also from the prior knowledge these models may have about these implicated beliefs\. We leave a thorough investigation of the interaction between LLMs’ pragmatic understanding and their prior knowledge to future work\.

## 7Cancellation Recognition and Belief Update Experiment

Next, we present the experiments and results for the cancellation recognition and belief update tasks\.

### 7\.1Human Topline and Controls

We use our crowdsourcing annotation results from Section[4\.3](https://arxiv.org/html/2607.25094#S4.SS3)to compute cancellation recognition and belief updatehuman accuracy topline\. For the cancellation accuracy, we compute the proportion of items for which the average per\-item z\-scored Likert rating decreases when the cancelling utterance is introduced\. For thebelief update accuracy, we compute the proportion of items for which the average per\-item z\-scored Likert rating is above0given the triggering utterance and subsequently decreases given the cancelling utterance\.

#### Form control

We evaluate the extent to which the form of the cancellation affects an LLM’s success in performing cancellation recognition\. To do so, we create a new set of stimuli,Implicature⊥, in which cancelling utterances have the formu⊥=\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\bot\}\}=“d\+¬bd\+\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}” whereddis a discourse marker \(e\.g\.,in fact,actually, etc\.\) and¬b\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}denotes an explicit negation of the implicated beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}triggered byu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\.

#### Update type control

We also evaluate the extent to which LLM belief update accuracy is affected by the type of follow\-up utterance, i\.e\., whether the utterance that followsu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}cancels, strengthens, or leavesb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}unchanged\. To do so, we create two new sets of stimuli: \(1\)Implicature\+, in which cancelling utterancesu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}are replaced with strengthening utterancesu\+\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\+\}\}, i\.e\., utterances that strengthenb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}by providing information reinforcing its plausibility, and \(2\)Implicature≈, in whichu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}is replaced with a randomly sampled utteranceu≈\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\approx\}\}of the same stimuli type that leaves the beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}unchanged\. We evaluate a model’s ability to perform belief strengthening by computing the proportion ofImplicature\+items for whichPℳ​\(b∣tb​\(c,⟨u,u\+⟩\)\)\>Pℳ​\(b∣tb​\(c,⟨u⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\+\}\}\\rangle\)\)\>P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\), and its ability to leave its belief unchanged by computing the proportion ofImplicature≈items for which

Pℳ​\(b∣tb​\(c,⟨u,u≈⟩\)\)\\displaystyle P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\approx\}\}\\rangle\)\)\>Pℳ​\(¬b∣tb​\(c,⟨u,u≈⟩\)\)\\displaystyle\>P\_\{\\mathcal\{M\}\}\(\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\approx\}\}\\rangle\)\)⇔\\displaystyle\\iffPℳ​\(b∣tb​\(c,⟨u⟩\)\)\\displaystyle P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)\>Pℳ​\(¬b∣tb​\(c,⟨u⟩\)\)\.\\displaystyle\>P\_\{\\mathcal\{M\}\}\(\\neg\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)\.
We provide examples for the different experimental controls in Table[2](https://arxiv.org/html/2607.25094#S7.T2)\. Additional details regarding the composition of these control datasets can be found in[Appendix˜E](https://arxiv.org/html/2607.25094#A5)\.

Control TypeFollow\-Up Utteranceu\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u\}Originalu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}: “Though they must be done by now\.”Strengtheningu\+\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\+\}\}: “Maybe go over and ask them?”Unchangingu≈\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\approx\}\}: “Thankfully, I submitted before my laptop broke\.”Negationu⊥\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\bot\}\}: “Though, to be clear, I’m not saying that you could help by doingXXforCC\.”Table 2:Control conditions for the belief update task, with fixed contextc=c=“Do you need help with anything?”, utteranceu=u=“I thinkCCwas looking for someone to doXX\.”and beliefb=b=“You can help by doingXXforCC\.”

### 7\.2Results

We present model performance on cancellation recognition in Figure[5](https://arxiv.org/html/2607.25094#S6.F5)\. Overall, LLMs perform cancellation recognition at levels close to human performance, with some models slightly surpassing human cancellation recognition accuracy—e\.g\., Qwen 3 32B achieves a cancellation accuracy of0\.940\.94on discourse implicatures whereas humans achieve an accuracy of0\.900\.90\. However, models perform substantially worse when presented with naturally\-occurring cancellations\. In this setting, the best model, Qwen 3 32B Thinking, achieves a cancellation recognition accuracy of0\.840\.84, compared to human performance of0\.920\.92\.

We also report results for belief updating in Figure[6](https://arxiv.org/html/2607.25094#S6.F6)\. When evaluating the joint tasks of implicature and cancellation recognition, model performance falls sharply, especially for conversational implicatures\. For instance, the model with highest belief update accuracy, Qwen 3 32B Thinking, achieves belief update accuracies of0\.760\.76and0\.400\.40on synthetic and naturally\-occurring conversational implicatures, respectively, whereas humans achieve0\.900\.90and0\.720\.72\. Full results can be found in Tables[7](https://arxiv.org/html/2607.25094#A6.T7)and[8](https://arxiv.org/html/2607.25094#A6.T8)in Appendix[F](https://arxiv.org/html/2607.25094#A6)\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x7.png)Figure 7:Cancellation recognition accuracy of models on the naturally\-occurring conversational implicatures using the original cancelling utterances \(u×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\) and the cancelling utterances with explicit negation \(u⊥\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\bot\}\}\)\.![Refer to caption](https://arxiv.org/html/2607.25094v1/x8.png)Figure 8:Update recognition accuracy of models on the naturally\-occurring conversational implicatures using follow\-up utterances of different types: cancelling \(u×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\), unchanging \(u≈\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\approx\}\}\), and strengthening \(u\+\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\+\}\}\)\.#### Explicit negation does not lead to perfect LLM cancellation recognition

We compare model cancellation recognition accuracy when using the originalImplicatureXstimuli and theImplicature⊥variant, which contain explicitly negating cancelling utterances\. We show the cancellation recognition accuracies on the naturally\-occurring conversational implicatures in Figure[7](https://arxiv.org/html/2607.25094#S7.F7)and on the full set of implicature types in Figure[19](https://arxiv.org/html/2607.25094#A6.F19)in Appendix[F](https://arxiv.org/html/2607.25094#A6)\. While the implicature cancellations with explicit negations lead to higher cancellation recognition accuracies, we find that in most cases — aside from scalar implicatures — the cancellation recognition accuracy falls short of its theoretical upper bound of one\. This suggests that factors beyond pragmatic competence, such as difficulties in negation interpretation, may also contribute to observed cancellation recognition errors\.

#### LLMs are better at maintaining existing beliefs than changing them

We show the belief update accuracies using the cancelling, strengthening and unchanging utterances on naturally\-occurring conversational implicatures in Figure[8](https://arxiv.org/html/2607.25094#S7.F8)and on the entire set of implicature types in Figure[20](https://arxiv.org/html/2607.25094#A6.F20)in Appendix[F](https://arxiv.org/html/2607.25094#A6)\. Overall, we observe that LLMs struggle to update their beliefs both with strengthening and cancelling utterances\. However, they have relatively little difficulty maintaining their beliefs when presented with an irrelevant follow\-up utterance\. For instance, while GPT\-5\.4 maintains its belief with an accuracy of0\.960\.96on naturally\-occurring conversational implicatures, its update accuracies when presented with strengthening and cancelling utterances are both at0\.320\.32\. This result suggests that current LLMs are better at maintaining beliefs than at changing them\.

## 8Conclusion

In conclusion, in this paper we study LLM understanding of communicative beliefs and belief updates via implicature and cancellation recognition\. To this end, we develop the first implicature cancellation dataset,ImplicatureX, and show that LLMs fall short of a human understanding of unspoken beliefs and belief updates\. Control experiments reveal key weaknesses; for instance, a reliance on prior biases, an inability to reconcile explicit updates and a dependence on update type\. Our results highlight a critical gap between current LLM capabilities and human\-level pragmatic reasoning which calls for further research into how models manage evolving communicative beliefs\.

## Limitations

We identify three main limitations with our work\. Firstly, although we thoroughly reviewed whether the implicatures and corresponding cancellations are, in fact, annotated as such by both experts and lay annotators, whether or not an utterance cancels a prior statement remains open to interpretation\. With five annotations per stimulus, we consider our results to be reliable, but we encourage future work to further explore this, particularly focusing on the examples for which our annotators demonstrate disagreement\. Expanding the number of annotations could clarify whether or not such disagreement is due to legitimate stimulus ambiguity or instead to noise in our data\.

Secondly, we identified that models’ implicature recognition performance is strongly affected by their prior beliefs: without knowing utteranceu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}for⟨\\langlesome,all⟩\\ranglestatements, they will promptly score corresponding beliefb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}highly\. As a result, we cannot say with certainty that their implicature recognition accuracy is not affected by this; we leave disentangling scalar implicature reasoning from prior beliefs to future work\.

Lastly, we elicited LLM judgements by extracting probabilities from their output distributions\. Alternatively, one could provide the LLM with utterances and cancelling utterances, and ask them to further expand the provided discourse\. Consistency vs\. inconsistency with the beliefs implicitly conveyed in such generated text could reveal whether the LLMsreallymanage to remain consistent with those beliefs\. We did not opt for this type of evaluation due to its lack of scalability, i\.e\., high\-quality human annotations of generated text would be needed to assess the nuances behind the LLM responses\.

## Acknowledgements

The authors would like to thank the reviewers for their valuable comments\. We would also like to thank Gaurav Kamath and Austin Kraft for their helpful comments on an earlier version of this paper\. This work was supported by the Fonds de Recherche du Québec – Nature et Technologies \(FRQNT\), the Natural Sciences and Engineering Research Council of Canada \(NSERC\), the IVADO Postdoctoral Research Funding Program, and the Canada CIFAR AI Chair Program\. We acknowledge material support from NVIDIA Corporation in the form of computational resources provided to Mila\.

## References

- Analyzing intention in utterances\.Artificial intelligence15\(3\),pp\. 143–178\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0004-3702%2880%2990042-9)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- C\. Anderson \(2018\)Essentials of linguistics\.McMaster University\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/978-3-476-05678-8)Cited by:[footnote 3](https://arxiv.org/html/2607.25094#footnote3)\.
- J\. Andreas and D\. Klein \(2016\)Reasoning about pragmatics with neural listeners and speakers\.InProceedings of the 2016 conference on empirical methods in natural language processing,pp\. 1173–1182\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.18653/v1/d16-1125)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Benotti \(2010\)Implicature as an interactive process\.Ph\.D\. Thesis,Université Henri Poincaré\-Nancy I\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.70675/1880a7d1z6e5fz4f02z832dz3b4a0d293ebc)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- R\. Breheny, N\. Katsos, and J\. Williams \(2006\)Are generalised scalar implicatures generated by default? An on\-line investigation into the role of context in generating pragmatic inferences\.Cognition100\(3\),pp\. 434–463\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cognition.2005.07.003)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- T\. Byrt, J\. Bishop, and J\. B\. Carlin \(1993\)Bias, prevalence and kappa\.Journal of clinical epidemiology46\(5\),pp\. 423–429\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0895-4356%2893%2990018-V)Cited by:[§4\.2](https://arxiv.org/html/2607.25094#S4.SS2.SSS0.Px2.p1.2)\.
- Y\. Cho and S\. mook Kim \(2024\)Pragmatic inference of scalar implicature by LLMs\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 4: Student Research Workshop\),pp\. 10–20\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.18653/v1/2024.acl-srw.2)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- H\. H\. Clark and S\. E\. Brennan \(1991\)Grounding in communication\.\.Perspectives on Socially Shared Cognition\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1037/10096-006)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- H\. H\. Clark and M\. A\. Krych \(2004\)Speaking while monitoring addressees for understanding\.Journal of memory and language50\(1\),pp\. 62–81\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jml.2003.08.004)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- H\. H\. Clark and D\. Wilkes\-Gibbs \(1986\)Referring as a collaborative process\.Cognition22\(1\),pp\. 1–39\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0010-0277%2886%2990010-7)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Cliff \(1993\)Dominance statistics: Ordinal analyses to answer ordinal questions\.\.Psychological bulletin114\(3\),pp\. 494\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1037/0033-2909.114.3.494)Cited by:[§4\.3](https://arxiv.org/html/2607.25094#S4.SS3.SSS0.Px2.p1.1)\.
- Y\. Cong \(2024\)Manner implicatures in large language models\.Scientific Reports14\(1\),pp\. 29113\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41598-024-80571-3)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- R\. Dale and E\. Reiter \(1995\)Computational interpretations of the Gricean maxims in the generation of referring expressions\.Cognitive science19\(2\),pp\. 233–263\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/0364-0213%2895%2990018-7)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- K\. v\. Deemter, A\. Gatt, I\. v\. d\. Sluis, and R\. Power \(2012\)Generation of referring expressions: Assessing the incremental algorithm\.Cognitive science36\(5\),pp\. 799–836\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/j.1551-6709.2011.01205.x)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Degen \(2015\)Investigating the distribution of some \(but not all\) implicatures using corpora and web\-based methods\.Semantics and Pragmatics8,pp\. 11–1\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.3765/sp.8.11)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p2.13),[§4\.3](https://arxiv.org/html/2607.25094#S4.SS3.SSS0.Px1.p2.1)\.
- A\. Fetzer \(2018\)The linguistic realization of contrastive discourse relations in context: Contextualization and discourse common ground\.Modélisation et utilisation du contexte\.External Links:[Link](https://www.openscience.fr/The-Linguistic-Realization-of-Contrastive-Discourse-Relations-in-Context)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- Gemma\-Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot \(2025\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§5](https://arxiv.org/html/2607.25094#S5.SS0.SSS0.Px4.p1.1)\.
- E\. J\. George and R\. Mamidi \(2020\)Conversational implicatures in english dialogue: Annotated dataset\.Procedia Computer Science171,pp\. 2316–2323\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.procs.2020.04.251)Cited by:[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p5.10)\.
- J\. J\. Godfrey, E\. C\. Holliman, and J\. McDaniel \(1992\)SWITCHBOARD: Telephone speech corpus for research and development\.InProceedings of the 1992 IEEE International Conference on Acoustics, Speech and Signal Processing \- Volume 1,ICASSP’92,USA,pp\. 517–520\.External Links:ISBN 0780305329,[Document](https://dx.doi.org/https%3A//doi.org/10.1109/icassp.1992.225858)Cited by:[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p2.13),[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p6.5)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The Llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§5](https://arxiv.org/html/2607.25094#S5.SS0.SSS0.Px4.p1.1)\.
- H\. P\. Grice \(1957\)Meaning\.The philosophical review66\(3\),pp\. 377–388\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.2307/2182440)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- H\. P\. Grice \(1975\)Logic and conversation\.InSpeech acts,pp\. 41–58\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1163/9789004368811%5F003)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§1](https://arxiv.org/html/2607.25094#S1.p2.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- D\. J\. Grodner, N\. M\. Klein, K\. M\. Carbary, and M\. K\. Tanenhaus \(2010\)“Some,” and possibly all, scalar inferences are not delayed: Evidence for immediate pragmatic enrichment\.Cognition116\(1\),pp\. 42–55\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cognition.2010.03.014)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- B\. J\. Grosz and C\. L\. Sidner \(1986\)Attention, intentions, and the structure of discourse\.Computational linguistics12\(3\),pp\. 175–204\.External Links:[Link](https://aclanthology.org/J86-3001/)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- I\. R\. Heim \(1982\)The semantics of definite and indefinite noun phrases\.University of Massachusetts Amherst\.External Links:[Link](https://www.proquest.com/openview/9533c80af7894f76eb795c9de2cfa18f/1?pq-origsite=gscholar&cbl=18750&diss=y)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Hu, S\. Floyd, O\. Jouravlev, E\. Fedorenko, and E\. Gibson \(2023\)A fine\-grained comparison of pragmatic language understanding in humans and language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 4194–4213\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.230)Cited by:[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p5.10)\.
- Y\. T\. Huang and J\. Snedeker \(2018\)Some inferences still take time: prosody, predictability, and the speed of scalar implicatures\.Cognitive psychology102,pp\. 105–126\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cogpsych.2018.01.004)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- J\. D\. Hwang, C\. Bhagavatula, R\. Le Bras, J\. Da, K\. Sakaguchi, A\. Bosselut, and Y\. Choi \(2021\)\(Comet\-\)Atomic 2020: on symbolic and neural commonsense knowledge graphs\.InProceedings of the AAAI conference on artificial intelligence,pp\. 6384–6392\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1609/aaai.v35i7.16792)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p2.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- E\. A\. Isaacs and H\. H\. Clark \(1987\)References in conversation between experts and novices\.\.Journal of experimental psychology: general116\(1\),pp\. 26\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1037/0096-3445.116.1.26)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- P\. Jeretic, A\. Warstadt, S\. Bhooshan, and A\. Williams \(2020\)Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 8690–8705\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.18653/v1/2020.acl-main.768)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Kabbara and J\. C\. K\. Cheung \(2022\)Investigating the performance of transformer\-based NLI models on presuppositional inferences\.InProceedings of the 29th International Conference on Computational Linguistics,N\. Calzolari, C\. Huang, H\. Kim, J\. Pustejovsky, L\. Wanner, K\. Choi, P\. Ryu, H\. Chen, L\. Donatelli, H\. Ji, S\. Kurohashi, P\. Paggio, N\. Xue, S\. Kim, Y\. Hahm, Z\. He, T\. K\. Lee, E\. Santus, F\. Bond, and S\. Na \(Eds\.\),Gyeongju, Republic of Korea,pp\. 779–785\.External Links:[Link](https://aclanthology.org/2022.coling-1.65/)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- H\. Kamp \(1981\)A theory of truth and semantic representation\.Proceedings of the Third Amsterdam Colloquium,pp\. 277–322\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1093/oso/9780195136975.003.0013)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Körner, A\. Tolzin, A\. Janson, J\. M\. Leimeister, and R\. Rummer \(2025\)Common ground improves learning with conversational agents\.Behaviour & Information Technology,pp\. 1–17\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1080/0144929x.2025.2541222)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Lascarides and N\. Asher \(1991\)Discourse relations and defeasible knowledge\.In29th Annual Meeting of the Association for Computational Linguistics,pp\. 55–62\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.3115/981344.981352)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Lewis \(1979\)Scorekeeping in a language game\.Journal of philosophical logic8\(1\),pp\. 339–359\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/bf00258436)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Louis, D\. Roth, and F\. Radlinski \(2020\)“I’d rather just go to bed”: Understanding indirect answers\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 7411–7425\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.18653/v1/2020.emnlp-main.601)Cited by:[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p5.10)\.
- I\. A\. Noveck \(2001\)When children are more logical than adults: Experimental investigations of scalar implicature\.Cognition78\(2\),pp\. 165–188\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/s0010-0277%2800%2900114-1)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Prasad, B\. Webber, A\. Lee, and A\. Joshi \(2019\)Penn Discourse Treebank Version 3\.0\.Linguistic Data Consortium,Philadelphia\.Note:LDC2019T05Web DownloadExternal Links:[Document](https://dx.doi.org/10.35111/qebf-gk47)Cited by:[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p3.4)\.
- Qwen\-Team, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. Qiu \(2025\)Qwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§5](https://arxiv.org/html/2607.25094#S5.SS0.SSS0.Px4.p1.1)\.
- R\. Rudinger, V\. Shwartz, J\. D\. Hwang, C\. Bhagavatula, M\. Forbes, R\. Le Bras, N\. A\. Smith, and Y\. Choi \(2020\)Thinking like a skeptic: Defeasible inference in natural language\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 4661–4675\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.18653/v1/2020.findings-emnlp.418)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p2.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Ruis, A\. Khan, S\. Biderman, S\. Hooker, T\. Rocktäschel, and E\. Grefenstette \(2023\)The goldilocks of pragmatic understanding: Fine\-tuning strategy matters for implicature resolution by LLMs\.Advances in Neural Information Processing Systems36,pp\. 20827–20905\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.52202/075280-0913)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.
- R\. Selten and M\. Warglien \(2007\)The emergence of simple languages in an experimental coordination game\.Proceedings of the National Academy of Sciences104\(18\),pp\. 7361–7366\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1073/pnas.0702077104)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- L\. Shi, C\. Ma, W\. Liang, X\. Diao, W\. Ma, and S\. Vosoughi \(2025\)Judging the judges: A systematic study of position bias in LLM\-as\-a\-judge\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,K\. Inui, S\. Sakti, H\. Wang, D\. F\. Wong, P\. Bhattacharyya, B\. Banerjee, A\. Ekbal, T\. Chakraborty, and D\. P\. Singh \(Eds\.\),Mumbai, India,pp\. 292–314\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.ijcnlp-long.18),ISBN 979\-8\-89176\-298\-5Cited by:[§5](https://arxiv.org/html/2607.25094#S5.p1.12)\.
- D\. Sperber and D\. Wilson \(1986\)Relevance: Communication and Cognition\.Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§1](https://arxiv.org/html/2607.25094#S1.p2.1)\.
- R\. Stalnaker \(1998\)On the representation of context\.Journal of logic, language and information7\(1\),pp\. 3–19\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1023/a%3A1008254815298)Cited by:[§1](https://arxiv.org/html/2607.25094#S1.p1.1),[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px1.p1.1)\.
- E\. S\. Veinott, J\. Olson, G\. M\. Olson, and X\. Fu \(1999\)Video helps remote work: Speakers who need to negotiate common ground benefit from seeing each other\.InProceedings of the SIGCHI conference on Human Factors in Computing Systems,pp\. 302–309\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1145/302979.303067)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- A\. C\. Wilson and D\. V\. Bishop \(2021\)“Second guessing yourself all the time about what they really mean…”: Cognitive differences between autistic and non\-autistic adults in understanding implied meaning\.Autism Research14\(1\),pp\. 93–101\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.31234/osf.io/f6ksu)Cited by:[§4\.1](https://arxiv.org/html/2607.25094#S4.SS1.p5.10)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§5](https://arxiv.org/html/2607.25094#S5.SS0.SSS0.Px4.p1.1)\.
- X\. Yang, U\. Minai, and R\. Fiorentino \(2018\)Context\-sensitivity and individual differences in the derivation of scalar implicature\.Frontiers in psychology9,pp\. 1720\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.3389/fpsyg.2018.01720)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Yue, S\. Song, X\. Cheng, and H\. Hu \(2024\)Do large language models understand conversational implicature \- A case study with a chinese sitcom\.InProceedings of the 23rd Chinese National Conference on Computational Linguistics \(Volume 1: Main Conference\),pp\. 1270–1285\.External Links:[Link](https://aclanthology.org/2024.ccl-1.98/)Cited by:[§2](https://arxiv.org/html/2607.25094#S2.SS0.SSS0.Px3.p1.1)\.

## Appendix AExpert Annotation

#### Additional Annotation Details

The expert annotation interface is illustrated in Figures[9](https://arxiv.org/html/2607.25094#A1.F9),[10](https://arxiv.org/html/2607.25094#A1.F10),[11](https://arxiv.org/html/2607.25094#A1.F11), and[12](https://arxiv.org/html/2607.25094#A1.F12)\. The expert annotators were first presented with a landing page containing general instructions along with a reminder of the notions of implicature and implicature cancellation \(Figure[9](https://arxiv.org/html/2607.25094#A1.F9)\)\. The annotation task was organized into several sub\-batches based on stimulus type and the stimuli were presented in order of increasing cognitive load: scalar implicatures, synthetic conversational implicatures, naturally\-occurring conversational implicatures, and discourse implicatures\.

Before each stimulus type, annotators are given stimulus\-type\-specific instructions \(Figure[10](https://arxiv.org/html/2607.25094#A1.F10)\) and representative examples to familiarize them with the stimuli and task format \(Figure[11](https://arxiv.org/html/2607.25094#A1.F11)\)\. For each annotation item \(Figure[12](https://arxiv.org/html/2607.25094#A1.F12)\), annotators are asked to judge \(i\) whether the proposed interpretation constitutes a plausible implicature and \(ii\) whether the provided cancellation is correct\. If the implicature is judged implausible, annotators are asked to provide an alternative implicature\. Similarly, annotators are asked to provide an alternative cancellation when the given cancellation is incorrect or when the implicature itself is deemed implausible\. We elicit such alternatives for all data types except for our naturally\-occurring conversational implicatures\.

Expert annotators were paid at their affiliated university’s teaching assistant hourly rate\. This study received approval from the Research Ethics Board at McGill University \(REB File \#: 21\-06\-019\)\.

#### Additional Annotation Results

We provide confusion matrices for our first round of expert annotation during which two expert annotators annotated the same batch of5353items\. We report the two\-by\-two confusion matrix for the implicature plausibility responses in[Table˜3](https://arxiv.org/html/2607.25094#A1.T3)and for the implicature cancellation correctness in[Table˜4](https://arxiv.org/html/2607.25094#A1.T4)\. For the cancellation correctness confusion matrix, we only consider the items for which there was agreement in terms of the implicature’s plausibility\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/expert_annot/front_page.png)Figure 9:Landing page shown to expert annotators at the beginning of the annotation task\. The landing page includes a reminder of the definitions of implicature and implicature cancellation\.![Refer to caption](https://arxiv.org/html/2607.25094v1/expert_annot/instructions.png)Figure 10:Instructions provided to expert annotators for the scalar implicature items\.![Refer to caption](https://arxiv.org/html/2607.25094v1/expert_annot/examples_provided.png)Figure 11:Example scalar implicatures shown to expert annotators prior to the annotation of the scalar implicature items\. Examples like these are shown before each implicature type\.![Refer to caption](https://arxiv.org/html/2607.25094v1/expert_annot/example_item.png)Figure 12:One of the scalar implicature items to annotate\. The context, utterance, candidate implicature, and candidate implicature cancellation are shown\. The expert annotator decides on the plausibility of the candidate implicature and on the correctness of the implicature cancellation\.A1A\_\{1\}A2A\_\{2\}PlausibleImplausiblePlausible434Implausible43

Table 3:Confusion matrix of annotatorsA1A\_\{1\}andA2A\_\{2\}on the implicature plausibility annotation task\.A1A\_\{1\}A2A\_\{2\}CorrectIncorrectCorrect353Incorrect32

Table 4:Confusion matrix of annotatorsA1A\_\{1\}andA2A\_\{2\}on the implicature cancellation correctness annotation task\.

## Appendix BCrowdsourcing Annotation

#### Additional Crowdsourcing Details

The crowdsourcing annotation interface is shown in Figures[13](https://arxiv.org/html/2607.25094#A2.F13)and[14](https://arxiv.org/html/2607.25094#A2.F14)\. Each annotator began the annotation from the same landing page \(Figure[13](https://arxiv.org/html/2607.25094#A2.F13)\) which contains general information about the study and its structure\. The crowdworkers were then shown six example stimuli \(Figure[14](https://arxiv.org/html/2607.25094#A2.F14)\): one scalar implicature, two discourse implicatures, two synthetic conversational implicatures and one naturally\-occurring conversational implicature\. They were asked to try again if they provided an unlikely Likert score rating \(e\.g\., 1, 2, or 3\) for an item which did not contain a cancellation or if they provided a likely Likert score rating \(e\.g\., 5, 6, or 7\) for an item which did contain a cancellation\.

After completing the training examples, the crowdworkers were presented with either4040or4242items depending on the batch assigned to them\. For the1616batches containing4040items,1010were attention checks \(five items which should be rated likely and five items which should be rated unlikely\),1515were items without the cancelling utterance and1515were items with the cancelling utterance\. For the two batches of4242items,1010were attention checks,1616were items without the cancelling utterance and1616were items with the cancelling utterance\. The1010attention checks were the same for all batches\. We ensured that no batch contained the same implicature more than once\.

All crowdworkers were paid 15 USD/hour via the Prolific platform666[https://www\.prolific\.com/](https://www.prolific.com/)\. Data ingestion was done through Proliferate777[https://docs\.proliferate\.alps\.science/](https://docs.proliferate.alps.science/)\. This study received approval from the Research Ethics Board at McGill University \(REB File \#: 21\-06\-019\)\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x9.png)Figure 13:Landing page of the crowdsourcing experiment, outlining the instructions and interface for participants\.![Refer to caption](https://arxiv.org/html/2607.25094v1/x10.png)Figure 14:Example stimulus from the experiment, demonstrating the Likert\-scale interface used to collect interpretation likelihood ratings\.
#### Additional Crowdsourcing Results

We provide the per\-implicature\-type disaggregated histograms of the average human likelihood z\-scores in[Figure˜15](https://arxiv.org/html/2607.25094#A2.F15)\. Overall, the disaggregated histograms reveal that, across all implicature types, cancelling utterances weaken the z\-scored human likelihoods of implicatures\.

We also provide disaggregated scatter plots with the human likelihood z\-scores with and without the cancelling utterance on the y\-axis and x\-axis respectively in[Figure˜16](https://arxiv.org/html/2607.25094#A2.F16)\. We also report the Pearson correlationrrfor each of the implicature types in their corresponding subplot\. Overall, we find no statistically significant correlation between an implicature’s likelihood with and without a cancelling utterance for scalar implicature, discourse implicatures and synthetic conversational implicatures\. However, we do find a statistically significant correlation for naturally\-occurring conversational implicatures \(r=0\.63r=0\.63,p<0\.001p<0\.001\) which suggests that “stronger” naturally\-occurring implicatures are more “difficult” to cancel\.

![Refer to caption](https://arxiv.org/html/2607.25094v1/x11.png)Figure 15:Disaggregated histograms of the average human likelihood z\-scores with and without the cancelling utterance\.![Refer to caption](https://arxiv.org/html/2607.25094v1/x12.png)Figure 16:Scatter plots of the average human likelihood z\-scores with and without the cancelling utterance\. The Pearson correlation coefficientrrand p\-valueppare reported at the top left\.

## Appendix CModel Details

Model NameHuggingFace or OpenAI IdentifierGemma 3 \(4B\)google/gemma\-3\-4b\-itGemma 3 \(12B\)google/gemma\-3\-12b\-itGemma 3 \(27B\)google/gemma\-3\-27b\-itLlama 3\.1 \(8B\)meta\-llama/Llama\-3\.1\-8B\-InstructLlama 3\.1 \(70B\)meta\-llama/Llama\-3\.1\-70B\-InstructLlama 3\.2 \(3B\)meta\-llama/Llama\-3\.2\-3B\-InstructLlama 3\.3 \(70B\)meta\-llama/Llama\-3\.3\-70B\-InstructQwen 2\.5 \(3B\)Qwen/Qwen2\.5\-3B\-InstructQwen 2\.5 \(7B\)Qwen/Qwen2\.5\-7B\-InstructQwen 2\.5 \(14B\)Qwen/Qwen2\.5\-14B\-InstructQwen 2\.5 \(32B\)Qwen/Qwen2\.5\-32B\-InstructQwen 2\.5 \(72B\)Qwen/Qwen2\.5\-72B\-InstructQwen 3 \(0\.6B\)Qwen/Qwen3\-0\.6BQwen 3 \(1\.7B\)Qwen/Qwen3\-1\.7BQwen 3 \(4B\)Qwen/Qwen3\-4BQwen 3 \(8B\)Qwen/Qwen3\-8BQwen 3 \(14B\)Qwen/Qwen3\-14BQwen 3 \(32B\)Qwen/Qwen3\-32BQwen 3 Thinking \(0\.6B\)Qwen/Qwen3\-0\.6B&\# reasoning tokens: 512Qwen 3 Thinking \(1\.7B\)Qwen/Qwen3\-1\.7B&\# reasoning tokens: 512Qwen 3 Thinking \(4B\)Qwen/Qwen3\-4B&\# reasoning tokens: 512Qwen 3 Thinking \(8B\)Qwen/Qwen3\-8B&\# reasoning tokens: 512Qwen 3 Thinking \(14B\)Qwen/Qwen3\-14B&\# reasoning tokens: 512Qwen 3 Thinking \(32B\)Qwen/Qwen3\-32B&\# reasoning tokens: 512GPT\-5\.2gpt\-5\.2\-2025\-12\-11GPT\-5\.2 Thinkinggpt\-5\.2\-2025\-12\-11&reasoning effort: mediumGPT\-5\.4gpt\-5\.4\-2026\-03\-05GPT\-5\.4 Thinkinggpt\-5\.4\-2026\-03\-05&reasoning effort: mediumTable 5:Model names as referenced in the paper and their corresponding HuggingFace or OpenAI identifiers\. We also provide the thinking budget for the models that leverage their reasoning capabilities\.
## Appendix DPrompt Templates

Model PromptSystem Prompt
You are an expert in pragmatic inference and in identifying the intended meaning of utterances\.Prompt Template
You will be given an utterance as well as a potential interpretation for this utterance\. Your task is to decide whether the interpretation provided for the utterance is true or false\. You will be given a numbered list containing true and false\. Read the utterance and the potential interpretation and pick the number corresponding to whether the interpretation is true or false\.\{scenario\}Interpretation:\{implicature\}\{question\}\{options\}Answer:Figure 17:Model prompt used for pragmatic inference evaluation\. Variables\{scenario\},\{implicature\},\{question\}, and\{options\}are filled per instance\.Example Instanceu\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}and some places, they, they, they really nail them for tax\.b\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}Some, but not all, places nail them for tax\.u×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}i think they all do that nowFilled Prompt \(withoutu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\)
You will be given an utterance as well as a potential interpretation for this utterance\. Your task is to decide whether the interpretation provided for the utterance is true or false\. You will be given a numbered list containing true and false\. Read the utterance and the potential interpretation and pick the number corresponding to whether the interpretation is true or false\.Utterance: and some places, they, they, they really nail them for tax\.Interpretation: Some, but not all, places nail them for tax\.Is the interpretation of the previous utterance true or false?1: True2: FalseAnswer:Filled Prompt \(withu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\)
You will be given an utterance as well as a potential interpretation for this utterance\. Your task is to decide whether the interpretation provided for the utterance is true or false\. You will be given a numbered list containing true and false\. Read the utterance and the potential interpretation and pick the number corresponding to whether the interpretation is true or false\.Utterance: and some places, they, they, they really nail them for tax\. i think they all do that nowInterpretation: Some, but not all, places nail them for tax\.Is the interpretation of the previous utterance true or false?1: True2: FalseAnswer:Figure 18:Example instance for the prompt in Figure[17](https://arxiv.org/html/2607.25094#A4.F17), shown without \(top\) and with \(bottom\) the cancellationu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}appended to the utterance in\{scenario\}\.
## Appendix EControl Datasets

To conduct our control experiments in Section[7](https://arxiv.org/html/2607.25094#S7), we created three additional datasets:Implicature⊥,Implicature\+andImplicature≈\. All of these datasets have the same size as the originalImplicatureXdataset\. They share the same context, triggering utterance and implicature as the original dataset but differ in their follow\-up utterance \(i\.e\., the utterance which follows the cancelling utterance\)\. We provide details regarding how their follow\-up utterances were created\.

TheImplicature⊥follow\-up utterance,u⊥\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\bot\}\}, is a cancelling utterance with a discourse marker and an explicit negation of the implicature\. We first determined which discourse markers best preserved coherence in each of the dataset splits and then applied formulaic negations to the implicature\. For instance, to explicitly negate a scalar implicature like “…some of them showed up\.” we use discourse markers like “in fact”, “actually” or “as a matter of fact” followed by the negation of the scalar implicature “all of them showed up\.” On the other hand, conversational implicatures suffer in coherence when negated using the same discourse markers\. Thus, in this case, we use the discourse marker “Though, I don’t mean to imply that” construction\. For instance, “I’m more behind than you\. Though, I don’t mean to imply that I can’t help you with this problem\.”

TheImplicature\+follow\-up utterance,u\+\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\+\}\}, is a strengthening utterance: an utterance designed to confirm the implicature triggered by the triggering utterance\. We manually created these utterances while controlling for coherence\. For instance, to strengthen a scalar implicature, we reinforce the fact that not the entire set of elements being discussed should be included e\.g\., “and some places, they, they, they really nail them for tax\. though some get that exemption i was talking about”\. For the other types of implicatures, we create similar utterances which implicitly reinforce the implicature\. For instance, in the case of synthetic conversational implicatures, the implicature “I cannot help you with this question\.” can be reinforced with the utterance “Is there no one else you can ask?”

TheImplicature≈follow\-up utterance,u≈\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\approx\}\}, is a neutral utterance: an utterance which is irrelevant to the triggering utterance and the implicature\. The neutral utterance is selected by randomly sampling a cancelling utterance from the same corresponding dataset split\.

## Appendix FAdditional Results

#### Full Results

We provide full implicature recognition, cancellation recognition and belief update accuracies in Tables[6](https://arxiv.org/html/2607.25094#A6.T6),[7](https://arxiv.org/html/2607.25094#A6.T7)and[8](https://arxiv.org/html/2607.25094#A6.T8)respectively\. We also include the full results for our control experiments in Tables[9](https://arxiv.org/html/2607.25094#A6.T9),[10](https://arxiv.org/html/2607.25094#A6.T10),[11](https://arxiv.org/html/2607.25094#A6.T11),[12](https://arxiv.org/html/2607.25094#A6.T12)\. In addition, the form and update type controls on all the implicature types are plotted in Figures[19](https://arxiv.org/html/2607.25094#A6.F19)and[20](https://arxiv.org/html/2607.25094#A6.F20)\.

#### Does length explain implicature and cancellation recognition difficulty?

To determine whether our results are confounded by length and the ability of LLMs to perform under longer prompts, we run a correlation analysis between different length variables and implicature and cancellation recognition probabilities\. In particular, we compute Kendall tau’s concordances between different space\-separated lengths,\|⋅\|:𝒱∗→ℕ\|\\cdot\|:\\mathcal\{V\}^\{\*\}\\rightarrow\\mathbb\{N\}, and the implicature and cancellation probabilities of Llama 3\.3 70B, the model which performed the best overall in this study\. We provide Kendall tau’s concordancesτ\\taubetweenPℳ​\(b∣tb​\(c,⟨u⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)and the space\-separated lengths ofc\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},u\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}andb\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}in Table[13](https://arxiv.org/html/2607.25094#A6.T13)\. We also provide Kendall tau’s concordancesτ\\taubetweenPℳ​\(b∣tb​\(c,⟨u,u×⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)\)and the space\-separated lengths ofc\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},u\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},b\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}andu×\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}in Table[14](https://arxiv.org/html/2607.25094#A6.T14)\.

Our results in Tables[13](https://arxiv.org/html/2607.25094#A6.T13)and[14](https://arxiv.org/html/2607.25094#A6.T14)indicate weak concordances between the tested length variables and the LLM\-induced probabilitiesPℳ​\(b∣tb​\(c,⟨u⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)andPℳ​\(b∣tb​\(c,⟨u,u×⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)\)\. In addition, in most cases, the computed concordances are not statistically significant\. The only statistically significant Kendall’s tau values betweenPℳ​\(b∣tb​\(c,⟨u⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)and the length variables are0\.270\.27\(context length of the naturally\-occurring implicatures\) and0\.280\.28\(utterance length of the scalar implicatures\)\. The only statistically significant Kendall’s tau betweenPℳ​\(b∣tb​\(c,⟨u,u×⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)\)and the length variables is0\.220\.22\(triggering utterance length of the scalar implicatures\)\. All other concordances are near zero and not statistically significant\. While a lack of concordance is not evidence that there is no relation between length and implicature and cancellation difficulty, we believe it suggests that the difficulty of the items inImplicatureXcannot be explained by length alone\.

ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)1\.0000\.8710\.7080\.700Gemma 3 \(12B\)1\.0000\.9030\.7220\.560Gemma 3 \(27B\)1\.0000\.8060\.6740\.460Llama 3Llama 3\.1 \(8B\)0\.9130\.6770\.2290\.160Llama 3\.1 \(70B\)0\.9780\.9350\.8060\.560Llama 3\.2 \(3B\)0\.4780\.4520\.2710\.100Llama 3\.3 \(70B\)0\.9570\.9350\.7710\.600Qwen 2\.5Qwen 2\.5 \(3B\)0\.5220\.4190\.0760\.080Qwen 2\.5 \(7B\)0\.7610\.6130\.2220\.160Qwen 2\.5 \(14B\)0\.8910\.8060\.4790\.340Qwen 2\.5 \(32B\)0\.9350\.8710\.5560\.380Qwen 2\.5 \(72B\)1\.0000\.8710\.6940\.440Qwen 3Qwen 3 \(0\.6B\)0\.0000\.0000\.0000\.000Qwen 3 \(1\.7B\)0\.8480\.8710\.7150\.600Qwen 3 \(4B\)1\.0000\.9030\.5900\.780Qwen 3 \(8B\)1\.0000\.8390\.3260\.420Qwen 3 \(14B\)0\.8910\.9030\.6740\.420Qwen 3 \(32B\)0\.9570\.8710\.7080\.480Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.9780\.6130\.4440\.500Qwen 3 \(1\.7B\)0\.4350\.7740\.2920\.400Qwen 3 \(4B\)1\.0000\.8390\.5210\.460Qwen 3 \(8B\)0\.8040\.7420\.3470\.380Qwen 3 \(14B\)0\.5870\.8390\.5280\.420Qwen 3 \(32B\)0\.9350\.9030\.7710\.420GPTGPT 5\.20\.8040\.8060\.6740\.500GPT 5\.40\.8910\.9030\.7640\.440GPT ThinkingGPT 5\.20\.6740\.8710\.7430\.460GPT 5\.40\.8260\.8390\.8060\.440Human \(avg\.\)1\.0000\.9030\.9100\.780

Table 6:Full implicature recognition accuracy results\.ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)0\.5650\.7100\.8260\.720Gemma 3 \(12B\)0\.3260\.5480\.7430\.560Gemma 3 \(27B\)0\.7170\.7420\.8330\.640Llama 3Llama 3\.1 \(8B\)0\.8700\.8390\.9440\.660Llama 3\.1 \(70B\)0\.9130\.7740\.9440\.820Llama 3\.2 \(3B\)0\.8480\.8060\.7290\.540Llama 3\.3 \(70B\)0\.7610\.6770\.8120\.600Qwen 2\.5Qwen 2\.5 \(3B\)0\.6740\.7740\.4380\.560Qwen 2\.5 \(7B\)0\.8910\.7100\.8190\.620Qwen 2\.5 \(14B\)0\.8260\.6770\.8470\.720Qwen 2\.5 \(32B\)0\.8700\.7100\.9170\.740Qwen 2\.5 \(72B\)0\.7610\.7100\.8820\.680Qwen 3Qwen 3 \(0\.6B\)0\.4570\.7420\.5560\.520Qwen 3 \(1\.7B\)0\.6090\.5160\.6880\.680Qwen 3 \(4B\)0\.3040\.6450\.8750\.660Qwen 3 \(8B\)0\.8260\.7100\.8400\.640Qwen 3 \(14B\)0\.6740\.6130\.7500\.640Qwen 3 \(32B\)0\.9780\.9350\.9100\.780Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.8260\.7420\.6600\.680Qwen 3 \(1\.7B\)0\.8700\.6130\.7360\.600Qwen 3 \(4B\)0\.8260\.7100\.8680\.760Qwen 3 \(8B\)0\.9570\.8710\.9380\.740Qwen 3 \(14B\)0\.9780\.8710\.9310\.880Qwen 3 \(32B\)0\.9780\.9030\.9510\.840GPTGPT 5\.20\.8910\.6130\.7150\.320GPT 5\.40\.8700\.6450\.7500\.380GPT ThinkingGPT 5\.20\.9130\.8710\.7850\.460GPT 5\.40\.9350\.8710\.8190\.360Human \(avg\.\)1\.0000\.9030\.9720\.920

Table 7:Full cancellation recognition accuracy results\.ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)0\.5650\.6450\.5900\.540Gemma 3 \(12B\)0\.3260\.4840\.5760\.280Gemma 3 \(27B\)0\.7170\.5810\.5690\.240Llama 3Llama 3\.1 \(8B\)0\.8040\.6450\.2220\.160Llama 3\.1 \(70B\)0\.8910\.7420\.7920\.460Llama 3\.2 \(3B\)0\.4570\.4520\.2710\.100Llama 3\.3 \(70B\)0\.7170\.6130\.6600\.360Qwen 2\.5Qwen 2\.5 \(3B\)0\.3700\.3870\.0690\.040Qwen 2\.5 \(7B\)0\.6740\.4840\.2150\.140Qwen 2\.5 \(14B\)0\.7170\.4840\.4380\.200Qwen 2\.5 \(32B\)0\.8040\.5810\.5210\.280Qwen 2\.5 \(72B\)0\.7610\.6130\.6110\.240Qwen 3Qwen 3 \(0\.6B\)0\.0000\.0000\.0000\.000Qwen 3 \(1\.7B\)0\.5650\.4840\.5350\.440Qwen 3 \(4B\)0\.3040\.5810\.5420\.480Qwen 3 \(8B\)0\.8260\.6130\.2920\.240Qwen 3 \(14B\)0\.6090\.5480\.6040\.340Qwen 3 \(32B\)0\.9350\.8710\.7010\.400Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.8040\.5810\.4030\.380Qwen 3 \(1\.7B\)0\.4130\.5160\.2570\.320Qwen 3 \(4B\)0\.8260\.6130\.4790\.360Qwen 3 \(8B\)0\.7610\.6770\.3400\.280Qwen 3 \(14B\)0\.5870\.7740\.5070\.380Qwen 3 \(32B\)0\.9130\.8390\.7640\.400GPTGPT 5\.20\.8040\.4840\.6460\.300GPT 5\.40\.8260\.6450\.7080\.280GPT ThinkingGPT 5\.20\.6520\.8060\.7150\.400GPT 5\.40\.8040\.8060\.7640\.340Human \(avg\.\)1\.0000\.9030\.9030\.720

Table 8:Full belief update accuracy results\.ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)1\.0000\.7741\.0000\.920Gemma 3 \(12B\)1\.0000\.8060\.9650\.660Gemma 3 \(27B\)0\.9570\.6130\.9510\.560Llama 3Llama 3\.1 \(8B\)1\.0000\.6770\.8540\.620Llama 3\.1 \(70B\)0\.9780\.7100\.8680\.580Llama 3\.2 \(3B\)1\.0000\.8060\.9790\.820Llama 3\.3 \(70B\)1\.0000\.7100\.8820\.640Qwen 2\.5Qwen 2\.5 \(3B\)0\.0870\.0320\.0210\.020Qwen 2\.5 \(7B\)0\.9570\.1940\.4650\.200Qwen 2\.5 \(14B\)0\.8910\.2900\.1670\.160Qwen 2\.5 \(32B\)0\.9570\.4840\.8060\.420Qwen 2\.5 \(72B\)1\.0000\.6770\.8330\.460Qwen 3Qwen 3 \(0\.6B\)0\.0000\.0000\.0000\.000Qwen 3 \(1\.7B\)1\.0000\.7420\.9100\.620Qwen 3 \(4B\)1\.0000\.7100\.9100\.580Qwen 3 \(8B\)0\.9780\.4520\.7150\.380Qwen 3 \(14B\)1\.0000\.3230\.5490\.260Qwen 3 \(32B\)0\.9350\.5160\.7290\.340Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.9130\.3550\.6110\.380Qwen 3 \(1\.7B\)0\.5000\.3230\.2920\.100Qwen 3 \(4B\)0\.9350\.4190\.3960\.140Qwen 3 \(8B\)0\.9130\.2580\.4860\.180Qwen 3 \(14B\)0\.9350\.2580\.5760\.260Qwen 3 \(32B\)0\.9570\.5480\.7570\.420GPTGPT 5\.20\.9780\.2900\.3890\.260GPT 5\.40\.9350\.3870\.5000\.260GPT ThinkingGPT 5\.20\.9350\.3230\.3400\.220GPT 5\.40\.9570\.4520\.4440\.200Human \(avg\.\)1\.0000\.9030\.9100\.780

Table 9:Full implicature recognition accuracy results for the prior common ground control experiment\.ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)0\.8910\.9680\.8120\.820Gemma 3 \(12B\)0\.8040\.8710\.8120\.760Gemma 3 \(27B\)0\.9780\.9030\.8540\.900Llama 3Llama 3\.1 \(8B\)1\.0001\.0000\.9380\.840Llama 3\.1 \(70B\)0\.9780\.9350\.9510\.940Llama 3\.2 \(3B\)0\.9781\.0000\.8330\.840Llama 3\.3 \(70B\)0\.9780\.8710\.8750\.840Qwen 2\.5Qwen 2\.5 \(3B\)0\.9570\.8060\.4790\.540Qwen 2\.5 \(7B\)1\.0000\.9030\.8680\.820Qwen 2\.5 \(14B\)0\.9570\.8710\.8890\.900Qwen 2\.5 \(32B\)1\.0000\.8710\.9650\.940Qwen 2\.5 \(72B\)0\.9780\.9030\.8890\.920Qwen 3Qwen 3 \(0\.6B\)0\.4130\.8060\.5420\.500Qwen 3 \(1\.7B\)0\.9130\.8710\.7010\.820Qwen 3 \(4B\)0\.8700\.9680\.8330\.880Qwen 3 \(8B\)1\.0000\.9680\.8470\.900Qwen 3 \(14B\)0\.9780\.8710\.8120\.680Qwen 3 \(32B\)1\.0000\.9680\.9240\.780Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.9570\.8390\.7920\.740Qwen 3 \(1\.7B\)0\.9780\.8710\.7850\.880Qwen 3 \(4B\)0\.9780\.9680\.8820\.960Qwen 3 \(8B\)1\.0000\.9680\.9100\.940Qwen 3 \(14B\)0\.9780\.9350\.9240\.920Qwen 3 \(32B\)0\.9780\.9030\.9580\.880GPTGPT 5\.20\.8910\.8390\.7500\.480GPT 5\.40\.9570\.8060\.7710\.500GPT ThinkingGPT 5\.20\.9350\.8390\.7990\.500GPT 5\.40\.9570\.8710\.8330\.500Human \(avg\.\)1\.0000\.9030\.9720\.920

Table 10:Full cancellation recognition accuracy results on theImplicature⊥items\.ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)0\.1740\.2260\.4240\.620Gemma 3 \(12B\)0\.0220\.1290\.3820\.540Gemma 3 \(27B\)0\.0650\.2260\.3890\.500Llama 3Llama 3\.1 \(8B\)0\.5220\.4840\.5350\.680Llama 3\.1 \(70B\)0\.2610\.1610\.5280\.600Llama 3\.2 \(3B\)0\.5220\.6130\.4580\.580Llama 3\.3 \(70B\)0\.0870\.0000\.2640\.460Qwen 2\.5Qwen 2\.5 \(3B\)0\.7610\.2260\.6040\.600Qwen 2\.5 \(7B\)0\.5650\.2260\.6740\.560Qwen 2\.5 \(14B\)0\.1960\.0970\.5210\.520Qwen 2\.5 \(32B\)0\.1300\.1610\.4650\.500Qwen 2\.5 \(72B\)0\.0220\.0320\.3190\.460Qwen 3Qwen 3 \(0\.6B\)0\.4780\.4190\.4440\.440Qwen 3 \(1\.7B\)0\.6300\.2260\.5210\.380Qwen 3 \(4B\)0\.0000\.0650\.5830\.520Qwen 3 \(8B\)0\.1960\.1940\.6250\.620Qwen 3 \(14B\)0\.2170\.1290\.6530\.620Qwen 3 \(32B\)0\.4780\.5480\.6250\.660Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.3700\.4840\.4930\.580Qwen 3 \(1\.7B\)0\.8040\.1940\.5280\.680Qwen 3 \(4B\)0\.0650\.0970\.4440\.380Qwen 3 \(8B\)0\.6740\.2900\.5760\.640Qwen 3 \(14B\)0\.6740\.3230\.5280\.580Qwen 3 \(32B\)0\.4130\.2580\.4440\.680GPTGPT 5\.20\.2610\.0650\.3190\.140GPT 5\.40\.1520\.0000\.2290\.200GPT ThinkingGPT 5\.20\.6300\.0970\.3060\.300GPT 5\.40\.4570\.1290\.2080\.320

Table 11:Full belief strengthening accuracy results on theImplicature\+items\.ModelScalarImplicatureDiscourseImplicatureSyntheticConversational ImplicatureNaturally\-OccurringConversational ImplicatureGemma 3Gemma 3 \(4B\)1\.0001\.0000\.7570\.860Gemma 3 \(12B\)0\.9130\.9680\.7920\.940Gemma 3 \(27B\)0\.8260\.9350\.8120\.960Llama 3Llama 3\.1 \(8B\)0\.6740\.8060\.8680\.900Llama 3\.1 \(70B\)0\.9130\.9350\.8470\.900Llama 3\.2 \(3B\)0\.6090\.7420\.8190\.940Llama 3\.3 \(70B\)0\.9130\.9680\.8470\.900Qwen 2\.5Qwen 2\.5 \(3B\)0\.7610\.8060\.9310\.920Qwen 2\.5 \(7B\)0\.6740\.8710\.8680\.980Qwen 2\.5 \(14B\)0\.6300\.8060\.7570\.900Qwen 2\.5 \(32B\)0\.5870\.9030\.7920\.920Qwen 2\.5 \(72B\)0\.7610\.9680\.7570\.960Qwen 3Qwen 3 \(0\.6B\)1\.0001\.0001\.0001\.000Qwen 3 \(1\.7B\)0\.7830\.9350\.8470\.900Qwen 3 \(4B\)1\.0000\.9680\.8260\.920Qwen 3 \(8B\)0\.8260\.9030\.8260\.940Qwen 3 \(14B\)0\.8701\.0000\.7990\.940Qwen 3 \(32B\)0\.8481\.0000\.8400\.940Qwen 3 ThinkingQwen 3 \(0\.6B\)0\.8700\.7740\.6740\.660Qwen 3 \(1\.7B\)0\.4570\.9030\.7710\.840Qwen 3 \(4B\)1\.0000\.9030\.8060\.880Qwen 3 \(8B\)0\.7390\.8710\.7710\.940Qwen 3 \(14B\)0\.6740\.9030\.7430\.920Qwen 3 \(32B\)0\.8700\.9350\.7710\.840GPTGPT 5\.20\.6960\.9680\.8120\.940GPT 5\.40\.7610\.9350\.7850\.940GPT ThinkingGPT 5\.20\.7170\.8710\.8470\.960GPT 5\.40\.8260\.9680\.8470\.960

Table 12:Full belief unchanging accuracy results on theImplicature≈items\.![Refer to caption](https://arxiv.org/html/2607.25094v1/x13.png)Figure 19:Cancellation recognition accuracy using original cancelling utterances fromImplicatureXand the explicit cancelling utterances fromImplicature⊥which contain negation\.![Refer to caption](https://arxiv.org/html/2607.25094v1/x14.png)Figure 20:Update belief accuracies given different update types\. Cancelling denotes the update belief as triggered by implicature cancellation while unchanging denotes the belief update accuracy usingImplicature≈items and strengthening denotes the belief update accuracy usingImplicature\+items\.Scalar Imp\.Discourse Imp\.Synthetic Conv\. Imp\.Naturally\-Occurring Conv\. Imp\.\|c\|\|\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}\|—−0\.06\-0\.06−0\.01\-0\.010\.27∗0\.27^\{\*\}\|b\|\|\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\|−0\.01\-0\.01−0\.26\-0\.260\.070\.07−0\.10\-0\.10\|u\|\|\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\|0\.28∗0\.28^\{\*\}−0\.13\-0\.130\.060\.06−0\.01\-0\.01Table 13:Kendall’s tauτ\\taubetweenPℳ​\(b∣tb​\(c,⟨u⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\\rangle\)\)and length variables for the dataset splits\. Values with∗indicate statistical significance \(p<0\.05p<0\.05\)\.Scalar Imp\.Discourse Imp\.Synthetic Conv\. Imp\.Naturally\-Occurring Conv\. Imp\.\|c\|\|\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\}\|—−0\.21\-0\.210\.020\.020\.160\.16\|b\|\|\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\|0\.080\.08−0\.00\-0\.000\.070\.070\.020\.02\|u\|\|\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\}\|0\.22∗0\.22^\{\*\}−0\.06\-0\.060\.010\.010\.140\.14\|u×\|\|\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\|0\.090\.09−0\.05\-0\.05−0\.07\-0\.07−0\.10\-0\.10Table 14:Kendall’s tauτ\\taubetweenPℳ​\(b∣tb​\(c,⟨u,u×⟩\)\)P\_\{\\mathcal\{M\}\}\(\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\\mid\\mathrm\{\\texttt\{t\}\}\_\{\{\\color\[rgb\]\{0,0\.4453125,0\.69921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.4453125,0\.69921875\}b\}\}\(\{\\color\[rgb\]\{0,0\.62109375,0\.44921875\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.62109375,0\.44921875\}c\},\\langle\{\\color\[rgb\]\{0\.80078125,0\.47265625,0\.65625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.80078125,0\.47265625,0\.65625\}u\},\{\\color\[rgb\]\{0\.90234375,0\.625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.90234375,0\.625,0\}u^\{\\times\}\}\\rangle\)\)and length variables for the different dataset splits\. Values with∗indicate statistical significance \(p<0\.05p<0\.05\)\.

Similar Articles

Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

arXiv cs.CL

This paper evaluates the abilities of large language models (LLMs) and multimodal LLMs for addressee detection, turn-change prediction, and next speaker prediction in multi-party meeting conversations. Results show text-based LLMs outperform supervised models and humans in next speaker prediction, while multimodal LLMs improve over text-only models in other tasks but remain below human performance.

Can LLMs Take Retrieved Information with a Grain of Salt?

arXiv cs.CL

This paper investigates how large language models adapt to the certainty of retrieved information, identifying systematic limitations in handling uncertainty. It proposes an interaction strategy that reduces obedience errors by 25% without modifying model weights.