Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry
Summary
This paper investigates whether large language models exhibit pragmatic cooperativity in multi-party collaborative tasks under conditions of epistemic asymmetry, formalizing Grice's cooperative principle and evaluating LLMs as both speakers and listeners. Results show that while LLMs display some pragmatic capabilities, they struggle with incomplete information and fail to recognize certain violations of Gricean maxims.
View Cached Full Text
Cached at: 07/14/26, 04:23 AM
# Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry
Source: [https://arxiv.org/html/2607.11053](https://arxiv.org/html/2607.11053)
Hannah VanderHoeven, Abhijnan Nath & Nikhil Krishnaswamy Situated Grounding and Natural Language \(SIGNAL\) Lab Department of Computer Science Colorado State University Fort Collins, CO 80523, USA \{hannah\.vanderhoeven,nkrishna\}@colostate\.edu
###### Abstract
Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning\. The increasing use of LLMs in collaborative and agentic pipelines raises questions about the extent to which they exhibit these pragmatic capabilities, especially in scenarios where they may not have access to the same information as their collaborators\. In this paper, we perform a novel investigation into the pragmatic reasoning capabilities of LLMs in a multi\-party collaborative task under partial information conditions\. We formalize a notion ofcollaborative epistemic asymmetrythat explicitly connects objective task success to Grice’s cooperative principle and empirically assess various LLMs’ abilities to act cooperatively as both speakers and listeners, including both prompting and post\-training strategies\. Our results show that while LLMs exhibit certain pragmatic capabilities in collaborative settings, and these can be elicited through prompting and post\-training, they still face challenges in pragmatic communication with incomplete information, and that certain failure modes do correlate with floutings of Grice’s maxims that go unrecognized\.
## 1Introduction
Successful collaboration rests on establishingcommon ground\(Clark & Brennan,[1991](https://arxiv.org/html/2607.11053#bib.bib10); Clark,[1996](https://arxiv.org/html/2607.11053#bib.bib9)\)\. Collaborative groups can achieve outcomes exceeding the abilities and understanding of any individual member\(Boyd,[2021](https://arxiv.org/html/2607.11053#bib.bib3)\), but only if they can effectively share perspectives\. Inherent in this is the adjustment of communications to becooperative—expressing no more or less than the needed information, at the right time, in the appropriate way\(Grice,[1975](https://arxiv.org/html/2607.11053#bib.bib23)\)\.
The dialogue capabilities of large language models \(LLMs\) have led them to be incorporated into workflows across many domains, often as “collaborators” with humans, or with other AI systems in “agentic” pipelines\(Butler et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib4); Maslej et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib41)\)\. In large part this has come about because they appear to be fluent, cooperative communicators; the implicit assumption is that LLMs and agentic systems driven by them will exchange information and interpret collaborator needs and instructions in a way that is roughly equivalent to the cooperativity exhibited by humans\(Zarrieß & Schlangen,[2019](https://arxiv.org/html/2607.11053#bib.bib69); Guo et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib24)\)\. However, evidence regarding the pragmatic competence of LLMs is mixed at best\(Nguyen,[2023](https://arxiv.org/html/2607.11053#bib.bib50); Jian & Siddharth,[2024](https://arxiv.org/html/2607.11053#bib.bib30); Park et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib53); Sravanthi et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib61); Ma et al\.,[2025a](https://arxiv.org/html/2607.11053#bib.bib38);[b](https://arxiv.org/html/2607.11053#bib.bib39); Eisenstein et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib12); Nath et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib49)\)\. This calls into question whether current LLMs truly possess the pragmatic capabilities required to be good collaborators, or if their apparent demonstrations of pragmatic competence are illusory\. This question is particularly relevant inmultipartycollaborations with more than 2 agents \(human or AI\), because each agent may labor under ”false assumptions” about other agents and how they communicate\.
This paper presents a first of its kind examination of common LLMs’ cooperative communication capabilities in a challenging multi\-agent, partially\-observable collaborative reasoning task\. Fig\.[1](https://arxiv.org/html/2607.11053#S1.F1)shows an overview of our approach\. Through a focus on agent\-agent collaborations, we address the following research questions\.RQ1:In multiagent collaborations, do common LLMs behave more likeliteralorpragmaticlisteners?RQ2:To what extent can greater pragmatic listening capabilities be elicited through prompting and encouraging exploration?RQ3:To what extent can speaker agents be made more robust against pragmatically suboptimal listeners through offline alignment?
Through a novel theoretical formulation ofcollaborative epistemic asymmetry\(CEA\), we explicitly connect Gricean maxims to quantitative task outcomes\. We perform an exploration of LLMs pragmatic reasoning capabilities through prompting, exploration, and post\-training strategies, demonstrate where failures in collaborative task performance can be connected to maxim flouting by the agents\. Our results show that while LLMs have some pragmatic speaking, listening, and reasoning capabilities in collaborative tasks, they still encounter challenges with cooperativity in the incomplete information setting and may flout cooperative maxims in ways that correlate with degraded collaborative task performance\.
Figure 1:High level overview of our experimental pipeline\.
## 2Related Work
LLMs have proven to be grammatically fluent language generators and linguistically competent “listeners”\(Kibria et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib35); Hardt,[2025](https://arxiv.org/html/2607.11053#bib.bib25); Ma et al\.,[2025a](https://arxiv.org/html/2607.11053#bib.bib38)\), yet tend to be pragmatically fragile, often missing meaning that goes beyond words\.Grice \([1975](https://arxiv.org/html/2607.11053#bib.bib23)\)’s influentialmaxims of conversationdefined how accurate, concise, relevant, and clear comminication enhances group trust and interaction efficiency\. These maxims have proven challenging to operationalize computationally\(Saygin & Cicekli,[2002](https://arxiv.org/html/2607.11053#bib.bib59)\), in part due to the depth of hard\-to\-quantify social nuances in communication\(Jurafsky,[2006](https://arxiv.org/html/2607.11053#bib.bib31); Floyd et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib15)\)\.
According to the Rational Speech Act \(RSA\) framework\(Goodman & Stuhlmüller,[2013](https://arxiv.org/html/2607.11053#bib.bib22)\), which is built on the assumption that speakers are rational and informative, the utterances of an optimally pragmatic speaker \(S1S\_\{1\}\) will maximize the probability that a listener’sliteralinterpretation \(L0L\_\{0\}\) will correctly infer their intended meaning\. The RSA framework has been used to show that humans reason recursively as pragmatic speakers and listeners, however similar applications have proven challenging to scale up in Natural Language Processing \(NLP\)\(Frank et al\.,[2016](https://arxiv.org/html/2607.11053#bib.bib17); Goodman & Frank,[2016](https://arxiv.org/html/2607.11053#bib.bib21)\)\. Instead, certain specific inferential problems have received the lion’s share of attention\(Jurafsky,[2006](https://arxiv.org/html/2607.11053#bib.bib31)\), such asdiscourse structure and coherence relations,reference resolution, andabduction\(Nath et al\.,[2024a](https://arxiv.org/html/2607.11053#bib.bib44);[b](https://arxiv.org/html/2607.11053#bib.bib45); Gan et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib18); Min et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib42); Zhu et al\.,[2025a](https://arxiv.org/html/2607.11053#bib.bib70)\)\.
Forcollaborators, inferring referents and intended actions from each other’s statements is key to task progression\(Jurafsky,[2006](https://arxiv.org/html/2607.11053#bib.bib31)\)\. This often necessitatesperspective\-takingto achieve common ground with interlocutors\(Filipi & Wales,[2004](https://arxiv.org/html/2607.11053#bib.bib14); Cantiani et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib5); Davidson et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib11)\)while remaining anchored in one’s own private knowledge\(Katsos et al\.,[2023](https://arxiv.org/html/2607.11053#bib.bib33)\)\. This rests on atheory of mind\(Premack & Woodruff,[1978](https://arxiv.org/html/2607.11053#bib.bib56)\), which is challenging for LLMs\(Ullman,[2023](https://arxiv.org/html/2607.11053#bib.bib62); Hu et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib28)\)\. Multiple RSA and MDP\-based frameworks describe how agents maintain beliefs about the world and each other’s beliefs and capabilities\(Langlois & Everitt,[2021](https://arxiv.org/html/2607.11053#bib.bib36); Nath & Krishnaswamy,[2025](https://arxiv.org/html/2607.11053#bib.bib43); Maeda et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib40)\), and retain robustness to non\-cooperative or adversarial speakers\(Gmytrasiewicz & Doshi,[2005](https://arxiv.org/html/2607.11053#bib.bib20); Gmytrasiewicz,[2020](https://arxiv.org/html/2607.11053#bib.bib19)\)\.Estienne et al\. \([2025](https://arxiv.org/html/2607.11053#bib.bib13)\)empirically improve performance in a partial\-information reference game by optimizing cooperativity as a gain function\. Dynamic epistemic logic \(DEL\) approaches explore necessary communicative strategies for collaborative inference\(Van Benthem & Pacuit,[2011](https://arxiv.org/html/2607.11053#bib.bib63); Van Benthem et al\.,[2014](https://arxiv.org/html/2607.11053#bib.bib64); Pacuit,[2017](https://arxiv.org/html/2607.11053#bib.bib52); Khebour et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib34)\)\. Prompting interventions can help LLMs reason more pragmatically, but remain limited in tasks with higher\-order prerequisites like perspective\-taking\(Wilf et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib67); Just et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib32); Ma et al\.,[2025b](https://arxiv.org/html/2607.11053#bib.bib39)\)\. In grounded multiagent LLM collaboration\(Andreas & Klein,[2016](https://arxiv.org/html/2607.11053#bib.bib1)\), recent work has addressed maze\-solving\(Davidson et al\.,[2025](https://arxiv.org/html/2607.11053#bib.bib11)\)and 3D block building\(Wu et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib68)\)including in partial\-information settings\(Nath et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib49)\)but few works directly analyze collaborative failures through a Gricean lens\. In our work, all agents are set up to seek collaborative success rather than adversarial outcomes, and so we focus onfloutingsof Gricean maxims such as accidental rule\-breaking or omission of important task\-relevant information, rather than intentionalviolationsto hide meaning\.
## 3Collaborative Epistemic Asymmetry: Actionability Requires Cooperativity
Our work unifies the above challenges into a novel investigation into the pragmatic reasoning capabilities of LLMs in amulti\-turn, multi\-agent collaborationunderpartial information and observability conditions\. Tasks of this nature, e\.g\.,Zhu et al\. \([2025b](https://arxiv.org/html/2607.11053#bib.bib71);[2026](https://arxiv.org/html/2607.11053#bib.bib72)\), simulate situations where parties with different background knowledge and capabilities must bring them together to solve a problem,but no one individual knows the solution at the outset\. We refer to this ascollaborative epistemic asymmetry\(CEA\)\. This is commonplace in real\-world scenarios like “jigsaw” problem solving in classrooms\(Perkins & Tagler,[2011](https://arxiv.org/html/2607.11053#bib.bib54)\), challenging medical diagnostics\(Poradzisz & Florczak,[2019](https://arxiv.org/html/2607.11053#bib.bib55)\), or even military exercises\. In these situations, collaborators must be cooperative in their utterances and instructions to communicate the necessary information for progress toward the shared goal\. This leads to a core insight of our work:in collaborative task environments with shared goals but epistemic asymmetry, the mostcooperativeutterances are those that are mostactionable\.We show this is a valid assumption\.
#### Definitions
Let𝒲\\mathcal\{W\}be theworld statesand𝒰\\mathcal\{U\}be theutterancesin a collaborative task under epistemic asymmetry\. LetG∈𝒲G\\in\\mathcal\{W\}be theshared goal, or world state in which the shared problem is solved\. No individual knowsGGexactly; each agentiihas a priorPi\(G\)P\_\{i\}\(G\)over possible goals, drawn from the partial informationdid\_\{i\}they have\.𝐝\\mathbf\{d\}is the agents’joint information state⋃i=1Ndi\\bigcup\_\{i=1\}^\{N\}d\_\{i\}and eachdid\_\{i\}is drawn from partitionℐi\\mathcal\{I\}\_\{i\}of𝒲\\mathcal\{W\}\.Collaborative epistemic asymmetry \(CEA\)can be parameterized by⟨𝒲,𝒰,G,𝐝,\{Pi\}⟩\\langle\\mathcal\{W\},\\mathcal\{U\},G,\\mathbf\{d\},\\\{P\_\{i\}\\\}\\rangle\. An utterance isactionableif it is \(1\)informative\(Shannon,[1948](https://arxiv.org/html/2607.11053#bib.bib60)\)and \(2\)goal\-oriented\(Van Rooy,[2004](https://arxiv.org/html/2607.11053#bib.bib65)\)\. This stipulative definition captures an intuitive notion of an actionable utterance in a collaborative scenario: an utterance enables progress toward a shared goal if and only if it both updates the listener’s model of the worldandreduces their uncertainty about which actions are goal\-directed\. In this setting,Grice \([1975](https://arxiv.org/html/2607.11053#bib.bib23)\)’s 4 cooperative maxims can be considered constraints on utterances\. An utteranceuufully iscooperativeif it satisfies all 4 constraints\.
###### Lemma 1\(Actionability in CEA Settings Requires Gricean Cooperativity\)\.
An utteranceu∈𝒰u\\in\\mathcal\{U\}isactionablefor listenerLLif it \(1\) results in a non\-trivial belief update toLL’s posterior over world states \(i\.e\.,uuis informative aboutww\) and \(2\) reducesLL’s uncertainty about the goalGG:
PL\(w\|u\)∝̸PL\(w\)∧H\(PL\(G\|u\)\)<H\(PL\(G\)\), whereH\(⋅\)denotes Shannon entropy\.\\displaystyle P\_\{L\}\(w\|u\)\\not\\propto P\_\{L\}\(w\)\\wedge H\(P\_\{L\}\(G\|u\)\)<H\(P\_\{L\}\(G\)\)\\text\{, where $H\(\\cdot\)$ denotes Shannon entropy\.\}\(1\)Under CEA, if utteranceuuis actionable, thenuuadheres to Grice’s cooperative maxims\. Flouting any single maxim is sufficient to degrade actionability, and necessary for certain failure modes\.
Proof SketchQualityFlouting: ifuudoes not reflect speakerSiS\_\{i\}’sdid\_\{i\},P\(u\|w\)P\(u\|w\)loses its dependence onwwand the posterior collapses to the prior, violating the first condition\.QuantityFlouting: ifuuis under\- or overinformative, entropy reduction towardGGis not guaranteed, violating the second condition and in certain cases, the first\.RelationFlouting: ifuuis irrelevant toGG, no posterior update reduces uncertainty towardGG, violating the second condition\.MannerFlouting: ifuuhas multiple ambiguous interpretations, entropy towardGGnecessarily increases, violating the second condition\.Thus under collaborative epistemic asymmetry, Gricean maxims are necessary conditions for resolving the asymmetry\. Any maxim flouting violates one of the two conditions for actionability as defined\. A full proof is presented in Appendix[A](https://arxiv.org/html/2607.11053#A1)\.
This is demonstrated in an example from our experimental setting \(Fig\.[2](https://arxiv.org/html/2607.11053#S3.F2)\)\. Here, each Director \(D\) utterance updates the posterior of the Builder \(B\) toward a refinement of which color block should go where \(satifying condition 1\)\. The set of possible moves B should make to build the structure is also narrowed \(satisfying condition 2\)\. Bactswhen both conditions are satisfied: he removes a block from the board in anticipation of correcting the block beneath it\. The CEA setting produces the pragmatic behaviors predicted by Lemma[1](https://arxiv.org/html/2607.11053#Thmlemma1)\. See Appendix[B](https://arxiv.org/html/2607.11053#A2)for an extended analysis based on the RSA framework\.

AgentUtterance / ActionD1The long\.BThis is yellow\.D1Yeah, that one’s yellow\.BSo this is yellow\.BWhich one’s this color?BGreen?D3Yeah, the bottom one is green\.BAction: REMOVErsfromlayer1layer~1
Figure 2:Still from Distributed Partial Information Puzzle \(DPIP\) Lego task dataset\(Zhu et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib72)\)with accompanying dialogue snippet at the point of execution\. ThreeDirectorsreceive separate side views of a single structure made of large Lego blocks, and have to instruct oneBuilderto place the correct blocks in the correct places on a board to build it\.Connecting cooperativity to actionability this way enables measurement of cooperativity in terms of task progress, providing an objective metric that is explicitly connected to satisfying the Gricean maxims\. We can therefore quantitatively assess when LLMs are appropriately Gricean by assessing when their utterances contribute toward shared task progress\.
## 4Methodology
We perform our examination in a setting derived from the Distributed Partial Information Puzzle \(DPIP\) Lego task\(Zhu et al\.,[2025b](https://arxiv.org/html/2607.11053#bib.bib71);[2026](https://arxiv.org/html/2607.11053#bib.bib72)\)\(Fig\.[2](https://arxiv.org/html/2607.11053#S3.F2)\)\. The DPIP Lego task is both a “jigsaw” task\(Hattie,[2012](https://arxiv.org/html/2607.11053#bib.bib26)\)and fundamentally areference game\(Frank & Goodman,[2012](https://arxiv.org/html/2607.11053#bib.bib16)\)that requires the Builder to selectbothobjects of specified color and size from a fixed inventory, andlocationsat which to place them\. All goal structures have a3×33\\times 3footprint \(1 unit = 1 square Lego block\) and are 3 layers high\. The Directors cannot show their privately\-held views to anyone else, and only the Builder can manipulate blocks and the board\. All participants may see the board\. This distinction in roles allows us to map theSpeakersas defined in Sec\.[3](https://arxiv.org/html/2607.11053#S3)to the Directors, and theListenerto the Builder\. We can thus meaningfully speak of “pragmatic” or “literal” Builders who interpret Director utterances accordingly, or “optimal” Directors who attempt to maximize cooperativity in their utterances\.
Zhu et al\. \([2026](https://arxiv.org/html/2607.11053#bib.bib72)\)present 10 annotated videos of humans performing the task, including speech transcriptions and action annotations\. They show that this task presents defined challenges for SOTA LLMs in inference tasks like structure and participant belief prediction\. However, the static dataset makes it challenging to assess counterfactual outcomes, such aswhat if the participants spoke more/less cooperatively?111This is a known challenge of fixed datasets in the domain of human\-LLM and LLM\-LLM interaction, as noted byNath et al\.\([2025a](https://arxiv.org/html/2607.11053#bib.bib46)\)\.
### 4\.1Task Environment
Inspired by the richness of the DPIP Lego task but cognizant of the challenges posed by its relatively sparse naturalistic data and fixed nature of the dataset, we adopt the associated CRAFT simulator and benchmark platform\(Nath et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib49)\)\.222[https://zenodo\.org/records/18626419](https://zenodo.org/records/18626419)CRAFTis a multi\-agent coordination benchmark which tests LLMs for grounded communication under partial observability, built upon a practical implementation of the Bounded Pragmatic Speaker \(BPS\) formalism\(Nguyen,[2023](https://arxiv.org/html/2607.11053#bib.bib50)\)\.CRAFTsimulates the DPIP Lego task—including 3D “Lego” structure generation, conversion to Director\-specific 2D partial views, and the Builder’s move execution and validation logic—using atext\-onlyagentic framework that supports both open\-weight and proprietary models in turn\-based synchronous communication protocols\(Li et al\.,[2023](https://arxiv.org/html/2607.11053#bib.bib37); Nath et al\.,[2025b](https://arxiv.org/html/2607.11053#bib.bib47);[c](https://arxiv.org/html/2607.11053#bib.bib48)\)for creating high\-quality “expert” trajectories\.
LLMs can be assigned the role of a Director \(with privately held information\) or the Builder\. All agents have access to the current board state, or the public portion of the world stateww\. Each turn, 3 Director instructions are generated in the context of the current board state, their private information, and the current dialogue history\. Due to the text\-based nature of CRAFT, Director utterances eschew demonstratives to indicate location in favor of specific descriptions using directional and relational terms\. Directors are randomly sampled with replacement\. The Builder chooses one instruction to follow\. CRAFT provides amove exploration toolthe Builder can use to assess move options before executing\. This separately interprets each Director instruction and returns its estimated impact toward goal completion\. This is in effect a computational approximation of what the pragmatic listenerL1L\_\{1\}should do\. The Builder then executes the instruction of the most positive value\.
The Builder canPLACEorREMOVEone block of specified color and size at a specified location, which updates the board state\. It also generates aconfirmationof its move in plain English\. The move mayfailif Builder attempts something illegal \(e\.g\., to place a block somewhere without a supporting block underneath it\)\. The Builder may also request toCLARIFYthe instructions without executing\. Moves \(including fail and clarify outcomes\) are appended to the dialogue history, which continues for a prespecified length\. Further technical details on CRAFT are inNath et al\. \([2026](https://arxiv.org/html/2607.11053#bib.bib49)\)\. Further details on our specific usage are in Appendix[C](https://arxiv.org/html/2607.11053#A3)\.
Each “game” consists of a pre\-generated goal structureGG, which is partitioned to populate the Directors’ private information\(d1,d2,d3\)\(d\_\{1\},d\_\{2\},d\_\{3\}\)\. One of 6 partial completion statuses is randomly pre\-assigned, such that at the beginning of the game \(start statew∈𝒲w\\in\\mathcal\{W\}\), the board can beempty, or the structure can have itsfirst 1 or 2 layers, or a Director’s wall \(D1,D2,D3\) pre\-completed consistent with the goal state\. Directors are assigned personality archetypes to increase lexical diversity in their utterances, which are specified using fixed roleplay prompts \(Appendix[F](https://arxiv.org/html/2607.11053#A6)\)\. These form a prior over𝒰\\mathcal\{U\}, which is sampled from to create the dialogue history\. Thus a game maps neatly to the CEA specification \(Sec\.[3](https://arxiv.org/html/2607.11053#S3)\)\. Lemma[1](https://arxiv.org/html/2607.11053#Thmlemma1)predicts thatfailed movesare necessary consequences of non\-actionable instructions, and flouting of Gricean maxims should adversely impact the correctness of valid instructions\.
### 4\.2Data Generation
LeveragingCRAFTand 100 pre\-generated goal structures, we generated “gold standard” games for sampling preference data\. For each structure, we ran 2 games ofT=20T=20turns\. Separate GPT\-4\.1\-mini instances roleplayed the Builder and the Directors\.333GPT\-4\.1\-mini was selected after initial experimentation that showed it to optimally balance lexical diversity, overall task progress in 20 turns, and cost\. Details are given in Appendix[H](https://arxiv.org/html/2607.11053#A8)\.The Builder agent used the move exploration tool to maximize progress\. From this data, we constructed a preference dataset𝒟=\{\(x,yw,yℓ\)\}i=1N\\mathcal\{D\}=\\\{\(x,y\_\{w\},y\_\{\\ell\}\)\\\}\_\{i=1\}^\{N\}where each sample consisted of a dialogue history \(context\)xxterminating before a given turntt, the most actionable Director utteranceywy\_\{w\}\(defined as the utterance inttthat the Builder acted upon\) and a suboptimal utteranceyℓy\_\{\\ell\}from the same turn\. This data \(6,201 samples\) was used to align open\-weight models to act as Directors\. A further 20 goal structures were held out solely for evaluation \(Sec\.[5](https://arxiv.org/html/2607.11053#S5)\)\.
### 4\.3Metrics
One of our core insights in this work is that in a CEA setting, a Director utterance beingactionablemeans it must be cooperative \(Sec\.[3](https://arxiv.org/html/2607.11053#S3)\), but this does not mean that the instruction or Builder action based on it is necessarilycorrectw\.r\.t\. the final goal structure\. Therefore we calculated a “correctness score” \(0–6\) for each turn based on how the utterances and actions therein furthered task completion and group common ground\. Starting from a floor of 0 for afailedmove \(non\-actionable instruction\), 1 point was assigned if the Builder requested clarification \(signaling unactionable but partially informative instructions\)\. A minimum of 2 points was assigned for a successful move, with 1 additional point each for1\)placing/removing the right block at/from the right position in the goal structure,2\)placing/removing the right block at/from the right position in the in a Director’s private view,3\)if all blocks currently in the goal structure were in the correct place444To avoid unduly penalizing for errors made in previous turns, a discount factor ofγ=0\.5\\gamma=0\.5was applied to this metric\. The correctness value was incremented by 1 if this metric wasTrue, and ifFalse, by\(1−\(γp−1\)\)\(1\-\(\\gamma^\{p\-1\}\)\)whereppis the number of consecutive previous turns where the overall structure remained incorrect due to an existing uncorrected error\., and4\)if utterances; in the l turn agree on the move, according to an LLM\-Judge \(prompt in Appendix[F](https://arxiv.org/html/2607.11053#A6)\)\.
We also calculated objective metrics of task progress:IoU\(IoU between blocks on the board and in the goal structure\),Completion %\(percentage of target blocks that are correctly placed\),Position Accuracy\(accuracy of blocks in layers independent of layer order\), andOverall Progress\(mean of the previous 3\)\.555Formulas are given in Appendix[D\.1](https://arxiv.org/html/2607.11053#A4.SS1)\.As some games started with a non\-empty board, we normalized metrics relative to their values att=0t=0and report deltas\.
## 5Experiments
We testedclosed models\(GPT\-4\.1\-mini,GPT\-4o\-mini\) acting as the Builder, while closed oropen\-weight models\(Qwen2\.5\-7B\-Instruct,Llama 3\.1\-8B\-Instruct\) could be Directors\.666Full prompts are in Appendix[F](https://arxiv.org/html/2607.11053#A6)and training hyperparameters are in Appendix[G](https://arxiv.org/html/2607.11053#A7)\.Goal structures were randomly pre\-split among the 6 possible start states except where noted\. Each game was played with pre\-assigned Director archetypes, forT=20T=20turns, to keep conditions consistent\.
#### Prompting Experiments
This set of experiments assessed the effect ofpromptingthe Builder LLM \(a closed model\) to act as aliterallistener \(L0L\_\{0\}\) or apragmaticlistener \(L1L\_\{1\}\) as in the RSA framework\. These states were invoked through simple prompt “infixes” following prior works\(Ward et al\.,[2023](https://arxiv.org/html/2607.11053#bib.bib66); Nath & Krishnaswamy,[2025](https://arxiv.org/html/2607.11053#bib.bib43)\)\. The prompt infix used for the “Literal Builder” stated “Read each director’s message at face value”, while the “Pragmatic Builder” was prompted to identify which Director’s utterance was the most informative and to think through a set of questions adapted fromChenail & Chenail \([2014](https://arxiv.org/html/2607.11053#bib.bib8)\), relating to how well the information presented aligns with Gricean maxims\. These conditions were compared to a baseline condition where no explicit instructions on how to interpret Director utterances were provided, aside from the default system prompt from CRAFT\. Each of the 20 test structures was run twice\. This addressedRQ1andRQ2\.
#### Tool\-Calling Experiments
We further assessed outcomes where the Builder acted with themove exploration toolvs\. without\. This tested how Builder exploration before commitment improved selection of actionable utterances and therefore outcomes, addressingRQ2\.
#### Post\-Training Experiments
In this set of experiments, open\-weight models acting as Directors werepost\-trainedusing SFT, DPO\(Rafailov et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib57)\)and IPO\(Azar et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib2)\)methods against the preference dataset𝒟\\mathcal\{D\}\(Sec\.[4\.2](https://arxiv.org/html/2607.11053#S4.SS2)\)\. Intuitively, if the Builder were behaving more like a literal listener \(L0L\_\{0\}\), which could impact task progress, then the Directors could be made more robust to this by aligning their utterances toward more actionable examples as captured in the preference dataset\. GPT\-4\.1\-mini was chosen for the Builder \(see Appendix[H](https://arxiv.org/html/2607.11053#A8)for more\)\. Each test structure was runoncein all these experiments, which addressedRQ3\. The main experiments we report here start from an empty board state\.
## 6Results
Table 1:Comparison \(Mean±SEM\{\}\_\{\\pm\\text\{SEM\}\}\) of Overall Progress and Completion %Δ\\Deltaacross Builders \(Directors: GPT\-4\.1\-mini,N=40N=40\)\.ppvalues compare to analogous base prompt condition given a pairedtt\-test\. Significant values underlined at threshold ofp=0\.05p=0\.05\.Table[1](https://arxiv.org/html/2607.11053#S6.T1)shows progress compared to game start under various prompting/tool\-calling conditions using GPT\-4\.1\-mini and GPT\-4o\-mini\. The first thing we see is that GPT\-4\.1\-mini substantially outperformed GPT\-4o\-mini, so we focus on this Builder agent for the remainder of this paper\. Using the move exploration tool helps performance under all prompting conditions, suggesting that allowing the Builder to explore more interpretations of Director utterances helps it more optimally reason about Directors’ own interpretations ofuuand the intended action\. Eliciting pragmatic reasoning from the GPT\-4\.1 Builder with thePragmatic Builderprompt infix results in the best overall performance with use of the exploration tool, attaining statistical significance\. This combination is also the most reliable Builder, with the best “worst case” outcomes \(\-0\.102 Overall ProgressΔ\\Deltaand \-0\.120 Completion %Δ\\Delta\)\. However, the Pragmatic Builder without tool\-calling is as bad as or worse than the Literal Builder\. This aggregate evidence suggests that in this CEA setting, OTS LLMs show no significant difference in behavior from aliteral listener\(RQ1\), but can be made to behave more like pragmatic listeners with acombinationof appropriate prompting and exploration \(RQ2\)\. The combination is important: the prompt tells the model to reason pragmatically, the tool gives it the machinery to do so\.
Figure 3:Main prompting and tool\-calling experimental results with GPT\-4\.1\-mini instances as Directors and Builder\. Shading reflects std\. error of the mean\.Fig\.[3](https://arxiv.org/html/2607.11053#S6.F3)breaks down Overall ProgressΔ\\Deltaover turns and by start state\. We see the same improvement with move exploration and the Pragmatic Builder\. The strongest effect is in games starting with an empty board, which show steady progress toward the goal helped by pragmatic prompting and exploration\. In other cases the group makes almost no progress or even regresses by undoing some of the pre\-completed structure, even though it was guaranteed to be correct\. We discuss this further in Sec\.[7](https://arxiv.org/html/2607.11053#S7)\.
Figure 4:Effects of Director post\-training by prompting condition with move exploration\.Fig\.[4](https://arxiv.org/html/2607.11053#S6.F4)shows post\-trained open models \(Qwen and Llama\) as Directors in games starting from an empty board, by Builder prompt condition\. Supervised fine\-tuning of the Director model hurts performance relative to the base model—“mean\-seeking”\(Chan et al\.,[2022](https://arxiv.org/html/2607.11053#bib.bib6)\)SFT may produce “safe” actionable\-sounding utterances decoupled from actual task state\. However, DPO alignment brings it back up when the Builder is explicitly Literal or, in particular, Pragmatic\. This shows the convergence between more optimal speakers and pragmatic listeners predicted by RSA, alignment toward optimal utterances may provide robustness against listener suboptimalities \(RQ3\), with Qwen as a stronger Director overall than Llama\. Llama with IPO in particular seems to repeatedly place and remove the same few blocks, creating a correction spiral that inhibits progress\. IPO’s stronger KL\-regularization may keep it closer to the underperforming SFT reference model, creating a feedback loop\. The relatively small preference dataset𝒟\\mathcal\{D\}may also be a factor here\.
## 7Discussion: Where Do Pragmatic Inference Failures Occur?
Results show broad trends demonstrating how prompting, exploration, and post\-training elicit pragmatic speaking and listening behaviors in LLM collaborators under epistemic asymmetry\. However, the overall ceiling on pragmatic capabilities appears to remain low \(22\.96%outright move failure rate with move exploration\)\. Where, then, do pragmatic inference failures come from? We examine specific cases from the test data where Gricean maxim flouting as operationally defined correlated with non\-actionable instructions\.
\(a\)
\(b\)
\(c\)
\(d\)
Figure 5:\([5\(a\)](https://arxiv.org/html/2607.11053#S7.F5.sf1)\) Mean executed move correctness vs\. correct grounding dimensions in Director utterances\.Circlesindicate the mean correctness ofnon\-failed moves only\. \([5\(b\)](https://arxiv.org/html/2607.11053#S7.F5.sf2)\) Move failure rate/mean correctness vs\. normalized message length by Director model\. \([5\(c\)](https://arxiv.org/html/2607.11053#S7.F5.sf3)\)PLACEandREMOVEoutcomes vs\. maxim of relation adherence\. \([5\(d\)](https://arxiv.org/html/2607.11053#S7.F5.sf4)\)REMOVEoutcomes vs\. maxim of manner adherence\. For failure rate analyses, BuilderCLARIFYrequests are considered failed moves\. \(Builder: GPT\-4\.1\-mini with Pragmatic prompting and move exploration\)\.#### Quality Flouting
If a director does not accurately describe their wall view, the maxim of quality is flouted\. An LLM\-Judge \(GPT\-4\.1\-mini, the same model used as the Builder in most experiments\) examined Director utterances for accurate representation of blocks in their wall view in terms of color, size, and layer position, and assigned 1 point for each dimension correct\. Fig\.[5\(a\)](https://arxiv.org/html/2607.11053#S7.F5.sf1)shows mean move correctness vs\. number of accurate spatial grounding \(SG\) dimensions in Director utterances\. We see that Director alignment correlates with improved actionable quality of utterances over SFT models\. Certain models like Qwen IPO or Llama DPO can provide actionable instructions that result in low move failure rates, although the moves may not be correct\.
#### Quantity Flouting
Fig\.[5\(b\)](https://arxiv.org/html/2607.11053#S7.F5.sf2)shows move failure rate and correctness score between 0 and 6 as a function of normalized Director utterance length \(in words\) using a rolling average with a window of 20% of total samples\. Particularly with open models, move failure rate trends upward and correctness declines as Director message length grows\. This shows sensitivity to flouting the maxim of quantity, and that overinformativeness \(as proxied by length\) adversely affects actionability as Lemma[1](https://arxiv.org/html/2607.11053#Thmlemma1)predicts\. Aligned Llama appears less sensitive to this, with its low move failure rate also reflected here, but at the cost of correctness\.
#### Relation Flouting
If a Director’s wall is fully completed, the wall’s “owner” should avoid modifying it, as doing so would have no bearing on progress toward goal stateGGand flout the maxim of relation\. Fig\.[5\(c\)](https://arxiv.org/html/2607.11053#S7.F5.sf3)showsPLACEandREMOVEoutcomes \(fail/succeed\) in cases where a Director’s wall is complete but the “owner” gives an instruction anyway\. The Builder mayignorethis maxim flouting and act upon it, or act upon utterances from a different Director \(which adheres to the maxim\)\. Flouting the relation maxim strongly correlates with move failure: the “wall owner” frequently pushes for changes to a complete wall, which are not actionable and fail\. This explains in part why games that started with non\-empty boards showed inferior overall progress \(Sec\.[6](https://arxiv.org/html/2607.11053#S6), Fig\.[3](https://arxiv.org/html/2607.11053#S6.F3)\)\.777To save limited compute resources, Llama and IPO directors were not run in start conditions beginning with fully completed walls, which was required to assess this definition of relation\.
#### Manner Flouting
If there are two instructions in a turn from the same Director, which flouts the maxim of manner, it forces the Builder to marginalize over their interpretations and increase uncertainty\. However, the Builder may ignore this flouting and act anyway, choosing one of the multiple interpretations or a different Director’s utterance entirely\. Fig\.[5\(d\)](https://arxiv.org/html/2607.11053#S7.F5.sf4)shows the effect of the occurrence of manner flouting on move outcomes\. In contrast to flouting other maxims, we see that a strong Builder model is very capable of ignoring manner flouting and making successful moves\. Manner flouting shows a very slight effect of increasing failure ofREMOVEs, particularly in SFT models, compared to cases where there were no duplicate director utterances in a turn \(manner adherence\)\.
## 8Conclusion
In this paper we performed a novel investigation of LLMs pragmatic speaking and listening capabilities under conditions of epistemic asymmetry\. Similar questions have been examined in human\-human collaborations and single AI\-human collaborations, but never to our knowledge in multiagent collaborations, where the asymmetric nature of information available to different agents creates unique conditions that challenge LLMs’ reasoning capabilities\. Other recent work \(e\.g\.,Eisenstein et al\. \([2026](https://arxiv.org/html/2607.11053#bib.bib12)\)\) has identified shortcomings in LLM reasoning in private information contexts; importantly, our work identifies unique challenges inmultiagentsettings with shared goal\-directednesswithout explicit goalknowledge\. While we make no specific claims about cooperativity in human\-LLM interactions, our results suggest that even in LLM\-LLM interactions rendered in human\-readable language, LLMs suffer from pragmatic failures in both speaking and listening, but also display some pragmatic listening ability to ignore speakers’ flouting of certain maxims\.
We formalized a notion of collaborative epistemic asymmetry \(CEA\) that explicitly connects the actionability of task\-centered utterances to consistency with Grice’s maxims of conversation, showing that under certain assumptions, flouting the maxims negatively impacts actionability and thus collaborative task progress\. Explicitly connecting Gricean maxims to objective task metrics allowed us to investigate the contributions of prompting, exploration, and alignment to LLMs’ cooperativity in collaboration and the challenges this setting poses\. Empirical evidence bolsters our theoretical formulation: flouting Gricean maxims under epistemic asymmetry is associated with adverse impacts on agent\-agent coordination and task performance\. Closed model Builder agents are better pragmatic listeners when there is not too much context \(such as a partially full board\) to incorporate into their reasoning\. Open model directors can be aligned to be more optimal in their utterances but they remain fragile; certain models give instructions that are superficially cooperative but have low goal\-directness, resulting in trivially correct moves, like placing and removing the same block\. Future work may explore novel optimization techniques to enable Builder/Listener models to adaptively seek information from specific speakers or in certain ways, online Director optimization using methods like PPO or GRPO, or intentional Griceanviolationsby introducing perturbed private information or adversarial agents\.
## Ethics Statement
Due to human diversity of expression and the human tendency to behave in ways that challenge even theoretically rigorous and empirically validated AI systems, our results have bearing on human\-LLM interactions and suggest that assumptions of current LLMs’ pragmatic competence in collaborative settings should be treated cautiously or even skeptically\. These conclusions would need to be tested in real multiparty human\-AI collaborations in order to assess the precise effects of human diversity on LLM performance, as well as susceptibility to factors such as bias against non\-normative modes of communication\. Additionally, because of the focus of our work on agent\-agent interactions in collaborative tasks with measurable outcomes, “flouting” as used in the context of our experiments focuses less on social nuance in human communication, such as irony, hyperbole, or the use metaphors to imply meaning, in favor of measuring instances where LLMs break the rules outlined by the maxims of conversation without direct intent to imply additional meaning\. These social nuances are often situation\- and culture\-dependent and so agentic AI settings remain suboptimal settings to study and make claims about them\.
## References
- Andreas & Klein \(2016\)Jacob Andreas and Dan Klein\.Reasoning about pragmatics with neural listeners and speakers\.In*Proceedings of the 2016 conference on empirical methods in natural language processing*, pp\. 1173–1182, 2016\.
- Azar et al\. \(2024\)Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello\.A general theoretical paradigm to understand learning from human preferences\.In*International Conference on Artificial Intelligence and Statistics*, pp\. 4447–4455\. PMLR, 2024\.
- Boyd \(2021\)Kenneth Boyd\.Group understanding\.*Synthese*, 198\(7\):6837–6858, 2021\.
- Butler et al\. \(2025\)J\. Butler, S\. Jaffe, R\. Janßen, N\. Baym, B\. Hecht, J\. Hofman, S\. Rintel, B\. Sarrafzadeh, A\. Sellen, M\. Vorvoreanu, and J\. Teevan\.Microsoft new future of work report 2025\.Technical Report MSR\-TR\-2025\-58, Microsoft Research, 2025\.URL[https://aka\.ms/nfw2025](https://aka.ms/nfw2025)\.
- Cantiani et al\. \(2024\)Anabela Cantiani, Ilja van Beest, Frans Cruijssen, Goos Kant, and Thorsten M Erle\.Perspective\-taking predicts success in coalition formation\.*European Journal of Social Psychology*, 54\(6\):1364–1377, 2024\.
- Chan et al\. \(2022\)Alan Chan, Hugo Silva, Sungsu Lim, Tadashi Kozuno, A Rupam Mahmood, and Martha White\.Greedification operators for policy optimization: Investigating forward and reverse kl divergences\.*Journal of Machine Learning Research*, 23\(253\):1–79, 2022\.
- Chen et al\. \(2024\)Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, et al\.From persona to personalization: A survey on role\-playing language agents\.*Transactions on Machine Learning Research*, 2024\.
- Chenail & Chenail \(2014\)Jan Chenail and Ronald Chenail\.Communicating qualitative analytical results following grice’s conversational maxims\.*The Qualitative Report*, pp\. 276–285, 10 2014\.doi:10\.46743/2160\-3715/2011\.1053\.
- Clark \(1996\)Herbert H Clark\.*Using language*\.Cambridge university press, 1996\.
- Clark & Brennan \(1991\)Herbert H Clark and Susan E Brennan\.Grounding in communication\.1991\.
- Davidson et al\. \(2025\)Tim R Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, and Ece Kamar\.The collaboration gap\.*arXiv preprint arXiv:2511\.02687*, 2025\.
- Eisenstein et al\. \(2026\)Jacob Eisenstein, Fantine Huot, Adam Fisch, Jonathan Berant, and Mirella Lapata\.Mt\-pingeval: Evaluating multi\-turn collaboration with private information games\.*arXiv preprint arXiv:2602\.24188*, 2026\.
- Estienne et al\. \(2025\)Lautaro Estienne, Gabriel Ben Zenou, Nona Naderi, Jackie CK Cheung, and Pablo Piantanida\.Collaborative rational speech act: Pragmatic reasoning for multi\-turn dialog\.In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng \(eds\.\),*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 22509–22523, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.doi:10\.18653/v1/2025\.emnlp\-main\.1145\.URL[https://aclanthology\.org/2025\.emnlp\-main\.1145/](https://aclanthology.org/2025.emnlp-main.1145/)\.
- Filipi & Wales \(2004\)Anna Filipi and Roger Wales\.Perspective\-taking and perspective\-shifting as socially situated and collaborative actions\.*Journal of pragmatics*, 36\(10\):1851–1884, 2004\.
- Floyd et al\. \(2025\)Sammy Floyd, Olessia Jouravlev, Moshe Poliak, Zachary Mineroff, Edward Gibson, and Evelina Fedorenko\.Three distinct components of pragmatic language use: Social conventions, intonation, and world knowledge–based causal reasoning\.*Proceedings of the National Academy of Sciences*, 122\(50\):e2424400122, 2025\.
- Frank & Goodman \(2012\)Michael C Frank and Noah D Goodman\.Predicting pragmatic reasoning in language games\.*Science*, 336\(6084\):998–998, 2012\.
- Frank et al\. \(2016\)Michael C Frank, Andrés Gómez Emilsson, Benjamin Peloquin, Noah D Goodman, and Christopher Potts\.Rational speech act models of pragmatic reasoning in reference games\.*psyarxiv*, 2016\.
- Gan et al\. \(2025\)Yujian Gan, Yuan Liang, Yanni Lin, Juntao Yu, and Massimo Poesio\.Improving llms’ learning of coreference resolution\.In*Proceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue*, pp\. 311–321, 2025\.
- Gmytrasiewicz \(2020\)Piotr Gmytrasiewicz\.How to do things with words: A bayesian approach\.*Journal of Artificial Intelligence Research*, 68:753–776, 2020\.
- Gmytrasiewicz & Doshi \(2005\)Piotr J Gmytrasiewicz and Prashant Doshi\.A framework for sequential planning in multi\-agent settings\.*Journal of Artificial Intelligence Research*, 24:49–79, 2005\.
- Goodman & Frank \(2016\)Noah D Goodman and Michael C Frank\.Pragmatic language interpretation as probabilistic inference\.*Trends in cognitive sciences*, 20\(11\):818–829, 2016\.
- Goodman & Stuhlmüller \(2013\)Noah D Goodman and Andreas Stuhlmüller\.Knowledge and implicature: Modeling language understanding as social cognition\.*Topics in cognitive science*, 5\(1\):173–184, 2013\.
- Grice \(1975\)Herbert P Grice\.Logic and conversation\.In*Speech acts*, pp\. 41–58\. Brill, 1975\.
- Guo et al\. \(2026\)Xudong Guo, Kaixuan Huang, Jiale Liu, Wenhui Fan, Natalia Vélez, Qingyun Wu, Huazheng Wang, Thomas L Griffiths, and Mengdi Wang\.Embodied llm agents learn to cooperate in organized teams\.*IEEE Transactions on Computational Social Systems*, 2026\.
- Hardt \(2025\)Daniel Hardt\.Sparks of pure competence in llms: the case of syntactic center embedding in english\.In*Proceedings of the Society for Computation in Linguistics 2025*, pp\. 346–355, 2025\.
- Hattie \(2012\)John Hattie\.*Visible learning for teachers: Maximizing impact on learning*\.Routledge, 2012\.
- Hu et al\. \(2021\)Edward J\. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\.Lora: Low\-rank adaptation of large language models, 2021\.URL[https://arxiv\.org/abs/2106\.09685](https://arxiv.org/abs/2106.09685)\.
- Hu et al\. \(2025\)Jennifer Hu, Felix Sosa, and Tomer Ullman\.Re\-evaluating theory of mind evaluation in large language models\.*Philosophical Transactions of the Royal Society B: Biological Sciences*, 380\(1932\), 2025\.
- Jaccard \(1912\)Paul Jaccard\.The distribution of the flora in the alpine zone\. 1\.*New phytologist*, 11\(2\):37–50, 1912\.
- Jian & Siddharth \(2024\)Mingyue Jian and N Siddharth\.Are llms good pragmatic speakers?In*NeurIPS 2024 Workshop on Behavioral Machine Learning*, 2024\.
- Jurafsky \(2006\)Daniel Jurafsky\.Pragmatics and computational linguistics\.*The handbook of pragmatics*, pp\. 578–604, 2006\.
- Just et al\. \(2025\)Hoang Anh Just, Mahavir Dabas, Lifu Huang, Ming Jin, and Ruoxi Jia\.Dipt: Enhancing llm reasoning through diversified perspective\-taking\.In*Findings of the Association for Computational Linguistics: NAACL 2025*, pp\. 6344–6374, 2025\.
- Katsos et al\. \(2023\)Napoleon Katsos, Blanche Gonzales de Linares, Ekaterina Ostashchenko, and Elspeth Wilson\.Perspective\-taking in deriving implicatures: The listener’s perspective is important too\.*Cognition*, 241:105582, 2023\.
- Khebour et al\. \(2024\)Ibrahim Khalil Khebour, Kenneth Lai, Mariah Bradford, Yifan Zhu, Richard A Brutti, Christopher Tam, Jingxuan Tu, Benjamin A Ibarra, Nathaniel Blanchard, Nikhil Krishnaswamy, et al\.Common ground tracking in multimodal dialogue\.In*Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\)*, pp\. 3587–3602, 2024\.
- Kibria et al\. \(2024\)Raihan Kibria, Sheikh Intiser Uddin Dipta, and Muhammad Abdullah Adnan\.On functional competence of llms for linguistic disambiguation\.In*Proceedings of the 28th Conference on Computational Natural Language Learning*, pp\. 143–160, 2024\.
- Langlois & Everitt \(2021\)Eric D Langlois and Tom Everitt\.How rl agents behave when their actions are modified\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 35, pp\. 11586–11594, 2021\.
- Li et al\. \(2023\)Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem\.Camel: Communicative agents for” mind” exploration of large language model society\.*Advances in neural information processing systems*, 36:51991–52008, 2023\.
- Ma et al\. \(2025a\)Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, and Barbara Plank\.Pragmatics in the era of large language models: A survey on datasets, evaluation, opportunities and challenges\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 8679–8696, 2025a\.
- Ma et al\. \(2025b\)Ziqiao Ma, Jing Ding, Xuejun Zhang, Dezhi Luo, Jiahe Ding, Sihan Xu, Yuchen Huang, Run Peng, and Joyce Chai\.Vision\-language models are not pragmatically competent in referring expression generation\.In*Second Conference on Language Modeling*, 2025b\.
- Maeda et al\. \(2026\)Kiyosu Maeda, William P McCarthy, Ching\-Yi Tsai, Jeffrey Mu, Haoliang Wang, Robert D Hawkins, Judith E Fan, and Parastoo Abtahi\.Gesturing toward abstraction: Multimodal convention formation in collaborative physical tasks\.*arXiv preprint arXiv:2602\.08914*, 2026\.
- Maslej et al\. \(2025\)Nestor Maslej, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Njenga Kariuki, Emily Capstick, Anka Reuel, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Russell Wald, Toby Walsh, Armin Hamrah, Lapo Santarlasci, Julia Betts Lotufo, Alexandra Rome, Andrew Shi, and Sukrut Oak\.Artificial intelligence index report 2025, 2025\.URL[https://arxiv\.org/abs/2504\.07139](https://arxiv.org/abs/2504.07139)\.
- Min et al\. \(2025\)Qingkai Min, Zitian Qu, Qipeng Guo, Xiangkun Hu, Zheng Zhang, and Yue Zhang\.Multi\-document event extraction using large and small language models\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 19265–19296, 2025\.
- Nath & Krishnaswamy \(2025\)Abhijnan Nath and Nikhil Krishnaswamy\.Learning “partner\-aware” collaborators in multi\-party collaboration\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Nath et al\. \(2024a\)Abhijnan Nath, Shadi Manafi Avari, Avyakta Chelle, and Nikhil Krishnaswamy\.Okay, let’s do this\! modeling event coreference with generated rationales and knowledge distillation\.In*Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pp\. 3931–3946, 2024a\.
- Nath et al\. \(2024b\)Abhijnan Nath, Videep Venkatesha, Mariah Bradford, Avyakta Chelle, Austin C Youngren, Carlos Mabrey, Nathaniel Blanchard, and Nikhil Krishnaswamy\.“any other thoughts, hedgehog?” linking deliberation chains in collaborative dialogues\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pp\. 5297–5314, 2024b\.
- Nath et al\. \(2025a\)Abhijnan Nath, Carine Graff, Andrei Bachinin, and Nikhil Krishnaswamy\.Frictional agent alignment framework: Slow down and don’t break things\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 11042–11089, 2025a\.
- Nath et al\. \(2025b\)Abhijnan Nath, Carine Graff, and Nikhil Krishnaswamy\.Collaborate, deliberate, evaluate: How llm alignment affects coordinated multi\-agent outcomes\.*arXiv preprint arXiv:2509\.05882*, 2025b\.
- Nath et al\. \(2025c\)Abhijnan Nath, Carine Graff, and Nikhil Krishnaswamy\.Let’s roleplay: Examining llm alignment in collaborative dialogues\.*arXiv e\-prints*, pp\. arXiv–2509, 2025c\.
- Nath et al\. \(2026\)Abhijnan Nath, Hannah VanderHoeven, and Nikhil Krishnaswamy\.Craft: Grounded multi\-agent coordination under partial information, 2026\.URL[https://arxiv\.org/abs/2603\.25268](https://arxiv.org/abs/2603.25268)\.
- Nguyen \(2023\)Khanh Xuan Nguyen\.Language models are bounded pragmatic speakers\.In*First Workshop on Theory of Mind in Communicating Agents*, 2023\.
- Ouyang et al\. \(2022\)Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al\.Training language models to follow instructions with human feedback\.*Advances in neural information processing systems*, 35:27730–27744, 2022\.
- Pacuit \(2017\)Eric Pacuit\.*Neighborhood semantics for modal logic*\.Springer, 2017\.
- Park et al\. \(2024\)Dojun Park, Jiwoo Lee, Seohyun Park, Hyeyun Jeong, Youngeun Koo, Soonha Hwang, Seonwoo Park, and Sungeun Lee\.Multiprageval: Multilingual pragmatic evaluation of large language models\.In*Proceedings of the 2nd GenBench Workshop on Generalisation \(Benchmarking\) in NLP*, pp\. 96–119, 2024\.
- Perkins & Tagler \(2011\)David V Perkins and Michael J Tagler\.Jigsaw classroom\.*Promoting student engagement*, 1:195–197, 2011\.
- Poradzisz & Florczak \(2019\)Michele Poradzisz and Kristine L Florczak\.Collaboration: Does it require pragmatic thought?*Nursing Science Quarterly*, 32\(4\):271–275, 2019\.
- Premack & Woodruff \(1978\)David Premack and Guy Woodruff\.Does the chimpanzee have a theory of mind?*Behavioral and brain sciences*, 1\(4\):515–526, 1978\.
- Rafailov et al\. \(2024\)Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn\.Direct preference optimization: Your language model is secretly a reward model\.*Advances in neural information processing systems*, 36:53728–53741, 2024\.
- Reimers & Gurevych \(2019\)Nils Reimers and Iryna Gurevych\.Sentence\-bert: Sentence embeddings using siamese bert\-networks\.In*Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing \(EMNLP\-IJCNLP\)*, pp\. 3982–3992, 2019\.
- Saygin & Cicekli \(2002\)Ayse Pinar Saygin and Ilyas Cicekli\.Pragmatics in human\-computer conversations\.*Journal of Pragmatics*, 34\(3\):227–258, 2002\.
- Shannon \(1948\)Claude Elwood Shannon\.A mathematical theory of communication\.*The Bell system technical journal*, 27\(3\):379–423, 1948\.
- Sravanthi et al\. \(2024\)Settaluri Sravanthi, Meet Doshi, Pavan Tankala, Rudra Murthy, Raj Dabre, and Pushpak Bhattacharyya\.Pub: A pragmatics understanding benchmark for assessing llms’ pragmatics capabilities\.In*Findings of the Association for Computational Linguistics: ACL 2024*, pp\. 12075–12097, 2024\.
- Ullman \(2023\)Tomer Ullman\.Large language models fail on trivial alterations to theory\-of\-mind tasks\.*arXiv preprint arXiv:2302\.08399*, 2023\.
- Van Benthem & Pacuit \(2011\)Johan Van Benthem and Eric Pacuit\.Dynamic logics of evidence\-based beliefs\.*Studia Logica*, 99\(1\):61, 2011\.
- Van Benthem et al\. \(2014\)Johan Van Benthem, David Fernández\-Duque, and Eric Pacuit\.Evidence and plausibility in neighborhood structures\.*Annals of Pure and Applied Logic*, 165\(1\):106–133, 2014\.
- Van Rooy \(2004\)Robert Van Rooy\.Utility, informativity and protocols\.*Journal of philosophical logic*, 33\(4\):389–419, 2004\.
- Ward et al\. \(2023\)Francis Ward, Francesca Toni, Francesco Belardinelli, and Tom Everitt\.Honesty is the best policy: defining and mitigating ai deception\.*Advances in neural information processing systems*, 36:2313–2341, 2023\.
- Wilf et al\. \(2024\)Alex Wilf, Sihyun Lee, Paul Pu Liang, and Louis\-Philippe Morency\.Think twice: Perspective\-taking improves large language models’ theory\-of\-mind capabilities\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 8292–8308, 2024\.
- Wu et al\. \(2024\)Guande Wu, Chen Zhao, Claudio Silva, and He He\.Your co\-workers matter: Evaluating collaborative capabilities of language models in blocks world, 2024\.URL[https://arxiv\.org/abs/2404\.00246](https://arxiv.org/abs/2404.00246)\.
- Zarrieß & Schlangen \(2019\)Sina Zarrieß and David Schlangen\.Know what you don’t know: Modeling a pragmatic speaker that refers to objects of unknown categories\.In*Proceedings of the 57th annual meeting of the association for computational linguistics*, pp\. 654–659, 2019\.
- Zhu et al\. \(2025a\)Lixing Zhu, Jun Wang, and Yulan He\.Llmlink: Dual llms for dynamic entity linking on long narratives with collaborative memorisation and prompt optimisation\.In*Proceedings of the 31st international conference on computational linguistics*, pp\. 11334–11347, 2025a\.
- Zhu et al\. \(2025b\)Yifan Zhu, Changsoo Jung, Kenneth Lai, Videep Venkatesha, Mariah Bradford, Jack Fitzgerald, Huma Jamil, Carine Graff, Sai Kiran Ganesh Kumar, Bruce Draper, et al\.Multimodal common ground annotation for partial information collaborative problem solving\.In*Proceedings of the 21st Joint ACL\-ISO Workshop on Interoperable Semantic Annotation \(ISA\-21\)*, pp\. 85–91, 2025b\.
- Zhu et al\. \(2026\)Yifan Zhu, Mariah Bradford, Kenneth Lai, Timothy Obiso, Videep Venkatesha, James Pustejovsky, and Nikhil Krishnaswamy\.Distributed partial information puzzles: Examining common ground construction under epistemic asymmetry\.*arXiv preprint arXiv:2603\.05450*, 2026\.
## Appendix AProof of Lemma[1](https://arxiv.org/html/2607.11053#Thmlemma1)
###### Proof\.
Let us first explicitly define the constraints on an utteranceuuentailed byGrice \([1975](https://arxiv.org/html/2607.11053#bib.bib23)\)’s cooperative maxims in the collaborative epistemic asymmetry \(CEA\) setting \(see Sec\.[3](https://arxiv.org/html/2607.11053#S3)\):
- •Quality: speakerSiS\_\{i\}should only assert what they believe to true givendid\_\{i\}, their own information state\.
- •Quantity: speakerSiS\_\{i\}should assert no more and no less than is required to communicate elements ofdid\_\{i\}\.
- •Relation: speakerSiS\_\{i\}should assert only what is relevant to the shared goalGG, of whichdid\_\{i\}is partial information\.
- •Manner: speakerSiS\_\{i\}should assert clearly and without ambiguity\.
We now show that flouting any single maxim violates either\(1\)the condition thatuuperform a non\-trivial belief update to listenerLL’s posterior over world states—PL\(w\|u\)∝̸PL\(w\)P\_\{L\}\(w\|u\)\\not\\propto P\_\{L\}\(w\)\)—or\(2\)the condition thatuumust reduceLL’s uncertainty about the goalGG—H\(PL\(G\|u\)\)<H\(PL\(G\)\)H\(P\_\{L\}\(G\|u\)\)<H\(P\_\{L\}\(G\)\)\.
1. 1\.Qualityflouting: Ifuudoes not reflectdid\_\{i\},SiS\_\{i\}’s actual private information, thenuuwas not generated fromPi\(u\|w,di,ℐi\)P\_\{i\}\(u\|w,d\_\{i\},\\mathcal\{I\}\_\{i\}\), but rather from some other policy distribution uncoupled fromdid\_\{i\}\. Sincedid\_\{i\}must be drawn from a partitionℐi\\mathcal\{I\}\_\{i\}of world states𝒲\\mathcal\{W\}, removinguu’s dependence ondid\_\{i\}likewise removes its dependence onww\. Thus, listenerLL’s posterior collapses: PL\(w\|u\)∝PL\(u\|w,di,ℐi\)⋅PL\(w\)⇒\\displaystyle P\_\{L\}\(w\|u\)\\propto P\_\{L\}\(u\|w,d\_\{i\},\\mathcal\{I\}\_\{i\}\)\\cdot P\_\{L\}\(w\)\\RightarrowPL\(w\|u\)∝PL\(u\)⋅PL\(w\)⇒\\displaystyle P\_\{L\}\(w\|u\)\\propto P\_\{L\}\(u\)\\cdot P\_\{L\}\(w\)\\RightarrowPL\(w\|u\)∝c⋅PL\(w\)⇒\\displaystyle P\_\{L\}\(w\|u\)\\propto c\\cdot P\_\{L\}\(w\)\\RightarrowPL\(w\|u\)∝PL\(w\)\\displaystyle P\_\{L\}\(w\|u\)\\propto P\_\{L\}\(w\)\(2\) Condition 1 for actionability fails\.
2. 2\.Relationflouting: Ifuuis true and well\-formed, but irrelevant toGG, even ifuucauses a non\-trivial posterior update overww\(satisfying the first condition for actionability\), the update has no bearing on progress towardGG\.PL\(G\|u\)=PL\(G\)P\_\{L\}\(G\|u\)=P\_\{L\}\(G\), and thus the uncertainty towardGGremains the same \(H\(PL\(G\|u\)\)=H\(PL\(G\)\)H\(P\_\{L\}\(G\|u\)\)=H\(P\_\{L\}\(G\)\)\)\. Condition 2 for actionability fails\.
3. 3\.Mannerflouting: Ifuuviolates the maxim of manner by being unclear or ambiguous, then it has multiple interpretations, e\.g\.,\{⟦u⟧1,…,⟦u⟧N\}\\\{\\llbracket u\\rrbracket\_\{1\},\\dots,\\llbracket u\\rrbracket\_\{N\}\\\}\(rendered hereafter as\{u1,…,uN\}\\\{u\_\{1\},\\ldots,u\_\{N\}\\\}for simplicity\)\. Each interpretation supports different posterior updates, and the listener must marginalize over them: PL\(w\|u\)=∑k=1NPL\(w\|uk\)⋅PL\(uk\|u\)\\displaystyle P\_\{L\}\(w\|u\)=\\sum\_\{k=1\}^\{N\}P\_\{L\}\(w\|u\_\{k\}\)\\cdot P\_\{L\}\(u\_\{k\}\|u\)\(3\) This creates a mixture of distributions whose entropy has a lower bound of the weighted combination of component entropies: H\(PL\(w\|u\)\)≥∑k=1NPL\(uk\|u\)⋅H\(PL\(w\|uk\)\)\\displaystyle H\(P\_\{L\}\(w\|u\)\)\\geq\\sum\_\{k=1\}^\{N\}P\_\{L\}\(u\_\{k\}\|u\)\\cdot H\(P\_\{L\}\(w\|u\_\{k\}\)\)\(4\) Therefore adoption of any one of the interpretations causesH\(PL\(G\|uk\)\)H\(P\_\{L\}\(G\|u\_\{k\}\)\)to rise relative toH\(PL\(G\)\)H\(P\_\{L\}\(G\)\), or if no interpretation is adopted,H\(PL\(G\)\)H\(P\_\{L\}\(G\)\)remains the same\. Condition 2 for actionability is thus violated\.
4. 4\.Quantityflouting: An utteranceuumight violate the maxim of quantity in two ways—by beingunderinformativeor beingoverinformative\. - •Ifuuisunderinformative, it means that speakerSiS\_\{i\}has withheld information relevant towwthatdid\_\{i\}contains\. Listener updates toPL\(w\|u\)P\_\{L\}\(w\|u\)are made based on less information thanSiS\_\{i\}has available\. This leaves residual entropy aboutGGthat could have been reduced ifuuwere fully informative\. - •Ifuuisoverinformative, it means thatuuincludes some information that is irrelevant toGG\. The listener cannot distinguish which parts ofuuare relevant or not without knowing what part ofuuis relevant already, but knowing that would renderuuunnecessary\. In either case, entropy reductionH\(PL\(G\|u\)\)<H\(PL\(G\)\)H\(P\_\{L\}\(G\|u\)\)<H\(P\_\{L\}\(G\)\)cannot be guaranteed and condition 2 for actionability fails\. Note that while flouting the maxims of quality, relation, and manner deterministically violate one of the conditions for actionabilty, flouting the maxim of quantity does not necessarily do the same\. Flouting the maxim of quantity only prohibits aguaranteethat entropy toward the goal is reduced\. An over\- or underinformative utterancemaystill be actionable in certain cases, even if actionability cannot be guaranteed in general\. Instead, the maxim of quantity can be characterized as asufficientcondition for reliable actionability: adhering to the maxim of quantity guarantees a reduction in entropy toward the goal that may not otherwise occur\. However, in specific cases, severeviolationsof the maxim of quantity—which underGrice \([1975](https://arxiv.org/html/2607.11053#bib.bib23)\)are often considered intentional—can also lead to violations of condition 1 for actionability \(non\-trivial posterior update\): - •Severeunderinformativeness: if speakerSiS\_\{i\}withholds so much ofdid\_\{i\}thatuuis essentially vacuous with respect toww, thenuucontains no information dependent onww\. Following the demonstration of flouting the maxim ofqualityabove, removinguu’s dependencewwcauses the listener’s posterior to collapse:PL\(w\|u\)∝PL\(w\)P\_\{L\}\(w\|u\)\\propto P\_\{L\}\(w\)\. Severe underinformativeness collapses a quantity violation to a quality violation\. Condition 1 for actionability is thus violated\. - •Severeoverinformativeness: ifuumixes genuine signal aboutwwwith ancillary noise, then the listener must determine which interpretation ofuuto use by marginalizing over them, as in the manner case\. This likewise creates a mixture of distributions and flattens the posterior, leading to a violation of condition 2 for actionability\. In the limit of extreme noise, this drivesPL\(w\|u\)→PL\(w\)P\_\{L\}\(w\|u\)\\rightarrow P\_\{L\}\(w\), causing a violation of condition 1 as well\.
Flouting or violating a single cooperative maxim result in violations of at least one condition for actionability in a CEA setting as defined \(Sec\.[3](https://arxiv.org/html/2607.11053#S3)\)\. ∎
###### Corollary 1\.
In collaborative, fully observable, symmetric information settings \(that is, settings where all agents have the equal access to the same set of observables\) and the goalGGis commonly known, condition 1 \(non\-trivial posterior update\) alone is necessary and sufficient for actionability\. Under symmetric information conditions, the joint information state𝐝\\mathbf\{d\}is not partitioned and every agent’sdi≅𝐝d\_\{i\}\\cong\\mathbf\{d\}\. All agents can then compute the same posterior updateP\(w\|𝐝\)P\(w\|\\mathbf\{d\}\)\. The goalGGbecomes identifiable with a specific region of the set of world states𝒲\\mathcal\{W\},w∗w^\{\*\}\. WithGGandw∗w^\{\*\}thus conflated,H\(PL\(G\)\)≈0H\(P\_\{L\}\(G\)\)\\approx 0before any utterance is made\. Entropy that does not exist cannot be reduced and so condition 2 for actionability loses its independent force under symmetric information conditions\. Thus Lemma[1](https://arxiv.org/html/2607.11053#Thmlemma1)depends on the collaborative epistemic asymmetry \(CEA\) setting\.
## Appendix BRSA Example from DPIP Transcript
### B\.1Pragmatic Repair and Grounding in Multi\-Agent Dialogue
We analyze a representative dialogue segment from the real human DPIP collaborative building task\(Zhu et al\.,[2025b](https://arxiv.org/html/2607.11053#bib.bib71);[2026](https://arxiv.org/html/2607.11053#bib.bib72)\), illustrating ambiguity resolution, perceptual misalignment, and multi\-agent pragmatic repair\. The action of interest occurs at timet=599\.6t=599\.6s \(REMOVErsfromlayer1layer~1\)\. Table[2](https://arxiv.org/html/2607.11053#A2.T2)presents a curated transcript leading up to this action\.
Table 2:Samples dialogue segment from the DPIP Lego task data\(Zhu et al\.,[2026](https://arxiv.org/html/2607.11053#bib.bib72)\)illustrating ambiguity resolution, perceptual misalignment, and multi\-agent grounding prior to action execution\.Δt\\Delta tis measured relative to the action att=599\.6t=599\.6s\.#### RSA Interpretation
We interpret this interaction through the lens of Rational Speech Acts \(RSA\), where the builder acts as a pragmatic listener maintaining a belief distributionP\(w\|u\)P\(w\|u\)over possible world stateswwgiven utterancesuu\.
Table 3:RSA\-based interpretation of key utterances in the dialogue\.
#### Analysis\.
This example illustrates a full pragmatic “belief repair” cycle\. Initially, the builder’s posteriorP\(w\|u\)P\(w\|u\)exhibits high entropy due to ambiguous and perceptually inconsistent references \(e\.g\., color disagreement\)\. The builder issues targeted clarification queries to maximize expected information gain, while multiple directors iteratively refine their utterances to increase informativeness\. Notably, the dialogue reveals a mismatch in perceptual grounding, where different agents map the same object to different color categories\. Through repeated corrections and confirmations, the agents align their internal representations, effectively converging to a shared belief state\. Once the posterior over world states becomes sufficiently concentrated, the builder executes the intended action\.
Formally, the interaction can be viewed as an iterative process where the speaker selects utterances according to:
u∗=argmaxulogPL\(w∗\|u\)−C\(u\),u^\{\*\}=\\arg\\max\_\{u\}\\log P\_\{L\}\(w^\{\*\}\|u\)\-C\(u\),while the listener updates beliefs via:
P\(w\|u\)∝P\(u\|w\)⋅P\(w\),P\(w\|u\)\\propto P\(u\|w\)\\cdot P\(w\),and issues clarification queries when uncertainty remains high\. The presence of multiple directors introduces a multi\-agent extension, where belief updates incorporate multiple utterances:
P\(w\|u1,u2,…\)∝∏iP\(w\|ui\)\.P\(w\|u\_\{1\},u\_\{2\},\\dots\)\\propto\\prod\_\{i\}P\(w\|u\_\{i\}\)\.
This interaction highlights that pragmatic communication in DPIP is not a single\-turn inference process but a multi\-step, collaborative grounding procedure under partial observability\.
## Appendix CCRAFT Simulator Usage
### C\.1Block Encoding and the World State
The environment is defined over a3×33\\times 3grid of positions, where each location is indexed by a coordinate pair\(i,j\)\(i,j\)withi,j∈\{0,1,2\}i,j\\in\\\{0,1,2\\\}\. Each grid cell contains an ordered stack of blocks, where the ordering corresponds to vertical layers\.
Blocks are represented as two\-character strings: the first character specifies the color—green, blue, red, yellow, or orange—and the second character indicates size, either small \(s\) or large \(l\)\. The complete set of admissible block types is given by:
ℬ=\{gs,gl,bs,bl,rs,rl,ys,yl,os,ol\}\\mathcal\{B\}=\\\{\\texttt\{gs\},\\texttt\{gl\},\\texttt\{bs\},\\texttt\{bl\},\\texttt\{rs\},\\texttt\{rl\},\\texttt\{ys\},\\texttt\{yl\},\\texttt\{os\},\\texttt\{ol\}\\\}
At any timestep, the world state is described by a functionS:𝒞→ℬ∗S:\\mathcal\{C\}\\to\\mathcal\{B\}^\{\*\}, where each coordinatec∈𝒞=\{\(i,j\)∣i,j∈\{0,1,2\}\}c\\in\\mathcal\{C\}=\\\{\(i,j\)\\mid i,j\\in\\\{0,1,2\\\}\\\}maps to a \(possibly empty\) ordered sequence of blocks\. Each stack is bounded to a maximum height of three layers\.
### C\.2Structure Generation
We construct target structures through a two\-phase procedure: assigning stack heights across the grid, followed by independently populating each layer with blocks\.
#### Stack Height Assignment
The grid positions are divided into two categories\. Seven*required*positions—all cells except\(1,1\)\(1,1\)and\(2,1\)\(2,1\)—are deterministically assigned a height of three layers\. The remaining two*optional*positions,\(1,1\)\(1,1\)and\(2,1\)\(2,1\), are assigned heights drawn independently and uniformly from\{0,1,2\}\\\{0,1,2\\\}\. This scheme ensures that the outer region of the grid forms a consistently tall structure, while the interior exhibits variability in depth\.
#### Layer Tiling
Given the height assignments, each layer is populated independently\. For a particular layer, the subset of positions requiring blocks is determined by the previously sampled heights\. These positions are then filled using a combination of small and large blocks\. Large blocks occupy two orthogonally adjacent cells within the same layer, forming a domino configuration, and are never placed across layers\. Small blocks occupy a single cell\.
For each candidate position, the generator probabilistically attempts to place a domino with an available orthogonal neighbor\. If no suitable neighbor exists or the attempt is unsuccessful, a small block is placed instead\. Block colors are sampled uniformly from the set of five colors\. To reduce repetitive patterns, the generator performs a limited number of retries to avoid assigning identical block types to the same position in consecutive layers\.
#### Complexity Classification
After generation, structures are categorized based on their total number of blocks\. Structures containing at most 22 blocks are labeled*simple*, those with 23–24 blocks are labeled*medium*, and those with more than 24 blocks are labeled*complex*\. Since the required positions always contribute 21 blocks \(seven positions each with three layers\), variation in complexity primarily arises from the optional positions and the frequency of large blocks, which may increase block counts when large block placements cover positions that would otherwise remain empty\.
## Appendix DProgress Measurement
### D\.1Metrics
Progress toward the target structureS∗S^\{\*\}is measured after each successful move and computes four complementary metrics over the normalized representations of the current stateStS\_\{t\}and targetS∗S^\{\*\}\.
#### Intersection over Union \(IoU\)
For each positionc∈𝒞c\\in\\mathcal\{C\}, letAc=\{b∈St\(c\)\}A\_\{c\}=\\\{b\\in S\_\{t\}\(c\)\\\}andBc=\{b∈S∗\(c\)\}B\_\{c\}=\\\{b\\in S^\{\*\}\(c\)\\\}be the multisets of blocks treated as sets\. The IoU score aggregates overlap across all positions:
IoU\(St,S∗\)=∑c∈𝒞\|Ac∩Bc\|∑c∈𝒞\|Ac∪Bc\|\\text\{IoU\}\(S\_\{t\},S^\{\*\}\)=\\frac\{\\sum\_\{c\\in\\mathcal\{C\}\}\|A\_\{c\}\\cap B\_\{c\}\|\}\{\\sum\_\{c\\in\\mathcal\{C\}\}\|A\_\{c\}\\cup B\_\{c\}\|\}
This metric is insensitive to block order within a stack and rewards partial position matches\.
#### Completion Percentage
This metric measures layer\-exact correctness — a block at positionccand layerkkcounts as correct only if it matchesS∗\(c\)\[k\]S^\{\*\}\(c\)\[k\]:
CP\(St,S∗\)=∑c∈𝒞∑k=0\|S∗\(c\)\|−1𝟏\[St\(c\)\[k\]=S∗\(c\)\[k\]\]∑c∈𝒞\|S∗\(c\)\|\\text\{CP\}\(S\_\{t\},S^\{\*\}\)=\\frac\{\\sum\_\{c\\in\\mathcal\{C\}\}\\sum\_\{k=0\}^\{\|S^\{\*\}\(c\)\|\-1\}\\mathbf\{1\}\[S\_\{t\}\(c\)\[k\]=S^\{\*\}\(c\)\[k\]\]\}\{\\sum\_\{c\\in\\mathcal\{C\}\}\|S^\{\*\}\(c\)\|\}
#### Position Accuracy
A coarser metric that rewards positions where the set of blocks matches exactly, regardless of layer order:
PA\(St,S∗\)=19∑c∈𝒞𝟏\[\{b:b∈St\(c\)\}=\{b:b∈S∗\(c\)\}\]\\text\{PA\}\(S\_\{t\},S^\{\*\}\)=\\frac\{1\}\{9\}\\sum\_\{c\\in\\mathcal\{C\}\}\\mathbf\{1\}\[\\\{b:b\\in S\_\{t\}\(c\)\\\}=\\\{b:b\\in S^\{\*\}\(c\)\\\}\]
#### Overall Progress
The scalar summary used for termination and trend analysis is the unweighted mean of the three metrics:
OP\(St,S∗\)=IoU\+CP\+PA3\\text\{OP\}\(S\_\{t\},S^\{\*\}\)=\\frac\{\\text\{IoU\}\+\\text\{CP\}\+\\text\{PA\}\}\{3\}
### D\.2Progress Tracking and Delta Computation
After each move, we compute a deltaΔt=OP\(St,S∗\)−OP\(St−1,S∗\)\\Delta\_\{t\}=\\text\{OP\}\(S\_\{t\},S^\{\*\}\)\-\\text\{OP\}\(S\_\{t\-1\},S^\{\*\}\)relative to the previous turn, for each metric\. A recent trend estimate is computed over a sliding window of size 3:
τ=1min\(t,3\)∑i=t−2tΔi\\tau=\\frac\{1\}\{\\min\(t,3\)\}\\sum\_\{i=t\-2\}^\{t\}\\Delta\_\{i\}
The game is considered*improving*ifτ\>−0\.05\\tau\>\-0\.05, allowing for small temporary regressions\. An estimated turns remaining is computed as⌈\(1−OPt\)/Δ¯⌉\\lceil\(1\-\\text\{OP\}\_\{t\}\)/\\bar\{\\Delta\}\\rceilwhereΔ¯\\bar\{\\Delta\}is the average per\-turn progress, though this estimate degrades in stagnation conditions whereΔ¯≈0\\bar\{\\Delta\}\\approx 0\.
### D\.3Director View Projections
Each director is assigned a fixed 2D projection of the 3D world, corresponding to a distinct face of the structure\. These projections are defined as follows:
#### D1 — Left Column View
D1 observes the cells\(0,0\)\(0,0\),\(1,0\)\(1,0\), and\(2,0\)\(2,0\)across all vertical layers, corresponding to the left\-facing side of the structure\.
#### D2 — Top Row View
D2 observes the cells\(0,0\)\(0,0\),\(0,1\)\(0,1\), and\(0,2\)\(0,2\)across all vertical layers, corresponding to the far\-facing side\.
#### D3 — Right Column View
D3 observes the cells\(0,2\)\(0,2\),\(1,2\)\(1,2\), and\(2,2\)\(2,2\)across all vertical layers, corresponding to the right\-facing side\.
Within each projection, blocks are displayed from left to right according to the director’s viewing orientation\. Each visible cell is encoded as a color–size pair, while empty cells are represented with colornone\. A large block is represented as size 2 only when both cells of its domino placement lie within the director’s field of view; otherwise, it appears as size 1 since only a single face is observable\.
#### Information Coverage
The views of D1 and D3 overlap at a single coordinate,\(0,0\)\(0,0\), which serves as a shared reference point between the two lateral perspectives\. In contrast, D2 uniquely observes interior cells such as\(1,1\)\(1,1\)and\(2,1\)\(2,1\), corresponding to the optional positions\. As a result, D2 plays a crucial role in conveying information about interior structure depth\. No individual director has sufficient information to reconstruct the full 3D state independently; effective collaboration requires each director to communicate information that is inaccessible to the others\.
## Appendix EDetails on Generated Data
### E\.1Descriptive Statistics
Table 4:Overall Progress and Completion %Δ\\Deltafrom Start over 100 test structures, by start condition completion category \(Builder: GPT\-4\.1\-mini, Director: GPT\-4\.1\-mini\)\.Table[4](https://arxiv.org/html/2607.11053#A5.T4)shows Overall Progress and Completion %Δ\\Deltafrom Start over the 2×\\times100 runs of the ”training” structures, broken down by start condition completion category\.
### E\.2Preference Data Selection
For creation of the preference dataset \(Sec\.[4\.2](https://arxiv.org/html/2607.11053#S4.SS2)\), selection of chosen utteranceywy\_\{w\}and rejected utteranceyℓy\_\{\\ell\}given dialogue contextxxand turnttwas performed using the following method:
1. 1\.The Builder’s confirmation \(Sec\.[4\.1](https://arxiv.org/html/2607.11053#S4.SS1)\) was searched for mentions of specific Directors \(D1–D3\); if only one match was found, that Director’s utterance was retained asywy\_\{w\}\.
2. 2\.Otherwise, we built keystrings of key words in the Builder’s confirmation and in the matching Director’s utterances888If no Director label was found in the Builder’s confirmation, this step checks againstallDirector utterances fromtt\.\(consisting of color, size, or spatial relation terms\), computed the IoU\(Jaccard,[1912](https://arxiv.org/html/2607.11053#bib.bib29)\)over the Builder keystring and each Director keystring, and retained asywy\_\{w\}the Director utterance whose keystring returns the highest value\.
3. 3\.In the event of a tie in IoU, we usedSentenceTransformers\-MiniLM\-L6\-v2\(Reimers & Gurevych,[2019](https://arxiv.org/html/2607.11053#bib.bib58)\)to construct embeddings of the Builder confirmation and the remaining candidate Director utterances and used the highest cosine similarity to break ties\.
4. 4\.All other utterances in the turn not selected according to this method were paired with theywy\_\{w\}utterance as an associatedyℓy\_\{\\ell\}\.
## Appendix FPrompts
This section describes prompt infixes used in the prompting experiments and for calculating LLM\-Judge\-based metrics\. All other prompts remain unchanged from the pre\-specified system prompts in the CRAFT simulator, reported inNath et al\. \([2026](https://arxiv.org/html/2607.11053#bib.bib49)\)\.
### F\.1Literal and Pragmatic Builder Prompts
The Literal and Pragmatic prompts used in the main experiments were both Variant 1, given below\. The other variants were used in experiments to test sensitivity to prompt wording \(see Appendix[I](https://arxiv.org/html/2607.11053#A9)\)\.
Literal Prompt Variant 1Read each director’s message at face value\. Use only the semantic meaning of the directors’ utterances and your prior beliefs in your interpretation\. Disregard director intent or informativeness and consider only the literal interpretation\.
Pragmatic Prompt Variant 1Infer what you can from the director’s message\.Before deciding on a move, consider the following and make inferences based on your own thinking:1\. Identify which director gave the most informative description in order to take an action and further the task\.2\. Do you think any of the dialogue is false or misleading?3\. Do you find the dialogue to be appropriate and focused on the task?4\. Do you find the meaning of any words used in the dialogue to be unclear or obscure?
Literal Prompt Variant 2Interpret each director’s messages with strict literalism\. Use only the semantic meaning of the directors’ utterances and your prior beliefs in your interpretation\. Treat the director statements as a series of literal propositions, independent of the speaker’s goals\.
Pragmatic Prompt Variant 2Infer what you can from the director’s message\.Before deciding on a move, consider the following and make inferences based on your own thinking:1\. Do you find any of the dialogue to be unnecessary in order for you to understand the perspectives of the directors?2\. Do find that any of the observations or conclusions made in the dialogue lack adequate evidence to support their ideas?3\. Do you find the dialogue to be appropriate and focused on the task?4\. Do you find the organization of the dialogue to be clear and easy to follow?
Literal Prompt Variant 3Restrict interpretation to the compositional semantics of each utterance\. Disregard pragmatic inference or speaker intent; evaluate statements solely as independent logical propositions\. Base interpretations exclusively on literal meaning and established prior knowledge\.
Pragmatic Prompt Variant 3Infer what you can from the director’s message\.Before deciding on a move, consider the following and make inferences based on context, intent, and shared knowledge1\. Do you find the dialogue to be informative enough to take an action and further the task?2\. Is any of this phrasing potentially misrepresentative?3\. Does the exchange remain focused, or does it drift into irrelevant details?4\. Is the structure of this conversation easy to navigate?
### F\.2Common Ground LLM\-Judge Prompt
Common Ground Prompt \(I\): Identity and Game StateYou are a Common Ground Analysis agent for a collaborative LEGO construction task\.Your job is to determine the TRUE alignment between three directors \(D1, D2, D3\) based on both their public messages and private internal thoughts\.CURRENT BOARD STATE:\{current\_board\_state\}CONVERSATION HISTORY:\{conversation\_history\}DIRECTOR INTERNAL THOUGHTS \(PRIVATE\):D1 Internal:\{d1\_internal\}D2 Internal:\{d2\_internal\}D3 Internal:\{d3\_internal\}DIRECTOR PUBLIC MESSAGES:D1 Message:\{d1\_public\} D2 Message:\{d2\_public\} D3 Message:\{d3\_public\}
Common Ground Prompt \(II\): Analysis InstructionsANALYSIS INSTRUCTIONS:1\. Look for discrepancies between what directors say publicly vs\. think privately2\. Identify blocks/positions where directors have genuine alignment vs\. surface agreement3\. Detect uncertainty, confusion, or hidden disagreements4\. Consider confidence levels expressed in internal thoughtsOUTPUT TWO SECTIONS:<analysis\>Provide a brief analysis \(3\-4 sentences\) about the true alignment state\. Highlight any:\- Surface agreements that mask underlying disagreement/uncertainty\- Positions where directors are genuinely aligned\- Areas of confusion or misunderstanding\- Confidence mismatches between directors</analysis\><aligned\_structure\>``` {{ "D1": {{ "row_0": [{{ "color":"[color]", "size":[size], "confidence":"high/medium/low" }}, {{ "color":"[color]", "size":[size], "confidence":"high/medium/low" }}, {{ "color":"[color]", "size":[size], "confidence":"high/medium/low" }}], "row_1": [...], "row_2": [...] }}, "D2": {{...}}, "D3": {{...}} }} ``` </aligned\_structure\>IMPORTANT NOTES:\- Use ”unknown” for positions not mentioned or unclear\- Use ”disputed” for positions where directors disagree\- Set confidence based on internal thoughts: ”high” = certain, ”medium” = somewhat sure, ”low” = uncertain/guessing\- The aligned structure should reflect what each director TRULY believes \(from internal thoughts\), not just what they said publicly\- Colors: ”red”, ”blue”, ”green”, ”yellow”, ”orange”, ”brown”, ”unknown”, ”disputed”
### F\.3Director Archetype Prompts
Assertive Archetype PromptYou are confident and direct\. You form hypotheses quickly from your data and share them, but you genuinely listen to other groups and update your thinking when their evidence is compelling\. You sometimes move faster than the evidence warrants but you’re not closed\-minded\.
Cautious Archetype PromptYou are methodical and prefer to verify before claiming\. You ask clarifying questions and often synthesize what others have said before adding your own interpretation\. You can make claims when evidence is strong enough — you’re not paralyzed, just careful\.
Observant Archetype PromptYou notice patterns and anomalies in your data that others might overlook\. You tend to flag inconsistencies and ask “does this match what you’re seeing?” rather than broadcasting conclusions\. You’re collaborative by nature and often connect dots across groups\.
Skeptical Archetype PromptYou question assumptions including your own\. When someone makes a claim you probe it — not to be difficult but because you want the group to get it right\. You’re comfortable with uncertainty and say so openly\.
Synthesizer Archetype PromptYou actively try to integrate what all groups are saying into a coherent picture\. You summarize, reconcile contradictions, and push the group toward a shared understanding\. You ask “how does your data fit your view and what the other directors have said?”
### F\.4Quality LLM\-Judge Prompt
Spatial Grounding Quality Judge PromptYou are evaluating the spatial grounding quality of a director agent in a collaborative construction task\. The director has a private view of one wall of a 3D target structure and must reason about what blocks are missing before instructing a builder\.TARGET VIEW \(what this director needs the structure to look like\):\{target\_view\} CURRENT BOARD STATE:\{board\_state\}DIRECTOR PUBLIC MESSAGE:\{public\_message\}EVALUATION:For each question below answer with “Yes” or “No”, and provide a brief one\-sentence justification\.Questions:1\. Does the director’s public message accurately identify at least one block that appears in their privately\-held target view of their wall?2\. Does the public message correctly interpret the size of this block \(small versus large\) as it appears in the target view?3\. Does the director’s public message accurately reference the correct layer that this block should be at, accounting for what is already stacked at that position on the current game board?Return your response as a JSON object with keys SG1 through SG3, each containing an “answer” field \(“Yes” or “No”\) and a “reason” field\. Return only valid JSON with no additional text\.
Due to budget constraints, the Spatial Grounding Quality Judge was run once over each Director utterance\. As validation, we ran the Judge prompt over a subsample of utterance \(consisting of 10 games with Llama variants as Directors\), and established that std\. error of the mean was very low \(Table[5](https://arxiv.org/html/2607.11053#A6.T5)\), validating the use of a single Judge evaluation\.
Table 5:Outcomes of Spatial Grounding Quality LLM\-Judge \(Mean±SEM\{\}\_\{\\pm\\text\{SEM\}\}\) run twice over a subsample of dialogues \(10 games\)\.
## Appendix GModel Training Hyperparameters
We use Qwen2\.5\-7B\-Instruct999[https://huggingface\.co/Qwen/Qwen2\.5\-7B\-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct)and Llama 3\.1\-8B\-Instruct101010[https://huggingface\.co/meta\-llama/Meta\-Llama\-3\-8B\-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct)as base models for our finetuning experiments\. Specifically, we trained the SFT\(Ouyang et al\.,[2022](https://arxiv.org/html/2607.11053#bib.bib51)\), DPO\(Rafailov et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib57)\)and IPO\(Azar et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib2)\)baselines using these two base models as the initial starting point\. Similar to prior work\(Rafailov et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib57)\), we use SFT\-trained reference models for preference\-alignment training in all DPO and IPO baselines\.
### Supervised Fine\-Tuning Baseline
Both SFT baselines—Qwen2\.5\-7B\-Instruct and Llama 3\.1\-8B\-Instruct—were fine\-tuned using parameter\-efficient fine\-tuning via LoRA\(Hu et al\.,[2021](https://arxiv.org/html/2607.11053#bib.bib27)\)\. Supervised learning loss was applied on the chosen utterancesywy\_\{w\}\(Sec\.[E\.2](https://arxiv.org/html/2607.11053#A5.SS2)\) with a held\-out validation split evaluated every 60 steps\. We trained for 3 epochs with a learning rate of2×10−52\\times 10^\{\-5\}, effective batch size of 48 \(batch size 12, gradient accumulation 4\), and maximum sequence length of 1024 tokens\. LoRA rank was set tor=32r=32with scalingα=16\\alpha=16and dropout 0\.05, applied to all attention projection layers with no quantization\. Full hyperparameters are summarized in Table[6](https://arxiv.org/html/2607.11053#A7.T6)\.
Table 6:SFT baseline hyperparameters\.
### DPO and IPO Baselines
For both initial base models, preference aligment baselines that require constrastive sample training—DPO\(Rafailov et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib57)\)and IPO\(Azar et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib2)\)—were initialized from the SFT checkpoint \(epoch 3\) rather than the base model, following standard practice\. We use both the chosen utterancesywy\_\{w\}and the rejected utterancesyℓy\_\{\\ell\}as contrastive utterance samples \(Sec\.[E\.2](https://arxiv.org/html/2607.11053#A5.SS2)\) and trained for 4 epochs with a learning rate of5×10−65\\times 10^\{\-6\}, effective batch size of 24 \(batch size 6, gradient accumulation 4\), and KL regularization coefficientβ=0\.1\\beta=0\.1\. LoRA configuration was identical to the SFT stage \(r=32r=32,α=16\\alpha=16, no dropout\) with 4\-bit quantization enabled to accommodate the larger gradient memory footprint of contrastive training\. Hyperparameters are summarized in Table[7](https://arxiv.org/html/2607.11053#A7.T7)\.
Table 7:Preference alignment baseline hyperparameters\.
## Appendix HSelection of Models and Baselines for Primary Experiments
### H\.1Director and Builder Model Selection
One primary goal in the selection of models for the Director and Builder roles was to maximize diversity of the Director utterances while not sacrificing task progress \(requiring a strong Builder model\)\. Table[8](https://arxiv.org/html/2607.11053#A8.T8)shows a lexical and embedding diversity comparison of Director utterances generated by GPT\-4\.1\-mini and GPT\-4o\-mini in a sample game using thebase promptcondition, and by GPT\-4\.1\-mini using the base prompt with Directors roleplaying a fixed, randomly assigned set of archetypes \(prompts in Appendix[F](https://arxiv.org/html/2607.11053#A6)\)\.
Table 8:Comparative lexical and embedding diversity over one game using the same goal structure, comparing GPT\-4\.1\-mini and GPT\-4o\-mini in the Director role \(Builder: GPT\-4\.1\-mini\)\.- •GPT\-4\.1\-mini shows improved lexical diversity with a lower SelfBLEU\-4 \(31\.2 vs\. 35\.3\) and higher Distinct\-1/2 scores\.
- •GPT\-4\.1\-mini also significantly reduces repetition, with a Repetition\-2 score of 0\.157 compared to 0\.230\.
- •GPT\-4o\-mini maintains a higher Embedding Diversity \(0\.134 vs\. 0\.0995\), suggesting more varied semantic positioning despite the higher lexical repetition\.
These outcomes, all assessed using the base prompt from CRAFT before the experimentation with prompting or post\-training, motivated the choice of GPT\-4\.1\-mini as the Director model in training data generation\. GPT\-4\.1\-mini produces meaningfully more diverse utterances, even when both models use Director archetype prompts to increase lexical diversity\. The archetype prompting is a common technique with precedent in the literature\(Li et al\.,[2023](https://arxiv.org/html/2607.11053#bib.bib37); Chen et al\.,[2024](https://arxiv.org/html/2607.11053#bib.bib7); Nath et al\.,[2025b](https://arxiv.org/html/2607.11053#bib.bib47);[c](https://arxiv.org/html/2607.11053#bib.bib48)\)that meaningfully increases both lexical and semantic diversity of utterances over the majority of metrics\.
Fig\.[6](https://arxiv.org/html/2607.11053#A8.F6)and Table[9](https://arxiv.org/html/2607.11053#A8.T9)show sample experimental outcomes with GPT\-4o\-mini in the role of the Builder\.
Table 9:Overall Progress and Completion %Δ\\Deltaby Builder and Prompting Condition \(N=40N=40\)Figure 6:Prompting and tool\-calling outcomes results with GPT\-4o\-mini as Builder and GPT\-4\.1\-mini as Directors\. Shading reflects std\. error of the mean\.Many times the GPT\-4o\-mini Builder shows no progress from the start condition, or Markov Chain\-like behavior where the entire game consists of placing and removing the same block, resulting in no net change\. Task progress is uniformly worse than when GPT\-4\.1\-mini played the Builder, motivating the choice of GPT\-4\.1\-mini as the Builder in the main experiments, including post\-training experiments\.
### H\.2Alignment Baseline Selection
Our choice to focus on offline Director alignment using DPO and IPO, as opposed to online methods like PPO, arose from both principled and practical concerns\. PPO and other online methods require continued environmental interaction during training, including generating rollouts, receiving rewards, and iterative updating\. For 7\-8B parameter models, including all combinations of prompting and exploration conditions in the rollout and reward\-generation process, this would have substantially expanded the required compute and cost budget, rendering such a choice infeasible\.
Additionally, online RL requires design of a suitable reward signal\. Natural choices in this task would be the move correctness score or one of the objective progress metrics like Overall ProgressΔ\\Delta, but these signals are delayed and sparse—occurring only after the end of a turn that includes multiple utterances and potential Builder exploration\. The offline preference dataset𝒟\\mathcal\{D\}meanwhile, isolates signal at a contrastive level that is more granular than simply per\-turn\. Defining a dense per\-utterance reward that accurately reflects Gricean cooperativity remains the topic of future work\. Therefore, we assessed the problem of exploring the effects of Director alignment on task performance rather than training a globally optimal Director policy, and our results show that training on examples of more vs\. less actionable utterancescanshift Director behavior in the right direction, but it remains far from optimal\. Offline methods arenotstrictly sufficient in this task\. DPO’s gains are fragile and SFT or IPO frequently impede task performance\.
## Appendix ISensitivity to Prompt Phrasing
Tables[10](https://arxiv.org/html/2607.11053#A9.T10)and[11](https://arxiv.org/html/2607.11053#A9.T11), and Figs\.[7](https://arxiv.org/html/2607.11053#A9.F7)–[9](https://arxiv.org/html/2607.11053#A9.F9)show results from tests of a subsample of structures \(1 per start condition\) using the Literal and Pragmatic Builder prompt variants \(Appendix[F](https://arxiv.org/html/2607.11053#A6)\)\.
Table 10:Overall ProgressΔ\\Deltafrom Start \(Builder: GPT\-4\.1\-mini, Director: GPT\-4\.1\-mini, 1 structure per start condition\)Table 11:Completion %Δ\\Deltafrom Start \(Builder: GPT\-4\.1\-mini, Director: GPT\-4\.1\-mini, 1 structure per start condition\)Figure 7:Prompting and tool\-calling outcomes results with GPT\-4\.1\-mini as Builder and GPT\-4\.1\-mini as Directors, using Literal and Pragmatic Builder prompt variant 1 \(same as versions used in main experiments\)\.Figure 8:Prompting and tool\-calling outcomes results with GPT\-4\.1\-mini as Builder and GPT\-4\.1\-mini as Directors, using Literal and Pragmatic Builder prompt variant 2 \(Appendix[F](https://arxiv.org/html/2607.11053#A6)\)\.Figure 9:Prompting and tool\-calling outcomes results with GPT\-4\.1\-mini as Builder and GPT\-4\.1\-mini as Directors, using Literal and Pragmatic Builder prompt variant 3 \(Appendix[F](https://arxiv.org/html/2607.11053#A6)\)\.These results show that trends are preserved with prompt rephrasing and that overall variance across rephrasings remains low\.Similar Articles
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
This paper investigates asymmetries in LLMs' pragmatic competence by comparing their performance as judges of linguistic appropriateness versus as generators of pragmatically appropriate language. The study finds that many models perform substantially better as pragmatic listeners than as speakers, suggesting misalignment between evaluation and generation capabilities.
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
This paper formalizes deliberative collaboration for LLM agents under partial observability, introduces a scalable benchmark across multiple domains, and systematically evaluates representative LLMs, finding that complex tasks remain challenging while deliberation can enable error correction.
LLMs know when they are wrong. I made a fix relating to Anthropic's new "global workspace" paper [R]
The author presents a method to make LLMs verbalize calibrated confidence by using a linear probe on mid-layer states and a small trained bridge to confidence logits, requiring only 200 labeled examples and no weight modification. This is linked to Anthropic's global workspace paper explaining the know-say gap.
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators
Study shows LLM agents can model counterparty preferences in negotiation but fail to turn that knowledge into strategic bargaining to improve outcomes, limiting their effectiveness in multi-turn negotiations.