"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo
Summary
This paper evaluates how LLMs perform in the game Taboo under lexical constraints, showing trade-offs between compliance and communicative effectiveness, and finding that models are weaker guessers than humans.
View Cached Full Text
Cached at: 07/02/26, 05:37 AM
# "Don’t Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo
Source: [https://arxiv.org/html/2607.00601](https://arxiv.org/html/2607.00601)
\\copyrightclause
Copyright for this paper by its authors\. Use permitted under Creative Commons License Attribution 4\.0 International \(CC BY 4\.0\)\.
\\conference
CLiC\-it 2026: Twelfth Italian Conference on Computational Linguistics, September 14 — 16, 2026, Palermo, Italy
\[ orcid=0009\-0004\-5198\-6970, email=sara\.candussio@phd\.units\.it, \]\\fnmark\[1\]
\[ orcid=0009\-0007\-3489\-9631, email=f\.padovani@rug\.nl, \]\\fnmark\[1\]
\[orcid=0009\-0006\-0518\-6504, email=d\.scalena@rug\.nl, \]\\fnmark\[1\]
\[orcid=0000\-0001\-5289\-0971, email=m\.nissim@rug\.nl \]
\\cortext
\[1\]Corresponding author\.\\fntext\[1\]These authors contributed equally\.
Francesca PadovaniCenter for Language and Cognition \(CLCG\), University of Groningen, The NetherlandsDaniel ScalenaUniversity of Milano \- Bicocca, ItalyMalvina Nissim
\(2026\)
###### Abstract
The game of Taboo requires describing a target word without using a set of forbidden words, so that other players can guess it\. This deceptively simple task combines strict lexical constraints with the need for communicatively effective descriptions, making it a compelling playground for examining how LLMs navigate competing demands at inference time\. We evaluate two open\-weight models under conditions that intervene at progressively deeper levels of the generative process, from prompting to generation\-time constraints to internal representations manipulations\. We assess their outputs through forbidden word violation detection, LLM\-as\-a\-judge measuring the degree to which generated descriptions successfully evoke the target concept for both human and machine guessers, and examining whether the strategies models adopt under constraint align with those of human players\. Our results show that compliance with the rules of the game and communicative effectiveness trade off differently across conditions, and that models remain substantially weaker than humans as guessers, suggesting that lexical grounding under constraint is an open challenge for current language models111Code and data can be found at[https://github\.com/DanielSc4/LMtaboo/](https://github.com/DanielSc4/LMtaboo/)\.
###### keywords:
taboo\\sepprompting\\sepconstrained generation\\sepword\-guessing\\sepSAEs
## 1Introduction
The game of Taboo requires a player to describe a target word without using a set of forbidden words, so that another player can guess it\. Beyond its appeal as aparlourgame, Taboo instantiates a linguistically interesting challenge: the speaker must suppress some lexically and conceptually salient words while simultaneously producing a description that is informative enough to identify the target\. This combination of constraint satisfaction and communicative effectiveness makes it a particularly suitable setting for probing language models in a setting that goes beyond standard instruction\-following benchmarks\. In particular, it remains unclear whether descriptions generated under constraint conform to the demands of the task — namely whether they are sufficiently detailed, salient enough to evoke the target concept, and communicatively effective enough for the game to be successfully completed\.
In this paper we investigate how LLMs play Taboo in Italian, evaluating two open\-weight models under conditions that intervene at progressively deeper levels of the generative process: from prompting, to generation\-time constraints, to the manipulation of internal representations\. To ground our evaluation in an actual game\-play, we conduct a human study in which the roles are reversed in both directions: human evaluators read model\-generated descriptions and attempt to guess the target word, and vice versa\. This bidirectional design allows us to directly assess how well descriptions produced succeed in conveying the target concept, and whether the descriptive choices adopted by models and humans are mutually interpretable\.
Our findings reveal that the relationship between rule compliance and description quality varies substantially across conditions, and that models fall considerably short of human performance in the guessing role, pointing to lexical grounding under constraint as an open challenge for current language models\.
## 2Related works
##### Language games as NLP benchmarks
Word\-based games have been exploited as playgrounds for NLP, offering well\-defined rules and measurable objectives suitable for benchmarking\. Within the Italian NLP community, language games have attracted growing attention as probes for linguistic competence: the TV gameLa Ghigliottinahas been tackled both with dedicated NLP systems\[sangati\-etal\-2020\-challenge\]and used to benchmark LLMs\[manna\-etal\-2024\-riddle\], while studies on rebus and crossword solving reveal consistent limitations in tasks requiring compositional, phonological, and lateral reasoning\[sarti\-etal\-2024\-non,sarti\-etal\-2024\-eurekarebus,ciaccio\-etal\-2025\-crossword,ciaccio2026cruciverbit\]\. Beyond Italian, benchmarks such as NYT Connections\[loredo\-lopez\-etal\-2025\-nyt\]and Codenames\[stephenson2025codenamesbenchmarklargelanguage\]consistently show a substantial gap between LLM and human performance on tasks requiring deliberate reasoning and theory of mind\. Word games have also been studied as training environments:horst\-etal\-2025\-playpenshow that LLMs can learn from dialogue game self\-play feedback, with Taboo among the training games\. Closest to our work,cywinski2025elicitinglatentknowledgellmsstudy Taboo specifically in the context of LLM interpretability, using the game as a test for eliciting latent knowledge\. Our work differs in scope: rather than focusing on a single intervention, we systematically compare constraint enforcement strategies at multiple levels of the generative process and we furthermore ground the evaluation in an actual human game\-play\.
##### Constrained text generation
Comprehensive evaluations of constrained text generation across open\-source LLMs show that models exhibit systematic deficiencies even on relatively simple lexical constraints\. Prompt\-based approaches suffer from position bias and struggle with morphologically complex forms, and stricter format constraints have been demonstrated to cause performance degradation\[yao2024collie,tam\-etal\-2024\-speak\]\.ciaccio\-etal\-2024\-controllableevaluate several Italian LLMs on their ability to generate sentences adhering to morpho\-syntactic specifications, finding systematic deficiencies across models and constraint types\.calderaro\-etal\-2025\-oulibenchfurther show that even state\-of\-the\-art proprietary models consistently fall short when faced with fine\-grained linguistic restrictions in Italian and that stricter requirements tend to degrade output quality more severely\. These findings motivate our decision to go beyond prompting and to opt for different types of interventions\.
##### Mechanistic interpretability and SAEs
Beyond prompting, a more direct approach to output control is logit masking, i\.e\. forcing forbidden token probabilities to−∞\-\\inftyat every decoding step, as implemented in frameworks such as that proposed bywillard2023efficientguidedgenerationlargeand its predecessorshokamp\-liu\-2017\-lexically\. At a deeper level, concept manipulation via internal model representations offers an alternative approach\. Rather than suppressing surface tokens, it targets the concepts encoded in the model’s latent space using Sparse AutoEncoders \(SAEs\)\[cunningham2023sparseautoencodershighlyinterpretable\]\. They decompose dense LLM activations into sparse, monosemantic features\[templeton2024scaling\]that can be ablated at inference time without modifying model weights\. Furthermore, a feature’s interpretability does not guarantee its effectiveness as a steering target: features that cleanly represent a concept may have little causal influence on the model’s output – a gap whose limits we help characterise in the context of lexical avoidance\.
Figure 1:An example of a Taboo card\. The word to be guessed is on top\. The five words below are forbidden from appearing in the description\.
## 3Methodology
### 3\.1Data and models
The datasets consists of 194 italian Taboo cards manually transcribed from a physical version of the board game\. Each example contains a target word and a list of forbidden words that cannot be used in the description\.
We experiment with two open\-weight models\.google/gemma\-3\-4b\-it\[gemma\_2025\]is a lightweight instruction\-tuned model selected for its strong general performance and for the availability of pre\-trained SAEs222google/gemma\-scope\-2\-4b\-it\[lieberum\-etal\-2024\-gemma\]compatible with our interpretability\-based experimental condition\.openai/gpt\-oss\-20b\[openai2025gptoss120bgptoss20bmodel\]is a Mixture\-of\-Experts reasoning model post\-trained with chain\-of\-thought reinforcement learning, chosen to assess the role of explicit reasoning in constrained description generation, evaluated with and without reasoning\. Generating an effective Taboo description requires implicitly planning ahead: the model must produce a description that is semantically informative while simultaneously avoiding a set of lexically related forbidden words\. We therefore expect explicit reasoning to play a meaningful role in this task\.
### 3\.2Forbidding experiments
We investigate three degrees of forbidding, each intervening at a different level of the model’s behaviour\. In the first condition, models are*prompted*with the target word and the list of forbidden words, analogously to how human players are instructed before their turn\. The second condition moves the intervention downstream,*constraining the model during generation*so as to directly block forbidden tokens from being produced\. The third condition operates at the level of internal representations,*manipulating the model’s latent space*to suppress the encoded concepts corresponding to the forbidden words\.
#### 3\.2\.1Prompting
The prompt is constructed programmatically for each card, embedding the list of forbidden words and the target word directly into the user message\. To ensure clean and comparable outputs, we additionally provide output instructions specifying that the model should return only the description, with no surrounding text, and within a limit of 30 words\. The prompt explicitly instructs the model to avoid not only the forbidden words verbatim, but also any word derived from the same morphological root, in order to fully comply with the rules of the game\. The full prompt structure is provided in TableLABEL:tab:prompt\_exampleof Appendix[A](https://arxiv.org/html/2607.00601#A1)\. Notably, the prompt does not explicitly instruct the model to avoid using the target word itself in the description, an error that would trivially make the model wrong at the game\. We deemed it preferable to assess the extent to which this occurs empirically, and defer this analysis to Section[3\.3](https://arxiv.org/html/2607.00601#S3.SS3)\. In Appendix[A](https://arxiv.org/html/2607.00601#A1)we additionally report results obtained with a relaxed prompt variant that omits the morphological constraint, discussing the impact of this simplification on model behaviour\.
#### 3\.2\.2Generation\-time constraints
##### A prioristems censorship
Before generation begins, we compute a static set of banned token from the forbidden word list\. For each forbidden word, we extract all stems and ban any vocabulary token whose surface form is a prefix of any such stem or lemma, or vice versa\. The banned token are then masked by forcing their logits to−∞\-\\inftyat every decoding step, making them impossible to sample regardless of context\. A key design choice is the use of morphological normalisation rather than exact string matching: the censorship extends beyond the literal forbidden words to their inflected and derived forms\. As a consequence, the model is forced to use synonyms or periphrases to express the forbidden elements\. Complete freedom is left to the reasoning process\.
##### Inference\-time constrained generation
A limitation of thea\-prioriapproach is that it cannot condition the ban on generation context, occasionally suppressing tokens that would be appropriate in the description but happen to share a morphological root with a forbidden word\. A practical example of its aggressiveness can be found in Appendix[B](https://arxiv.org/html/2607.00601#A2)\. We address this by constraining the reasoning process dynamically\. Rather than banning tokens upfront, we allow the model to generate freely and intervene only when a complete word is detected; its stem is then compared against the forbidden ones\. If a match is found, generation backtracks to the first token of the offending word, which is added to a persistent position\-level mask, and generation resumes from that position with the triggering token suppressed\. Compared to thea prioriapproach, it operates on complete surface forms in place of subword fragments, avoiding spurious bans and higher computational overhead due to potential backtracing\.
#### 3\.2\.3Model Manipulations
We also explore a representation\-based approach that intervenes directly on the model’s residual stream at generation time\. The intervention operates in two stages: feature identification and steered generation\.
##### Feature identification
For each forbidden wordwiw\_\{i\}, we identify the SAE feature most strongly associated with it by running the model on the short natural prompt "La parola vietata èwiw\_\{i\}" \(translated: "The forbidden word iswiw\_\{i\}"\) and extracting the residual\-stream activations at layer 29 ofgoogle/gemma\-3\-4b\-it333Layer 29 corresponds to∼85%\\sim 85\\%network depth, where representations are richer in semantic content \(see Appendix[C](https://arxiv.org/html/2607.00601#A3)\)\.\. These activations are passed through a pre\-trained JumpReLU SAE fromgoogle/gemma\-scope\-2\-4b\-it\[lieberum\-etal\-2024\-gemma\], using a residual\-post hook at width 65k\. The feature with the highest activation across the token span corresponding towiw\_\{i\}is selected as its representative\. This process is applied to all forbidden words in a card, yielding a set of feature indices\{f1,…,fk\}\\\{f\_\{1\},\\ldots,f\_\{k\}\\\}\.
##### Steered generation
At generation time, we register a forward hook on the same residual\-stream layer\. At each decoding step, we compute the mean decoder vectord=1k∑j𝐰fidecd=\\frac\{1\}\{k\}\\sum\_\{j\}\\mathbf\{w\}^\{\\text\{dec\}\}\_\{f\_\{i\}\}across the selected features and apply an activation\-norm\-scaled intervention:
𝐡′=𝐡\+α⋅‖𝐡‖2⋅𝐝\\mathbf\{h\}^\{\\prime\}=\\mathbf\{h\}\+\\alpha\\cdot\\\|\\mathbf\{h\}\\\|\_\{2\}\\cdot\\mathbf\{d\}
steering the residual stream*away*from the direction associated with the forbidden concepts\. The model is otherwise prompted without any mention of forbidden words, receiving only the target word and the standard output instructions\. The entire process operates at inference time without modifying model weights\.
### 3\.3Evaluation
##### Exact amount of violations
We evaluate constraint compliance in the description generation through a number of metrics\.Accuracymeasures the percentage of valid descriptions that fully comply with the rules of the game, i\.e\. outputs that contain neither verbatim forbidden words nor any morphologically related form\.Share of Exact Violationsreports the percentage of descriptions in which the model uses one of the forbidden words verbatim\.Share of Morphological Violationsreports the percentage of descriptions in which the model produces a word sharing the same stem or lemma as a forbidden word444Detected using theit\_core\_news\_smspaCy model\[honnibal2020spacy\]and thenltkItalian Snowball stemmer\[bird\-loper\-2004\-nltk\]\.\.Share of no<eot\>applies exclusively toopenai/gpt\-oss\-20bin reasoning mode and it measures the proportion of cases in which the model exhausts its 2048\-token reasoning budget without producing a final description\. Finally, as anticipated in Section[3\.2\.1](https://arxiv.org/html/2607.00601#S3.SS2.SSS1), the prompt does not explicitly instruct the model to avoid using the target word itself in the description\.Share of Target Leaked Wordsreports the percentage of outputs in which the model nevertheless commits this error, inadvertently revealing the answer it was supposed to help guess\.
##### LLM\-as\-a\-judge
To assess the communicative effectiveness of the generated descriptions, we employ an LLM\-as\-a\-judge evaluation\[bavaresco\-etal\-2025\-llms\]as a complement to human evaluation\. We use bothgoogle/gemma\-3\-4b\-itandopenai/gpt\-oss\-20bas automatic judges, evaluating descriptions generated by each model with both judges — that is, each model judges its own outputs as well as those of the other\. Each judge is presented with the generated description alone, with no access to the forbidden words or the target word, and asked to guess the word being described\. Instead of generating a free\-form answer, we extract the judge’s next\-token probability distribution immediately after reading the description and measure how much probability mass is assigned to the first token of the target word\. This allows for a fine\-grained, continuous assessment of how strongly the description evokes the target concept, rather than a binary correct/incorrect judgment\.
We derive two metrics from this distribution:Pass@kkmeasures whether the target word’s first token falls within the top\-kkmost likely tokens,k∈\{1,2,3,5,10\}k\\in\\\{1,2,3,5,10\\\}; theRaw token likelihoodreports the actual probability assigned to the target\.
##### Human Evaluation
We conduct a human evaluation study involving 8 annotators, structured into two phases on disjoint sets of cards\.
In theguessing phase\(80 cards\), annotators are presented with model\-generated descriptions and asked to guess the target word as if playing an actual game of Taboo, albeit without the interactive component\. Each annotator sees 20 cards, with system conditions distributed randomly across annotators to ensure balanced coverage of all evaluated systems, and each card evaluated under two different conditions by two different annotators\.
In thedescription phase\(100 cards, disjoint from the guessing phase\), each annotator receives 20 cards\. For 5 of these, they produce a spontaneous, oral\-style description; for the other 5, they write a deliberate, carefully worded description\. This design mirrors the*without thinking*and*with thinking*regimes observed in model experiments respectively\. These 10 descriptions of unique items produced by each annotator are then presented to the other annotators for guessing, with the oral description setting being the closest to the actual Taboo game\. The last 10 cards are held constant across all annotators: every participant describes the same items, enabling future analysis of inter\-annotator variability in description strategies, which we leave for future work\.
## 4Results
### 4\.1Models are quite good at following given constraints
Figure 2:Compliance breakdown by method and model\. Bars are partitioned into mutually exclusive categories summing to 100%; outputs with multiple error types are grouped into Mixed\. The Baseline frequently violates taboo rules, while prompting substantially reduces errors\. Constrained decoding achieves near\-perfect compliance by construction, with residual failures almost exclusively due to target\-word leaks\. Forgpt\-oss\-20b, enabling reasoning further improves compliance\.Figure[2](https://arxiv.org/html/2607.00601#S4.F2)provides an overview of model performance across forbidding experiments, compared against abaselinein which models generate target word descriptions without any forbidden word constraint\.555Prompt in Appendix[A](https://arxiv.org/html/2607.00601#A1); full results in Tables[4](https://arxiv.org/html/2607.00601#A5.T4)–[6](https://arxiv.org/html/2607.00601#A5.T6)in Appendix[E](https://arxiv.org/html/2607.00601#A5)\.Baseline accuracy varies considerably across models, ranging from 24\.7% foropenai/gpt\-oss\-20bwithout reasoning to 34\.0% forgoogle/gemma\-3\-4b\-it, reflecting how readily each model would spontaneously use the forbidden words for descriptive purposes, even without knowing they are banned\.
Under the prompting constraint, accuracy rises markedly for bothgoogle/gemma\-3\-4b\-it\(91\.2%\) andopenai/gpt\-oss\-20bwith reasoning \(97\.4%\)\. Forgoogle/gemma\-3\-4b\-it, a small proportion of exact and morphological violations persist;openai/gpt\-oss\-20bwith reasoning exhibits only morphological violations, suggesting that chain\-of\-thought is effective at suppressing verbatim forbidden words but less reliable on morphologically related forms\. When reasoning is disabled, accuracy drops to 72\.2%, with a higher share of both violation types, confirming that the reasoning process plays an active role in enforcing lexical constraints\.
In the constrained generation setting, all three models achieve high accuracy, withgoogle/gemma\-3\-4b\-itreaching near\-perfect compliance under both a priori \(98\.5%\) and online \(96\.4%\) settings\.openai/gpt\-oss\-20bwith reasoning performs similarly under a priori constraints \(92\.3%\), though accuracy drops in the online condition \(87\.1%\), accompanied by a higher rate of target word leaks\. Without reasoning, the model’s accuracy reaches 82\.5% and 82\.0% respectively, with a markedly higher proportion of leaked target words in both cases\.
Finally, the SAE\-based approach \(Section[2](https://arxiv.org/html/2607.00601#S2)\), evaluated ongoogle/gemma\-3\-4b\-itonly, reaches 55\.7%, well below both prompting \(91\.2%\) and constrained generation \(98\.5% and 96\.4%\), confirming that representation\-based interventions alone cannot match explicit lexical guidance, though the result remains well above the unconstrained baseline \(34\.0%\)\.
### 4\.2The type of constraining strategy shapes description informativeness
Figure 3:Compliance vs\. pass@10 across models, methods, and evaluators\. All three proposed methods \(Prompt, Constrained, SAE\) substantially increase compliance over the Baseline, confirming their effectiveness at enforcing the Taboo constraint\. Pass@10 remains broadly stable across methods, suggesting compliance gains do not come at the cost of description quality\. Gemma, in the guesser role, consistently achieves higher pass@10 than GPT both in\- and out\-of\-distribution\. Finally, enabling reasoning in gpt\-oss\-20b does not yield consistent improvements in its description effectiveness\.Having established that constraining strategy shapes compliance, we now ask whether it also affects how informative the resulting descriptions are as measured by the ability of a judge model to identify the target word\.
The two judge models agree on the broad ordering of conditions but differ substantially in sensitivity:openai/gpt\-oss\-20bassigns near\-zero probability to the target word across almost all conditions, with pass@10 never exceeding 10%;google/gemma\-3\-4b\-itis consistently more generous, reaching up to 24\.7% on the same descriptions \(Figure[3](https://arxiv.org/html/2607.00601#S4.F3)\)\.
Despite this gap, both judges agree that the constraining strategy matters\. Descriptions produced under inference\-time constrained generation are in most conditions easier to guess than prompt\-based ones: forgoogle/gemma\-3\-4b\-itas clue\-giver, constrained online generation yields pass@10 of 17\.0% and 8\.8% against 11\.3% and 3\.6% for prompting under thegoogle/gemma\-3\-4b\-itandopenai/gpt\-oss\-20bjudge respectively\. A likely explanation is that forcing the model away from forbidden tokens at decoding time pushes it toward less obvious but more semantically precise paraphrases, whereas an instructed model tends toward overly cautious, informationally sparse descriptions\.
SAE\-based steering presents the most striking dissociation: compliance is low \(56\.7% strict accuracy\) yet communicative effectiveness is competitive, with pass@10 reaching 21\.1% under thegoogle/gemma\-3\-4b\-itjudge, also exceeding constrained generation for the same clue\-giver\. This suggests that blocking surface tokens at decoding time imposes a descriptive cost that representation\-level intervention avoids: the model loses access to forbidden words but retains the underlying conceptual associations needed to produce an informative description\.
### 4\.3Models are weak guessers
Figure 4:Pass@kkguessing accuracy on 10 shared Taboo cards, for all clue\-givers evaluated byopenai/gpt\-oss\-20b\(left\) andgoogle/gemma\-3\-4b\-it\(right\)\. Each cell reports the fraction of cards for which the judge ranked the target word within its topkkguesses\.We generated descriptions for 10 shared Taboo cards usinggoogle/gemma\-3\-4b\-it,openai/gpt\-oss\-20b\(with and without reasoning\), Claude Sonnet\-4\.6 \(max effort\)\[anthropic2026claudesonnet46\], Gemini\-3\.1 Pro\[googledeepming2026gemini31pro\]666Proprietary models are used in their chat version\., along with human\-generated ones\. These are passed toopenai/gpt\-oss\-20bandgoogle/gemma\-3\-4b\-itin order to be guessed and the results are reported in Figure[4](https://arxiv.org/html/2607.00601#S4.F4)\.openai/gpt\-oss\-20bassigns near\-zero probability to the correct target across virtually all clue givers, reaching at most 1\.2% at pass@10 on human descriptions\.google/gemma\-3\-4b\-itis substantially more capable, yet still requires between 5 and 10 guesses to approach a decent accuracy: on human\-written descriptions, its pass@1 is 0\.0% while pass@10 reaches 30\.0%\. The same asymmetry holds when models guess from model\-written descriptions, with Claude Sonnet\-4\.6 and Gemini\-3\.1\-Pro as clue givers yielding comparable pass@10 of 25\.0% and 30\.0% respectively undergoogle/gemma\-3\-4b\-itjudge, whileopenai/gpt\-oss\-20bremains at or below 20\.0% across all clue givers\.
### 4\.4Humans versus models; models versus humans
##### Humans as guessers\.
We asked human annotators to read model\-generated descriptions and guess the target word\. Overall, annotators correctly identified the target in 30\.77% of cases \(weighted average across annotators and conditions\), with performance varying considerably by constraining strategy: constrained generation yields the highest human accuracy \(38\.10%\), followed by the SAE\-based condition \(31\.03%\), with prompting producing the lowest \(23\.44%\)\. This ordering mirrors the pattern observed in Section[4\.3](https://arxiv.org/html/2607.00601#S4.SS3)for model judges, suggesting that the informativeness advantage of constrained generation generalises beyond automatic evaluation\.
Beyond aggregate accuracy, qualitative inspection reveals recurring failure patterns in model\-generated descriptions, including referential failures, category mismatches, and code\-switching under the SAE condition, which we discuss in Appendix[D](https://arxiv.org/html/2607.00601#A4)\.
##### Humans as descriptors\.
We also collected human\-generated descriptions across three conditions that differ in prior exposure and time constraints \(seedescription phasein Section[3\.3](https://arxiv.org/html/2607.00601#S3.SS3)\):spontaneous\(no preparation, 5 cards\),deliberate\(time to reflect before writing, 5 cards\), andshared\(all annotators describe privately the same ten words, enabling direct comparison with models\)\. For thespontaneousanddeliberateconditions, we assessed how accurately other annotators could guess the target word from the description:deliberatedescriptions achieve 85% accuracy, whilespontaneousdescriptions yield a higher rate of 95%\. The gap between these two reflects the interactive nature of oral description, where players could refine in real time, versus the deliberate setting, which more closely mirrors the constraints imposed on models\.
If we ask the models to guess the human\-described words, we can spot a clear ordering: shared\>\>deliberate\>\>spontaneous at all values ofkk\(Table[1](https://arxiv.org/html/2607.00601#S4.T1)\)\. At pass@10, shared descriptions yield 15\.6%, deliberate 13\.8%, and spontaneous only 5\.1%\. As observed throughout,openai/gpt\-oss\-20balmost never guesses correctly \(pass@10≤\\leq2\.5%\), whereasgoogle/gemma\-3\-4b\-itreaches up to 30\.0% on shared descriptions, yet still requiring between 5 and 10 guesses to rival what a human achieves at pass@1\.
This contrast is revealing: humans benefit from spontaneous descriptions that, though fragmented, are iteratively refined towards the referent in real time\. When presented to models as static text \(without the interactive context that makes them effective for human guessers\) such descriptions yield substantially lower pass@kkrates, while models are better suited to process the structured, fixed outputs of the deliberate and shared conditions\.
Table 1:Human clue\-giver evaluation\. Accuracy computed on strict criterion \(no exact \+ no morphological violations\)\.ConditionJudgepass@kkAccuracy@1@2@3@5@10loosestrictSpontaneous\(n=39n=39\)gpt\-oss\-20b0\.0%0\.0%0\.0%0\.0%0\.0%97\.4%82\.1%gemma\-3\-4b\-it0\.0%2\.6%5\.1%7\.7%10\.3%Deliberate\(n=40n=40\)gpt\-oss\-20b0\.0%0\.0%0\.0%0\.0%2\.5%100\.0%97\.5%gemma\-3\-4b\-it0\.0%5\.0%7\.5%10\.0%25\.0%Shared\(n=80n=80\)gpt\-oss\-20b0\.0%0\.0%0\.0%1\.2%1\.2%100\.0%91\.2%gemma\-3\-4b\-it0\.0%7\.5%12\.5%18\.8%30\.0%
## 5Discussion
Our results reveal a consistent tension between constraint adherence and descriptive informativeness that cuts across all conditions\. The three families of constraining strategies we examine operate at qualitatively different levels \(instruction, decoding, and internal representation\) and this difference determines not only how reliably the model avoids forbidden words, but also how it describes the target item\. Prompting achieves high compliance at the cost of informationally sparse output; constrained generation forces the model toward more semantically precise paraphrases, at the cost of occasionally leaking the target word; SAE\-based steering sits at the opposite extreme, preserving conceptual associations at the cost of surface compliance\. This trade\-off suggests that lexical avoidance and communicative effectiveness are not independently controllable, and that the mechanism through which a constraint is imposed shapes the output in ways that go beyond mere compliance\.
A second thread concerns the asymmetry between humans and models as players\. On the descriptor side, reasoning does help models:openai/gpt\-oss\-20bwith chain\-of\-thought reaches 97\.4% accuracy versus 72\.2% without, a substantial gain\. For humans, the picture is more nuanced: spontaneous oral descriptions, despite being unplanned, yield higher guessing accuracy \(95\.0%\) than deliberate written ones \(85\.0%\), likely because the interactive oral setting allows real\-time refinement\. This suggests that planning confers different benefits depending on the medium and the player: for models it primarily enforces compliance, whereas for humans it is the interactivity of the setting, rather than deliberationper se, that drives performance\. On the guesser side, the asymmetry is starker: both judge models struggle to identify target words, withgoogle/gemma\-3\-4b\-itrequiring between 5 and 10 guesses to approach the accuracy a human achieves at pass@1\. This gap suggests that models do not share the same network of salient lexical associations that makes Taboo intuitive for humans, and that guessing from indirect descriptions remains a genuinely hard task for current language models\. This gap is also reminiscent of the distinction between fast, associative System 1 processing and deliberate System 2 reasoning\[kahneman2011thinking\]: humans navigate Taboo by rapidly activating and suppressing salient lexical associates, a mode of processing that current language models do not appear to replicate, at least not in the guessing direction\.
Our study has however several limitations\. The benchmark is currently Italian\-only, and it is unclear how results generalise to languages with different morphological profiles or lexical structures\. The shared evaluation set covers only 10 words, limiting the statistical power of cross\-model comparisons in Section[4\.3](https://arxiv.org/html/2607.00601#S4.SS3)\. SAE\-based steering was evaluated ongoogle/gemma\-3\-4b\-itonly, and extending it to other architectures remains open\. The two judge models differ substantially in sensitivity, raising questions about which better reflects human guessing behaviour, a question our human evaluation begins to address but does not fully resolve\.
## 6Conclusions and Future Work
We presented a framework that operationalises lexical constraint following as a communicative game, and used it to evaluate three families of constraining strategies across multiple models and human annotators\. Our results show that compliance and communicative effectiveness are not independently controllable: the mechanism through which a constraint is imposed \(instruction, decoding, or internal representation\) shapes the output in ways that go beyond mere rule adherence, with constrained generation and SAE\-based steering producing qualitatively different trade\-offs\. Beyond the specific findings, our work supports the broader argument that games constitute a natural and underexplored testbed for language model capabilities: they impose well\-defined rules, require communicative grounding, and admit quantitative evaluation while remaining ecologically valid\. LMtaboo in particular is inherently interactive, yet we have evaluated models only in a single\-turn, single\-agent setting\. Extending the benchmark to multi\-turn and multi\-agent configurations \(where models iteratively refine descriptions based on guesser feedback, or negotiate meaning across turns\) would more faithfully reflect the dynamics of the original game and probe capabilities, such as pragmatic adaptation and collaborative grounding, that current evaluations leave largely untested\.
###### Acknowledgements\.
The authors sincerely thank Valentina, Frida, Daniel, and Arianna for their invaluable contribution to the annotation process\. The work of Daniel Scalena has been partially funded by MUR under the grant ReGAInS,Dipartimenti di Eccellenza 2023\-2027of the Department of Informatics, Systems and Communication at the University of Milano\-Bicocca\. The work of Sara Candussio has been funded by Fondo Sociale Europeo Plus of Regione Autonoma Friuli Venezia Giulia\. We also thank the Center for Information Technology of the University of Groningen for providing access to the Hábrók high\-performance computing cluster used for part of the experiments\.
## Declaration on Generative AI
During the preparation of this work, the authors usedClaude\(Anthropic\) in order toimprove writing style,grammar and spelling check,paraphrase and reword,drafting content\(including experimental code\)\. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the publication’s content\.
## References
Impact of Prompt Relaxation In the main experiments we adopt a prompt that explicitly forbids the use of words sharing the same morphological root as any of the forbidden words, as this most faithfully reflects the rules of the Taboo game\. To assess the sensitivity of our results to this design choice, we also tested a relaxed variant that omits this morphological constraint, instructing the model only to avoid the forbidden words verbatim\. For Gemma, this relaxed variant yields a marginal decrease in*Loose Accuracy*\(96\.4% vs\. 97\.4%\) and a marginal improvement in*Strict Accuracy*\(93\.8% vs\. 92\.8%\), with the number of target word leaks remaining identical \(3\)\. The difference is negligible and can be attributed to noise rather than any systematic effect\. For GPT\-OSS 20B with increased repetition penalty, removing the morphological constraint reduces description failures from 1 to 0, but slightly decreases*Strict Accuracy*from 99\.5% to 99\.0%, while*Loose Accuracy*remains perfect \(100\.0%\)\. Target word leaks increase marginally from 2 to 3\.
Overall, softening the morphological constraint does not lead to substantial changes in model behaviour across either configuration, suggesting that the models’ compliance strategies are largely robust to this variation in prompt formulation\.
## Appendix AAppendix A: prompt details, generation parameters and repetition penalty
This appendix reports the prompt templates used in our experiments, together with the generation hyperparameters adopted for each model and condition\.
### A\.1Prompt Templates
Two prompt variants were used\. Thebaselineprompt asks the model to describe the target word freely, with no mention of forbidden words\. Thepromptedcondition additionally instructs the model to avoid the forbidden words and their morphological relatives\.
User Message — BaselineDescrivi la seguente parola in modo che qualcuno possa indovinarla\.Parola da descrivere:\[target word\]Scrivi in output solamente la descrizione senza nessun testo di contorno\. Considera che hai un limite di 30 parole, quindi cerca di dare una descrizione quanto più breve possibile\.Descrizione:Describe the following word so that someone can guess it\.Word to describe:\[target word\]Output only the description with no surrounding text\. You have a limit of 30 words, so try to give as brief a description as possible\.Description:User Message — Prompted conditionDescrivi la seguente parola in modo che qualcuno possa indovinarla, ma NON usare nessuna di queste parole o nessuna parola con la stessa radice:\[fw1, fw2, …\]\.Parola da descrivere:\[target word\]Scrivi in output solamente la descrizione senza nessun testo di contorno\. Considera che hai un limite di 30 parole, quindi cerca di dare una descrizione quanto più breve possibile\.Descrizione:Describe the following word so that someone can guess it, but do NOT use any of these words or any word with the same root:\[fw1, fw2, …\]\.Word to describe:\[target word\]Output only the description with no surrounding text\. You have a limit of 30 words, so try to give as brief a description as possible\.Description:Table 2:Prompt templates for thebaseline\(top\) andprompted\(bottom\) conditions\. English translation in italics\.
### A\.2Generation parameters and repetition penalty
All models were run in greedy decoding mode \(do\_sample=False\)\. Table[3](https://arxiv.org/html/2607.00601#A1.T3)summarises the generation parameters used per model and condition\.
Modelmax\_new\_tokensrep\. penaltygemma\-3\-4b\-it801\.0gpt\-oss\-20b\(reasoning\)20481\.2gpt\-oss\-20b\(no reasoning\)801\.0Table 3:Generation hyperparameters per model and condition\.Foropenai/gpt\-oss\-20bin the reasoning condition, the assistant turn was left open so the model could produce a chain\-of\-thought trace before committing to a final answer\. In the no\-reasoning condition, the assistant turn was instead prefilled with a channel token that routes the model directly to its final\-answer channel, bypassing the reasoning trace\. A repetition penalty of1\.21\.2was applied exclusively in the reasoning condition to mitigate degenerate repetition loops that can arise during long generations; no penalty was applied otherwise\.
## Appendix BAppendix B: a priori constrained generation vs online
Both thea prioriand the inference\-time \(online\) approaches enforce the taboo constraint by operating on stems and lemmas of the forbidden words, using the same morphological pipeline \(spaCyit\_core\_news\_sm\+nltkItalian Stemmer\)\. They differ fundamentally in*when*and*how broadly*the constraint is applied\.
##### A priori censorship
Before generation begins, the method computes a static set of banned token IDs: for each forbidden word, all stems and lemmas are extracted, and every vocabulary token whose surface form shares a prefix with any of them \(in either direction\) is masked by forcing its logit to−∞\-\\inftyfor the entire generation\. This static mask is context\-free: a token is banned regardless of whether its occurrence in the output would actually constitute a violation\.
The consequences can be severe\. Consider the target wordaccipicchia\(an Italian exclamation of surprise\), whose forbidden list includesper la miseria\(a common interjection meaning "what a misery"\)\. Because the phrase contains the function wordsper\("for"\) andla\("the"\), the stemmer extracts them as forbidden stems, causing21712171vocabulary tokens to be masked — including common prepositions, determiners, and unrelated content words that merely share a surface prefix withperorla\(e\.g\. tokens such as\_permitting,\_lancé,larghezza\)\. Across the full dataset of 194 items, thea prioriapproach bans a cumulative total of90 95790\\,957tokens \(on average469469per item, out of a vocabulary of262 145262\\,145\)\.
##### Inference\-time constrained generation
The online approach bans*no tokens*before generation begins\. Instead, it generates freely token by token, accumulates tokens into complete words, and only when a word boundary is detected checks whether the completed word matches any forbidden stem or lemma\. If a match is found, the method backtracks to the first token of the offending word and re\-samples, blocking only that specific token at that position\. Foraccipicchia, the online approach evaluates each generated word against the forbidden forms at runtime, and only intervenes if a word such asmeravigliaorcaspitais actually produced — leavingper,la, and all other innocent tokens freely available\. Across the same 194 items, the model produces on average19\.319\.3words per output, all of which are checked at runtime — yet zero violations are recorded, confirming that the online constraint intervenes surgically only when needed rather than preemptively suppressing large portions of the vocabulary\.
## Appendix CAppendix C: SAE implementation details
##### SAE configuration
We use JumpReLU SAE fromgoogle/gemma\-scope\-2\-4b\-it\[lieberum\-etal\-2024\-gemma\], applied to the residual stream ofgoogle/gemma\-3\-4b\-it, which has 34 transformers layers in its text module\. The Gemma Scope 2 checkpoints provides SAEs at four depths: layer 9, 17, 22 and 29 \(approximately25%25\\%,50%50\\%,65%65\\%and85%85\\%of network depth\)\. We select layer 29, as later layers are known to carry richer semantic representation compared to early layers\[skean2025layer\], making them better suited for identifying concept\-level features associated with specific words\. The SAE uses a dictionary width of 65k features and medium L0 sparsity, with weights loaded inbfloat16and kept frozen throughout\. Feature identification runs once per forbidden word per card and is cached to avoid redundant forward passes across items sharing the same forbidden word\. The feature prompt template"La parola vietata è:\{word\}"is tokenized with special tokens; the token span ofwiw\_\{i\}is located by matching the tokenized word \(with and without a leading space\) against the full input sequence, falling back to the final token if no march is found\.
## Appendix DAppendix E: Failure Patterns in Model\-Generated Descriptions
Beyond aggregate accuracy, qualitative inspection of the guessing phase reveals several recurring failure patterns in model\-generated descriptions\.
##### Code\-switching
While descriptions generated under the prompting condition remain consistently in Italian, constrained generation introduces a small but noticeable rate of language shifts \(6\.35%\), which becomes far more pronounced under the SAE condition \(44\.83%\), with English, French, and Spanish being the most frequent alternates\. This suggests that intervening on internal representations disrupts the model’s linguistic behaviour in ways that explicit lexical guidance does not\.
##### Referential failure
A significant proportion of descriptions are nonsensical or semantically detached from the target word, making guessing effectively impossible\. Out of 160 ground\-truth entries evaluated, 25 \(15\.62%\) were flagged as containing an inappropriate description\. A striking example is the wordsemolino\(semolina\), described under two conditions by the same model as follows:
- •gemma\_sae:“Piccola protuberanza dura e rotonda, spesso presente sulla pelle, solitamente causata da irritazione o crescita di foruncoli\.”\(“A small, hard, round bump, often found on the skin, usually caused by irritation or the growth of pimples\.”\)
- •gemma\_constrained:“Piccola goccia di liquido, spesso salata, che si stacca da una fonte, come una lacrima o una pioggia leggera\.”\(“A small drop of liquid, often salty, that detaches from a source, such as a tear or light rain\.”\)
Neither description bears any semantic relation to the target word: the generated text is fluent and well\-formed, but lacks any associative link to the intended referent\.
##### Category mismatch
In several cases, the model describes a semantically related but categorically distinct concept — for instance, foregrounding the activity rather than the agent, or the domain rather than the practitioner\. An example is the wordfalegname\(carpenter\), described bygpt\_think\_constrainedas:
- •gpt\_think\_constrained:“Arte che trasforma il legno: taglia, intaglia, costruisce mobili e strutture con seghe, scalpelli e martelli\.”\(“The art of transforming wood: cutting, carving, building furniture and structures with saws, chisels and hammers\.”\)
By framing the description around the craft \(arte\) rather than the practitioner, the model effectively steers the guesser away from the intended referent\.
## Appendix ETables
Note on evaluation metrics\.The metrics reported in the Eval section of each table are independent counters and do not sum to 100%\. A single output may simultaneously trigger multiple error categories \(e\.g\., a target leak and a filter failure\), so the categories are not mutually exclusive\.
PromptingConstrainedManip\.BaselineSectionMetrica priorionlineSAEsno forb\.EvalTrue accuracy91\.2098\.5096\.4055\.7034\.00% exact viol\.2\.600\.000\.0031\.4052\.10% only morphol\. viol\.4\.600\.000\.0011\.9012\.30% no<eot\>0\.000\.000\.000\.000\.00% target leaked1\.501\.503\.601\.001\.50gpt\-oss judgepass@10\.000\.000\.000\.000\.00pass@20\.000\.500\.500\.000\.50pass@31\.001\.502\.601\.003\.60pass@51\.503\.605\.203\.105\.70pass@103\.607\.708\.805\.709\.30min4e\-64e\-64e\-61e\-64e\-6p251\.1e\-53\.0e\-52\.6e\-51\.0e\-52\.7e\-5median1\.9e\-56\.9e\-56\.2e\-51\.9e\-51\.2e\-4avg1\.9e\-33\.4e\-34\.5e\-31\.8e\-35\.0e\-3p751\.3e\-43\.5e\-46\.1e\-41\.6e\-41\.3e\-3max0\.11430\.15640\.17650\.05760\.1437gemma judgepass@10\.000\.000\.000\.000\.00pass@21\.502\.104\.604\.103\.10pass@34\.104\.607\.707\.207\.20pass@56\.209\.809\.8013\.9011\.90pass@1011\.3017\.5017\.0021\.1024\.70min00000p2500000median00000avg6\.9e\-401\.4e\-44\.5e\-41\.9e\-4p754e\-63e\-62e\-61\.2e\-52e\-6max0\.12723\.2e\-31\.8e\-24\.7e\-21\.8e\-2Table 4:Evaluation results –gemma\-3\-4b\-itPromptingConstrainedBaselineSectionMetrica priorionlineno forb\.EvalTrue accuracy72\.2082\.5082\.0024\.70% exact viol\.16\.000\.000\.0055\.20% only morphol\. viol\.8\.700\.000\.5014\.90% no<eot\>0\.000\.000\.000\.00% target leaked3\.1017\.5017\.5012\.90gpt\-oss judgepass@10\.000\.000\.000\.00pass@20\.000\.000\.000\.00pass@30\.500\.500\.501\.00pass@51\.000\.500\.502\.60pass@103\.602\.102\.106\.20min2e\-62e\-61e\-61e\-6p251\.1e\-51\.1e\-51\.1e\-51\.2e\-5median2\.5e\-51\.6e\-51\.5e\-53\.2e\-5avg8\.6e\-48\.7e\-59\.0e\-51\.8e\-3p751\.4e\-43\.0e\-52\.6e\-53\.9e\-4max0\.06353\.3e\-35\.6e\-30\.1416gemma judgepass@10\.000\.000\.000\.00pass@22\.106\.706\.703\.10pass@34\.1011\.3013\.907\.70pass@510\.3017\.5019\.1011\.30pass@1018\.6024\.2024\.7017\.50min0000p250000median0000avg4\.4e\-42\.8e\-31\.1e\-31\.1e\-3p754e\-63\.7e\-54\.3e\-54e\-6max0\.03660\.24560\.04980\.1358Table 5:Evaluation results –gpt\-oss\-20b \(no reasoning\)PromptingConstrainedBaselineSectionMetrica priorionlineno forb\.EvalTrue accuracy97\.4092\.3087\.1034\.00% exact viol\.1\.000\.001\.0045\.90% only morphol\. viol\.0\.500\.000\.5017\.50% no<eot\>1\.002\.102\.100\.00% target leaked2\.107\.7012\.404\.10gpt\-oss judgepass@10\.000\.500\.000\.00pass@20\.001\.000\.500\.50pass@30\.001\.500\.500\.50pass@50\.502\.101\.004\.60pass@101\.505\.704\.606\.70min1e\-62e\-62e\-61e\-6p251\.0e\-51\.1e\-51\.0e\-51\.4e\-5median1\.8e\-52\.7e\-52\.2e\-53\.8e\-5avg3\.4e\-43\.5e\-31\.9e\-32\.3e\-3p756\.1e\-51\.6e\-41\.9e\-42\.5e\-4max0\.02690\.27790\.19160\.1281gemma judgepass@10\.500\.000\.000\.00pass@24\.605\.204\.104\.10pass@36\.707\.707\.709\.80pass@511\.3013\.9010\.3017\.50pass@1016\.5019\.6020\.6021\.60min0000p250000median0000avg2\.2e\-37\.0e\-43\.6e\-44\.7e\-4p755e\-68e\-68e\-69e\-6max0\.28700\.07460\.02280\.0309Table 6:Evaluation results –gpt\-oss\-20bSimilar Articles
Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs
This paper provides causal evidence that large language models acquire negative linguistic knowledge (what not to say) through statistical preemption, a mechanism from Construction Grammar, by showing that manipulating competing-form frequencies via fine-tuning shifts preemption behavior in predicted directions.
Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry
This paper investigates whether large language models exhibit pragmatic cooperativity in multi-party collaborative tasks under conditions of epistemic asymmetry, formalizing Grice's cooperative principle and evaluating LLMs as both speakers and listeners. Results show that while LLMs display some pragmatic capabilities, they struggle with incomplete information and fail to recognize certain violations of Gricean maxims.
Exploring the Capability Boundaries of LLMs in Mastering Chinese Chouxiang Language
This paper introduces Mouse, a specialized benchmark for evaluating LLMs on Chinese Chouxiang Language tasks across six NLP domains, revealing that current state-of-the-art models have significant limitations with this subcultural internet language despite performing well on contextual understanding tasks.
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
This paper presents a multi-dimensional analysis of human-like behaviors in LLMs, examining prevalence, effects, and controllability across 21,000 conversations from four models, finding that behaviors vary by model and user factors, with implications for responsible design.
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
This paper evaluates nine LLMs on their ability to accurately communicate probabilistic predictions in natural language, finding that models are consistent but miscalibrated, particularly for uncertainty tasks.