Tag
This paper evaluates LLMs' ability to recognize unspoken beliefs (implicatures) and their updates through implicature cancellation, introducing the expert-annotated ImplicatureX dataset. Results show LLMs lag behind humans, especially in natural scenarios.
Proposes SAGE, a neuro-symbolic framework combining language models with cognitive models for pragmatic reasoning, demonstrated on three case studies including referential expression generation and implicatures.
The article proposes a 'Genie coefficient' to measure how well AI agents understand user intent, arguing that current benchmarks fail to capture the gap between what users ask and what they mean.
This paper studies language models' failure to act on communicative intent despite robust internal representations. Using linear probes, the authors show intent is decodable from hidden states but often not reflected in outputs, and steering a late-layer direction can recover the intended behavior.
This paper proposes demographic-conditioned fusion embeddings to model perspectivist social meaning in language, showing consistent improvements over text-only baselines by integrating annotator demographics into NLP systems.
Introduces KARMA, a framework that trains a reward model on Reddit conversations to improve LLMs' context-sensitive conversational behavior via reinforcement learning, finding that the best reward model for predicting karma does not yield the best downstream alignment.
Introduces DRInQ, a benchmark for evaluating conversational implicature in question utterances, revealing that LLMs often fail to recover intended implications at inference time despite being able to generate plausible pragmatic scenarios.