Tag
Introduces DRInQ, a benchmark for evaluating conversational implicature in question utterances, revealing that LLMs often fail to recover intended implications at inference time despite being able to generate plausible pragmatic scenarios.