Tag
This paper reevaluates the imperfective paradox benchmark for large language models, identifying conceptual and evaluation mis-specifications, and introduces lexically matched minimal pairs to reveal sufficiency bias in models' semantic reasoning.
Stanford NLP announces that a linguistics-informed NLP paper titled 'The Imperfective Paradox in LLMs' has won a Best Paper Award at ACL 2026, encouraging more detailed linguistic evaluation of large language models.