Tag
This paper introduces TatBLiMP, the first linguistic minimal pairs benchmark for the Tatar language, evaluating 16 morphosyntactic phenomena across models from from-scratch Tatar models to frontier multilingual LLMs.
This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.
This paper investigates whether large language models have a localized causal mechanism for handling the animacy concept, using circuit discovery on minimal pairs; they find an animacy circuit that is distributed and only partially generalizes.
Introduces ALEE, a framework that uses Abstract Meaning Representations to generate English minimal pairs with controlled semantic shifts and translates them for evaluating text embeddings across 275+ languages, revealing persistent gaps in cross-lingual semantic representation.
This paper evaluates whether Sign Language Recognition models exhibit phonological sensitivity by probing them with minimal pairs of signs, revealing architectural trade-offs and emergent but limited phonological perception.