Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning
Summary
Wazobia Eval is a benchmark designed to evaluate AI models on emotion understanding, sarcasm detection, and cultural reasoning in Nigerian Pidgin, aiming to advance NLP capabilities for this language.
View Cached Full Text
Cached at: 08/25/26, 04:09 AM
# Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning Source: [https://arxiv.org/abs/2608.21369](https://arxiv.org/abs/2608.21369) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
This paper introduces VIVID, the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese, comprising 1,636 idioms and proverbs. Evaluation of eight state-of-the-art models reveals significant gaps, with Vietnamese-specialized models drastically underperforming multilingual systems and even top models achieving less than 50% correctness on average.
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models
SpeechEQ introduces a benchmark and dataset for evaluating emotional intelligence in speech-language models, covering 15 EQ subscales across 2,265 dialogues. Experiments reveal current models struggle with paralinguistic cues, exhibiting text-reliant shortcuts and other limitations.
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
Introduces CAREBench, a benchmark grounded in appraisal theory to evaluate LLMs' emotion understanding through cognitive appraisal reasoning, revealing that current models struggle with reasoning and positive emotion recognition despite matching humans on some downstream tasks.
Evaluation Awareness in Language Models: Representation, Verbalization, and Control
This paper provides a systematic study of evaluation awareness in language models, showing that models internalize evaluation context, leading to a disconnect between internal representation, verbalization, and steering behavior, with implications for benchmark reliability.
Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
Introduces Inspect India Evals, an open-source framework for evaluating LLMs in Indian linguistic and cultural contexts, with six benchmarks testing multilingual ability, bias, safety, and cultural knowledge. Tests on five models show Sarvam-M 24B and Gemma 2 27B lead.