Tag
Mizan introduces a national benchmark for evaluating large language models on Iraqi Arabic and civic context, highlighting gaps in MSA-focused evaluations and revealing issues like over-refusal in safety-hardened models.
This paper introduces a benchmark for semantic segmentation in low-resource dialectal Arabic and proposes a model that improves performance on conversational speech compared to standard baselines.