Tag
Introduces BavGround, a benchmark for evaluating LLMs' regional cultural grounding and dialect competence in Bavarian across English, German, and Bavarian, finding that models struggle with dialectal and localized cultural knowledge.
This paper presents a unified poly-dialectal neural machine translation system for 12 Bangla regional dialects, introducing the largest multi-dialect parallel corpus to date and achieving state-of-the-art BLEU scores with a fine-tuned BanglaT5 model using DoRA.
This paper introduces VialectBench, a benchmark evaluating LLM robustness to six Vietnamese dialect groups across four tasks, finding average performance drops of 2.82% and no model fully dialect-invariant.
This paper investigates methods to steer Arabic LLMs toward dialect-specific generation by identifying sparse neuron populations and extracting dialect activation directions, enabling dialect control at inference time without fine-tuning.
Tests on Qwen3.5-35B-A3B show that AAVE-coded prompts cause MoE models to respond differently, with refusal layers masking dialect-conditioned safety failures that become visible when refusal is weakened.