When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs
Summary
This paper introduces a benchmark dataset of 1,516 expert-verified Bangla sentences for disambiguating culturally entangled homographs (words that are both names and common nouns). It shows that LLMs suffer from dominant-meaning bias and proposes contrastive chain-of-thought prompting and distillation to reduce this bias.
View Cached Full Text
Cached at: 07/21/26, 06:46 AM
# When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs
Source: [https://arxiv.org/html/2607.17828](https://arxiv.org/html/2607.17828)
Md\. Asaduzzaman Shuvo
United International University, Bangladesh Emails:\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishashuvo221104@bscse\.uiu\.ac\.bd
###### Abstract
Many Bangla words are at once personal names and culturally loaded common nouns:\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliমায়া \(Maya\) is both a girl’s name and a word for affectionate compassion\. Choosing the right reading demands cultural knowledge that is scarce in the pretraining data of modern language models\. We introduceCulturally Entangled Homograph \(CEH\)disambiguation and build a Bangla benchmark of 1,516 expert\-verified sentences \(3,032 labelled occurrences\) in which one word appears twice with two distinct readings, each labelled with a culturally grounded category and an explanation of the reasoning behind it\. Across open\- and closed\-source models, we find a systematic*dominant\-meaning bias*: models default to the common\-noun sense and overlook the name\. A Bangla\-specific model fails under every prompting regime we test, showing that language\-specific pretraining alone does not confer cultural grounding\. We further show that contrastive chain\-of\-thought prompting can sharply reduce this bias without training, and that distilling cultural explanations teaches small \(1–3B\) models to reason toward the correct reading rather than memorise labels, cutting dominant\-meaning bias from as high as 100% to under 5% and turning the failed Bangla\-specific model into our strongest system\. We release the code and the dataset\.\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1Dataset and code are available at[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishhttps://github\.com/ashuvo25/BanglaCEH](https://github.com/ashuvo25/BanglaCEH)\.
\\fontspec\_if\_language:nTF
ENG\\addfontfeatureLanguage=English
When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low\-Resource LLMs
Md\. Asaduzzaman ShuvoUnited International University, BangladeshEmails:\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishashuvo221104@bscse\.uiu\.ac\.bd
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1Introduction
Word\-sense ambiguity is a long\-standing challenge in NLP, but a particularly difficult and understudied variant arises when a single word functions simultaneously as a personal name and as a culturally loaded common noun\. In Bangla this is pervasive: parents routinely name children after words denoting prized emotions, virtues, or spiritual states, so one surface form carries both an individuating \(name\) reading and an abstract \(concept\) reading\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 1:The Culturally Entangled Homograph \(CEH\) task\.\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliমায়া \(Maya\) appears twice in one Bangla sentence first as a*name*, then as the*emotion*it is named after and the model must disambiguate the two using cultural knowledge alone\.The word\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliমায়া \(Maya\) is at once a common girl’s name and a word for deep affectionate compassion;\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliআরিফ \(Arif\) is both a boy’s name and, in Sufi theology, one who possesses intuitive knowledge of the divine\. Resolving the intended reading demands cultural and pragmatic knowledge beyond surface lexical statistics\.
We term this phenomenon theCulturally Entangled Homograph \(CEH\)and argue it is a sharp diagnostic of genuine cultural grounding\. Unlike conventional word\-sense disambiguation, a CEH cannot be resolved by frequency or distributional cues: both readings coexist in the same cultural space, and disambiguation hinges on recognising when a form*names*a person versus*invokes*the concept it was drawn from making CEH especially demanding for models trained on high\-resource, Western\-centric corpora\.
LLMs are known to encode a Western\-dominance bias and to struggle with low\-resource cultures\(Belay et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib2); Yu et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib23)\); language\-specific competence does not entail cultural competence, and in\-language prompting alone does not supply cultural context\(Belay et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib2)\)\. For Bangla, recent work has produced capable models and benchmarks\(Nahin et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib13); Raihan and Zampieri,[2025](https://arxiv.org/html/2607.17828#bib.bib18); Joy and Shatabda,[2026](https://arxiv.org/html/2607.17828#bib.bib8)\), but evaluation has centred on knowledge, reasoning, and classification, leaving cultural grounding in everyday lexical usage largely unexamined\. Related studies show that fluent Bangla output can mask systematic register failures\(Shuvo et al\.,[2026](https://arxiv.org/html/2607.17828#bib.bib20)\)and persistent hallucination\(Adib et al\.,[2026](https://arxiv.org/html/2607.17828#bib.bib1)\)symptoms of shallow rather than grounded understanding\.
We introduce the CEH task and build a benchmark of Bangla sentences in which one word appears twice with two distinct readings, each annotated with one of six culturally grounded categories and an explanation of the underlying cultural reasoning\. Evaluating open\- and closed\-source models, we uncover a systematic*dominant\-meaning bias*: models default to the common\-noun reading and overlook the personal\-name reading \(Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1](https://arxiv.org/html/2607.17828#S1.F1)\)\. A Bangla\-specific model fails under every regime we test, confirming that language\-specific pretraining alone does not confer cultural grounding\. We further show that contrastive chain\-of\-thought prompting can sharply reduce this bias without training, though its effect is strongly model\-dependent, and that distilling cultural reasoning into small models teaches them to reason toward the correct reading rather than memorise labels\.
Contributions\.\(i\) We formalise CEH disambiguation, a task isolating cultural grounding from surface lexical competence in a low\-resource language\. \(ii\) We release a Bangla CEH benchmark with token\-level cultural categories and cultural\-reasoning explanations\. \(iii\) We evaluate open\- and closed\-source models under zero\-shot, few\-shot, contrastive chain\-of\-thought, and knowledge\-distillation regimes, revealing a consistent dominant\-meaning bias and the failure of a Bangla\-specific model across all settings\. \(iv\) We show contrastive reasoning and distilled supervision mitigate this bias, and analyse when and why they succeed or fail\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2Related Work
Word\-sense disambiguation \(WSD\) is a long\-standing NLP problem\(Navigli and Ponzetto,[2010](https://arxiv.org/html/2607.17828#bib.bib14); Raganato et al\.,[2017](https://arxiv.org/html/2607.17828#bib.bib17)\), and remains especially hard for morphologically rich, low\-resource languages where annotated corpora are scarce\(Habtamu and Gizachew,[2024](https://arxiv.org/html/2607.17828#bib.bib5); Masethe et al\.,[2024](https://arxiv.org/html/2607.17828#bib.bib11)\); recent work has extended WSD evaluation to under\-resourced settings through hybrid sense annotation\(Goworek et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib4)\)and dedicated Arabic and Urdu resources\(Khalilia et al\.,[2024](https://arxiv.org/html/2607.17828#bib.bib9); Saeed et al\.,[2019](https://arxiv.org/html/2607.17828#bib.bib19)\), while studies probing whether large language models truly grasp word senses report that surface fluency often masks shallow lexical understanding\(Meconi et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib12); Ortega\-Martín et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib15)\)\. Our CEH task differs from classical WSD in that the competing readings a personal name versus the culturally loaded concept it derives from are not lexically distinct senses but culturally entangled ones, requiring pragmatic and cultural knowledge rather than dictionary sense inventories\. This connects our work to a rapidly growing literature on the cultural competence of LLMs, which finds that models encode a Western\-dominance bias and underperform on the values and everyday knowledge of low\-resource cultures\(Belay et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib2); Yu et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib23)\), and that language\-specific ability does not guarantee cultural grounding\(Belay et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib2)\)\. For Bangla specifically, recent efforts have produced dedicated language models and knowledge benchmarks\(Nahin et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib13); Raihan and Zampieri,[2025](https://arxiv.org/html/2607.17828#bib.bib18); Joy and Shatabda,[2026](https://arxiv.org/html/2607.17828#bib.bib8)\), yet cultural grounding in everyday lexical usage remains underexplored, and fluent Bangla generation can still hide systematic register and factuality failures\(Shuvo et al\.,[2026](https://arxiv.org/html/2607.17828#bib.bib20); Adib et al\.,[2026](https://arxiv.org/html/2607.17828#bib.bib1)\)\. Finally, our mitigation strategy builds on chain\-of\-thought prompting\(Wei et al\.,[2022](https://arxiv.org/html/2607.17828#bib.bib22)\)and on knowledge distillation of reasoning into smaller models\(Magister et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib10); Hsieh et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib6)\), where a teacher model’s rationales are distilled into a compact student; unlike prior distillation work that targets mathematical or commonsense reasoning, we distil*cultural*reasoning, teaching small models to justify why a given occurrence is a name or a concept rather than merely to reproduce the label\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3Dataset Construction
Task formulation\.Each instance in our benchmark is a single Bangla sentence in which one word form appears*twice*with two distinct readings\. Given the sentence and the target word, a model must assign each of the two occurrences one of six culturally grounded categories Name, Concept, Emotion, Emotional State, Collective State, and Spiritual and justify each assignment\. Crucially, the two readings are not arbitrary lexical senses but*culturally entangled*ones: the same form functions once as a personal name and once as the emotion, concept, or spiritual state from which that name is conventionally drawn\. Word selection\.We do not rely on a fixed name registry or lexical database\. Instead, a native Bangla speaker manually curated words such as\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliমায়া \(Maya\),\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliআশা \(Asha\), and\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliগগন \(Gagan\) that are both common Bangla given names and culturally loaded common nouns\. This expert\-driven selection deliberately targets forms with a genuine dual reading, which a generic name list would not isolate, at the cost of coverage reflecting a single annotator’s naming knowledge Section \([Limitations](https://arxiv.org/html/2607.17828#Sx1)\)\. Multi\-model generation\.For each selected word, we generated sentences and accompanying cultural explanations using three state\-of\-the\-art models as complementary generators: Claude Fable 5, Kimi K3, and Gemini 3\.5 Flash, producing 690, 480, and 344 instances respectively\. Multiple generators reduce the stylistic and distributional bias of a single teacher and yield greater lexical and structural diversity\.
Word\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliগগন \(Gagan\)CategoryName↔\\leftrightarrowConceptSentence\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliসমুদ্রসৈকতে সূর্যাস্ত দেখতে দেখতেগগন\-র মনে হলো, প্রাচীন শ্লোকে ’গগন’ দিয়ে ঠিক এই আকাশ; নভোমণ্ডল\-কেই বোঝানো হতো।GlossWatching the sunset, it struckGaganthat ancient verse used ‘Gagan’ for exactly this the celestial firmament\.Occ\. 1Namea person, the grammatical subjectOcc\. 2Conceptthe celestial firmamentCultural noteBengali naming turns words for nature, deities, and virtues into personal names;\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliগগন is thus once a living identity and once a cultural concept\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 1:An example CEH instance\. The same form\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliগগন \(Gagan\) appears twice with a name reading and a concept reading\. The full bilingual schema is in Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishC](https://arxiv.org/html/2607.17828#A3)\.Cross\-model and human verification\.We applied a two\-stage verification protocol\. First, outputs were cross\-checked round\-robin by a different model \(Claude verified by Kimi, Kimi by Gemini, Gemini by Claude\), flagging inconsistent labels or explanations without any model auditing its own output\. Second, every instance was manually reviewed by two native Bangla\-speakers \(profile in appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishA](https://arxiv.org/html/2607.17828#A1)\), who corrected labels, explanations, and culturally inaccurate reasoning\. Human verification is the final authority; the cross\-model stage only surfaces candidate errors for the reviewer\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 2:CEH statistics entanglement categories \(left\) and token\-level label distribution \(right\)\.Annotations\.Each verified instance contains token\-level labels for both occurrences, a natural\-language explanation of why each occurrence carries its label, and a concise cultural\-entanglement note describing the naming convention that links the two readings\. These explanations form the supervision signal for our knowledge\-distillation
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 2:Overview of our methodology, spanning dataset creation, prompt\-based evaluation \(zero\-shot, few\-shot, C\-CoT\), knowledge\-distillation fine\-tuning with LoRA, and evaluation on held\-out CEH sentences using automatic, human, and LLM\-based judges\.experiments Section \([\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2607.17828#S4)\)\. An example instance is shown in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1](https://arxiv.org/html/2607.17828#S3.T1); the complete field schema and a raw JSON record are provided in Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishC](https://arxiv.org/html/2607.17828#A3)\. Statistics\.The final benchmark contains 1,516 verified instances, each contributing two labelled occurrences \(3,032 labelled tokens\)\. Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2](https://arxiv.org/html/2607.17828#S3.T2)reports the distribution over entanglement categories and over the six labels\. Name and Concept are the most frequent labels, reflecting the prominence of name–concept entanglement in Bangla naming practice, while Emotion and Spiritual form a rarer long tail that we find models handle least reliably in Section \([\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5](https://arxiv.org/html/2607.17828#S5)\)\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4Methodology
We evaluate the CEH task under four regimes of increasing supervision: zero\-shot prompting, few\-shot prompting, Cultural Chain\-of\-Thought \(C\-CoT\) prompting, and knowledge\-distillation \(KD\) fine\-tuning\. The first two establish how much cultural grounding models possess without task\-specific adaptation; the latter two are our proposed interventions for mitigating the dominant\-meaning bias\. All regimes are evaluated on the same held\-out test set to enable direct comparison\. We split the benchmark into 78% training, 10% validation, and 12% test partitions with a fixed random seed\. Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2](https://arxiv.org/html/2607.17828#S3.F2)provides an overview of the full pipeline\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.1Prompting Baselines
#### Zero\-shot\.
The model receives the sentence and an instruction to assign each of the two occurrences of the target word a label from the six categories, with no examples\. This measures a model’s out\-of\-the\-box cultural grounding\.Few\-shot\.We prepend a small number of labelled in\-context examples before the query, testing whether demonstrations alone without weight updates help the model recognise the entangled readings\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.2Cultural Chain\-of\-Thought
Standard chain\-of\-thought prompting elicits intermediate reasoning before a final answer\(Wei et al\.,[2022](https://arxiv.org/html/2607.17828#bib.bib22)\), but it does not counteract a model’s prior tendency to collapse both occurrences onto the dominant sense\. We introduceCultural Chain\-of\-Thought \(C\-CoT\)\(Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishD](https://arxiv.org/html/2607.17828#A4)\), a contrastive variant that explicitly forces the model to argue*both*candidate readings for each occurrence before committing\. For every occurrence, the model must \(i\) state why the form could be aName, \(ii\) state why it could be the competing cultural category \(Concept,Emotion,Spiritual, orState\), and \(iii\) decide which reading the sentential and cultural context supports\. By requiring the suppressed reading to be articulated before a decision is made, C\-CoT directly targets the dominant\-meaning bias rather than merely eliciting free\-form rationale\. C\-CoT requires no training and is applied at inference time\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.3KD Fine\-Tuning
Our final regime tests whether the cultural reasoning behind each label can be*taught*to a small model rather than merely prompted\. We fine\-tune compact open\-source models on the human\-verified cultural explanations in our benchmark, following a knowledge\-distillation formulation in which a student model learns to reproduce the reasoning traces of stronger teacher models\(Hsieh et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib6); Magister et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib10)\)\. Crucially, the training target is the*full cultural explanation followed by the label*, not the label alone: the model first generates the reasoning that justifies each occurrence’s reading and only then emits the final labels\. This encourages the model to internalise the pattern of cultural reasoning rather than memorise a surface mapping from word to label\. We use parameter\-efficient fine\-tuning via Low\-Rank Adaptation\(Hu et al\.,[2021](https://arxiv.org/html/2607.17828#bib.bib7)\)in its quantized form, QLoRA\(Dettmers et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib3)\), which back\-propagates gradients through a frozen 4\-bit base model into low\-rank adapters, enabling training of billion\-parameter models on a single consumer GPU\. Loss is computed only over the target reasoning and label tokens, with the prompt masked, so that supervision is concentrated on the cultural reasoning the model must learn to produce\. We select the checkpoint with the lowest validation loss and apply early stopping to prevent overfitting\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.4Models
We evaluate a mix of open\- and closed\-source models spanning a range of parameter scales\. The open\-source set comprises Qwen2\.5\-1\.5B, Llama\-3\.2\-3B, Gemma\-3\-1B, Ministral\-3B, and the Bangla\-specific TituLLM; the last is included specifically to test whether language\-targeted pretraining confers cultural grounding\. We include GPT\-4o\-mini as a strong closed\-source reference\. The prompting baselines are evaluated across all models, while C\-CoT and knowledge\-distillation fine\-tuning are applied to a representative subset spanning small general\-purpose and Bangla\-specific models\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.5Evaluation Metrics
We evaluate along two tracks\. The label track measures disambiguation accuracy through Exact Match \(both occurrences correct\), Average Score \(Average Score is the percentage of total points earned over the maximum possible points, see Appendix\([\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishF\.1](https://arxiv.org/html/2607.17828#A6.SS1)\)\), together with Macro\- and Micro\-averaged F1 over the six labels\. We additionally report a Hallucination rate the fraction of outputs producing a label outside the valid set and, as our central diagnostic, a Dominant\-Bias rate measuring how often a model misses a Name label and defaults to the entangled cultural reading\. The reasoning track evaluates the quality of generated cultural explanations against the gold explanations using chrF\(Popović,[2015](https://arxiv.org/html/2607.17828#bib.bib16)\)and BERTScore\(Zhang et al\.,[2019](https://arxiv.org/html/2607.17828#bib.bib24)\)\. In addition, we assess outputs through a Human\-as\-a\-Judge and LLM\-as\-a\-Judge evaluation\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5Result Analysis
Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3](https://arxiv.org/html/2607.17828#S5.T3)reports performance across all four regimes\. We organise our analysis around four findings: the severity of the dominant\-meaning bias under prompting, the model\-dependent behaviour of C\-CoT, the decisive effect of knowledge\-distillation fine\-tuning, and the striking reversal of the Bangla\-specific model\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 3:CEH results across four regimes\.↑\\uparrowhigher is better;↓\\downarrowlower is better\. EM = Exact Match, Avg = Average Score, Ma\-F1 / Mi\-F1 = Macro / Micro F1, BERT = BERTScore\-F1, Hall\. = Hallucination, Bias = Dominant\-Bias \(all %\)\. C\-CoT and QLoRA\-KD were run on closed\-source prompting reference\. For each model, the best value per metric is inbold\.### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.1Zero\-Shot Prompting
Under zero\-shot prompting, every open\-source model shows a pronounced dominant\-meaning bias, defaulting to the common\-noun reading and misclassifying the personal\-name occurrence\. Bias ranges from 62\.5% \( Ministral\) to 100% \(Gemma, TituLLM\), with exact match below 11% for all open models\. The failure modes differ: Ministral and Llama produce valid but wrong labels, while others collapse entirely TituLLM hallucinates on 100% of inputs and Gemma shows complete \(100%\) bias, indicating the name reading is not even represented as a candidate\. The closed\-source GPT\-4o\-mini is markedly stronger 43\.87% exact match, 0\.65% hallucination yet retains a non\-trivial 16\.35% bias\. This open–closed gap holds across every metric, confirming CEH disambiguation is genuinely hard: even a capable proprietary model without task\-specific adaptation misreads the name as the entangled concept in roughly one in six cases\. Uniformly low Macro\-F1 \(0\.005–0\.467\) further shows zero\-shot failure is disproportionate on rarer cultural labels, not merely the frequent name/concept pairing\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.2Few\-Shot Prompting
Adding a few in\-context examples improves most models over zero\-shot, but the gains are inconsistent and leave the core bias largely unresolved\. Gemma benefits most dramatically, rising from 0\.91% to 31\.29% exact match as its dominant\-meaning bias falls from 100% to 39\.5% evidence that demonstrations can teach a model to represent the name reading it previously ignored\. Qwen and Llama show more modest gains \(exact match 0\.65%→\\rightarrow8\.39% and 10\.97%→\\rightarrow22\.58%\), with bias dropping into the low\-20% range for both\. Two models resist this trend: Ministral*degrades*under few\-shot prompting exact match falls from 9\.0% to 5\.8% with bias barely moving \(62\.5%→\\rightarrow60\.6%\), the demonstrations disrupting rather than guiding its predictions while TituLLM remains effectively non\-functional, hallucinating on 100% of inputs in both regimes despite a nominal rise to 0\.31% exact match, confirming its failure is one of task representation rather than a shortage of examples\. GPT\-4o\-mini again leads, reaching 49\.68% exact match with bias reduced to 2\.88%, its strongest prompting result\. Overall, few\-shot prompting narrows but does not close the gap: even the best open\-source configuration \(Gemma, 31\.29%\) trails GPT\-4o\-mini by a wide margin, and dominant\-meaning bias remains substantial for every open model motivating the reasoning\- and training\-based interventions we examine next\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.3Cultural Chain\-of\-Thought
C\-CoT produces the largest prompting\-time improvement we observe, but only for models able to follow its contrastive structure\. On Qwen2\.5\-1\.5B, it drives dominant\-meaning bias from 63\.46% \(zero\-shot\) to 2\.88%, eliminates hallucination entirely \(0%\), and raises exact match to 30\.97% a level no other prompting configuration reaches for this model\. Forcing the model to articulate the suppressed name\-reading before committing counteracts the bias without parameter updates, supporting our hypothesis that the dominant\-meaning failure is one of reasoning order rather than missing knowledge\. The benefit does not transfer uniformly: on Llama\-3\.2\-3B, C\-CoT is counter\-productive dominant\-meaning bias*rises*to 75% and hallucination climbs to 43\.23%, both worse than the zero\- and few\-shot baselines\. The degradation coincides with a sharp rise in unparseable outputs: Llama’s extended contrastive reasoning often fails to terminate in the required label format, so otherwise valid predictions are discarded\. C\-CoT’s effectiveness is thus contingent on sustaining format adherence over long reasoning chains an important limitation for smaller and lower\-resource models\. TituLLM remains largely non\-functional under C\-CoT \(89\.12% hallucination, 1\.14% exact match\); contrastive reasoning cannot compensate for a model that does not reliably represent the task\. Together, these results show C\-CoT can nearly eliminate dominant\-meaning bias when a model can execute it, but not a universal remedy motivating the training\-based approach we examine next\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.4QLoRA\-KD Fine\-Tuning
The QLoRA\-KD regime yields the strongest results by a wide margin\. All three fine\-tuned models exceed 82% exact match and reduce dominant\-meaning bias to at most 4\.2% an order\-of\-magnitude improvement over their best prompting configuration\. Qwen2\.5\-1\.5B reaches 85\.16% exact match with 0% dominant\-meaning bias and 1\.65% wrong predictions, while Llama\-3\.2\-3B, which C\-CoT had degraded, recovers to 82\.42% exact match and 2\.8% bias\. This confirms that the dominant\-meaning failure is not an inherent limitation of these models but a gap that targeted supervision can close\. Crucially, the improvement extends to reasoning quality, not merely label accuracy\. For Qwen, BERTScore rises from 0\.178 under C\-CoT to 0\.784 under QLoRA\-KD and chrF from 24\.3 to 79\.35, indicating that the fine\-tuned model generates cultural explanations closely aligned with the gold reasoning rather than producing correct labels for spurious reasons\. Because the distillation target pairs each label with its cultural justification, the model learns the*pattern*of cultural reasoning that licenses a reading rather than a surface word\-to\-label mapping precisely the capability that prompting failed to elicit\.
These gains are obtained with parameter\-efficient adaptation on models of only 1–3B parameters\. Notably, Qwen2\.5 fine\-tuned model surpasses the strongest closed\-source prompting baseline \(GPT\-4o\-mini few\-shot, 49\.68% exact match, 2\.88% bias\) by a large margin\. We stress that this is not a like\-for\-like comparison GPT\-4o\-mini is evaluated only under prompting, not fine\-tuned but it establishes that a small, openly available model, once taught cultural reasoning, can substantially outperform a much larger proprietary model prompted on the same task\. The residual gap between Micro\- and Macro\-F1 \( 0\.926 vs\. 0\.737 for Qwen\) indicates that errors now concentrate almost entirely in the rarer cultural labels, which we identify as the primary remaining challenge \(full per\-label results and confusion matrices in Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishH](https://arxiv.org/html/2607.17828#A8)\)\. An ablation study is provided in Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishB\.1](https://arxiv.org/html/2607.17828#A2.SS1), confirming the importance of reasoning supervision\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.5Human & LLM Judge Evaluation
ModelHall\.%↓\\downarrowBias%↓\\downarrowExact%↑\\uparrowP%↑\\uparrowHuman\-as\-a\-JudgeQwen2\.5\-1\.5B1\.00\.28415Llama\-3\.2\-3B1\.02\.08216titulm\-1b2\.03\.08710LLM\-as\-a\-Judge \(GPT\-5\.4\)Qwen2\.5\-1\.5B1\.00\.38514Llama\-3\.2\-3B1\.12\.58216titulm\-1b2\.03\.58710\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 4:Human and LLM \(GPT\-5\.4\) judge evaluation of the fine\-tuned \(QLoRA\-KD\) models\.↑\\uparrowhigher is better;↓\\downarrowlower is better and P = Partial\.To verify that the fine\-tuned models reach correct labels through culturally sound reasoning, we evaluate their outputs with two\-stage independent judges: two human judges \(native Bangla speakers\) and an LLM judge GPT\-5\.4 \(the template is in Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishE](https://arxiv.org/html/2607.17828#A5)\)\. Each judge scores every output for hallucination, dominant\-meaning bias, and exact\- and partial\-match accuracy; results are reported in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2607.17828#S5.T4)\. The two judges agree closely, differing by at most 1 point on any metric, which both corroborates our automatic evaluation and supports the use of an LLM judge as a scalable proxy for human assessment\. Consistent with the automatic results, all three fine\-tuned models exhibit near\-zero hallucination and dominant\-meaning bias under both judges, and TituLLM attains the highest exact\-match rating \(87%\), confirming that its post\-distillation predictions are not only accurate but judged culturally faithful by a human expert\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6Discussion
Our results separate two mechanisms for resolving culturally entangled homographs: surfacing knowledge already latent in a model’s parameters, and acquiring knowledge it lacks\. Prompting, including contrastive C\-CoT, engages only the former it helps only insofar as cultural grounding is already present and can be reorganised through reasoning order, which is why C\-CoT nearly eliminates the dominant\-meaning bias on Qwen yet leaves TituLLM, which lacks such grounding, effectively unchanged\. Knowledge\-distillation fine\-tuning engages the latter: supervising on cultural explanations rather than labels alone induces the reasoning that connects a name to its source concept, so predictions follow from reasoning rather than memorised surface mappings\. Under this regime accuracy and reasoning quality improve together rather than in isolation the models do not merely predict the right label but justify it with culturally faithful explanations, a coupling our human and LLM judges independently corroborate\.
These gains proved sensitive to how the adaptation was configured\. Because the benchmark is small relative to adapter capacity, an initial higher\-capacity setting overfit within a few epochs; reducing the adapter rank and learning rate and strengthening regularisation delayed overfitting and lowered the validation minimum, and combined with early stopping this configuration transferred unchanged across all three models, indicating that the recipe is not narrowly tuned to a single architecture \(Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishG](https://arxiv.org/html/2607.17828#A7)\)\.
The most consequential result is the reversal of the Bangla\-specific model: non\-functional under every prompting regime, yet best\-in\-class after distillation\. This suggests that monolingual pretraining supplies a useful substrate more faithful tokenisation of Bangla script and denser exposure to the relevant names and morphology but not task competence itself; the substrate instead allows limited reasoning supervision can be absorbed far more efficiently than by a general model with a weaker lexical prior\. For low\-resource, culturally rich languages that lack large instruction\-tuned systems, this points to a reproducible recipe: pair a compact language\-specific model with a small set of cultural reasoning traces and adapt it with parameter\-efficient fine\-tuning\. Although our experiments focus on Bangla, the proposed framework is language\-agnostic and could be evaluated on other languages with similar name–concept entanglement\.
Two limitations qualify these conclusions\. First, the benefit of C\-CoT is model\-dependent: it helps only when a model can sustain the required output format across long reasoning chains, and degrades performance where it cannot, as observed for Llama\. Second, the persistent gap between Micro\- and Macro\-averaged F1 after fine\-tuning shows that the rarer cultural labels remain the hardest, locating the current bottleneck in long\-tail data scarcity rather than in the method, and marking targeted expansion of these categories as the clearest direction for future work\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English7Conclusion
We introduced Culturally Entangled Homograph \(CEH\) disambiguation, a task in which the same Bangla word must be read once as a personal name and once as the cultural concept it is drawn from, and released a benchmark of annotated instances paired with cultural\-reasoning explanations\. Across open\- and closed\-source models we found a systematic dominant\-meaning bias models default to the common\-noun sense and overlook the name and showed that language\-specific pretraining alone does not fix it: the Bangla\-specific model failed under every prompting regime\. Yet the same model became the strongest system once fine\-tuned on distilled cultural reasoning, reaching near\-perfect disambiguation with the bias eliminated\. Our results suggest that cultural grounding is not conferred by monolingual pretraining alone, but can be improved through supervision with reasoning\-annotated examples\. This provides a practical approach for improving cultural disambiguation in low\-resource languages\. Evaluating its applicability to other languages with similar name–concept entanglement remains an important direction for future work\.
## Limitations
Our benchmark has several limitations\. First, word selection was performed manually by a single native Bangla speaker rather than drawn from a name registry or lexical database; while this deliberately targets forms with a genuine dual reading, it means coverage reflects one annotator’s naming knowledge and may under\-represent regional or less common names\. Second, the sentences and cultural explanations were initially produced by large language models and subsequently verified by two native Bangla\-speaking experts\. Although every instance was manually reviewed, model\-generated text may still exhibit subtle stylistic regularities that a fully human\-authored corpus would not\. Third, although every instance was independently reviewed by two native Bangla\-speaking annotators and the resulting inter\-annotator agreement was high \(Cohen’sκ=0\.95\\kappa=0\.95\), the annotation process still relied on a relatively small annotator pool\. While disagreements were resolved through discussion using a shared annotation guideline, involving more annotators from diverse regional and linguistic backgrounds would further improve the robustness and generalizability of the benchmark\. Finally, our fine\-tuning experiments use parameter\-efficient adaptation on small models under limited compute, so the reported gains may not transfer directly to larger models or full\-parameter fine\-tuning\.
## References
- Adib et al\. \(2026\)Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Ishmam Tashdeed, and Md Taukir Azam Chowdhury\. 2026\.Benhallueval: A multi\-task hallucination evaluation framework for large language models on bengali\.*arXiv preprint arXiv:2605\.31483*\.
- Belay et al\. \(2025\)Tadesse Destaw Belay, Ahmed Haj Ahmed, Alvin C Grissom Ii, Iqra Ameer, Grigori Sidorov, Olga Kolesnikova, and Seid Muhie Yimam\. 2025\.Culemo: Cultural lenses on emotion\-benchmarking llms for cross\-cultural emotion understanding\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 18894–18909\.
- Dettmers et al\. \(2023\)Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer\. 2023\.[Qlora: Efficient finetuning of quantized llms](https://proceedings.neurips.cc/paper_files/paper/2023/file/1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf)\.In*Advances in Neural Information Processing Systems*, volume 36, pages 10088–10115\. Curran Associates, Inc\.
- Goworek et al\. \(2025\)Roksana Goworek, Harpal Singh Karlcut, Hamza Shezad, Nijaguna Darshana, Abhishek Mane, Syam Bondada, Raghav Sikka, Ulvi Mammadov, Rauf Allahverdiyev, Sriram Satkirti Purighella, Paridhi Gupta, Muhinyia Ndegwa, Bao Khanh Tran, and Haim Dubossarsky\. 2025\.[SenWiCh: Sense\-annotation of low\-resource languages for WiC using hybrid methods](https://doi.org/10.18653/v1/2025.sigtyp-1.7)\.In*Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP*, pages 61–74, Vienna, Austria\. Association for Computational Linguistics\.
- Habtamu and Gizachew \(2024\)Robbel Habtamu and Beakal Gizachew\. 2024\.State\-of\-the\-art approaches to word sense disambiguation: A multilingual investigation\.In*Pan\-African Conference on Artificial Intelligence*, pages 176–202, Cham\. Springer Nature Switzerland\.
- Hsieh et al\. \(2023\)Cheng\-Yu Hsieh, Chun\-Liang Li, Chih\-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen\-Yu Lee, and Tomas Pfister\. 2023\.[Distilling step\-by\-step\! outperforming larger language models with less training data and smaller model sizes](https://doi.org/10.18653/v1/2023.findings-acl.507)\.In*Findings of the Association for Computational Linguistics: ACL 2023*, pages 8003–8017, Toronto, Canada\. Association for Computational Linguistics\.
- Hu et al\. \(2021\)Edward J\. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\. 2021\.[Lora: Low\-rank adaptation of large language models](https://arxiv.org/abs/2106.09685)\.*Preprint*, arXiv:2106\.09685\.
- Joy and Shatabda \(2026\)Saman Sarker Joy and Swakkhar Shatabda\. 2026\.[BnMMLU: Measuring massive multitask language understanding in Bengali](https://doi.org/10.18653/v1/2026.findings-acl.593)\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 12211–12230, San Diego, California, United States\. Association for Computational Linguistics\.
- Khalilia et al\. \(2024\)Mohammed Khalilia, Sanad Malaysha, Reem Suwaileh, Mustafa Jarrar, Alaa Aljabari, Tamer Elsayed, and Imed Zitouni\. 2024\.[ArabicNLU 2024: The first Arabic natural language understanding shared task](https://doi.org/10.18653/v1/2024.arabicnlp-1.30)\.In*Proceedings of the Second Arabic Natural Language Processing Conference*, pages 361–371, Bangkok, Thailand\. Association for Computational Linguistics\.
- Magister et al\. \(2023\)Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn\. 2023\.[Teaching small language models to reason](https://doi.org/10.18653/v1/2023.acl-short.151)\.In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 1773–1781, Toronto, Canada\. Association for Computational Linguistics\.
- Masethe et al\. \(2024\)Hlaudi Daniel Masethe, Mosima Anna Masethe, Sunday Olusegun Ojo, Fausto Giunchiglia, and Pius Adewale Owolawi\. 2024\.[Word sense disambiguation for morphologically rich low\-resourced languages: A systematic literature review and meta\-analysis](https://doi.org/10.3390/info15090540)\.*Information*, 15\(9\)\.
- Meconi et al\. \(2025\)Domenico Meconi, Simone Stirpe, Federico Martelli, Leonardo Lavalle, and Roberto Navigli\. 2025\.[Do large language models understand word senses?](https://doi.org/10.18653/v1/2025.emnlp-main.1720)In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 33897–33916, Suzhou, China\. Association for Computational Linguistics\.
- Nahin et al\. \(2025\)Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker, Quazi Sarwar Muhtaseem, Md Kowsher, Apu Chandraw Shill, Md Ibrahim, Mehadi Hasan Menon, Tareq Al Muntasir, and Firoj Alam\. 2025\.[TituLLMs: A family of Bangla LLMs with comprehensive benchmarking](https://doi.org/10.18653/v1/2025.findings-acl.1279)\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 24922–24940, Vienna, Austria\. Association for Computational Linguistics\.
- Navigli and Ponzetto \(2010\)Roberto Navigli and Simone Paolo Ponzetto\. 2010\.Babelnet: Building a very large multilingual semantic network\.In*Proceedings of the 48th annual meeting of the association for computational linguistics*, pages 216–225\.
- Ortega\-Martín et al\. \(2023\)Miguel Ortega\-Martín, Óscar García\-Sierra, Alfonso Ardoiz, Jorge Álvarez, Juan Carlos Armenteros, and Adrián Alonso\. 2023\.Linguistic ambiguity analysis in chatgpt\.*arXiv preprint arXiv:2302\.06426*\.
- Popović \(2015\)Maja Popović\. 2015\.chrf: character n\-gram f\-score for automatic mt evaluation\.In*Proceedings of the tenth workshop on statistical machine translation*, pages 392–395\.
- Raganato et al\. \(2017\)Alessandro Raganato, Jose Camacho\-Collados, and Roberto Navigli\. 2017\.[Word sense disambiguation: A unified evaluation framework and empirical comparison](https://aclanthology.org/E17-1010/)\.In*Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers*, pages 99–110, Valencia, Spain\. Association for Computational Linguistics\.
- Raihan and Zampieri \(2025\)Nishat Raihan and Marcos Zampieri\. 2025\.[TigerLLM \- a family of Bangla large language models](https://doi.org/10.18653/v1/2025.acl-short.69)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 887–896, Vienna, Austria\. Association for Computational Linguistics\.
- Saeed et al\. \(2019\)Ali Saeed, Rao Muhammad Adeel Nawab, Mark Stevenson, and Paul Rayson\. 2019\.[A sense annotated corpus for all\-words urdu word sense disambiguation](https://doi.org/10.1145/3314940)\.*ACM Trans\. Asian Low\-Resour\. Lang\. Inf\. Process\.*, 18\(4\)\.
- Shuvo et al\. \(2026\)Md\. Asaduzzaman Shuvo, Mahedi Hasan, Md\. Tashin Parvez, Azizul Haque Noman, and Md\. Shafayet Hossain Ovi\. 2026\.[Polite on the surface, broken in practice: A curated dataset for fixing generation and register failures in low\-resource bangla text generation](https://arxiv.org/abs/2605.22487)\.*Preprint*, arXiv:2605\.22487\.
- Thakur et al\. \(2025\)Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes\. 2025\.[Judging the judges: Evaluating alignment and vulnerabilities in LLMs\-as\-judges](https://aclanthology.org/2025.gem-1.33/)\.In*Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics \(GEM²\)*, pages 404–430, Vienna, Austria and virtual meeting\. Association for Computational Linguistics\.
- Wei et al\. \(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou\. 2022\.[Chain\-of\-thought prompting elicits reasoning in large language models](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf)\.In*Advances in Neural Information Processing Systems*, volume 35, pages 24824–24837\. Curran Associates, Inc\.
- Yu et al\. \(2025\)Haeun Yu, Seogyeong Jeong, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, and Isabelle Augenstein\. 2025\.Entangled in representations: Mechanistic investigation of cultural biases in large language models\.*arXiv preprint arXiv:2508\.08879*\.
- Zhang et al\. \(2019\)Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi\. 2019\.Bertscore: Evaluating text generation with bert\.*arXiv preprint arXiv:1904\.09675*\.
- Zheng et al\. \(2023\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others\. 2023\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.volume 36, pages 46595–46623\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix AAnnotator Profile
The dataset was verified by two native Bangla\-speaking annotators with graduate\-level backgrounds in Computer Science and prior experience in Bangla NLP research\. Both annotators independently reviewed the labels and cultural reasoning associated with each instance\. Disagreements were resolved through discussion using a shared annotation guideline\. The resulting inter\-annotator agreement, measured using Cohen’sκ\\kappa, was 0\.95, indicating almost perfect agreement\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix BResult Analysis Extend
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishB\.1Ablation Study
Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5](https://arxiv.org/html/2607.17828#A2.T5)examines the contribution of reasoning\-aware knowledge distillation by comparing zero\-shot inference, label\-only distillation, and the full QLoRA\-KD framework\. Zero\-shot prompting performs poorly on the CEH task, achieving only 0\.65% Exact Match with high dominant\-meaning bias \(63\.46%\) and hallucination \(52\.9%\), indicating that the pretrained model lacks sufficient cultural grounding\. Introducing label\-only knowledge distillation substantially improves Exact Match to 41\.33% and Macro\-F1 to 48\.12%, demonstrating that supervision from teacher predictions is beneficial\. However, the model still exhibits considerable dominant\-meaning bias \(34\.56%\) and a high hallucination rate \(36\.11%\), suggesting that labels alone are insufficient for learning the underlying cultural distinctions\. In contrast, the full QLoRA\-KD framework, which distills both cultural reasoning and labels, achieves 85\.16% Exact Match and 73\.7% Macro\-F1 while reducing dominant\-meaning bias to 0% and hallucination rate 1\.65%\. These results show that the largest gains are obtained by supervising the reasoning process rather than the prediction labels alone\. These results show that supervising the reasoning process rather than the prediction labels alone enables substantially more reliable cultural disambiguation\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 5:Ablation study\.↑\\uparrowindicates higher is better and↓\\downarrowindicates lower is better\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix CDataset Schema and Example Record
Each instance in the CEH benchmark is stored as a JSON record with the fields listed below\. The body of the paper \(Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1](https://arxiv.org/html/2607.17828#S3.T1)\) shows a condensed view; here we give the complete bilingual schema and a full raw record\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\{ "id": 1516, "word\_bangla": "\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliগগন", "word\_roman": "Gagan", "category": "Intra\-sentential \(Name↔\\leftrightarrowConcept\)", "input": \{ "narrative\_bangla": "\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengali…গগন\-র মনে হলো…", "label\_1": "Name", "label\_2": "Concept" \}, "token\_labels": \{ "token\_1": "Name", "token\_2": "Concept" \}, "cultural\_entanglement\_note": "Gagan carries the concept of the celestial firmament in Bengali cultural imagination while remaining a common given name\."
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 3:A raw CEH benchmark record \(condensed\)\. Bangla fields render in native script; the full schema additionally includes bilingual justifications, cultural\-entanglement paragraphs, and full explanations omitted here for space\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix DCultural Chain\-of\-Thought Prompt
Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2607.17828#A4.F4)shows the full Cultural Chain\-of\-Thought \(C\-CoT\) prompt template\. Unlike standard chain\-of\-thought, it forces the model to argue*both*the name and the competing cultural reading for each occurrence before committing to a label, directly targeting the dominant\-meaning bias\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\[System\] You are an expert in Bangla cultural linguistics\. Some Bangla words are simultaneously a person’s NAME and a culturally meaningful concept, emotion, or spiritual state\. Explain WHY each occurrence carries its meaning, then give the final labels\. \[User\] The Bangla word ‘\\fontspec\_if\_script:nTFbeng\\addfontfeatureScript=Bengali\\fontspec\_if\_language:nTFBEN\\addfontfeatureLanguage=Bengaliগগন’ \(Gagan\) appears TWICE in this sentence with two different meanings\. Sentence: \{narrative\} Reason contrastively for EACH occurrence: For Occurrence 1: \- Argument A: why it could be a Name\. \- Argument B: why it could be an Emotion / Concept / Spiritual / State\. \- Decision: which reading the context supports\. For Occurrence 2: \- Argument A: why it could be a Name\. \- Argument B: why it could be an Emotion / Concept / Spiritual / State\. \- Decision: which reading the context supports\. After reasoning, end with EXACTLY this block: FINAL: Occurrence 1: <label\> Occurrence 2: <label\> Valid labels: Name, Emotion, Concept, Spiritual, Emotional State, Collective State
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 4:The Cultural Chain\-of\-Thought \(C\-CoT\) prompt template\. The model must articulate a Name argument and a competing cultural\-category argument for each occurrence before producing the final labels, counteracting the dominant\-meaning bias\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix ELLM\-as\-a\-Judge Prompt
Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5](https://arxiv.org/html/2607.17828#A5.F5)shows the prompt given to the LLM judge \(GPT\-5\.4\) for evaluating the fine\-tuned models\. To mitigate known judge biases\(Zheng et al\.,[2023](https://arxiv.org/html/2607.17828#bib.bib25); Thakur et al\.,[2025](https://arxiv.org/html/2607.17828#bib.bib21)\), the judge model is distinct from all models under evaluation, and the human annotators were given an identical rubric to ensure the two judges measure the same quantities\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishYou are an expert evaluator of Bangla cultural linguistics, judging a model’s performance on the Culturally Entangled Homograph \(CEH\) task, where one Bangla word appears twice in a sentence with two different meanings \(e\.g\., once as a NAME and once as a CONCEPT, EMOTION, SPIRITUAL, or STATE reading\)\. You are given a batch of items, each with the sentence, gold labels, gold reasoning, and the model output\. Valid labels:Name, Emotion, Concept, Spiritual, Emotional State, Collective State\. Per\-item judgement: \-\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishmatch\_1: does the model’s Occurrence\-1 label equal the gold label? \-\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishmatch\_2: does the model’s Occurrence\-2 label equal the gold label? \-\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishbias: did the model label a gold NAME occurrence as a non\-name \(cultural\) reading? \-\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishhallucinated: did the model output a label outside the valid set, or fail to produce a parseable label? Aggregation \(over N items\): Exact% = 100×\\times\(match\_1 AND match\_2\) / N Partial% = 100×\\times\(exactly one match\) / N Bias% = 100×\\times\(bias\) / \(items with a gold Name\) Hall\.% = 100×\\times\(hallucinated\) / N Output:return ONLY a JSON object, rounded to one decimal \{"model", "hallucination\_pct", "bias\_pct", "exact\_pct", "partial\_pct"\}
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 5:The LLM\-as\-a\-judge prompt \(GPT\-5\.4\)\. The judge scores each output for label match, dominant\-meaning bias, and hallucination, then aggregates into the percentage metrics reported in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2607.17828#S5.T4)\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix FFormula
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishF\.1Average Score
Each CEH instance contains two target occurrences\. A prediction receives one point for each correctly classified occurrence \(maximum two points per instance\)\. The Average Score is computed as
Average Score=Total points earnedNumber of instances×2×100\.\\text\{Average Score\}=\\frac\{\\text\{Total points earned\}\}\{\\text\{Number of instances\}\\times 2\}\\times 100\.
Thus, Average Score reflects occurrence\-level accuracy, whereas Exact Match requires both occurrences in an instance to be classified correctly\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix GHyperparameter Settings
Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6](https://arxiv.org/html/2607.17828#A7.T6)lists the final QLoRA fine\-tuning configuration, applied without modification across all fine\-tuned models\. We arrived at these values by monitoring validation loss: an initial higher\-capacity configuration \(rank 32,α=64\\alpha=64, learning rate2×10−42\\times 10^\{\-4\}, dropout 0\.05, weight decay 0\.01\) overfit within three epochs, so we reduced adapter capacity and learning rate and increased regularisation until the validation curve stabilised\.
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 6:Final QLoRA\-KD fine\-tuning hyperparameters, applied identically across Qwen2\.5\-1\.5B, Llama\-3\.2\-3B, and TituLLM\-1B\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix HPer\-Label Analysis
\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 7:Per\-label precision \(P\), recall \(R\), and F1 for the three fine\-tuned \(QLoRA\-KD\) models\.Nameachieves near\-perfect F1 across all models; the rarest labels \(Emotion,Spiritual\) remain hardest\.To locate where residual errors concentrate after fine\-tuning, we report per\-label precision, recall, and F1 for the three QLoRA\-KD models in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English7](https://arxiv.org/html/2607.17828#A8.T7)\. Two patterns are consistent across all models\. First,Nameis recovered almost perfectly \(F1≥0\.98\\geq 0\.98everywhere\), confirming at the label level that the dominant\-meaning bias central to our study is eliminated by distillation models no longer default away from the name reading\. Second, the errors that remain fall almost entirely on the two rarest labels,EmotionandSpiritual\(F1 as low as 0\.59 and 0\.62\), while the more frequent categories \(Name,Concept,Emotional State\) are handled reliably\. This concentration of error in the long tail accounts for the gap between Micro\- and Macro\-averaged F1, and indicates that the remaining challenge is one of data scarcity in culturally rarer readings rather than a limitation of the method itself\.Similar Articles
Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation
This paper introduces BLADE, a culturally aligned instruction-tuning dataset of 4,196 interaction pairs for fixing honorific failures and pragmatic gaps in multilingual Bangla generation. Fine-tuning models like DeepSeek-8B and LLaMA-3.2-3B on this dataset yields substantial improvements in structural fidelity and honorific alignment.
When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models
This paper introduces CulturalNB, a dataset of Bengali cultural question-answer pairs, and evaluates nine LLMs for cross-lingual cultural bias. Findings show that English prompting increases global narrative substitution and reduces local perspectives, revealing that cultural failures in LLMs are grounding and prioritization issues, not just missing knowledge.
When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates
This paper introduces SemCog Bench, a curated benchmark of 1,858 Arabic-Hebrew word pairs with sentence-level annotations, to evaluate LLMs' ability to distinguish true cognates from false friends and loanwords. Results show high accuracy on true cognates but sharp drops on false friends, highlighting a key limitation in cross-lingual semantic reasoning.
MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media
This paper introduces MultiSoc-4D, a benchmark for diagnosing instruction-induced label collapse in LLMs annotating Bengali social media. It reveals that LLMs systematically prefer fallback labels, leading to under-detection of minority categories like hate speech and sarcasm.
BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts
Researchers introduce BenSyc, the first benchmark for evaluating conversational sycophancy in Bengali social contexts, finding that LLMs struggle to distinguish empathetic support from validation and escalation, achieving only ~61% Macro-F1.