Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

arXiv cs.CL Papers

Summary

This paper investigates why LLMs underperform in Arabic medical tasks, showing via mechanistic analysis that knowledge exists internally but fails to surface, then proposes TLoRA, a targeted low-rank adaptation method that outperforms full-network LoRA on medical QA and introduces a new Arabic clinical dialogue benchmark.

arXiv:2608.00207v1 Announce Type: new Abstract: Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic insight motivates a targeted adaptation strategy: rather than fine-tuning the full network, we propose Targeted Low-Rank Adaptation (TLoRA), restricted to the layer window where cross-lingual representations diverge, upstream of the output layers where the failure manifests. We evaluate TLoRA on multiple-choice medical QA, where our approach outperforms full-network LoRA, zero-shot, and few-shot baselines. We further evaluate it on short-answer generation and multi-turn clinical dialogue, where it performs competitively without the need for task-specific finetuning. We additionally introduce AraClinicDialog, a clinician-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented-language medical LLMs.
Original Article
View Cached Full Text

Cached at: 08/04/26, 07:40 AM

# Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
Source: [https://arxiv.org/html/2608.00207](https://arxiv.org/html/2608.00207)
Chaimae Abouzahir1, Musa Khan1, Hala Ali\-Hassan1, Congbo Ma1, Khaled Saleh2, Yousra Sadqi2, Jihad Mallat2, Walid Al\-Eisawi1, Nizar Habash1, Farah E\. Shamout1 1New York University Abu Dhabi 2Cleveland Clinic Abu Dhabi ca2627@nyu\.edu

###### Abstract

Large Language Models \(LLMs\) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data\. We systematically investigate this assumption via tuned lens probing and causal activation patching, and we find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output\. This mechanistic insight motivates a targeted adaptation strategy: rather than fine\-tuning the full network, we propose Targeted Low\-Rank Adaptation \(TLoRA\), restricted to the layer window where cross\-lingual representations diverge, upstream of the output layers where the failure manifests\. We evaluated TLoRA on multiple\-choice medical QA, where our approach outperforms full\-network LoRA, zero\-shot, and few\-shot baselines\. We further evaluated it on short\-answer generation and multi\-turn clinical dialogue, where it performed competitively without the need for task\-specific finetuning\. We additionally introduce AraClinicDialog, a clinician\-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects\. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented\-language medical LLMs\.

\[ Path = fonts/, Extension = \.otf, UprightFont = \*\-regular, BoldFont = \*\-bold, ItalicFont = \*\-italic, BoldItalicFont = \*\-bolditalic \] \[ Path = fonts/, Extension = \.otf, UprightFont = \*\-regular, BoldFont = \*\-bold, ItalicFont = \*\-italic, BoldItalicFont = \*\-bolditalic \] \[ Path = fonts/, Extension = \.otf, UprightFont = \*\-regular, ItalicFont = \*\-italic \]

Bridging the English\-Arabic Medical Knowledge Gap: Targeted Low\-Rank Adaptation via Causal Layer Selection

Chaimae Abouzahir1, Musa Khan1, Hala Ali\-Hassan1, Congbo Ma1, Khaled Saleh2,Yousra Sadqi2, Jihad Mallat2, Walid Al\-Eisawi1, Nizar Habash1, Farah E\. Shamout11New York University Abu Dhabi2Cleveland Clinic Abu Dhabica2627@nyu\.edu

## 1Introduction

The rise of Large Language Models \(LLMs\) has led to rapid progress in healthcare, with systems demonstrating strong performance across medical question answering, clinical language understanding, and real\-world diagnostic predictionWanget al\.\([2023b](https://arxiv.org/html/2608.00207#bib.bib58)\); Nazi and Peng \([2024](https://arxiv.org/html/2608.00207#bib.bib57)\); Jianget al\.\([2023b](https://arxiv.org/html/2608.00207#bib.bib59)\)\. However, this progress has been defined and measured mainly in English\(Jinet al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib25); Joshiet al\.,[2020](https://arxiv.org/html/2608.00207#bib.bib24)\)\. Pretraining corpora remain heavily English\-dominated, even for nominally multilingual models\(Touvronet al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib23); Scaoet al\.,[2022](https://arxiv.org/html/2608.00207#bib.bib22)\), and the benchmarks used to evaluate medical capability are similarly English\-centric\(Qiuet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib71); Palet al\.,[2022](https://arxiv.org/html/2608.00207#bib.bib14); Jinet al\.,[2019](https://arxiv.org/html/2608.00207#bib.bib55)\)\. Outside English, prior work shows that LLM performance degrades on medical tasksWanget al\.\([2024](https://arxiv.org/html/2608.00207#bib.bib49)\); Jinet al\.\([2023](https://arxiv.org/html/2608.00207#bib.bib25)\)\. Importantly, the mechanisms underlying this degradation remain poorly understood and insufficiently addressed, particularly in low\-resource and underrepresented languages\.

Arabic, spoken by over 400 million people, remains underrepresented in medical LLM training and evaluation due to the scarcity of high\-quality domain\-specific data\(Statista,[2023](https://arxiv.org/html/2608.00207#bib.bib74); Daoudet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib8)\)\. Beyond data availability, Arabic is morphologically rich, and its diglossic structure spans Modern Standard Arabic \(MSA\) and dialectal variants that lack written standards and resources\(Habash,[2010](https://arxiv.org/html/2608.00207#bib.bib66); Moaiadet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib65)\)\. These factors make Arabic a particularly challenging setting for cross\-lingual generalization in medical LLMs\. Despite growing efforts in Arabic medical NLP, performance gaps persist and their causes remain underspecified\(Abouzahiret al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib69)\)\.

Existing approaches to improving Arabic medical performance largely treat this as a data problem\. Domain\-specific efforts, such as BiMediX\(Pieriet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib72)\), rely on translating English medical data and fine\-tuning on the result\. However, this approach inherits English\-centric clinical norms and biases while lacking native Arabic grounding, and recent work shows that fine\-tuning on Arabic medical data can even degrade performance on several benchmarks\(Saadiet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib54)\)\. At the same time, general\-purpose Arabic LLMs, including Jais and ALLaM, are not designed for medical reasoning and perform poorly on medical domain benchmarks\(Daoudet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib5)\)\. Broader multilingual adaptation methods, such as representation alignment and encoder\-bridging approaches\(Zhaoet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib7); Yoonet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib35); Huanget al\.,[2024b](https://arxiv.org/html/2608.00207#bib.bib6)\), offer only limited gains and assume the bottleneck is data quantity or language coverage, leaving the model’s internal failure mechanism unexamined\.

Crucially, all of these approaches treat the model as a black box: none ask where in the model the failure occurs or why Arabic queries fail to elicit knowledge the model demonstrably possesses\. We identify that Arabic medical failure in LLMs is a knowledge\-routing breakdown rather than a knowledge deficit\. More specifically, the model answers correctly in English on approximately 30% of questions where it fails on an identical Arabic query \(Mistral\-Small\-3\.2\-24B on MedAraBench; see §[3](https://arxiv.org/html/2608.00207#S3)\)\. Using tuned lens probing, causal activation patching, and KL divergence profiling, we localize this breakdown to a specific layer window and derive a targeted adaptation strategy, whose design follows directly from the mechanistic evidence\. Our contributions are as follows:

- •To the best of our knowledge, we provide the first mechanistic analysis of Arabic medical failure in LLMs, identifying aknowledge\-routing failurespecific to a layer window that motivates our adaptation design\.
- •We proposelayer\-targeted LoRA, denoted as TLoRA, with a cross\-lingual alignment objective whose window and probe layer are both derived from mechanistic evidence\.
- •We presentAraClinicDialog, a new clinician\-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects\.

## 2Related Work

### 2\.1LLMs in Healthcare

Clinical question answering benchmarks have become the primary measure of progress for medical LLMs, with proprietary systems such as GPT\-4 and Med\-PaLM 2 now achieving near\-expert performance and medical\-domain open models such as Meditron increasingly approaching the same ceiling\(Singhalet al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib19); Noriet al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib11); Singhalet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib51); Chenet al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib20)\)\. These benchmarks, however, are designed around English\-language clinical data and evaluation conventions, establishing a performance standard whose underlying assumptions are English\-first\. Efforts to extend evaluation beyond English exist but remain constrained, relying on translation from other source languages\(Alonsoet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib67)\)\.

While MCQA benchmarks offer a scalable and reproducible measure of medical knowledge, they remain an insufficient basis for assessing clinical competence, as the format is susceptible to artifacts and structural cues\(Cocchieriet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib18)\)\. Recent work has therefore extended evaluation to open\-ended generation and clinical dialogue\. For example, Med\-PaLM 2 introduced physician preference judgments over long\-form answers, AMIE evaluated diagnostic reasoning through blinded OSCE\-style consultations, and HealthBench and MedHELM formalized rubric\-based scoring over free\-form outputs\(Singhalet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib51); Tuet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib21); Aroraet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib64); Bediet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib61)\)\. Like their MCQA counterparts, however, these frameworks are developed and validated primarily in English\.

Cross\-lingual degradation in medical LLMs has been documented consistently across languages and evaluation paradigms\. In MCQA settings, MedExpQA reports roughly a ten\-point accuracy drop across European languages\(Alonsoet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib67)\), while multilingual medical models such as Apollo and Apollo\-MoE still show a fifteen\-point gap on Arabic after multilingual adaptation\.\(Wanget al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib49); Zhenget al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib48)\)\. Beyond MCQA, Jin et al\.Jinet al\.\([2024](https://arxiv.org/html/2608.00207#bib.bib50)\)document correctness, consistency, and verifiability disparities in open\-ended healthcare queries\. In each case, the degradation is attributed to data scarcity or insufficient multilingual pretraining, leaving open whether the failure instead reflects an inability to retrieve knowledge already present in the model\.

### 2\.2Arabic Medical LLMs

At the intersection of Arabic language and clinical medicine, LLM development has been largely absent\. Models trained specifically on Arabic such as Jais and Allam demonstrate strong performance across standard NLP tasks but underperform on clinical benchmarks\(Daoudet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib5)\)\. A first wave of dedicated work has begun to address this with new benchmarks establishing evaluation grounded in native Arabic medical material\(Daoudet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib5),[2026](https://arxiv.org/html/2608.00207#bib.bib8)\), and BiMediX representing the first purpose\-built Arabic medical LLM through bilingual fine\-tuning on translated clinical data\(Pieriet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib72)\)\.

Model adaptation efforts share some limitations\. First, training and evaluation data are derived primarily from English translation pipelines rather than native Arabic clinical sources\. All existing work frames the Arabic medical capability gap as a problem of data coverage and bilingual supervision rather than one of representational access\. One studyAbouzahiret al\.\([2026](https://arxiv.org/html/2608.00207#bib.bib69)\)confirms through cross\-lingual empirical analysis that the gap is consistent and domain\-specific, yet stops short of a mechanistic account\.

### 2\.3Cross\-Lingual Transfer and Representation Alignment

The mechanistic basis for this gap has been established in the broader multilingual literature\. Multilingual LLMs process non\-English inputs through an English\-centered latent pathway, with cross\-lingual prediction failures concentrating precisely at the middle layers where alignment to English breaks down\(Wendleret al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib46); Schutet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib47)\)\. This failure is causal, patching English hidden states at those layers recovers correct predictions in the majority of failure cases, and disproportionately affects knowledge retrieval rather than abstract reasoning, implicating representational access to stored knowledge as the primary bottleneck\(Ravisankaret al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib44); Huet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib40); Iferganet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib42)\)\. This misalignment is measurable through alignment scores between parallel representations\(Kargaranet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib39); Hämmerlet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib38)\), and is especially severe in typologically distant, morphologically complex languages, making Arabic a critical and underexplored test case\.

The dynamics of cross\-lingual knowledge transfer in domain adaptation have only recently been studied\. For example, Kobayashi et al\.Kobayashiet al\.\([2025](https://arxiv.org/html/2608.00207#bib.bib26)\)show that English biomedical corpora support low\-resource Japanese medical adaptation but that transfer depends sensitively on corpus composition, while Zhao et al\.Zhaoet al\.\([2026](https://arxiv.org/html/2608.00207#bib.bib37)\)trace how domain facts are memorized and generalized during multilingual medical adaptation, finding that transfer remains challenging even under high\-quality bilingual training\. Both studies, however, analyze transfer at the behavioral level, through performance curves and learning dynamics, and examine a relatively high\-resource language pair\. Our work extends this line to Arabic medicine, a typologically distant and morphologically complex setting, and connects behavioral failure to the representational mechanism: whether Arabic hidden states remain aligned enough to access English\-anchored medical knowledge\.

![Refer to caption](https://arxiv.org/html/2608.00207v1/x1.png)Figure 1:Mechanistic motivation for targeted adaptation of Mistral\-Small\-3\.2\-24B:\(a\)tuned lens probing,\(b\)causal activation patching, and\(c\)cross\-lingual KL divergence profile\.Existing adaptation methods fall into three families, none of which addresses the representational root cause identified above\. Relearning methods\(Cuiet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib33); Wanget al\.,[2023a](https://arxiv.org/html/2608.00207#bib.bib32)\)improve target\-language performance but erode cross\-lingual structure, inducing catastrophic forgetting in zero\-shot generation\(Vuet al\.,[2022](https://arxiv.org/html/2608.00207#bib.bib31)\)and proximity\-dependent forgetting around injected medical concepts\(Liu and Niehues,[2025](https://arxiv.org/html/2608.00207#bib.bib30); Zhouet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib29)\)\. External bridge methods route non\-English reasoning through cross\-lingual prompts, trainable bridging parameters, or auxiliary multilingual encoders\(Huanget al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib36); Yoonet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib35); Huanget al\.,[2024b](https://arxiv.org/html/2608.00207#bib.bib6)\), avoiding full relearning but introducing external components that require additional training and do not leverage the model’s existing English domain representations\.

To the best of our knowledge, no prior method combines a mechanistic diagnosis of domain\-specific knowledge access failure with a targeted adaptation strategy that uses the model’s own English hidden states as an alignment anchor, in Arabic or in any other language\.

## 3Mechanistic Motivation

We initially evaluate Mistral\-Small\-3\.2\-24B on the MedAraBench test set\(Daoudet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib8)\), where it achieves 57\.8% in English and 50\.8% in Arabic\. The English set is a machine translation of the original Arabic questions via the Google Translate API, constructed specifically for this diagnostic comparison\. Among English\-correct questions, 29\.6% fail on the identical Arabic query\. This implies that the model demonstrably possesses the knowledge but fails to surface it\. We ask if this reflects absent knowledge or a failure of access\.

To answer this, we use three complementary mechanistic probes\. Tuned\-lens probing\(Belroseet al\.,[2023](https://arxiv.org/html/2608.00207#bib.bib2)\)and causal activation patching\(Menget al\.,[2022](https://arxiv.org/html/2608.00207#bib.bib3)\)test whether the correct answer is encoded mid\-network and whether replacing Arabic representations with English ones recovers it, respectively\. A cross\-lingual KL divergence profile\(Wendleret al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib46)\)then locates where English and Arabic diverge, setting the adaptation window used in §[4\.1](https://arxiv.org/html/2608.00207#S4.SS1)\.

##### Tuned Lens\.

We applied tuned lens probing to track the correct answer probability layer\-by\-layer across parallel English and Arabic forward passes\. Figure[1](https://arxiv.org/html/2608.00207#S2.F1)\(a\) shows the correct answer accumulates through intermediate Arabic layers, reaching English levels by mid\-network, then collapses before the output; the both\-wrong quadrant remains flat throughout\. This rules out a general decoding deficit such that the knowledge is encoded but not committed\.

##### Causal Activation Patching\.

We apply causal activation patching to verify this causally, replacing Arabic hidden states with English counterparts one layer at a time and measuring gap recovery\. Figure[1](https://arxiv.org/html/2608.00207#S2.F1)\(b\) shows that recovery exceeds 80% at a single injection layerLpatch=24L\_\{\\mathrm\{patch\}\}=24, with peak recovery reaching 171%\. Layers below this threshold yield near\-zero recovery\. This suggests that the failure is both causal and precisely localized\.

##### KL Divergence Profile\.

We compute the cross\-lingual KL divergence profile to identify the upstream source of the breakdown\. Figure[1](https://arxiv.org/html/2608.00207#S2.F1)\(c\) points to divergence rising sharply atLKL=34L\_\{\\mathrm\{KL\}\}=34, after which English and Arabic trajectories become irreconcilable\. Patching and KL profiling together identify two mechanistically salient boundary layers,Lpatch=24L\_\{\\text\{patch\}\}=24andLKL=34L\_\{\\text\{KL\}\}=34, marking, respectively, the onset of causal recoverability and the onset of sharp cross\-lingual divergence, and defining the boundary signals for the candidate adaptation windows in §[4\.1](https://arxiv.org/html/2608.00207#S4.SS1)\.

## 4Methodology

Our method has two mechanistically\-grounded design choices\. The first identifies which layers to apply LoRA to, and the second determines at which layer to apply the cross\-lingual alignment\. Both are determined from the model’s own internal signals, the causal boundaryLpatchL\_\{\\text\{patch\}\}and the divergence onsetLKLL\_\{\\text\{KL\}\}identified in §[3](https://arxiv.org/html/2608.00207#S3), before any training begins and fixed thereafter\.

### 4\.1Mechanistically\-Guided Window Selection

The two boundary layers identified in §[3](https://arxiv.org/html/2608.00207#S3)partition the network into three contiguous regions, generating a small, exhaustive set of candidate adaptation windows shown in Table[1](https://arxiv.org/html/2608.00207#S4.T1)\.

WindowLayersRegionW1W\_\{1\}L1–LpatchL\_\{\\text\{patch\}\}below causal boundaryW2W\_\{2\}LpatchL\_\{\\text\{patch\}\}–LmaxL\_\{\\text\{max\}\}causal windowW3W\_\{3\}L1–LKLL\_\{\\text\{KL\}\}below divergence onsetW4W\_\{4\}LKLL\_\{\\text\{KL\}\}–LmaxL\_\{\\text\{max\}\}active zone onlyW5W\_\{5\}L1–LmaxL\_\{\\text\{max\}\}full modelTable 1:Candidate LoRA adaptation windows\.Rather than searching arbitrarily over layer subsets, this partition ensures candidates correspond to meaningful network regions defined by causal evidence\. We run an independent learning\-rate sweep for each window under the full training objective \(Appendix[D](https://arxiv.org/html/2608.00207#A4)\) and select the window with the best held\-out performance\. Becauseβ∗\\beta^\{\*\}is derived from initialisation losses before training \(see §[4\.3](https://arxiv.org/html/2608.00207#S4.SS3)\), and LoRA weights are zero at initialisation regardless of window,β∗\\beta^\{\*\}is identical across all five sweeps, ensuring a fair comparison\. The causal boundaryLpatchL\_\{\\text\{patch\}\}is defined as:

Lpatch=min⁡\{ℓ:recovery​\(ℓ\)≥τpatch\}L\_\{\\text\{patch\}\}\\;=\\;\\min\\bigl\\\{\\ell:\\mathrm\{recovery\}\(\\ell\)\\geq\\tau\_\{\\text\{patch\}\}\\bigr\\\}\(1\)withτpatch=0\.5\\tau\_\{\\text\{patch\}\}=0\.5\.

### 4\.2KL Probe Layer

Independently of window selection, we determine the layer at which to apply cross\-lingual alignment pressure from the model’s own divergence profile\. We run parallel\-bilingual data forward passes, and at each layerℓ\\ellproject the final\-token hidden state through the frozen RMSNorm and unembedding matrix to obtain vocabulary distributionsp^ℓAr\\hat\{p\}\_\{\\ell\}^\{\\text\{Ar\}\}andp^ℓEn\\hat\{p\}\_\{\\ell\}^\{\\text\{En\}\}\. The per\-layer divergence is:

KL​\(ℓ\)=𝔼\(xAr,xEn\)​\[DKL​\(p^ℓEn∥p^ℓAr\)\]\\mathrm\{KL\}\(\\ell\)\\;=\\;\\mathbb\{E\}\_\{\(x^\{\\text\{Ar\}\},\\,x^\{\\text\{En\}\}\)\}\\\!\\left\[D\_\{\\text\{KL\}\}\\\!\\left\(\\hat\{p\}\_\{\\ell\}^\{\\text\{En\}\}\\,\\big\\\|\\,\\hat\{p\}\_\{\\ell\}^\{\\text\{Ar\}\}\\right\)\\right\]\(2\)
The profile partitions naturally into three zones: neutral, ramp, and active, defined by the meanμ\\muand standard deviationσ\\sigmaofKL​\(ℓ\)\\mathrm\{KL\}\(\\ell\)across layers\. The probe layer is the first layer to enter the active zone:

LKL=min⁡\{ℓ:KL​\(ℓ\)≥μ\+σ\}L\_\{\\text\{KL\}\}\\;=\\;\\min\\bigl\\\{\\ell:\\mathrm\{KL\}\(\\ell\)\\geq\\mu\+\\sigma\\bigr\\\}\(3\)

### 4\.3Training Objective

We train with a combined cross\-entropy and cross\-lingual alignment loss:

ℒ=ℒCE\+β∗​ℒalign\\mathcal\{L\}\\;=\\;\\mathcal\{L\}\_\{\\text\{CE\}\}\\;\+\\;\\beta^\{\*\}\\,\\mathcal\{L\}\_\{\\text\{align\}\}\(4\)
whereℒalign\\mathcal\{L\}\_\{\\text\{align\}\}is the KL divergence between English and Arabic logit\-lens distributions atLKLL\_\{\\text\{KL\}\}:

ℒalign=DKL​\(p^LKLEn∥p^LKLAr\)\\mathcal\{L\}\_\{\\text\{align\}\}\\;=\\;D\_\{\\text\{KL\}\}\\\!\\left\(\\hat\{p\}\_\{L\_\{\\text\{KL\}\}\}^\{\\text\{En\}\}\\,\\big\\\|\\,\\hat\{p\}\_\{L\_\{\\text\{KL\}\}\}^\{\\text\{Ar\}\}\\right\)\(5\)
We calibrateβ\\betafrom the model at initialization\. A single forward pass measures the initial lossesℒCE\(0\)\\mathcal\{L\}\_\{\\text\{CE\}\}^\{\(0\)\}andℒalign\(0\)\\mathcal\{L\}\_\{\\text\{align\}\}^\{\(0\)\}, and we set:

β∗=ℒCE\(0\)ℒalign\(0\)\\beta^\{\*\}\\;=\\;\\frac\{\\mathcal\{L\}\_\{\\text\{CE\}\}^\{\(0\)\}\}\{\\mathcal\{L\}\_\{\\text\{align\}\}^\{\(0\)\}\}\(6\)
The alignment term requires parallel Arabic–English pairs at every training step\. Arabic provides the student distribution, while English, computed without gradient, provides the teacher\. The CE term trains on Arabic medical MCQs only\.

## 5Experimental Setup

### 5\.1Implementation Details

Unless otherwise specified, all models are evaluated zero\-shot, with task instructions provided through prompting \(Appendix[A](https://arxiv.org/html/2608.00207#A1)\)\. For all trained adaptation methods, the MedAraBench training split \(17,860 training examples and 1,987 validation examples; stratified 90/10 split\) is the sole training source\. For TLoRA and LoRA v2, the alignment objective uses the same Arabic training examples machine\-translated into English via Google Translate API to form parallel pairs\. The task and alignment losses are optimized jointly at every training step\.

Applying the layer\-selection criteria \(§[3](https://arxiv.org/html/2608.00207#S3)\) to Mistral\-Small\-3\.2\-24B identifies two boundaries:Lpatch=24L\_\{\\text\{patch\}\}=24andLKL=34L\_\{\\text\{KL\}\}=34\(μ=0\.88\\mu=0\.88,σ=1\.04\\sigma=1\.04,τ=1\.92\\tau=1\.92\)\. These separate the network into five candidate adaptation windows \(Table[3](https://arxiv.org/html/2608.00207#S6.T3)\), from which the optimal window is selected via held\-out performance on the training split\.

### 5\.2Evaluation Tasks

We evaluate across three tasks of increasing complexity: multiple\-choice question answering \(MCQA\), short answer generation, and multi\-turn clinical dialogue\. MCQA is the primary task our method is optimized for, while the remaining two serve as generalization probes, testing whether adaptation gains transfer without catastrophic forgetting\.

#### 5\.2\.1Multiple\-Choice Question Answering

We evaluate on four native Arabic MCQA benchmarks: MedArabiQ\(Daoudet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib5)\), MedAraBench\(Daoudet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib8)\), and the biology and medicine subsets of ArabicMMLU\(Kotoet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib9)\)and AraSTEM\(Boussahaet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib10)\)\. All are derived from regional medical examinations in MSA\. We retain only questions with at least four answer options for consistency\. Models predict the correct letter and are evaluated using exact\-match accuracy\. Full dataset statistics are in Appendix[S1](https://arxiv.org/html/2608.00207#A1.T1)\.

In\-DomainOut\-of\-Domain \(OOD\)CategoryModelMedAraBenchMedArabiQAraSTEMArabicMMLURandomRandom Baseline24\.020\.021\.129\.7Closed\-SourceGeneral\-PurposeGPT\-5\.2\(OpenAI,[2025](https://arxiv.org/html/2608.00207#bib.bib76)\)71\.078\.988\.071\.6Gemini\-2\.5\-Flash\(Comaniciet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib96)\)70\.279\.085\.273\.8Claude\-Opus\-4\.6\(Anthropic,[2026](https://arxiv.org/html/2608.00207#bib.bib77)\)74\.382\.189\.474\.1Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.3\(Jianget al\.,[2023a](https://arxiv.org/html/2608.00207#bib.bib82)\)27\.225\.226\.631\.3Llama\-3\.1\-8B\-Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib95)\)37\.732\.637\.838\.6Mistral\-Small\-3\.2\-24B\(Mistral AI,[2025](https://arxiv.org/html/2608.00207#bib.bib78)\)52\.651\.562\.655\.2Gemma\-3\-27B\-IT\(Teamet al\.,[2025b](https://arxiv.org/html/2608.00207#bib.bib81)\)51\.647\.364\.253\.8Llama\-3\.3\-70B\-Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib95)\)46\.840\.055\.051\.0DeepSeek\-V3\.2\(DeepSeek\-AIet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib83)\)63\.466\.376\.966\.0Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat\(Anwaret al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib79)\)37\.930\.540\.949\.3ALLaM\-7B\-Instruct\-Preview\(Bariet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib84)\)43\.836\.846\.146\.2Aya\-Expanse\-8B\(Danget al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib85)\)39\.634\.743\.142\.6SILMA\-9B\-Instruct\-v1\.0\(SILMA AI,[2026](https://arxiv.org/html/2608.00207#bib.bib80)\)42\.234\.042\.843\.9Falcon\-H1\-7B\-Instruct\(Zuoet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib86)\)46\.740\.048\.645\.6Fanar\-1\-9B\-Instruct\(Teamet al\.,[2025a](https://arxiv.org/html/2608.00207#bib.bib87)\)40\.833\.644\.347\.0Medical\-DomainMedGemma\-27B\-Text\-IT\(Sellergrenet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib88)\)51\.951\.563\.851\.3Meditron\-3\-70B\(Sallinenet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib89)\)55\.950\.566\.455\.9Llama\-3\-Med42\-70B\(Christopheet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib90)\)39\.029\.062\.851\.9AdaptationMethodsMistral \+ Few\-Shot \(k=5\)\(Brownet al\.,[2020](https://arxiv.org/html/2608.00207#bib.bib91)\)53\.746\.361\.557\.6Mistral \+ AUTOCAP\(Zhanget al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib92)\)59\.754\.068\.459\.6Mistral \+ English Translation57\.854\.071\.556\.8Mistral \+ MindMerger\(Huanget al\.,[2024a](https://arxiv.org/html/2608.00207#bib.bib93)\)24\.022\.020\.325\.1BiMedix \(Zero\-shot\)\(Pieriet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib72)\)28\.628\.037\.529\.3Mistral \+ LoRA\(Huet al\.,[2021](https://arxiv.org/html/2608.00207#bib.bib94)\)42\.642\.045\.536\.3Mistral \+ LoRA v2 \(with KL\)61\.955\.060\.760\.3Mistral \+ TLoRA \(Ours\)62\.160\.065\.161\.6

Table 2:Multiple\-Choice QA results across four Arabic medical benchmarks\. Exact\-match accuracy \(%\) is reported on MedAraBench \(in\-domain\) and MedArabiQ, AraSTEM\-medicine, and ArabicMMLU\-biology \(out\-of\-domain benchmarks\)\. Bold indicates the best result within each group per column, and underline points to second best result\. The Mistral variant used as backbone for the adaptation methods is Mistral\-Small\-3\.2\-24B\-Instruct\-2506\.
#### 5\.2\.2Short Answer Generation

We repurpose MedAraBench for free\-text generation\. Starting from4,9894,989test examples, we remove MCQA artifacts \(e\.g\., “All of the above”\) and rewrite remaining questions into standalone open\-ended queries via an LLM\-assisted pipeline, yielding 940 examples\. We refer to this subset asMedAraBench\-OE\(Open\-Ended\)\. Full preprocessing details are in Appendix[B\.1](https://arxiv.org/html/2608.00207#A2.SS1)\.

Standard n\-gram metrics are poorly suited for Arabic medical generation due to morphological variability and short reference answers\. We therefore use LLM\-as\-a\-judge \(GPT\-5\.2\) as our primary metric, validated against human judgements on a 100\-example subset: LLM\-judge accuracy achieves Pearson=0\.978=0\.978and Spearman=0\.982=0\.982with human labels, substantially outperforming BERTScore\-F1 \(Pearson=0\.802=0\.802, Spearman=0\.715=0\.715\)\. BERTScore\-F1 \(AraBERTv2\) is reported as a secondary metric\. Full reliability analysis is in Appendix[C](https://arxiv.org/html/2608.00207#A3)\.

#### 5\.2\.3Multi\-Turn Clinical Dialogue

We introduceAraClinicDialog, constructed by three native Arabic\-speaking physicians from a multi\-specialty hospital\. FollowingAroraet al\.\([2025](https://arxiv.org/html/2608.00207#bib.bib64)\), physicians authored 100 clinical cases spanning nine organ systems, expanded into multi\-turn dialogues using Claude\-Opus\-4\.6 and verified by clinicians\. Each physician authored an independent reference response for the final turn, while a fourth clinician assessed agreement, achieving 87% inter\-annotator agreement\. Full construction details are in Appendix[B\.2](https://arxiv.org/html/2608.00207#A2.SS2)\.

Dataset construction proceeded in four stages: \(a\) Case template authoring: each physician wrote a clinical scenario spanning nine organ systems, yielding 100 cases \(Appendix[S4](https://arxiv.org/html/2608.00207#A2.T4)\)\. \(b\) Dialogue generation: each template was expanded into a 3–8 turn MSA patient\-assistant dialogue using Claude\-Opus\-4\.6, with the final assistant turn withheld \(Appendix[S9](https://arxiv.org/html/2608.00207#A2.F9)\)\. \(c\) Reference collection: two clinicians per case independently authored the withheld final turn via a structured form\. Audio responses were transcribed with Whisper\-Large\-V3 and manually reviewed \(Appendix[S6](https://arxiv.org/html/2608.00207#A2.T6)\)\. \(d\) Dialect translation: all 100 MSA dialogues were translated into four regional Arabic dialects, Emirati, Jordanian, Moroccan and Egyptian, using GPT\-5\.2 and reviewed by two native speakers per dialect \(Appendix[S8](https://arxiv.org/html/2608.00207#A2.T8)\)\.

Models are evaluated on the final dialogue turn using correct/incorrect judgement against the clinician\-written reference, anchored in a primary reasoning objective authored by clinicians, and assessed by LLM\-as\-a\-judge\. This framing is consistent with Task 2 and targets factual correctness of the concluding clinical response rather than dialogue quality overall\. Given the validated alignment between LLM\-judge and human labels established in §[5\.2\.2](https://arxiv.org/html/2608.00207#S5.SS2.SSS2), we use the same metrics\.

![Refer to caption](https://arxiv.org/html/2608.00207v1/x2.png)Figure 2:Short answer generation performance across model categories, evaluated on MedArabench\-OE\. For each model, we report BERTScore\-F1 \(araBERTv2\) and LLM\-as\-a\-judge Correct %\. The star \(⋆\\star\) denotes our proposed method\.![Refer to caption](https://arxiv.org/html/2608.00207v1/x3.png)Figure 3:Multi\-Turn Clinical Dialogue performance across model categories, evaluated on AraClinicDialog \(MSA and dialect variants\)\. For each model, we report BERTScore\-F1 \(AraBERTv2\) and LLM\-as\-a\-judge Correct %\. The star \(⋆\\star\) denotes our proposed method\.

### 5\.3Baselines

Adaptation baselines include few\-shot prompting \(k=5k=5\)\(Brownet al\.,[2020](https://arxiv.org/html/2608.00207#bib.bib91)\), AUTOCAP\(Zhanget al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib92)\), purpose\-built Arabic medical model BiMedix\(Pieriet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib72)\), full LoRA\(Huet al\.,[2021](https://arxiv.org/html/2608.00207#bib.bib94)\), and MindMerger\(Huanget al\.,[2024a](https://arxiv.org/html/2608.00207#bib.bib93)\)\. The latter two are trained on the MedAraBench training split, while the remaining baselines are prompt\-based\. MindMerger is evaluated on MCQA only, as it struggled to produce the required output format on generation and dialogue tasks\.

## 6Results

### 6\.1MCQA

Table[2](https://arxiv.org/html/2608.00207#S5.T2)shows results across the four Arabic medical MCQA benchmarks\. Closed\-source models form a clear upper tier, with macro\-averages between 77\.1% and 80\.0%\. Among open\-source general\-purpose models, performance varies substantially: Mistral\-Small\-3\.2\-24B \(55\.5%\) and DeepSeek\-V3\.2 \(68\.2%\) substantially outperform smaller models, though scale alone is not predictive: Llama\-3\.3\-70B underperforms Mistral\-Small\-3\.2\-24B despite its larger size\. Notably, Arabic/multilingual models do not consistently outperform general open\-source models, averaging in the low\-to\-mid 40% range; Arabic\-specific pretraining alone does not resolve the knowledge\-access gap, consistent with prior work\(Daoudet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib5)\)\.

Among adaptation methods, MindMerger degrades substantially below the zero\-shot Mistral baseline, and few\-shot prompting offers no reliable gain\. LoRA v2 improves over zero\-shot \(59\.5% vs\. 55\.5%\) but applies adaptation uniformly across all 40 layers\. Our method, TLoRA, achieves the highest macro\-average among adaptation methods \(62\.2%\), outperforming LoRA \(w/o KL\) by 20\.6 points and LoRA v2 \(with KL\) by 2\.7 points despite training fewer parameters\. TLoRA and LoRA v2 converge to a similar ceiling on MedAraBench \(in\-domain\): 95% confidence intervals overlap substantially \(TLoRA \[60\.8, 63\.4\] vs\. LoRA v2 \[60\.5, 63\.2\],p=0\.74p=0\.74\)\. The out\-of\-domain comparisons are more diagnostic, where TLoRA’s advantage is clearest on MedArabiQ and ArabicMMLU\. This advantage holds under controlled conditions: fixing the learning rate to the same value across all windows and removing the KL alignment term both preserve the L1–34 ranking, ruling out optimisation and loss formulation as confounds \(Appendix[F\.2](https://arxiv.org/html/2608.00207#A6.SS2),[F\.3](https://arxiv.org/html/2608.00207#A6.SS3)\)\. This advantage holds under controlled conditions: fixing the learning rate to the same value across all windows and removing the KL alignment term both preserve the L1–34 ranking, ruling out optimisation and loss formulation as confounds \(Appendix[F\.2](https://arxiv.org/html/2608.00207#A6.SS2),[F\.3](https://arxiv.org/html/2608.00207#A6.SS3)\)\.

MCQAGen\.Dial\.WindowMedAraBenchMedarabiQMMLU\-BioAraSTEMBERTScoreLLM JudgeBERTScoreLLM JudgeZero\-shot52\.651\.555\.262\.660\.030\.158\.444\.0Full LoRA \(W5W\_\{5\}: L1–40\)61\.955\.060\.360\.753\.825\.556\.068\.0Targeted \(W1W\_\{1\}: L1–24\)60\.756\.060\.263\.253\.429\.057\.272\.0Targeted \(W2W\_\{2\}: L24–40\)53\.542\.052\.958\.552\.638\.148\.376\.0Targeted \(W3W\_\{3\}: L1–34\)62\.160\.061\.665\.157\.929\.752\.564\.0Targeted \(W4W\_\{4\}: L34–40\)47\.333\.051\.354\.650\.626\.548\.375\.0

Table 3:Window ablation\. All targeted variants useℒCE\+β∗​ℒalign\\mathcal\{L\}\_\{\\text\{CE\}\}\+\\beta^\{\*\}\\mathcal\{L\}\_\{\\text\{align\}\}with probe layerLKL=34L\_\{\\text\{KL\}\}=34and bilingual training data\. Short Answer Generation \(Gen\.\) and Multi\-Turn Clinical Dialogue \(Dial\.\) scores are BERTScore\-F1 and LLM\-judge accuracy \(%\)\. Bold indicates the best result among trained variants for each column, while underline indicates the second\-best result\.
### 6\.2Short Answer Generation

Figure[2](https://arxiv.org/html/2608.00207#S5.F2)reports results on MedAraBench\-OE, a generalization task no adaptation method was optimized for\. Closed\-source models lead, with LLM\-judge accuracy between 60\.5% and 67\.6%; smaller open\-source and Arabic/multilingual models largely fail to produce coherent free\-text responses\. Among adaptation methods, MCQA fine\-tuning generally degrades generation performance relative to zero\-shot: Full LoRA drops from 30\.1% to 25\.5%\. Our method retains 29\.7% \[26\.8%, 32\.6%\], statistically indistinguishable from zero\-shot \(30\.1% \[27\.2%, 33\.0%\],p=0\.86p=0\.86\), and the closest to zero\-shot among trained variants, suggesting that targeted adaptation does not induce the generation\-forgetting observed in broader fine\-tuning\.

TLoRA is selected by MCQA performance while TLoRA \(optimal\), also shown in Figure[2](https://arxiv.org/html/2608.00207#S5.F2), instead uses the window best suited to this task and reaches higher generation accuracy \(38\.1%\)\. Ablation results indicate that the KL alignment term contributes specifically to this preservation: removing it while keeping the same window reduces generation BERTScore substantially, whereas MCQA accuracy is largely unchanged \(Appendix Table[S20](https://arxiv.org/html/2608.00207#A6.T20)\)\. Complete results for Short Answer Generation on MedAraBench\-OE are provided in Appendix[E\.1](https://arxiv.org/html/2608.00207#A5.SS1)\.

### 6\.3Multi\-Turn Clinical Dialogue

Figure[3](https://arxiv.org/html/2608.00207#S5.F3)reports results on AraClinicDialog’s MSA and dialect variants\. Unlike Short Answer Generation, adaptation yields substantial gains on dialogue over zero\-shot\. Both full LoRA and TLoRA improve, with our method remaining competitive despite training exclusively on MCQA data\. AUTOCAP, strong on MCQA, collapses on dialogue, suggesting its gains are format\-sensitive\. Closed\-source models remain the upper bound \(complete per\-dialect results are reported in Appendix[E\.2](https://arxiv.org/html/2608.00207#A5.SS2), Tables[S14](https://arxiv.org/html/2608.00207#A5.T14)–[S18](https://arxiv.org/html/2608.00207#A5.T18)\)\.

### 6\.4Ablations

Table[3](https://arxiv.org/html/2608.00207#S6.T3)reports performance across all five candidate adaptation windows\. On MCQA, L1–34 achieves the highest macro\-average, outperforming full\-network LoRA and all other windows\. The result supports the mechanistic hypothesis: restricting adaptation to layers below the divergence onset outperforms both broader and narrower windows\. On dialogue, no trained variant falls below zero\-shot\. The L1–34 advantage is consistent across loss functions and learning\-rate conditions\. Detailed comparisons against CE\-only and fixed\-LR baselines are provided in Appendix[F\.2](https://arxiv.org/html/2608.00207#A6.SS2),[F\.3](https://arxiv.org/html/2608.00207#A6.SS3)\. Sensitivity of results to the KL probe layer threshold is reported in Appendix[F\.1](https://arxiv.org/html/2608.00207#A6.SS1)\. We adopt L1–34 as our method for all subsequent comparisons\.

## 7Discussion

##### MCQA\.

The performance ordering across windows is directionally consistent with the mechanistic hypothesis: gains scale with the degree to which the adaptation window covers the identified routing failure region, rather than with the number of layers adapted\. CE\-only training across all windows \(Appendix[F\.2](https://arxiv.org/html/2608.00207#A6.SS2)\) shows that L1–34 retains its MCQA lead withoutℒalign\\mathcal\{L\}\_\{\\text\{align\}\}, confirming that window placement is the primary driver, while the alignment term contributes selectively to generation quality rather than classification accuracy\. TLoRA also narrows the access gap identified in Section[3](https://arxiv.org/html/2608.00207#S3), reducing the English\-correct/Arabic\-incorrect failure rate from 29\.6% \(zero\-shot\) to 19\.0%, slightly ahead of Full LoRA’s 20\.2%\. The same diagnostic pipeline applied to a second model family, Llama\-3\.1\-8B\-Instruct, identifies a different window that matches or exceeds full\-network LoRA on MCQA \(Appendix[I](https://arxiv.org/html/2608.00207#A9)\), suggesting the approach is not specific to Mistral\.

MindMerger’s degradation below the zero\-shot baseline is architecturally informative: having been fine\-tuned on bilingual medical terminology prior to MedAraBench training, domain mismatch can be excluded as an explanation\. Encoder augmentation with a frozen backbone intervenes at the input representation level, leaving the late\-layer routing deficit entirely unaddressed\. Few\-shot prompting is similarly ineffective for the same underlying reason: surface\-level context cannot recover knowledge that is representationally accessible but fails to route to the output in Arabic\.

##### Generation and Dialogue\.

MCQA fine\-tuning broadly degrades generation, likely through format overfitting; our method is least affected, suggesting targeted adaptation preserves more general capability\. On dialogue, the pattern is reversed: adaptation helps substantially\. The high dialogue scores are partly explained by the evaluation design: the judge is anchored in clinician\-authored reasoning objective field \(Appendix Figure[S5](https://arxiv.org/html/2608.00207#A2.T5)\), rewarding clinical accuracy over fluency\. Manual inspection confirms responses were short and frequently code\-switched, performing well factually while likely underperforming on communicative quality\.

## 8Conclusion

We present TLoRA, a mechanistically\-grounded approach to Arabic medical adaptation: diagnosing where a model fails directly informs where adaptation should intervene\. Tuned lens probing and causal activation patching identify a knowledge\-routing failure localised to a specific layer window; restricting LoRA to that window outperforms full\-network adaptation and all mechanistically\-motivated alternatives\. Results on generation and dialogue suggest targeted adaptation preserves general capability where broader fine\-tuning does not\. We conclude that targeted adaptation informed by mechanistic diagnosis offers a tractable path toward closing the Arabic medical NLP gap, with gains likely to compound as higher\-quality native Arabic clinical data becomes available\.

## Limitations

First, our mechanistic analysis and proposed TLoRA method are validated exclusively on the Mistral\-Small\-3\.2\-24B model family\. Future work will extend our analysis to a broader range of decoder\-only models, encoder\-decoder architectures, and larger\-scale language models to verify the generalizability of the observed localized routing failure across diverse model checkpoints and configurations\. Second, while AraClinicDialog contributes a new multi\-dialectal Arabic clinical dialogue benchmark covering modern standard arabic and four major dialects, it does not yet encompass all regional Arabic varieties\. In future work, we plan to expand the dataset to include more dialects and increase its scale to better cover long\-tail clinical cases, rare diseases, and low\-frequency medical scenarios\.

## Ethical Considerations

This study investigates cross\-lingual knowledge routing and mechanistic adaptation methods for improving Arabic medical language understanding in LLMs\. The research does not involve human subjects or the use of proprietary, private, or sensitive data\. All datasets, pretrained models, and external resources used in this work comply with their respective licenses and terms of use\. The proposed methodology is intended to improve access to medical knowledge in under\-represented languages such as Arabic and can contribute to more equitable multilingual medical NLP systems\. It is not intended as a substitute for clinicians, medical advice, diagnosis, or treatment\. Future work should further evaluate the robustness, safety, and generalizability of these methods across broader medical settings and languages\.

## Acknowledgements

This work was supported by the Meem Foundation and the New York University Abu Dhabi \(NYUAD\) Center for Interdisciplinary Data Science and AI \(CIDSAI\), funded by Tamkeen under the NYUAD Research Institute Award CG016\. The research was carried out on NYUAD’s High Performance Computing resources \(Jubail\)\.

## References

- Cross\-lingual empirical evaluation of large language models for Arabic medical tasks\.InProceedings of the 1st Workshop on Linguistic Analysis for Health \(HeaLing 2026\),V\. Danilova, M\. Kurfalı, Y\. Söderfeldt, J\. Reed, and A\. Burchell \(Eds\.\),Rabat, Morocco,pp\. 158–171\.External Links:[Link](https://aclanthology.org/2026.healing-1.13/),[Document](https://dx.doi.org/10.18653/v1/2026.healing-1.13),ISBN 979\-8\-89176\-367\-8Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.00207#S2.SS2.p2.1)\.
- A\. Alasmari, S\. Alhumoud, and W\. Alshammari \(2024\)AraMed: Arabic medical question answering using pretrained transformer language models\.InProceedings of the 6th Workshop on Open\-Source Arabic Corpora and Processing Tools \(OSACT\) with Shared Tasks on Arabic LLMs Hallucination and Dialect to MSA Machine Translation @ LREC\-COLING 2024,H\. Al\-Khalifa, K\. Darwish, H\. Mubarak, M\. Ali, and T\. Elsayed \(Eds\.\),Torino, Italia,pp\. 50–56\.External Links:[Link](https://aclanthology.org/2024.osact-1.6/)Cited by:[§B\.1](https://arxiv.org/html/2608.00207#A2.SS1.p1.1)\.
- I\. Alonso, M\. Oronoz, and R\. Agerri \(2024\)MedExpQA: multilingual benchmarking of large language models for medical question answering\.Artificial Intelligence in Medicine155,pp\. 102938\.External Links:ISSN 0933\-3657,[Link](http://dx.doi.org/10.1016/j.artmed.2024.102938),[Document](https://dx.doi.org/10.1016/j.artmed.2024.102938)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p3.1)\.
- Anthropic \(2026\)Introducing Claude Opus 4\.6\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-6](https://www.anthropic.com/news/claude-opus-4-6)Accessed: 2026\-05\-03Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.6.2)\.
- M\. Anwar, A\. Freihat, G\. Ibrahim, M\. Awad, A\. A\. M\. A\. Sadallah, G\. Gosal, G\. Ramakrishnan, S\. Chandran, B\. Mishra, R\. Joshi, A\. Frikha, E\. Goffinet, A\. Maiti, A\. El Filali, S\. Al Barri, S\. Ghosh, R\. Pal, P\. Mullah, A\. Shukla, S\. Siddiki, S\. Kamboj, O\. Pandit, S\. Sahu, A\. El Badawy, A\. Mohamed, A\. Chamma, P\. Nakov,et al\.\(2025\)Jais 2: A family of Arabic\-centric open large language models\.Technical ReportMBZUAI, Inception, Cerebras\.Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.13.2)\.
- R\. K\. Arora, J\. Wei, R\. S\. Hicks, P\. Bowman, J\. Quiñonero\-Candela, F\. Tsimpourlas, M\. Sharman, M\. Shah, A\. Vallone, A\. Beutel, J\. Heidecke, and K\. Singhal \(2025\)HealthBench: evaluating large language models towards improved human health\.External Links:2505\.08775,[Link](https://arxiv.org/abs/2505.08775)Cited by:[§B\.2\.2](https://arxiv.org/html/2608.00207#A2.SS2.SSS2.p1.1),[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p2.1),[§5\.2\.3](https://arxiv.org/html/2608.00207#S5.SS2.SSS3.p1.1)\.
- M\. S\. Bari, Y\. Alnumay, N\. A\. Alzahrani, N\. M\. Alotaibi, H\. A\. Alyahya, S\. AlRashed, F\. A\. Mirza, S\. Z\. Alsubaie, H\. A\. Alahmed, G\. Alabduljabbar, R\. Alkhathran, Y\. Almushayqih, R\. Alnajim, S\. Alsubaihi, M\. A\. Mansour, S\. A\. Hassan, Dr\. M\. Alrubaian, A\. Alammari, Z\. Alawami, A\. Al\-Thubaity, A\. Abdelali, J\. Kuriakose, A\. Abujabal, N\. Al\-Twairesh, A\. Alowisheq, and H\. Khan \(2025\)ALLam: large language models for arabic and english\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MscdsFVZrN)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.14.2)\.
- S\. Bedi, H\. Cui, M\. Fuentes,et al\.\(2026\)Holistic evaluation of large language models for medical tasks with MedHELM\.Nature Medicine32,pp\. 943–951\.External Links:[Document](https://dx.doi.org/10.1038/s41591-025-04151-2)Cited by:[Appendix C](https://arxiv.org/html/2608.00207#A3.p1.1),[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p2.1)\.
- N\. Belrose, I\. Ostrovsky, L\. McKinney, Z\. Furman, L\. Smith, D\. Halawi, S\. Biderman, and J\. Steinhardt \(2023\)Eliciting latent predictions from transformers with the tuned lens\.External Links:2303\.08112,[Document](https://dx.doi.org/10.48550/arXiv.2303.08112),[Link](https://arxiv.org/abs/2303.08112)Cited by:[§3](https://arxiv.org/html/2608.00207#S3.p2.1)\.
- B\. E\. A\. Boussaha, L\. Al Qadi, M\. Farooq, S\. Alsuwaidi, G\. Campesan, A\. Alzubaidi, M\. Alyafeai, and H\. Hacid \(2025\)3LM: bridging Arabic, STEM, and code through benchmarking\.InProceedings of The Third Arabic Natural Language Processing Conference,K\. Darwish, A\. Ali, I\. Abu Farha, S\. Touileb, I\. Zitouni, A\. Abdelali, S\. Al\-Ghamdi, S\. Alkhereyf, W\. Zaghouani, S\. Khalifa, B\. AlKhamissi, R\. Almatham, I\. Hamed, Z\. Alyafeai, A\. Alowisheq, G\. Inoue, K\. Mrini, and W\. Alshammari \(Eds\.\),Suzhou, China,pp\. 42–63\.External Links:[Link](https://aclanthology.org/2025.arabicnlp-main.4/),[Document](https://dx.doi.org/10.18653/v1/2025.arabicnlp-main.4),ISBN 979\-8\-89176\-352\-4Cited by:[Table S1](https://arxiv.org/html/2608.00207#A1.T1.1.5.1),[§5\.2\.1](https://arxiv.org/html/2608.00207#S5.SS2.SSS1.p1.1)\.
- T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei \(2020\)Language models are few\-shot learners\.External Links:2005\.14165,[Link](https://arxiv.org/abs/2005.14165)Cited by:[§5\.3](https://arxiv.org/html/2608.00207#S5.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.22.2)\.
- Z\. Chen, A\. H\. Cano, A\. Romanou, A\. Bonnet, K\. Matoba, F\. Salvi, M\. Pagliardini, S\. Fan, A\. Köpf, A\. Mohtashami, A\. Sallinen, A\. Sakhaeirad, V\. Swamy, I\. Krawczuk, D\. Bayazit, A\. Marmet, S\. Montariol, M\. Hartley, M\. Jaggi, and A\. Bosselut \(2023\)MEDITRON\-70b: scaling medical pretraining for large language models\.External Links:2311\.16079,[Link](https://arxiv.org/abs/2311.16079)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p1.1)\.
- C\. Christophe, P\. K\. Kanithi, T\. Raha, S\. Khan, and M\. A\. Pimentel \(2024\)Med42\-v2: a suite of clinical llms\.External Links:2408\.06142,[Link](https://arxiv.org/abs/2408.06142)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.21.2)\.
- A\. Cocchieri, L\. Ragazzi, G\. Tagliavini, and G\. Moro \(2026\)ReMedQA: are we done with medical multiple\-choice benchmarks?\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 2706–2738\.External Links:[Link](https://aclanthology.org/2026.eacl-long.124/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.124),ISBN 979\-8\-89176\-380\-7Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p2.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen, L\. Marris, S\. Petulla, C\. Gaffney, A\. Aharoni, N\. Lintz, T\. C\. Pais, H\. Jacobsson, I\. Szpektor, N\. Jiang, K\. Haridasan, A\. Omran, N\. Saunshi, D\. Bahri, G\. Mishra, E\. Chu, T\. Boyd, B\. Hekman, A\. Parisi, C\. Zhang, K\. Kawintiranon, T\. Bedrax\-Weiss, O\. Wang, Y\. Xu, O\. Purkiss, U\. Mendlovic, I\. Deutel, N\. Nguyen, A\. Langley, F\. Korn, L\. Rossazza, A\. Ramé, S\. Waghmare, H\. Miller, N\. Byrd, A\. Sheshan, R\. Hadsell, S\. Bhardwaj, P\. Janus, T\. Rissa, D\. Horgan, A\. Abdagic, L\. Belenki, J\. Allingham, A\. Singh, T\. Guidroz, S\. Srinivasan, H\. Schmit, K\. Chiafullo, A\. Elisseeff, N\. Jha, P\. Kolhar, L\. Berrada, F\. Ding, X\. Si, S\. B\. Mallick, F\. Och, S\. Erell, E\. Ni, T\. Latkar, S\. Yang, P\. Sirkovic, Z\. Feng, R\. Leland, R\. Hornung, G\. Wu, C\. Blundell, H\. Alvari, P\. Huang, C\. Yip, S\. Deur, L\. Liu, G\. Surita, P\. Duque, D\. Damen, J\. Jia, A\. Guez, M\. Mircea, A\. Sinha, A\. Magni, P\. Stradomski, T\. Marian, V\. Galić, W\. Chen, H\. Husain, A\. Singhal, D\. Grewe, F\. Aubet, S\. Song, L\. Blanco, L\. Rechis, L\. Ho, R\. Munoz, K\. Zheng, J\. Hamrick, K\. Mather, H\. Taitelbaum, E\. Rutherford, Y\. Lei, K\. Chen, A\. Shukla, E\. Moreira, E\. Doi, B\. Isik, N\. Shabat, D\. Rogozińska, K\. Kolipaka, J\. Chang, E\. Vušak, S\. Venkatachary, S\. Noghabi, T\. Bharti, Y\. Jun, A\. Zaks, S\. Green, J\. Challagundla, W\. Wong, M\. Mohammad, D\. Hirsch, Y\. Cheng, I\. Naim, L\. Proleev, D\. Vincent, A\. Singh, M\. Krikun, D\. Krishnan, Z\. Ghahramani, A\. Atias, R\. Aggarwal, C\. Kirov, D\. Vytiniotis, C\. Koh, A\. Chronopoulou, P\. Dogra, V\. Ion, G\. Tyen, J\. Lee, F\. Weissenberger, T\. Strohman, A\. Balakrishna, J\. Rae, M\. Velic, R\. de Liedekerke, O\. Elyada, W\. Yuan, C\. Liu, L\. Shani, S\. Kishchenko, B\. Alessio, Y\. Li, R\. Song, S\. Kwei, O\. Jankowski, A\. Pappu, Y\. Namiki, Y\. Ma, N\. Tripuraneni, C\. Cherry, M\. Ikonomidis, Y\. Ling, C\. Ji, B\. Westberg, A\. Wright, D\. Yu, D\. Parkinson, S\. Ramaswamy, J\. Connor, S\. H\. Yeganeh, S\. Grover, G\. Kenwright, L\. Litchev, C\. Apps, A\. Tomala, F\. Halim, A\. Castro\-Ros, Z\. Li, A\. Boral, P\. Sho, M\. Yarom, E\. Malmi, D\. Klinghoffer, R\. Lin, A\. Ansell, P\. K\. S, S\. Zhao, S\. Zuo, A\. Santoro, H\. Cheng, S\. Demmessie, Y\. Liu, N\. Brichtova, A\. Culp, N\. Braun, D\. Graur, W\. Ng, N\. Mehta, A\. Phillips, P\. Sundberg, V\. Godbole, F\. Liu, Y\. Katariya, D\. Rim, M\. Seyedhosseini, S\. Ammirati, J\. Valfridsson, M\. Malihi, T\. Knight, A\. Toor, T\. Lampe, A\. Ittycheriah, L\. Chiang, C\. Yeung, A\. Fréchette, J\. Rao, H\. Wang, H\. Srivastava, R\. Zhang, R\. Rhodes, A\. Brand, D\. Weesner, I\. Figotin, F\. Gimeno, R\. Fellinger, P\. Marcenac, J\. Leal, E\. Marcus, V\. Cotruta, R\. Cabrera, S\. Luo, D\. Garrette, V\. Axelrod, S\. Baltateanu, D\. Barker, D\. Chen, H\. Toma, B\. Ingram, J\. Riesa, C\. Kulkarni, Y\. Zhang, H\. Liu, C\. Wang, M\. Polacek, W\. Wu, K\. Hui, A\. N\. Reyes, Y\. Su, M\. Barnes, I\. Malhi, A\. Siddiqui, Q\. Feng, M\. Damaschin, D\. Pighin, A\. Steiner, S\. Yang, R\. S\. Boppana, S\. Ivanov, A\. Kandoor, A\. Shah, A\. Mujika, D\. Huang, C\. A\. Choquette\-Choo, M\. Patel, T\. Yu, T\. Creswell, Jerry, Liu, C\. Barros, Y\. Razeghi, A\. Roy, P\. Culliton, B\. Xiong, J\. Pan, T\. Strohmann, T\. Powell, B\. Seal, D\. DeCarlo, P\. Shyam, K\. Katircioglu, X\. Wang, C\. Hardin, I\. Odisho, J\. Broder, O\. Chang, A\. Nair, A\. Shtefan, M\. O’Brien, M\. Agarwal, S\. Potluri, S\. Goyal, A\. Jhindal, S\. Thakur, Y\. Stuken, J\. Lyon, K\. Toutanova, F\. Feng, A\. Wu, B\. Horn, A\. Wang, A\. Cullum, G\. Taubman, D\. Shrivastava, C\. Shi, H\. Tomlinson, R\. Patel, T\. Tu, A\. M\. Oflazer, F\. Pongetti, M\. Yang, A\. A\. Taïga, V\. Perot, N\. W\. Pierse, F\. Han, Y\. Drori, I\. Iturrate, A\. Chakrabarti, L\. Yeung, D\. Dopson, Y\. Chen, A\. Kulshreshtha, T\. Guo, P\. Pham, T\. Schuster, J\. Chen, A\. Polozov, J\. Xing, H\. Zhou, P\. Kacham, D\. Kukliansky, A\. Miech, S\. Yaroshenko, E\. Chi, S\. Douglas, H\. Fei, M\. Blondel, P\. Myla, L\. Madmoni, X\. Wu, D\. Keysers, K\. Kjems, I\. Albuquerque, L\. Yu, J\. D’sa, M\. Plantan, V\. Ionescu, J\. S\. Elias, A\. Gupta, M\. R\. Vuyyuru, F\. Alcober, T\. Zhou, K\. Ji, F\. Hartmann, S\. Puttagunta, H\. Song, E\. Amid, A\. Stefanoiu, A\. Lee, P\. Pucciarelli, E\. Wang, A\. Raul, S\. Petrov, I\. Tian, V\. Anklin, N\. Nti, V\. Gomes, M\. Schumacher, G\. Vesom, A\. Panagopoulos, K\. Bousmalis, D\. Andor, J\. Jacob, Y\. Zhang, B\. Rosgen, M\. Kecman, M\. Tung, A\. Belias, N\. Goodman, P\. Covington, B\. Wieder, N\. Saxena, E\. Davoodi, M\. Huang, S\. Maddineni, V\. Roulet, F\. Campbell\-Ajala, P\. G\. Sessa, Xintian, Wu, G\. Lai, P\. Collins, A\. Haig, V\. Sakenas, X\. Xu, M\. Giustina, L\. E\. Shafey, P\. Charoenpanit, S\. Garg, J\. Ainslie, B\. Severson, M\. G\. Arenas, S\. Pathak, S\. Rajayogam, J\. Feng, M\. Bakker, S\. Li, N\. Wichers, J\. Rogers, X\. Geng, Y\. Li, R\. Jagerman, C\. Jia, N\. Olmert, D\. Sharon, M\. Mauger, S\. Mariserla, H\. Ma, M\. Mohabey, K\. Kim, A\. Andreev, S\. Pollom, J\. Love, V\. Jain, P\. Agrawal, Y\. Schroecker, A\. Fortin, M\. Warmuth, J\. Liu, A\. Leach, I\. Blok, G\. P\. Girirajan, R\. Aharoni, B\. Uria, A\. Sozanschi, D\. Goldberg, L\. Ionita, M\. T\. Ribeiro, M\. Zlocha, V\. Birodkar, S\. Lachgar, L\. Yuan, H\. Choudhury, M\. Ginsberg, F\. Zheng, G\. Dibb, E\. Graves, S\. Lokhande, G\. Rasskin, G\. Muraru, C\. Quick, S\. Tata, P\. Sermanet, A\. Chawla, I\. Karo, Y\. Wang, S\. Zhang, O\. Keller, A\. Dragan, G\. Su, I\. Chou, X\. Liu, Y\. Tao, S\. Prabhakara, M\. Wilson, R\. Liu, S\. Wang, G\. Evans, D\. Du, A\. Castaño, G\. Prasad, M\. E\. Mahdy, S\. Gerlach, M\. Reid, J\. Kahn, A\. Zait, T\. S\. Pillai, T\. Ulrich, G\. Wang, J\. Wassenberg, E\. Farkash, K\. Yalasangi, C\. Wang, M\. Bauza, S\. Bucher, T\. Liu, J\. Yan, G\. Leung, V\. Sindhwani, P\. Barnes, A\. Singh, I\. Jurin, J\. Chang, N\. K\. Bhumihar, S\. Eiger, G\. Citovsky, B\. Withbroe, Z\. Li, S\. Xue, N\. D\. Santo, G\. Stoyanov, Y\. Raimond, S\. Zheng, Y\. Gao, V\. Listík, S\. Kwasiborski, R\. Saputro, A\. Ozturel, G\. Mallya, K\. Majmundar, R\. West, P\. Caron, J\. Wei, L\. Castrejon, S\. Vikram, D\. Ramachandran, N\. Dhawan, J\. Park, S\. Smoot, G\. van den Driessche, Y\. Blau, C\. Malik, W\. Liang, R\. Hirsch, C\. N\. dos Santos, E\. Weinstein, A\. van den Oord, S\. Lall, N\. FitzGerald, Z\. Jiang, X\. Yang, D\. Webster, A\. Elqursh, A\. Pope, G\. Rotival, D\. Raposo, W\. Zhu, J\. Dean, S\. Alabed, D\. Tran, A\. Gupta, Z\. Gleicher, J\. Austin, E\. Rosseel, M\. Umekar, D\. Das, Y\. Sun, K\. Chen, K\. Misiunas, X\. Zhou, Y\. Di, A\. Loo, J\. Newlan, B\. Li, V\. Ramasesh, Y\. Xu, A\. Chen, S\. Gandhe, R\. Soricut, N\. Gupta, S\. Hu, S\. El\-Sayed, X\. Garcia, I\. Brusilovsky, P\. Chen, A\. Bolt, L\. Huang, A\. Gurney, Z\. Zhang, A\. Pritzel, J\. Wilkiewicz, B\. Seybold, B\. K\. Shamanna, F\. Fischer, J\. Dean, K\. Gill, R\. Mcilroy, A\. Bhowmick, J\. Selier, A\. Yang, D\. Cheng, V\. Magay, J\. Tan, D\. Varma, C\. Walder, T\. Kocisky, R\. Nakashima, P\. Natsev, M\. Kwong, I\. Gog, C\. Zhang, S\. Dieleman, T\. Jimma, A\. Ryabtsev, S\. Brahma, D\. Steiner, D\. Du, A\. Žužul, M\. Žanić, M\. Raghavachari, W\. Gierke, Z\. Zheng, D\. Petrova, Y\. Dauphin, Y\. Liu, I\. Kessler, S\. Hand, C\. Duvarney, S\. Kim, H\. Lee, L\. Hussenot, J\. Hui, J\. Smith, D\. Jain, J\. Xia, G\. S\. Tomar, K\. Amiri, D\. Phan, F\. Fuchs, T\. Weyand, N\. Tomasev, A\. Cordell, X\. Liu, J\. Mallinson, P\. Joshi, A\. Crawford, A\. Suggala, S\. Chien, N\. Fernando, M\. Sanchez\-Vargas, D\. Williams, P\. Crone, X\. Luo, I\. Karpov, J\. Shan, T\. Thurk, R\. Strudel, P\. Voigtlaender, P\. Patil, T\. Dozat, A\. Khodaei, S\. Singla, P\. Ambroszczyk, Q\. Wu, Y\. Chang, B\. Roark, C\. Hegde, T\. Ding, A\. Filos, Z\. Wu, A\. S\. Pinto, S\. Liu, S\. Khanna, A\. Pandey, S\. Mcloughlin, Q\. Li, S\. Haves, A\. Zhou, E\. Buchatskaya, I\. Leal, P\. de Boursac, N\. Akazawa, N\. Anderson, T\. Chen, K\. Somandepalli, C\. Liang, S\. Goenka, S\. Winkler, A\. Grushetsky, Y\. Ding, J\. Smith, F\. Ye, J\. Pont\-Tuset, E\. Li, R\. Li, T\. Golany, D\. Wegner, T\. Jiang, O\. Barak, Y\. Shangguan, E\. Vértes, R\. Wong, J\. Bornschein, A\. Tudor, M\. Bevilacqua, T\. Schaul, A\. S\. Rawat, Y\. Zhao, K\. Axiotis, L\. Meng, C\. McLean, J\. Lai, J\. Beattie, N\. Kushman, Y\. Liu, B\. Kutzman, F\. Lang, J\. Ye, P\. Netrapalli, P\. Mishra, M\. Khan, M\. Goel, R\. Willoughby, D\. Tian, H\. Zhuang, J\. Chen, Z\. Tsai, T\. Kementsietsidis, A\. Khare, J\. Keeling, K\. Xu, N\. Waters, F\. Altché, A\. Popat, B\. Mittal, D\. Saxton, D\. E\. Badawy, M\. Mathieu, Z\. Zheng, H\. Zhou, N\. Ranka, R\. Shin, Q\. Duan, T\. Salimans, I\. Mihailescu, U\. Shaham, M\. Chang, Y\. Assael, N\. Dikkala, M\. Izzard, V\. Cohen\-Addad, C\. Graves, V\. Feinberg, G\. Chung, D\. Strouse, D\. Karmon, S\. Sharifzadeh, Z\. Ashwood, K\. Pham, J\. Blanton, A\. Vasiloff, J\. Barber, M\. Geller, A\. Zhou, F\. Zubach, T\. Huang, L\. Zhang, H\. Gupta, M\. Young, J\. Proskurnia, R\. Votel, V\. Gabeur, G\. Barcik, A\. Tripathi, H\. Yu, G\. Yan, B\. Changpinyo, F\. Pavetić, A\. Coyle, Y\. Fujii, J\. G\. Mendez, T\. Zhou, H\. Rajamani, B\. Hechtman, E\. Cao, D\. Juan, Y\. Tan, V\. Dalibard, Y\. Du, N\. Clay, K\. Yao, W\. Jia, D\. Vijaykumar, Y\. Zhou, X\. Bai, W\. Hung, S\. Pecht, G\. Todorov, N\. Khadke, P\. Gupta, P\. Lahoti, A\. Autef, K\. Duddu, J\. Lee\-Thorp, A\. Bykovsky, T\. Misiunas, S\. Flennerhag, S\. Thangaraj, J\. McGiffin, Z\. Nado, M\. Kunesch, A\. Noever, A\. Hertz, M\. Liang, V\. Stone, E\. Palmer, S\. Daruki, A\. Pramanik, S\. Põder, A\. Kyker, M\. Khan, E\. Sluzhaev, M\. Ritter, A\. Ruderman, W\. Zhou, C\. Nagpal, K\. Vodrahalli, G\. Necula, P\. Barham, E\. Pavlick, J\. Hartford, I\. Shafran, L\. Zhao, M\. Mikuła, T\. Eccles, H\. Shimokawa, K\. Garg, L\. Vilnis, H\. Chen, I\. Shumailov, K\. Lee, A\. Abdelhamed, M\. Xie, V\. Cohen, E\. Hlavnova, D\. Malkin, C\. Sitawarin, J\. Lottes, P\. Coquinot, T\. Yu, S\. Kumar, J\. Zhang, A\. Mahendru, Z\. Ahmed, J\. Martens, T\. Chen, A\. Boag, D\. Peng, C\. Devin, A\. Klimovskiy, M\. Phuong, D\. Vainstein, J\. Xie, B\. Ramabhadran, N\. Howard, X\. Yu, G\. Goswami, J\. Cui, S\. Shleifer, M\. Pinto, C\. Yeh, M\. Yang, S\. Javanmardi, D\. Ethier, C\. Lee, J\. Orbay, S\. Kotecha, C\. Bromberg, P\. Shaw, J\. Thornton, A\. G\. Rosenthal, S\. Gu, M\. Thomas, I\. Gemp, A\. Ayyar, A\. Ushio, A\. Selvan, J\. Wee, C\. Liu, M\. Majzoubi, W\. Yu, J\. Abernethy, T\. Liechty, R\. Pan, H\. Nguyen, Qiong, Hu, S\. Perrin, A\. Arora, E\. Pitler, W\. Wang, K\. Shivakumar, F\. Prost, B\. Limonchik, J\. Wang, Y\. Gao, T\. Cour, S\. Buch, H\. Gui, M\. Ivanova, P\. Neubeck, K\. Chan, L\. Kim, H\. Chen, N\. Goyal, D\. Chung, L\. Liu, Y\. Su, A\. Petrushkina, J\. Shen, A\. Joulin, Y\. Xu, S\. X\. Lin, Y\. Kulizhskaya, C\. Chelba, S\. Vasudevan, E\. Collins, V\. Bashlovkina, T\. Lu, D\. Fritz, J\. Park, Y\. Zhou, C\. Su, R\. Tanburn, M\. Sushkov, M\. Rasquinha, J\. Li, J\. Prendki, Y\. Li, P\. LV, S\. Sharma, H\. Fitoussi, H\. Huang, A\. Dai, P\. Dao, M\. Burrows, H\. Prior, D\. Qin, G\. Pundak, L\. L\. Sjoesund, A\. Khurshudov, Z\. Zhu, A\. Webson, E\. Kemp, T\. Tan, S\. Agrawal, S\. Sargsyan, L\. Cheng, J\. Stephan, T\. Kwiatkowski, D\. Reid, A\. Byravan, A\. H\. Michaely, N\. Heess, L\. Zhou, S\. Goenka, V\. Carpenter, A\. Levskaya, B\. Wang, R\. Roberts, R\. Leblond, S\. Chikkerur, S\. Ginzburg, M\. Chang, R\. Riachi, Chuqiao, Xu, Z\. Borsos, M\. Pliskin, J\. Pawar, M\. Lustman, H\. Kirkwood, A\. Anand, A\. Chaudhary, N\. Kalb, K\. Milan, S\. Augenstein, A\. Goldie, L\. Prince, K\. Raman, Y\. Sun, V\. Xia, A\. Cohen, Z\. Huo, J\. Camp, S\. Ellis, L\. Zilka, D\. V\. Torres, L\. Patel, S\. Arora, B\. Chan, J\. Adler, K\. Ayoub, J\. Liang, F\. Jamil, J\. Jiang, S\. Baumgartner, H\. Sun, Y\. Karov, Y\. Akulov, H\. Zheng, I\. Cai, C\. Fantacci, J\. Rubin, A\. R\. Acha, M\. Wang, N\. D’Souza, R\. Sathyanarayana, S\. Dai, S\. Rowe, A\. Simanovsky, O\. Goldman, Y\. Kuang, X\. Pan, A\. Rosenberg, T\. Rojas\-Esponda, P\. Dutta, A\. Zeng, I\. Jurenka, G\. Farquhar, Y\. Bansal, S\. Iqbal, B\. Roelofs, G\. Joung, P\. Beak, C\. Ryu, R\. Poplin, Y\. Wu, J\. Alayrac, S\. Buthpitiya, O\. Ronneberger, C\. Habtegebriel, W\. Li, P\. Cavallaro, A\. Wei, G\. Bensky, T\. Denk, H\. Ganapathy, J\. Stanway, P\. Joshi, F\. Bertolini, J\. Lo, O\. Ma, Z\. Charles, G\. Sampemane, H\. Sahni, X\. Chen, H\. Askham, D\. Gaddy, P\. Young, J\. Tan, M\. Eyal, A\. Bražinskas, L\. Zhong, Z\. Wu, M\. Epstein, K\. Bailey, A\. Hard, K\. Lee, S\. Goldshtein, A\. Ruiz, M\. Badawi, M\. Lochbrunner, J\. Kearns, A\. Brown, F\. Pardo, T\. Weber, H\. Yang, P\. Jiang, B\. Akin, Z\. Fu, M\. Wainwright, C\. Zou, M\. Gaba, P\. Manzagol, W\. Kan, Y\. Song, K\. Zainullina, R\. Lin, J\. Ko, S\. Deshmukh, A\. Jindal, J\. Svensson, D\. Tyam, H\. Zhao, C\. Kaeser\-Chen, S\. Baird, P\. Moradi, J\. Hall, Q\. Guo, V\. Tsang, B\. Liang, F\. Pereira, S\. Ganesh, I\. Korotkov, J\. Adamek, S\. Thiagarajan, V\. Tran, C\. Chen, C\. Tar, S\. Jain, I\. Dasgupta, T\. Bilal, D\. Reitter, K\. Zhao, G\. Vezzani, Y\. Gehman, P\. Mehta, L\. Beltrone, X\. Dotiwalla, S\. Guadarrama, Z\. Abbas, S\. Karp, P\. Georgiev, C\. Ferng, M\. Brockschmidt, L\. Peng, C\. Hirnschall, V\. Verma, Y\. Bi, Y\. Xiao, A\. Dabush, K\. Xu, P\. Wallis, R\. Parker, Q\. Wang, Y\. Xu, I\. Safarli, D\. Tewari, Y\. Zhang, S\. Kim, A\. Gesmundo, M\. Thomas, S\. Levi, A\. Chowdhury, K\. Rao, P\. Garst, S\. Conway\-Rahman, H\. Ran, K\. McKinney, Z\. Xiao, W\. Yu, R\. Agrawal, A\. Stjerngren, C\. Ionescu, J\. Chen, V\. Sharma, J\. Chiu, F\. Liu, K\. Franko, C\. Sanford, X\. Cai, P\. Michel, S\. Ganapathy, J\. Labanowski, Z\. Garrett, B\. Vargas, S\. Sun, B\. Gale, T\. Buschmann, G\. Desjardins, N\. Ghelani, P\. Jain, M\. Verma, C\. Asawaroengchai, J\. Eisenschlos, J\. Harlalka, H\. Kazawa, D\. Metzler, J\. Howland, Y\. Jian, J\. Ades, V\. Shah, T\. Gangwani, S\. Lee, R\. Ring, S\. M\. Hernandez, D\. Reich, A\. Sinha, A\. Sathe, J\. Kovac, A\. Gill, A\. Kannan, A\. D’olimpio, M\. Sevenich, J\. Whang, B\. Kim, K\. C\. Sim, J\. Chen, J\. Zhang, S\. Lall, Y\. Matias, B\. Jia, A\. Friesen, S\. Nasso, A\. Thapliyal, B\. Perozzi, T\. Yu, A\. Shekhawat, S\. Huda, P\. Grabowski, E\. Wang, A\. Sreevatsa, H\. Dib, M\. Hassen, P\. Schuh, V\. Milutinovic, C\. Welty, M\. Quinn, A\. Shah, B\. Wang, G\. Barth\-Maron, J\. Frye, N\. Axelsson, T\. Zhu, Y\. Ma, I\. Giannoumis, H\. Sedghi, C\. Ye, Y\. Luan, K\. Aydin, B\. Chandra, V\. Sampathkumar, R\. Huang, V\. Lavrenko, A\. Eleryan, Z\. Hong, S\. Hansen, S\. M\. Carthy, B\. Samanta, D\. Ćevid, X\. Wang, F\. Li, M\. Voznesensky, M\. Hoffman, A\. Terzis, V\. Sehwag, G\. Fidel, L\. He, M\. Cai, Y\. He, A\. Feng, M\. Nikoltchev, S\. Phatale, J\. Chase, R\. Lawton, M\. Zhang, T\. Ouyang, M\. Tragut, M\. H\. Manshadi, A\. Narayanan, J\. Shen, X\. Gao, T\. Bolukbasi, N\. Roy, X\. Li, D\. Golovin, L\. Panait, Z\. Qin, G\. Han, T\. Anthony, S\. Kudugunta, V\. Patraucean, A\. Ray, X\. Chen, X\. Yang, T\. Bhatia, P\. Talluri, A\. Morris, A\. Ražnatović, B\. Brownfield, J\. An, S\. Peng, P\. Kane, C\. Zheng, N\. Duduta, J\. Kessinger, J\. Noraky, S\. Liu, K\. Rong, P\. Veličković, K\. Rush, A\. Goldin, F\. Wei, S\. M\. R\. Garlapati, C\. Pantofaru, O\. Kwon, J\. Ni, E\. Noland, J\. D\. Trapani, F\. Beaufays, A\. G\. Roy, Y\. Chow, A\. Turker, G\. Cideron, L\. Mei, J\. Clark, Q\. Dou, M\. Bošnjak, R\. Leith, Y\. Du, A\. Yazdanbakhsh, M\. Nasr, C\. Kwak, S\. S\. Sheth, A\. Kaskasoli, A\. Anand, B\. Lakshminarayanan, S\. Jerome, D\. Bieber, C\. Chu, A\. Senges, T\. Shen, M\. Sridhar, N\. Ndebele, B\. Beyret, S\. Mohamed, M\. Chen, M\. Freitag, J\. Guo, L\. Liu, P\. Roit, H\. Chen, S\. Yan, T\. Stone, J\. Co\-Reyes, J\. Cole, S\. Scellato, S\. Azizi, H\. Hashemi, A\. Jin, A\. Iyer, M\. Valentine, A\. György, A\. Ahuja, D\. H\. Diaz, C\. Lee, N\. Clement, W\. Kong, D\. Garmon, I\. Watts, K\. Bhatia, K\. Gupta, M\. Miecnikowski, H\. Vallet, A\. Taly, E\. Loper, S\. Joshi, J\. Atwood, J\. Chick, M\. Collier, F\. Iliopoulos, R\. Trostle, B\. Gunel, R\. Leal\-Cavazos, A\. M\. Hrafnkelsson, M\. Guzman, X\. Ju, A\. Forbes, J\. Emond, K\. Chauhan, B\. Caine, L\. Xiao, W\. Zeng, A\. Moufarek, D\. Murphy, M\. Meng, N\. Gupta, F\. Riedel, A\. Das, E\. Lawal, S\. Narayan, T\. Sosea, J\. Swirhun, L\. Friso, B\. Neyshabur, J\. Lu, S\. Girgin, M\. Wunder, E\. Yvinec, A\. Pyne, V\. Carbune, S\. Rijhwani, Y\. Guo, T\. Doshi, A\. Briukhov, M\. Bain, A\. Hitron, X\. Wang, A\. Gupta, K\. Chen, C\. Du, W\. Zhang, D\. Shah, A\. Akula, M\. Dylla, A\. Kachra, W\. Kuo, T\. Zou, L\. Wang, L\. Xu, J\. Zhu, J\. Snyder, S\. Menon, O\. Firat, I\. Mordatch, Y\. Yuan, N\. Ponomareva, R\. Blevins, L\. Moore, W\. Wang, P\. Chen, M\. Scholz, A\. Dwornik, J\. Lin, S\. Li, D\. Antognini, T\. I, X\. Song, M\. Miller, U\. Kalra, A\. Raveret, O\. Akerlund, F\. Wu, A\. Nystrom, N\. Godbole, T\. Liu, H\. DeBalsi, J\. Zhao, B\. Liu, A\. Caciularu, L\. Lax, U\. Khandelwal, V\. Langston, E\. Bailey, S\. Lattanzi, Y\. Wang, N\. Kovelamudi, S\. Mondal, G\. Guruganesh, N\. Hua, O\. Roval, P\. Wesołowski, R\. Ingale, J\. Halcrow, T\. Sohn, C\. Angermueller, B\. Raad, E\. Stickgold, E\. Lu, A\. Kosik, J\. Xie, T\. Lillicrap, A\. Huang, L\. L\. Zhang, D\. Paulus, C\. Farabet, A\. Wertheim, B\. Wang, R\. Joshi, C\. Ko, Y\. Wu, S\. Agrawal, L\. Lin, X\. Sheng, P\. Sung, T\. Breland\-King, C\. Butterfield, S\. Gawde, S\. Singh, Q\. Zhang, R\. Apte, S\. Shetty, A\. Hutter, T\. Li, E\. Salesky, F\. Lebron, J\. Kanerva, M\. Paganini, A\. Nguyen, R\. Vallu, J\. Peter, S\. Velury, D\. Kao, J\. Hoover, A\. Bortsova, C\. Bishop, S\. Jakobovits, A\. Agostini, A\. Agarwal, C\. Liu, C\. Kwong, S\. Tavakkol, I\. Bica, A\. Greve, A\. GP, J\. Marcus, L\. Hou, T\. Duerig, R\. Moroshko, D\. Lacey, A\. Davis, J\. Amelot, G\. Wang, F\. Kim, T\. Strinopoulos, H\. Wan, C\. L\. Lan, S\. Krishnan, H\. Tang, P\. Humphreys, J\. Bai, I\. H\. Shtacher, D\. Machado, C\. Pang, K\. Burke, D\. Liu, R\. Aravamudhan, Y\. Song, E\. Hirst, A\. Singh, B\. Jou, L\. Bai, F\. Piccinno, C\. K\. Fu, R\. Alazard, B\. Meiri, D\. Winter, C\. Chen, M\. Zhang, J\. Heitkaemper, J\. Lambert, J\. Lee, A\. Frömmgen, S\. Rogulenko, P\. Nair, P\. Niemczyk, A\. Bulyenov, B\. Xu, H\. Shemtov, M\. Zadimoghaddam, S\. Toropov, M\. Wirth, H\. Dai, S\. Gollapudi, D\. Zheng, A\. Kurakin, C\. Lee, K\. Bullard, N\. Serrano, I\. Balazevic, Y\. Li, J\. Schalkwyk, M\. Murphy, M\. Zhang, K\. Sequeira, R\. Datta, N\. Agrawal, C\. Sutton, N\. Attaluri, M\. Chiang, W\. Farhan, G\. Thornton, K\. Lin, T\. Choma, H\. Nguyen, K\. Dasgupta, D\. Robinson, I\. Comşa, M\. Riley, A\. Pillai, B\. Mustafa, B\. Golan, A\. Zandieh, J\. Lespiau, B\. Porter, D\. Ross, S\. Rajayogam, M\. Agarwal, S\. Venugopalan, B\. Shahriari, Q\. Yan, H\. Xu, T\. Tobin, P\. Dubov, H\. Shi, A\. Recasens, A\. Kovsharov, S\. Borgeaud, L\. Dery, S\. Vasanth, E\. Gribovskaya, L\. Qiu, M\. Mahdieh, W\. Skut, E\. Nielsen, C\. Zheng, A\. Yu, C\. G\. Bostock, S\. Gupta, A\. Archer, C\. Rawles, E\. Davies, A\. Svyatkovskiy, T\. Tsai, Y\. Halpern, C\. Reisswig, B\. Wydrowski, B\. Chang, J\. Puigcerver, M\. H\. Taege, J\. Li, E\. Schnider, X\. Li, D\. Dena, Y\. Xu, U\. Telang, T\. Shi, H\. Zen, K\. Kastner, Y\. Ko, N\. Subramaniam, A\. Kumar, P\. Blois, Z\. Dai, J\. Wieting, Y\. Lu, Y\. Zeldes, T\. Xie, A\. Hauth, A\. Ţifrea, Y\. Li, S\. El\-Husseini, D\. Abolafia, H\. Zhou, W\. Ding, S\. Ghalebikesabi, C\. Guía, A\. Maksai, Á\. Weisz, S\. Arik, N\. Sukhanov, A\. Świetlik, X\. Jia, L\. Yu, W\. Wang, M\. Brand, D\. Bloxwich, S\. Kirmani, Z\. Chen, A\. Go, P\. Sprechmann, N\. Kannen, A\. Carin, P\. Sandhu, I\. Edkins, L\. Nooteboom, J\. Gupta, L\. Maggiore, J\. Azizi, Y\. Pritch, P\. Yin, M\. Gupta, D\. Tarlow, D\. Smith, D\. Ivanov, M\. Babaeizadeh, A\. Goel, S\. Kambala, G\. Chu, M\. Kastelic, M\. Liu, H\. Soltau, A\. Stone, S\. Agrawal, M\. Kim, K\. Soparkar, S\. Tadepalli, O\. Bunyan, R\. Soh, A\. Kannan, D\. Kim, B\. J\. Chen, A\. Halumi, S\. Roy, Y\. Wang, O\. Sercinoglu, G\. Gibson, S\. Bhatnagar, M\. Sano, D\. von Dincklage, Q\. Ren, B\. Mitrevski, M\. Olšák, J\. She, C\. Doersch, Jilei, Wang, B\. Liu, Q\. Tan, T\. Yakar, T\. Warkentin, A\. Ramirez, C\. Lebsack, J\. Dillon, R\. Mathews, T\. Cobley, Z\. Wu, Z\. Chen, J\. Simon, S\. Nath, T\. Sainath, A\. Bendebury, R\. Julian, B\. Mankalale, D\. Ćurko, P\. Zacchello, A\. R\. Brown, K\. Sodhia, H\. Howard, S\. Caelles, A\. Gupta, G\. Evans, A\. Bulanova, L\. Katzen, R\. Goldenberg, A\. Tsitsulin, J\. Stanton, B\. Schillings, V\. Kovalev, C\. Fry, R\. Shah, K\. Lin, S\. Upadhyay, C\. Li, S\. Radpour, M\. Maggioni, J\. Xiong, L\. Haas, J\. Brennan, A\. Kamath, N\. Savinov, A\. Nagrani, T\. Yacovone, R\. Kappedal, K\. Andriopoulos, L\. Lao, Y\. Li, G\. Rozhdestvenskiy, K\. Hashimoto, A\. Audibert, S\. Austin, D\. Rodriguez, A\. Ruoss, G\. Honke, D\. Karkhanis, X\. Xiong, Q\. Wei, J\. Huang, Z\. Leng, V\. Premachandran, S\. Bileschi, G\. Evangelopoulos, T\. Mensink, J\. Pavagadhi, D\. Teplyashin, P\. Chang, L\. Xue, G\. Tanzer, S\. Goldman, K\. Patel, S\. Li, J\. Wiesner, I\. Zheng, I\. Stewart\-Binks, J\. Han, Z\. Li, L\. Luo, K\. Lenc, M\. Lučić, F\. Xue, R\. Mullins, A\. Guseynov, C\. Chang, I\. Galatzer\-Levy, A\. Zhang, G\. Bingham, G\. Hu, A\. Hartman, Y\. Ma, J\. Griffith, A\. Irpan, C\. Radebaugh, S\. Yue, L\. Fan, V\. Ungureanu, C\. Sorokin, H\. Teufel, P\. Li, R\. Anil, D\. Paparas, T\. Wang, C\. Lin, H\. Peng, M\. Shum, G\. Petrovic, D\. Brady, R\. Nguyen, K\. Macherey, Z\. Li, H\. Singh, M\. Yenugula, M\. Iinuma, X\. Chen, K\. Kopparapu, A\. Stern, S\. Dave, C\. Thekkath, F\. Perot, A\. Kumar, F\. Li, Y\. Xiao, M\. Bilotti, M\. H\. Bateni, I\. Noble, L\. Lee, A\. Vázquez\-Reina, J\. Salazar, X\. Yang, B\. Wang, E\. Gruzewska, A\. Rao, S\. Raghuram, Z\. Xu, E\. Ben\-David, J\. Mei, S\. Dalmia, Z\. Zhang, Y\. Liu, G\. Bansal, H\. Pankov, S\. Schwarcz, A\. Burns, C\. Chan, S\. Sanghai, R\. Liang, E\. Liang, A\. He, A\. Stuart, A\. Narayanan, Y\. Zhu, C\. Frank, B\. Fatemi, A\. Sabne, O\. Lang, I\. Bhattacharya, S\. Settle, M\. Wang, B\. McMahan, A\. Tacchetti, L\. B\. Soares, M\. Hadian, S\. Cabi, T\. Chung, N\. Putikhin, G\. Li, J\. Chen, A\. Tarango, H\. Michalewski, M\. Kazemi, H\. Masoom, H\. Sheftel, R\. Shivanna, A\. Vadali, R\. Comanescu, D\. Reid, J\. Moore, A\. Neelakantan, M\. Sander, J\. Herzig, A\. Rosenberg, M\. Dehghani, J\. Choi, M\. Fink, R\. Hayes, E\. Ge, S\. Weng, C\. Ho, J\. Karro, K\. Krishna, L\. N\. Thiet, A\. Skerry\-Ryan, D\. Eppens, M\. Andreetto, N\. Sarma, S\. Bonacina, B\. K\. Ayan, M\. Nawhal, Z\. Shan, M\. Dusenberry, S\. Thakoor, S\. Gubbi, D\. D\. Nguyen, R\. Tsarfaty, S\. Albanie, J\. Mitrović, M\. Gandhi, B\. Chen, A\. Epasto, G\. Stephanov, Y\. Jin, S\. Gehman, A\. Amini, J\. Weber, F\. Behbahani, S\. Xu, M\. Allamanis, X\. Chen, M\. Ott, C\. Sha, M\. Jastrzebski, H\. Qi, D\. Greene, X\. Wu, A\. Toki, D\. Vlasic, J\. Shapiro, R\. Kotikalapudi, Z\. Shen, T\. Saeki, S\. Xie, A\. Cassirer, S\. Bharadwaj, T\. Kiyono, S\. Bhojanapalli, E\. Rosenfeld, S\. Ritter, J\. Mao, J\. G\. Oliveira, Z\. Egyed, B\. Bandemer, E\. Parisotto, K\. Kinoshita, J\. Pluto, P\. Maniatis, S\. Li, Y\. Guo, G\. Ghiasi, J\. Tarbouriech, S\. Chatterjee, J\. Jin, Katrina, Xu, J\. Palomaki, S\. Arnold, M\. Sewak, F\. Piccinini, M\. Sharma, B\. Albrecht, S\. Purser\-haskell, A\. Vaswani, C\. Chen, M\. Wisniewski, Q\. Cao, J\. Aslanides, N\. M\. Phu, M\. Sieb, L\. Agubuzu, A\. Zheng, D\. Sohn, M\. Selvi, A\. Andreassen, K\. Subudhi, P\. Eruvbetine, O\. Woodman, T\. Mery, S\. Krause, X\. Ren, X\. Ma, J\. Luo, D\. Chen, W\. Fan, H\. Griffiths, C\. Schuler, A\. Li, S\. Zhang, J\. Sarr, S\. Luo, R\. Patana, M\. Watson, D\. Naboulsi, M\. Collins, S\. Sidhwani, E\. Hoogeboom, S\. Silver, E\. Caveness, X\. Zhao, M\. Rodriguez, M\. Deines, L\. Bai, P\. Griffin, M\. Tagliasacchi, E\. Xue, S\. R\. Babbula, B\. Pang, N\. Ding, G\. Shen, E\. Peake, R\. Crocker, S\. S\. Raghvendra, D\. Swisher, W\. Han, R\. Singh, L\. Wu, V\. Pchelin, T\. Munkhdalai, D\. Alon, G\. Bacon, E\. Robles, J\. Bulian, M\. Johnson, G\. Powell, F\. T\. Ferreira, Y\. Li, F\. Benzing, M\. Velimirović, H\. Soyer, W\. Kong, Tony, Nguyên, Z\. Yang, J\. Liu, J\. van Amersfoort, D\. Gillick, B\. Sun, N\. Rauschmayr, K\. Zhang, S\. Zhan, T\. Zhou, A\. Frolov, C\. Yang, D\. Vnukov, L\. Rouillard, H\. Li, A\. Mandhane, N\. Fallen, R\. Venkataraman, C\. H\. Hu, J\. Brennan, J\. Lee, J\. Chang, M\. Sundermeyer, Z\. Pan, R\. Ke, S\. Tong, A\. Fabrikant, W\. Bono, J\. Gu, R\. Foley, Y\. Mao, M\. Delakis, D\. Bhaswar, R\. Frostig, N\. Li, A\. Zipori, C\. Hope, O\. Kozlova, S\. Mishra, J\. Djolonga, C\. Schiff, M\. A\. Merey, E\. Briakou, P\. Morgan, A\. Wan, A\. Hassidim, R\. Skerry\-Ryan, K\. Sengupta, M\. Jasarevic, P\. Kallakuri, P\. Kunkle, H\. Brennan, T\. Lieber, H\. Mansoor, J\. Walker, B\. Zhang, A\. Xie, G\. Žužić, A\. Chukwuka, A\. Druinsky, D\. Cho, R\. Yao, F\. Naeem, S\. Butt, E\. Kim, Z\. Jia, M\. Jordan, A\. Lelkes, M\. Kurzeja, S\. Wang, J\. Zhao, A\. Over, A\. Chakladar, M\. Prasetya, N\. Jha, S\. Ganapathy, Y\. Cong, P\. Shroff, C\. Saroufim, S\. Miryoosefi, M\. Hammad, T\. Nasir, W\. Xi, Y\. Gao, Y\. Maeng, B\. Hora, C\. Cheng, P\. Haghani, Y\. Lewenberg, C\. Lu, M\. Matysiak, N\. Raisinghani, H\. Wang, L\. Baugher, R\. Sukthankar, M\. Giang, J\. Schultz, N\. Fiedel, M\. Chen, C\. Lee, T\. Dey, H\. Zheng, S\. Paul, C\. Smith, A\. Ly, Y\. Wang, R\. Bansal, B\. Perz, S\. Ricco, S\. Blank, V\. Keshava, D\. Sharma, M\. Chow, K\. Lad, K\. Jalan, S\. Osindero, C\. Swanson, J\. Scott, A\. Ilić, X\. Li, S\. R\. Jonnalagadda, A\. S\. Soudagar, Y\. Xiong, B\. Batsaikhan, D\. Jarrett, N\. Kumar, M\. Shah, M\. Lawlor, A\. Waters, M\. Graham, R\. May, S\. Ramos, S\. Lefdal, Z\. Cankara, N\. Cano, B\. O’Donoghue, J\. Borovik, F\. Liu, J\. Grimstad, M\. Alnahlawi, K\. Tsihlas, T\. Hudson, N\. Grigorev, Y\. Jia, T\. Huang, T\. P\. Igwe, S\. Lebedev, X\. Tang, I\. Krivokon, F\. Garcia, M\. Tan, E\. Jia, P\. Stys, S\. Vashishth, Y\. Liang, B\. Venkatraman, C\. Gu, A\. Kementsietsidis, C\. Zhu, J\. Jung, Y\. Bai, M\. J\. Hosseini, F\. Ahmed, A\. Gupta, X\. Yuan, S\. Ashraf, S\. Nigam, G\. Vasudevan, P\. Awasthi, A\. M\. Gilady, Z\. Mariet, R\. Eskander, H\. Li, H\. Hu, G\. Garrido, P\. Schlattner, G\. Zhang, R\. Saxena, P\. Dević, K\. Muralidharan, A\. Murthy, Y\. Zhou, M\. Choi, A\. Wongpanich, Z\. Wang, P\. Shah, Y\. Xu, Y\. Huang, S\. Spencer, A\. Chen, J\. Cohan, J\. Wang, J\. Tompson, J\. Wu, R\. Haroun, H\. Li, B\. Huergo, F\. Yang, T\. Yin, J\. Wendt, M\. Bendersky, R\. Chaabouni, J\. Snaider, J\. Ferret, A\. Jindal, T\. Thompson, A\. Xue, W\. Bishop, S\. M\. Phal, A\. Sharma, Y\. Sung, P\. Radhakrishnan, M\. Shomrat, R\. Ingle, R\. Vij, J\. Gilmer, M\. D\. Istin, S\. Sobell, Y\. Lu, E\. Nottage, D\. Sadigh, J\. Willcock, T\. Zhang, S\. Xu, S\. Brown, K\. Lee, G\. Wang, Y\. Zhu, Y\. Tay, C\. Kim, A\. Gutierrez, A\. Sharma, Y\. Xian, S\. Seo, C\. Cui, E\. Pochernina, C\. Baetu, K\. Jastrzębski, M\. Ly, M\. Elhawaty, D\. Suh, E\. Sezener, P\. Wang, N\. Yuen, G\. Tucker, J\. Cai, Z\. Yang, C\. Wang, A\. Muzio, H\. Qian, J\. Yoo, D\. Lockhart, K\. R\. McKee, M\. Guo, M\. Mehrotra, A\. Mendonça, S\. V\. Mehta, S\. Ben, C\. Tekur, J\. Mu, M\. Zhu, V\. Krakovna, H\. Lee, A\. Maschinot, S\. Cevey, H\. Choe, A\. Bai, H\. Srinivasan, D\. Gasaway, N\. Young, P\. Siegler, D\. Holtmann\-Rice, V\. Piratla, K\. Baumli, R\. Yogev, A\. Hofer, H\. van Hasselt, S\. Grant, Y\. Chervonyi, D\. Silver, A\. Hogue, A\. Agarwal, K\. Wang, P\. Singh, F\. Flynn, J\. Lipschultz, R\. David, L\. Bellot, Y\. Yang, L\. Le, F\. Graziano, K\. Olszewska, K\. Hui, A\. Maurya, N\. Parotsidis, W\. Chen, T\. Oguntebi, J\. Kelley, A\. Baddepudi, J\. Mauerer, G\. Shaw, A\. Siegman, L\. Yang, S\. Shetty, S\. Roy, Y\. Song, W\. Stokowiec, R\. Burnell, O\. Savant, R\. Busa\-Fekete, J\. Miao, S\. Ghosh, L\. MacDermed, P\. Lippe, M\. Dektiarev, Z\. Behrman, F\. Mentzer, K\. Nguyen, M\. Wei, S\. Verma, C\. Knutsen, S\. Dasari, Z\. Yan, P\. Mitrichev, X\. Wang, V\. Shejwalkar, J\. Austin, S\. Sunkara, N\. Potti, Y\. Virin, C\. Wright, G\. Liu, O\. Riva, E\. Pot, G\. Kochanski, Q\. Le, G\. Balasubramaniam, A\. Dhar, Y\. Liao, A\. Bloniarz, D\. Shukla, E\. Cole, J\. Lee, S\. Zhang, S\. Kafle, S\. Vashishtha, P\. Mahmoudieh, G\. Chen, R\. Hoffmann, P\. Srinivasan, A\. D\. Lago, Y\. B\. Shalom, Z\. Wang, M\. Elabd, A\. Sharma, J\. Oh, S\. Kothawade, M\. Le, M\. Monteiro, S\. Yang, K\. Alarakyia, R\. Geirhos, D\. Mincu, H\. Garnes, H\. Kobayashi, S\. Mariooryad, K\. Krasowiak, Zhixin, Lai, S\. Mourad, M\. Wang, F\. Bu, O\. Aharoni, G\. Chen, A\. Goyal, V\. Zubov, A\. Bapna, E\. Dabir, N\. Kothari, K\. Lamerigts, N\. D\. Cao, J\. Shar, C\. Yew, N\. Kulkarni, D\. Mahaarachchi, M\. Joshi, Z\. Zhu, J\. Lichtarge, Y\. Zhou, H\. Muckenhirn, V\. Selo, O\. Vinyals, P\. Chen, A\. Brohan, V\. Mehta, S\. Cogan, R\. Wang, T\. Geri, W\. Ko, W\. Chen, F\. Viola, K\. Shivam, L\. Wang, M\. C\. Elish, R\. A\. Popa, S\. Pereira, J\. Liu, R\. Koster, D\. Kim, G\. Zhang, S\. Ebrahimi, P\. Talukdar, Y\. Zheng, P\. Poklukar, A\. Mikhalap, D\. Johnson, A\. Vijayakumar, M\. Omernick, M\. Dibb, A\. Dubey, Q\. Hu, A\. Suman, V\. Aggarwal, I\. Kornakov, F\. Xia, W\. Lowe, A\. Kolganov, T\. Xiao, V\. Nikolaev, S\. Hemingray, B\. Li, J\. Iljazi, M\. Rybiński, B\. Sandhu, P\. Lu, T\. Luong, R\. Jenatton, V\. Govindaraj, Hui, Li, G\. Dulac\-Arnold, W\. Park, H\. Wang, A\. Modi, J\. Pouget\-Abadie, K\. Greller, R\. Gupta, R\. Berry, P\. Ramachandran, J\. Xie, L\. McCafferty, J\. Wang, K\. Gupta, H\. Lim, B\. Bratanič, A\. Brock, I\. Akolzin, J\. Sproch, D\. Karliner, D\. Kim, A\. Goedeckemeyer, N\. Shazeer, C\. Schmid, D\. Calandriello, P\. Bhatia, K\. Choromanski, C\. Montgomery, D\. Dua, A\. Ramalho, H\. King, Y\. Gao, L\. Nguyen, D\. Lindner, D\. Pitta, O\. Johnson, K\. Salama, D\. Ardila, M\. Han, E\. Farnese, S\. Odoom, Z\. Wang, X\. Ding, N\. Rink, R\. Smith, H\. T\. Lehri, E\. Cohen, N\. Vats, T\. He, P\. Gopavarapu, A\. Paszke, M\. Patel, W\. V\. Gansbeke, L\. Loher, L\. Castro, M\. Voitovich, T\. von Glehn, N\. George, S\. Niklaus, Z\. Eaton\-Rosen, N\. Rakićević, E\. Jue, S\. Perel, C\. Zhang, Y\. Bahat, A\. Pouget, Z\. Xing, F\. Huot, A\. Shenoy, T\. Bos, V\. Coriou, B\. Richter, N\. Noy, Y\. Wang, S\. Ontanon, S\. Qin, G\. Makarchuk, D\. Hassabis, Z\. Li, M\. Sharma, K\. Venkatesan, I\. Kemaev, R\. Daniel, S\. Huang, S\. Shah, O\. Ponce, Warren, Chen, M\. Faruqui, J\. Wu, S\. Andačić, S\. Payrits, D\. McDuff, T\. Hume, Y\. Cao, M\. Tessler, Q\. Wang, Y\. Wang, I\. Rendulic, E\. Agustsson, M\. Johnson, T\. Lando, A\. Howard, S\. G\. S\. Padmanabhan, M\. Daswani, A\. Banino, M\. Kilgore, J\. Heek, Z\. Ji, A\. Caceres, C\. Li, N\. Kassner, A\. Vlaskin, Z\. Liu, A\. Grills, Y\. Hou, R\. Sukkerd, G\. Cheon, N\. Shetty, L\. Markeeva, P\. Stanczyk, T\. Iyer, Y\. Gong, S\. Gao, K\. Gopalakrishnan, T\. Blyth, M\. Reynolds, A\. Bhoopchand, M\. Bilenko, D\. Gharibian, V\. Zayats, A\. Faust, A\. Singh, M\. Ma, H\. Jiao, S\. Vijayanarasimhan, L\. Aroyo, V\. Yadav, S\. Chakera, A\. Kakarla, V\. Meshram, K\. Gregor, G\. Botea, E\. Senter, D\. Jia, G\. Kovacs, N\. Sharma, S\. Baur, K\. Kang, Y\. He, L\. Zhuo, M\. Kostelac, I\. Laish, S\. Peng, L\. O’Bryan, D\. Kasenberg, G\. R\. Rao, E\. Leurent, B\. Zhang, S\. Stevens, A\. Salazar, Y\. Zhang, I\. Lobov, J\. Walker, A\. Porter, M\. Redshaw, H\. Ke, A\. Rao, A\. Lee, H\. Lam, M\. Moffitt, J\. Kim, S\. Qiao, T\. Koo, R\. Dadashi, X\. Song, M\. Sundararajan, P\. Xu, C\. Kawamoto, Y\. Zhong, C\. Barbu, A\. Reddy, M\. Verzetti, L\. Li, G\. Papamakarios, H\. Klimczak\-Plucińska, M\. Cassin, K\. Kavukcuoglu, R\. Swavely, A\. Vaucher, J\. Zhao, R\. Hemsley, M\. Tschannen, H\. Ge, G\. Menghani, Y\. Yu, N\. Ha, W\. He, X\. Wu, M\. Song, R\. Sterneck, S\. Zinke, D\. A\. Calian, A\. Marsden, A\. C\. Ruiz, M\. Hessel, A\. Gueta, B\. Lee, B\. Farris, M\. Gupta, Y\. Li, M\. Saleh, V\. Misra, K\. Xiao, P\. Mendolicchio, G\. Buttimore, V\. Krayvanova, N\. Nayakanti, M\. Wiethoff, Y\. Pande, A\. Mirhoseini, N\. Lao, J\. Liu, Y\. Hua, A\. Chen, Y\. Malkov, D\. Kalashnikov, S\. Gupta, K\. Audhkhasi, Y\. Zhai, S\. Kopalle, P\. Jain, E\. Ofek, C\. Meyer, K\. Baatarsukh, H\. Strejček, J\. Qian, J\. Freedman, R\. Figueira, M\. Sokolik, O\. Bachem, R\. Lin, D\. Kharrat, C\. Hidey, P\. Xu, D\. Duan, Y\. Li, M\. Ersoy, R\. Everett, K\. Cen, R\. Santamaria\-Fernandez, A\. Taubenfeld, I\. Mackinnon, L\. Deng, P\. Zablotskaia, S\. Viswanadha, S\. Goel, D\. Yates, Y\. Deng, P\. Choy, M\. Chen, A\. Sinha, A\. Mossin, Y\. Wang, A\. Szlam, S\. Hao, P\. K\. Rubenstein, M\. Toksoz\-Exley, M\. Aperghis, Y\. Zhong, J\. Ahn, M\. Isard, O\. Lacombe, F\. Luisier, C\. Anastasiou, Y\. Kalley, U\. Prabhu, E\. Dunleavy, S\. Bijwadia, J\. Mao\-Jones, K\. Chen, R\. Pasumarthi, E\. Wood, A\. Dostmohamed, N\. Hurley, J\. Simsa, A\. Parrish, M\. Pajarskas, M\. Harvey, O\. Skopek, Y\. Kochinski, J\. Rey, V\. Rieser, D\. Zhou, S\. J\. Lee, T\. Acharya, G\. Li, J\. Jiang, X\. Zhang, B\. Gipson, E\. Mahintorabi, M\. Gelmi, N\. Khajehnouri, A\. Yeh, K\. Lee, L\. Matthey, L\. Baker, T\. Pham, H\. Fu, A\. Pak, P\. Gupta, C\. Vasconcelos, A\. Sadovsky, B\. Walker, S\. Hsiao, P\. Zochbauer, A\. Marzoca, N\. Velan, J\. Zeng, G\. Baechler, D\. Driess, D\. Jain, Y\. Huang, L\. Tao, J\. Maggs, N\. Levine, J\. Schneider, E\. Gemzer, S\. Petit, S\. Han, Z\. Fisher, D\. Zelle, C\. Biles, E\. Ie, A\. Fadeeva, C\. Liu, J\. V\. Franco, A\. Collister, H\. Zhang, R\. Wang, R\. Zhao, L\. Kieliger, K\. Shuster, R\. Zhu, B\. Gong, L\. Chan, R\. Sun, S\. Basu, R\. Zimmermann, J\. Hayes, A\. Bapna, J\. Snoek, W\. Yang, P\. Datta, J\. A\. Abdallah, K\. Kilgour, L\. Li, S\. Mah, Y\. Jun, M\. Rivière, A\. Karmarkar, T\. Spalink, T\. Huang, L\. Gonzalez, D\. Tran, A\. Nowak, J\. Palowitch, M\. Chadwick, E\. Talius, H\. Mehta, T\. Sellam, P\. Fränken, M\. Nicosia, K\. He, A\. Kini, D\. Amos, S\. Basu, H\. Jobe, E\. Shaw, Q\. Xu, C\. Evans, D\. Ikeda, C\. Yan, L\. Jin, L\. Wang, S\. Yadav, I\. Labzovsky, R\. Sampath, A\. Ma, C\. Schumann, A\. Siddhant, R\. Shah, J\. Youssef, R\. Agarwal, N\. Dabney, A\. Tonioni, M\. Ambar, J\. Li, I\. Guyon, B\. Li, D\. Soergel, B\. Fang, G\. Karadzhov, C\. Udrescu, T\. Trinh, V\. Raunak, S\. Noury, D\. Guo, S\. Gupta, M\. Finkelstein, D\. Petek, L\. Liang, G\. Billock, P\. Sun, D\. Wood, Y\. Song, X\. Yu, T\. Matejovicova, R\. Cohen, K\. Andra, D\. D’Ambrosio, Z\. Deng, V\. Nallatamby, E\. Songhori, R\. Dangovski, A\. Lampinen, P\. Botadra, A\. Hillier, J\. Cao, N\. Baddi, A\. Kuncoro, T\. Yoshino, A\. Bhagatwala, M\. Ranzato, R\. Schaeffer, T\. Liu, S\. Ye, O\. Sarvana, J\. Nham, C\. Kuang, I\. Gao, J\. Baek, S\. Mittal, A\. Wahid, A\. Gergely, B\. Ni, J\. Feldman, C\. Muir, P\. Lamblin, W\. Macherey, E\. Dyer, L\. Kilpatrick, V\. Campos, M\. Bhutani, S\. Fort, Y\. Ahmad, A\. Severyn, K\. Chatziprimou, O\. Ferludin, M\. Dimarco, A\. Kusupati, J\. Heyward, D\. Bahir, K\. Villela, K\. Millican, D\. Marcus, S\. Bahargam, C\. Unlu, N\. Roth, Z\. Wei, S\. Gopal, D\. Ghoshal, E\. Lee, S\. Lin, J\. Lees, D\. Lee, A\. Hosseini, C\. Fan, S\. Neel, M\. Wu, Y\. Altun, H\. Cai, E\. Piqueras, J\. Woodward, A\. Bissacco, S\. Haykal, M\. Bordbar, P\. Sundaram, S\. Hodkinson, D\. Toyama, G\. Polovets, A\. Myers, A\. Sinha, T\. Levinboim, K\. Krishnakumar, R\. Chhaparia, T\. Sholokhova, N\. B\. Gundavarapu, G\. Jawahar, H\. Qureshi, J\. Hu, N\. Momchev, M\. Rahtz, R\. Wu, A\. P\. S, K\. Dhamdhere, M\. Guo, U\. Gupta, A\. Eslami, M\. Schain, M\. Blokzijl, D\. Welling, D\. Orr, L\. Bolelli, N\. Perez\-Nieves, M\. Sirotenko, A\. Prasad, A\. Kar, B\. D\. B\. Pigem, T\. Terzi, G\. Weisz, D\. Ghosh, A\. Mavalankar, D\. Madeka, K\. Daugaard, H\. Adam, V\. Shah, D\. Berman, M\. Tran, S\. Baker, E\. Andrejczuk, G\. Chole, G\. Raboshchuk, M\. Mirzazadeh, T\. Kagohara, S\. Wu, C\. Schallhart, B\. Orlando, C\. Wang, A\. Rrustemi, H\. Xiong, H\. Liu, A\. Vezer, N\. Ramsden, S\. Chang, S\. Mudgal, Y\. Li, N\. Vieillard, Y\. Hoshen, F\. Ahmad, A\. Slone, A\. Hua, N\. Potikha, M\. Rossini, J\. Stritar, S\. Prakash, Z\. Wang, X\. Dong, A\. Nazari, E\. Nehoran, K\. Tekelioglu, Y\. Li, K\. Badola, T\. Funkhouser, Y\. Li, V\. Yerram, R\. Ganeshan, D\. Formoso, K\. Langner, T\. Shi, H\. Li, Y\. Yamamori, A\. Panda, A\. Saade, A\. S\. Scarpati, C\. Breaux, C\. Carey, Z\. Zhou, C\. Hsieh, S\. Bridgers, A\. Butryna, N\. Gupta, V\. Tulsyan, S\. Woo, E\. Eltyshev, W\. Grathwohl, C\. Parks, S\. Benjamin, R\. Panigrahy, S\. Dodhia, D\. D\. Freitas, C\. Sauer, W\. Song, F\. Alet, J\. Tolins, C\. Paduraru, X\. Zhou, B\. Albert, Z\. Zhang, L\. Shu, M\. Bansal, S\. Nguyen, A\. Globerson, O\. Xiao, J\. Manyika, T\. Hennigan, R\. Rong, J\. Matak, A\. Bakalov, A\. Sharma, D\. Sinopalnikov, A\. Pierson, S\. Roller, G\. Brown, M\. Gao, T\. Fukuzawa, A\. Ghafouri, K\. Vassigh, I\. Barr, Z\. Wang, A\. Korsun, R\. Jayaram, L\. Ren, T\. Zaman, S\. Khan, Y\. Lunts, D\. Deutsch, D\. Uthus, N\. Katz, M\. Samsikova, A\. Khalifa, N\. Sethi, J\. Sun, L\. Tang, U\. Alon, X\. Luo, D\. Yu, A\. Nayyar, B\. Petrini, W\. Truong, V\. Hellendoorn, N\. Chinaev, C\. Alberti, W\. Wang, J\. Hu, V\. Mirrokni, A\. Balashankar, A\. Aharon, A\. Mehta, A\. Iscen, J\. Kready, L\. Manning, A\. Mohananey, Y\. Chen, A\. Tripathi, A\. Wu, I\. Petrovski, D\. Hwang, M\. Baeuml, S\. Chandrakaladharan, Y\. Liu, R\. Coaguila, M\. Chen, S\. Ma, P\. Tafti, S\. Tatineni, T\. Spitz, J\. Ye, P\. Vicol, M\. Rosca, A\. Puigdomènech, Z\. Yahav, S\. Ghemawat, H\. Lin, P\. Kirk, Z\. Nabulsi, S\. Brin, B\. Bohnet, K\. Caluwaerts, A\. S\. Veerubhotla, D\. Zheng, Z\. Dai, P\. Petrov, Y\. Xu, R\. Mehran, Z\. Xu, L\. Zintgraf, J\. Choi, S\. A\. Hombaiah, R\. Thoppilan, S\. Reddi, L\. Lew, L\. Li, K\. Webster, K\. Sawhney, L\. Lamprou, S\. Shakeri, M\. Lunayach, J\. Chen, S\. Bagri, A\. Salcianu, Y\. Chen, Y\. Donchev, C\. Magister, S\. Nørly, V\. Rodrigues, T\. Izo, H\. Noga, J\. Zou, T\. Köppe, W\. Zhou, K\. Lee, X\. Long, D\. Eisenbud, A\. Chen, C\. Schenck, C\. M\. To, P\. Zhong, E\. Taropa, M\. Truong, O\. Levy, D\. Martins, Z\. Zhang, C\. Semturs, K\. Zhang, A\. Yakubovich, P\. Moreno, L\. McConnaughey, D\. Lu, S\. Redmond, L\. Weerts, Y\. Bitton, T\. Refice, N\. Lacasse, A\. Conmy, C\. Tallec, J\. Odell, H\. Forbes\-Pollard, A\. Socala, J\. Hoech, P\. Kohli, A\. Walton, R\. Wang, M\. Sazanovich, K\. Zhu, A\. Kapishnikov, R\. Galt, M\. Denton, B\. Murdoch, C\. Sikora, K\. Mohamed, W\. Wei, U\. First, T\. McConnell, L\. C\. Cobo, J\. Qin, T\. Avrahami, D\. Balle, Y\. Watanabe, A\. Louis, A\. Kraft, S\. Ariafar, Y\. Gu, E\. Rives, C\. Yoon, A\. Rusu, J\. Cobon\-Kerr, C\. Hahn, J\. Luo, Yuvein, Zhu, N\. Ahuja, R\. Benenson, R\. L\. Kaufman, H\. Yu, L\. Hightower, J\. Zhang, D\. Ni, L\. A\. Hendricks, G\. Wang, G\. Yona, L\. Jain, P\. Barrio, S\. Bhupatiraju, S\. Velusamy, A\. Dafoe, S\. Riedel, T\. Thomas, Z\. Yuan, M\. Bellaiche, S\. Panthaplackel, K\. Kloboves, S\. Jauhari, C\. Akbulut, T\. Davchev, E\. Gladchenko, D\. Madras, A\. Chuklin, T\. Hill, Q\. Yuan, M\. Madhavan, L\. Leonhard, D\. Scandinaro, Q\. Chen, N\. Niu, A\. Douillard, B\. Damoc, Y\. Onoe, F\. Pedregosa, F\. Bertsch, C\. Leichner, J\. Pagadora, J\. Malmaud, S\. Ponda, A\. Twigg, O\. Duzhyi, J\. Shen, M\. Wang, R\. Garg, J\. Chen, U\. Evci, J\. Lee, L\. Liu, K\. Kojima, M\. Yamaguchi, A\. Rajendran, A\. Piergiovanni, V\. K\. Rajendran, M\. Fornoni, G\. Ibagon, H\. Ragan, S\. M\. Khan, J\. Blitzer, A\. Bunner, G\. Sun, T\. Kosakai, S\. Lundberg, N\. Elue, K\. Guu, S\. Park, J\. Park, A\. Narayanaswamy, C\. Wu, J\. Mudigonda, T\. Cohn, H\. Mu, R\. Kumar, L\. Graesser, Y\. Zhang, R\. Killam, V\. Zhuang, M\. Giménez, W\. A\. Jishi, R\. Ley\-Wild, A\. Zhai, K\. Osawa, D\. Cedillo, J\. Liu, M\. Upadhyay, M\. Sieniek, R\. Sharma, T\. Paine, A\. Angelova, S\. Addepalli, C\. Parada, K\. Majumder, A\. Lamp, S\. Kumar, X\. Deng, A\. Myaskovsky, T\. Sabolić, J\. Dudek, S\. York, F\. de Chaumont Quitry, J\. Nie, D\. Cattle, A\. Gunjan, B\. Piot, W\. Khawaja, S\. Bang, S\. Wang, S\. Khodadadeh, R\. R, P\. Rawlani, R\. Powell, K\. Lee, J\. Griesser, G\. Oh, C\. Magalhaes, Y\. Li, S\. Tokumine, H\. N\. Vogel, D\. Hsu, A\. BC, D\. Jindal, M\. Cohen, Z\. Yang, J\. Yuan, D\. de Cesare, T\. Bruguier, J\. Xu, M\. Roy, A\. Jacovi, D\. Belov, R\. Arya, P\. Meadowlark, S\. Cohen\-Ganor, W\. Ye, P\. Morris\-Suzuki, P\. Banzal, G\. Song, P\. Ponnuramu, F\. Zhang, G\. Scrivener, S\. Zaiem, A\. R\. Rochman, K\. Han, B\. Ghazi, K\. Lee, S\. Drath, D\. Suo, A\. Girgis, P\. Shenoy, D\. Nguyen, D\. Eck, S\. Gupta, L\. Yan, J\. Carreira, A\. Gulati, R\. Sang, D\. Mirylenka, E\. Cooney, E\. Chou, M\. Ling, C\. Fan, B\. Coleman, G\. Tubone, R\. Kumar, J\. Baldridge, F\. Hernandez\-Campos, A\. Lazaridou, J\. Besley, I\. Yona, N\. Bulut, Q\. Wellens, A\. Pierigiovanni, J\. George, R\. Green, P\. Han, C\. Tao, G\. Clark, C\. You, A\. Abdolmaleki, J\. Fu, T\. Chen, A\. Chaugule, A\. Chandorkar, A\. Rahman, W\. Thompson, P\. Koanantakool, M\. Bernico, J\. Ren, A\. Vlasov, S\. Vassilvitskii, M\. Kula, Y\. Liang, D\. Kim, Y\. Huang, C\. Ye, D\. Lepikhin, and W\. Helmholz \(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.5.2)\.
- Y\. Cui, Z\. Yang, and X\. Yao \(2024\)Efficient and effective text encoding for chinese llama and alpaca\.External Links:2304\.08177,[Link](https://arxiv.org/abs/2304.08177)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- J\. Dang, S\. Singh, D\. D’souza, A\. Ahmadian, A\. Salamanca, M\. Smith, A\. Peppin, S\. Hong, M\. Govindassamy, T\. Zhao, S\. Kublik, M\. Amer, V\. Aryabumi, J\. A\. Campos, Y\. Tan, T\. Kocmi, F\. Strub, N\. Grinsztajn, Y\. Flet\-Berliac, A\. Locatelli, H\. Lin, D\. Talupuru, B\. Venkitesh, D\. Cairuz, B\. Yang, T\. Chung, W\. Ko, S\. S\. Shi, A\. Shukayev, S\. Bae, A\. Piktus, R\. Castagné, F\. Cruz\-Salinas, E\. Kim, L\. Crawhall\-Stein, A\. Morisot, S\. Roy, P\. Blunsom, I\. Zhang, A\. Gomez, N\. Frosst, M\. Fadaee, B\. Ermis, A\. Üstün, and S\. Hooker \(2024\)Aya expanse: combining research breakthroughs for a new multilingual frontier\.External Links:2412\.04261,[Link](https://arxiv.org/abs/2412.04261)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.15.2)\.
- M\. A\. Daoud, C\. Abouzahir, L\. Kharouf, W\. Al\-Eisawi, N\. Habash, and F\. E\. Shamout \(2025\)MedArabiQ: benchmarking large language models on arabic medical tasks\.InProceedings of the 10th Machine Learning for Healthcare Conference,M\. Agrawal, K\. Deshpande, M\. Engelhard, S\. Joshi, S\. Tang, and I\. Urteaga \(Eds\.\),Proceedings of Machine Learning Research, Vol\.298\.External Links:[Link](https://proceedings.mlr.press/v298/daoud25a.html)Cited by:[Table S1](https://arxiv.org/html/2608.00207#A1.T1.1.3.1),[§1](https://arxiv.org/html/2608.00207#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.00207#S2.SS2.p1.1),[§5\.2\.1](https://arxiv.org/html/2608.00207#S5.SS2.SSS1.p1.1),[§6\.1](https://arxiv.org/html/2608.00207#S6.SS1.p1.1)\.
- M\. A\. Daoud, L\. Kharouf, O\. E\. Hajj, D\. E\. Samad, M\. Al\-Omari, J\. Mallat, K\. Saleh, N\. Habash, and F\. E\. Shamout \(2026\)MedAraBench: large\-scale arabic medical question answering dataset and benchmark\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=1BXojAgNrg)Cited by:[Table S1](https://arxiv.org/html/2608.00207#A1.T1.1.2.1),[Appendix D](https://arxiv.org/html/2608.00207#A4.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2608.00207#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.00207#S2.SS2.p1.1),[§3](https://arxiv.org/html/2608.00207#S3.p1.1),[§5\.2\.1](https://arxiv.org/html/2608.00207#S5.SS2.SSS1.p1.1)\.
- DeepSeek\-AI, A\. Liu, A\. Mei, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Lu, C\. Zhao, C\. Deng, C\. Xu, C\. Ruan, D\. Dai, D\. Guo, D\. Yang, D\. Chen, E\. Li, F\. Zhou, F\. Lin, F\. Dai, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Xu, H\. Li, H\. Liang, H\. Wei, H\. Zhang, H\. Luo, H\. Ji, H\. Ding, H\. Tang, H\. Cao, H\. Gao, H\. Qu, H\. Zeng, J\. Huang, J\. Li, J\. Xu, J\. Hu, J\. Chen, J\. Xiang, J\. Yuan, J\. Cheng, J\. Zhu, J\. Ran, J\. Jiang, J\. Qiu, J\. Li, J\. Song, K\. Dong, K\. Gao, K\. Guan, K\. Huang, K\. Zhou, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Wang, L\. Zhao, L\. Yin, L\. Guo, L\. Luo, L\. Ma, L\. Wang, L\. Zhang, M\. S\. Di, M\. Y\. Xu, M\. Zhang, M\. Zhang, M\. Tang, M\. Zhou, P\. Huang, P\. Cong, P\. Wang, Q\. Wang, Q\. Zhu, Q\. Li, Q\. Chen, Q\. Du, R\. Xu, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. Yin, R\. Xu, R\. Shen, R\. Zhang, S\. H\. Liu, S\. Lu, S\. Zhou, S\. Chen, S\. Cai, S\. Chen, S\. Hu, S\. Liu, S\. Hu, S\. Ma, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. Zhou, T\. Ni, T\. Yun, T\. Pei, T\. Ye, T\. Yue, W\. Zeng, W\. Liu, W\. Liang, W\. Pang, W\. Luo, W\. Gao, W\. Zhang, X\. Gao, X\. Wang, X\. Bi, X\. Liu, X\. Wang, X\. Chen, X\. Zhang, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yu, X\. Li, X\. Yang, X\. Li, X\. Chen, X\. Su, X\. Pan, X\. Lin, X\. Fu, Y\. Q\. Wang, Y\. Zhang, Y\. Xu, Y\. Ma, Y\. Li, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Qian, Y\. Yu, Y\. Zhang, Y\. Ding, Y\. Shi, Y\. Xiong, Y\. He, Y\. Zhou, Y\. Zhong, Y\. Piao, Y\. Wang, Y\. Chen, Y\. Tan, Y\. Wei, Y\. Ma, Y\. Liu, Y\. Yang, Y\. Guo, Y\. Wu, Y\. Wu, Y\. Cheng, Y\. Ou, Y\. Xu, Y\. Wang, Y\. Gong, Y\. Wu, Y\. Zou, Y\. Li, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Z\. F\. Wu, Z\. Z\. Ren, Z\. Zhao, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Gou, Z\. Ma, Z\. Yan, Z\. Shao, Z\. Huang, Z\. Wu, Z\. Li, Z\. Zhang, Z\. Xu, Z\. Wang, Z\. Gu, Z\. Zhu, Z\. Li, Z\. Zhang, Z\. Xie, Z\. Gao, Z\. Pan, Z\. Yao, B\. Feng, H\. Li, J\. L\. Cai, J\. Ni, L\. Xu, M\. Li, N\. Tian, R\. J\. Chen, R\. L\. Jin, S\. S\. Li, S\. Zhou, T\. Sun, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Song, X\. Zhou, Y\. X\. Zhu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Z\. Huang, Z\. Xu, Z\. Zhang, D\. Ji, J\. Liang, J\. Guo, J\. Chen, L\. Xia, M\. Wang, M\. Li, P\. Zhang, R\. Chen, S\. Sun, S\. Wu, S\. Ye, T\. Wang, W\. L\. Xiao, W\. An, X\. Wang, X\. Sun, X\. Wang, Y\. Tang, Y\. Zha, Z\. Zhang, Z\. Ju, Z\. Zhang, and Z\. Qu \(2025\)DeepSeek\-v3\.2: pushing the frontier of open large language models\.External Links:2512\.02556,[Link](https://arxiv.org/abs/2512.02556)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.12.2)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.11.2),[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.8.2)\.
- N\. Y\. Habash \(2010\)Introduction to Arabic natural language processing\.1 edition,Synthesis Lectures on Human Language Technologies,Springer Cham\.External Links:ISBN 978\-3\-031\-01011\-8,[Document](https://dx.doi.org/10.1007/978-3-031-02139-8),[Link](https://doi.org/10.1007/978-3-031-02139-8)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p2.1)\.
- K\. Hämmerl, J\. Libovický, and A\. Fraser \(2024\)Understanding cross\-lingual Alignment—A survey\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 10922–10943\.External Links:[Link](https://aclanthology.org/2024.findings-acl.649/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.649)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1)\.
- E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2022\)LoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[Appendix D](https://arxiv.org/html/2608.00207#A4.SS0.SSS0.Px1.p1.3)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2021\)LoRA: low\-rank adaptation of large language models\.External Links:2106\.09685,[Link](https://arxiv.org/abs/2106.09685)Cited by:[§5\.3](https://arxiv.org/html/2608.00207#S5.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.27.2)\.
- P\. Hu, S\. Liu, C\. Gao, X\. Huang, X\. Han, J\. Feng, C\. Deng, and S\. Huang \(2025\)Large language models are cross\-lingual knowledge\-free reasoners\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 1525–1542\.External Links:[Link](https://aclanthology.org/2025.naacl-long.72/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.72),ISBN 979\-8\-89176\-189\-6Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1)\.
- H\. Huang, T\. Tang, D\. Zhang, X\. Zhao, T\. Song, Y\. Xia, and F\. Wei \(2023\)Not all languages are created equal in LLMs: improving multilingual capability by cross\-lingual\-thought prompting\.InThe 2023 Conference on Empirical Methods in Natural Language Processing,External Links:[Link](https://openreview.net/forum?id=E4ebDehO3O)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- Z\. Huang, W\. Zhu, G\. Cheng, L\. Li, and F\. Yuan \(2024a\)MindMerger: efficient boosting llm reasoning in non\-english languages\.External Links:2405\.17386,[Link](https://arxiv.org/abs/2405.17386)Cited by:[§5\.3](https://arxiv.org/html/2608.00207#S5.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.25.2)\.
- Z\. Huang, W\. Zhu, G\. Cheng, L\. Li, and F\. Yuan \(2024b\)MindMerger: efficiently boosting LLM reasoning in non\-english languages\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=Oq32ylAOu2)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- M\. Ifergan, L\. Choshen, R\. Aharoni, I\. Szpektor, and O\. Abend \(2025\)Beneath the surface of consistency: exploring cross\-lingual knowledge representation sharing in LLMs\.InFindings of the Association for Computational Linguistics: NAACL 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 4630–4644\.External Links:[Link](https://aclanthology.org/2025.findings-naacl.475/),ISBN 979\-8\-89176\-195\-7Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2023a\)Mistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.7.2)\.
- L\. Y\. Jiang, X\. C\. Liu, N\. P\. Nejatian,et al\.\(2023b\)Health system\-scale language models are all\-purpose prediction engines\.Nature619,pp\. 357–362\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06160-y)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- Q\. Jin, B\. Dhingra, Z\. Liu, W\. Cohen, and X\. Lu \(2019\)PubMedQA: a dataset for biomedical research question answering\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2567–2577\.External Links:[Link](https://aclanthology.org/D19-1259/),[Document](https://dx.doi.org/10.18653/v1/D19-1259)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- Y\. Jin, M\. Chandra, G\. Verma, Y\. Hu, M\. D\. Choudhury, and S\. Kumar \(2023\)Better to ask in english: cross\-lingual evaluation of large language models for healthcare queries\.External Links:2310\.13132,[Link](https://arxiv.org/abs/2310.13132)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- Y\. Jin, M\. Chandra, G\. Verma, Y\. Hu, M\. De Choudhury, and S\. Kumar \(2024\)Better to ask in english: cross\-lingual evaluation of large language models for healthcare queries\.InProceedings of the ACM Web Conference 2024,WWW ’24,New York, NY, USA,pp\. 2627–2638\.External Links:ISBN 9798400701719,[Link](https://doi.org/10.1145/3589334.3645643),[Document](https://dx.doi.org/10.1145/3589334.3645643)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p3.1)\.
- P\. Joshi, S\. Santy, A\. Budhiraja, K\. Bali, and M\. Choudhury \(2020\)The state and fate of linguistic diversity and inclusion in the NLP world\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 6282–6293\.External Links:[Link](https://aclanthology.org/2020.acl-main.560/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.560)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- A\. H\. Kargaran, A\. Modarressi, N\. Nikeghbal, J\. Diesner, F\. Yvon, and H\. Schuetze \(2025\)MEXA: multilingual evaluation of English\-centric LLMs via cross\-lingual alignment\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 27001–27023\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1385/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1385),ISBN 979\-8\-89176\-256\-5Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1)\.
- K\. Kobayashi, Z\. Wan, F\. Cheng, Y\. Tsuta, X\. Zhao, J\. Jiang, J\. Huang, Z\. Huang, Y\. Oda, R\. Yokota, Y\. Arase, D\. Kawahara, A\. Aizawa, and S\. Kurohashi \(2025\)Leveraging high\-resource English corpora for cross\-lingual domain adaptation in low\-resource Japanese medicine via continued pre\-training\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 11469–11488\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.615/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.615),ISBN 979\-8\-89176\-335\-7Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p2.1)\.
- F\. Koto, H\. Li, S\. Shatnawi, J\. Doughman, A\. Sadallah, A\. Alraeesi, K\. Almubarak, Z\. Alyafeai, N\. Sengupta, S\. Shehata, N\. Habash, P\. Nakov, and T\. Baldwin \(2024\)ArabicMMLU: assessing massive multitask language understanding in Arabic\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5622–5640\.External Links:[Link](https://aclanthology.org/2024.findings-acl.334/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.334)Cited by:[Table S1](https://arxiv.org/html/2608.00207#A1.T1.1.4.1),[§5\.2\.1](https://arxiv.org/html/2608.00207#S5.SS2.SSS1.p1.1)\.
- D\. Liu and J\. Niehues \(2025\)Conditions for catastrophic forgetting in multilingual translation\.InProceedings of the 5th Workshop on Multilingual Representation Learning \(MRL 2025\),D\. I\. Adelani, C\. Arnett, D\. Ataman, T\. A\. Chang, H\. Gonen, R\. Raja, F\. Schmidt, D\. Stap, and J\. Wang \(Eds\.\),Suzhuo, China,pp\. 347–359\.External Links:[Link](https://aclanthology.org/2025.mrl-main.23/),[Document](https://dx.doi.org/10.18653/v1/2025.mrl-main.23),ISBN 979\-8\-89176\-345\-6Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in GPT\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 17359–17372\.External Links:[Document](https://dx.doi.org/10.52202/068431-1262),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html)Cited by:[§3](https://arxiv.org/html/2608.00207#S3.p2.1)\.
- Mistral AI \(2025\)Mistral\-Small\-3\.2\-24B\-Instruct\-2506\.Note:[https://huggingface\.co/mistralai/Mistral\-Small\-3\.2\-24B\-Instruct\-2506](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506)Hugging Face model card\. Accessed: 2026\-05\-03Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.9.2)\.
- Y\. A\. Moaiad, M\. Alobed, M\. Alsakhnini, and A\. M\. Momani \(2024\)Challenges in natural arabic language processing\.Edelweiss Applied Science and Technology\.External Links:[Link](https://api.semanticscholar.org/CorpusID:274076822)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p2.1)\.
- Z\. A\. Nazi and W\. Peng \(2024\)Large language models in healthcare and medical domain: a review\.External Links:2401\.06775,[Link](https://arxiv.org/abs/2401.06775)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- H\. Nori, N\. King, S\. M\. McKinney, D\. Carignan, and E\. Horvitz \(2023\)Capabilities of gpt\-4 on medical challenge problems\.External Links:2303\.13375,[Link](https://arxiv.org/abs/2303.13375)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p1.1)\.
- OpenAI \(2025\)Introducing GPT\-5\.2\.Note:[https://openai\.com/index/introducing\-gpt\-5\-2/](https://openai.com/index/introducing-gpt-5-2/)Accessed: 2026\-05\-03Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.4.2)\.
- A\. Pal, L\. K\. Umapathi, and M\. Sankarasubbu \(2022\)MedMCQA: a large\-scale multi\-subject multi\-choice dataset for medical domain question answering\.InProceedings of the Conference on Health, Inference, and Learning,G\. Flores, G\. H\. Chen, T\. Pollard, J\. C\. Ho, and T\. Naumann \(Eds\.\),Proceedings of Machine Learning Research, Vol\.174,pp\. 248–260\.External Links:[Link](https://proceedings.mlr.press/v174/pal22a.html)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- S\. Pieri, S\. S\. Mullappilly, F\. S\. Khan, R\. M\. Anwer, S\. Khan, T\. Baldwin, and H\. Cholakkal \(2024\)BiMediX: bilingual medical mixture of experts LLM\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 16984–17002\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.989/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.989)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.00207#S2.SS2.p1.1),[§5\.3](https://arxiv.org/html/2608.00207#S5.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.26.2)\.
- P\. Qiu, C\. Wu, X\. Zhang,et al\.\(2024\)Towards building multilingual language model for medicine\.Nature Communications15,pp\. 8384\.External Links:[Document](https://dx.doi.org/10.1038/s41467-024-52417-z)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- K\. Ravisankar, H\. Han, S\. Wiegreffe, and M\. Carpuat \(2026\)Can you map it to English? the role of cross\-lingual alignment in the multilingual performance of LLMs\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 4854–4872\.External Links:[Link](https://aclanthology.org/2026.eacl-long.225/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.225),ISBN 979\-8\-89176\-380\-7Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1)\.
- N\. Saadi, T\. Raha, C\. Christophe, M\. A\. Pimentel, R\. Rajan, and P\. Kanithi \(2025\)Bridging language barriers in healthcare: a study on arabic LLMs\.InWorkshop on Large Language Models and Generative AI for Health at AAAI 2025,External Links:[Link](https://openreview.net/forum?id=Fpez4GovLN)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p3.1)\.
- A\. Sallinen, A\. Solergibert, M\. Zhang, G\. Boyé, M\. Dupont\-Roc, X\. Theimer\-Lienhard, E\. Boisson, B\. Bernath, H\. Hadhri, A\. Tran,et al\.\(2025\)Llama\-3\-meditron: an open\-weight suite of medical llms based on llama\-3\.1\.InWorkshop on Large Language Models and Generative AI for Health at AAAI 2025,Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.20.2)\.
- T\. L\. Scao, A\. Fan, C\. Akiki,et al\.\(2022\)BLOOM: a 176b\-parameter open\-access multilingual language model\.External Links:2211\.05100Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- L\. Schut, Y\. Gal, and S\. Farquhar \(2025\)Do multilingual LLMs think in english?\.InICLR 2025 Workshop on Building Trust in Language Models and Applications,External Links:[Link](https://openreview.net/forum?id=I8BOtOPcOv)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1)\.
- A\. Sellergren, S\. Kazemzadeh, T\. Jaroensri, A\. Kiraly, M\. Traverse, T\. Kohlberger, S\. Xu, F\. Jamil, C\. Hughes, C\. Lau, J\. Chen, F\. Mahvar, L\. Yatziv, T\. Chen, B\. Sterling, S\. A\. Baby, S\. M\. Baby, J\. Lai, S\. Schmidgall, L\. Yang, K\. Chen, P\. Bjornsson, S\. Reddy, R\. Brush, K\. Philbrick, M\. Asiedu, I\. Mezerreg, H\. Hu, H\. Yang, R\. Tiwari, S\. Jansen, P\. Singh, Y\. Liu, S\. Azizi, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Riviere, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Buchatskaya, J\. Alayrac, D\. Lepikhin, V\. Feinberg, S\. Borgeaud, A\. Andreev, C\. Hardin, R\. Dadashi, L\. Hussenot, A\. Joulin, O\. Bachem, Y\. Matias, K\. Chou, A\. Hassidim, K\. Goel, C\. Farabet, J\. Barral, T\. Warkentin, J\. Shlens, D\. Fleet, V\. Cotruta, O\. Sanseviero, G\. Martins, P\. Kirk, A\. Rao, S\. Shetty, D\. F\. Steiner, C\. Kirmizibayrak, R\. Pilgrim, D\. Golden, and L\. Yang \(2026\)MedGemma technical report\.External Links:2507\.05201,[Link](https://arxiv.org/abs/2507.05201)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.19.2)\.
- SILMA AI \(2026\)SILMA\-9B\-Instruct\-v1\.0\.Note:[https://huggingface\.co/silma\-ai/SILMA\-9B\-Instruct\-v1\.0](https://huggingface.co/silma-ai/SILMA-9B-Instruct-v1.0)Hugging Face model card\. Accessed: 2026\-05\-03Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.16.2)\.
- K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl, P\. Payne, M\. Seneviratne, P\. Gamble, C\. Kelly, A\. Babiker, N\. Schärli, A\. Chowdhery, P\. Mansfield, D\. Demner\-Fushman, B\. Agüera y Arcas, D\. Webster, G\. S\. Corrado, Y\. Matias, K\. Chou, J\. Gottweis, N\. Tomasev, Y\. Liu, A\. Rajkomar, J\. Barral, C\. Semturs, A\. Karthikesalingam, and V\. Natarajan \(2023\)Large language models encode clinical knowledge\.Nature620\(7972\),pp\. 172–180\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06291-2),[Link](https://doi.org/10.1038/s41586-023-06291-2)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p1.1)\.
- K\. Singhal, T\. Tu, J\. Gottweis,et al\.\(2025\)Toward expert\-level medical question answering with large language models\.Nature Medicine31,pp\. 943–950\.External Links:[Document](https://dx.doi.org/10.1038/s41591-024-03423-7)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p2.1)\.
- Statista \(2023\)The most spoken languages worldwide in 2023\.Note:[https://www\.statista\.com/statistics/266808/the\-most\-spoken\-languages\-worldwide/](https://www.statista.com/statistics/266808/the-most-spoken-languages-worldwide/)Accessed: 2026\-04\-20Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p2.1)\.
- F\. Team, U\. Abbas, M\. S\. Ahmad, F\. Alam, E\. Altinisik, E\. Asgari, Y\. Boshmaf, S\. Boughorbel, S\. Chawla, S\. Chowdhury, F\. Dalvi, K\. Darwish, N\. Durrani, M\. Elfeky, A\. Elmagarmid, M\. Eltabakh, M\. Fatehkia, A\. Fragkopoulos, M\. Hasanain, M\. Hawasly, M\. Husaini, S\. Jung, J\. K\. Lucas, W\. Magdy, S\. Messaoud, A\. Mohamed, T\. Mohiuddin, B\. Mousi, H\. Mubarak, A\. Musleh, Z\. Naeem, M\. Ouzzani, D\. Popovic, A\. Sadeghi, H\. T\. Sencar, M\. Shinoy, O\. Sinan, Y\. Zhang, A\. Ali, Y\. E\. Kheir, X\. Ma, and C\. Ruan \(2025a\)Fanar: an arabic\-centric multimodal generative ai platform\.External Links:2501\.13944,[Link](https://arxiv.org/abs/2501.13944)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.18.2)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot \(2025b\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.10.2)\.
- H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar, A\. Rodriguez, A\. Joulin, E\. Grave, and G\. Lample \(2023\)LLaMA: open and efficient foundation language models\.External Links:2302\.13971,[Link](https://arxiv.org/abs/2302.13971)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- T\. Tu, M\. Schaekermann, A\. Palepu,et al\.\(2025\)Towards conversational diagnostic artificial intelligence\.Nature642,pp\. 442–450\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-08866-7)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p2.1)\.
- T\. Vu, A\. Barua, B\. Lester, D\. Cer, M\. Iyyer, and N\. Constant \(2022\)Overcoming catastrophic forgetting in zero\-shot cross\-lingual generation\.External Links:2205\.12647,[Link](https://arxiv.org/abs/2205.12647)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- H\. Wang, C\. Liu, N\. Xi, Z\. Qiang, S\. Zhao, B\. Qin, and T\. Liu \(2023a\)HuaTuo: tuning llama model with chinese medical knowledge\.External Links:2304\.06975,[Link](https://arxiv.org/abs/2304.06975)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- X\. Wang, N\. Chen, J\. Chen, Y\. Wang, G\. Zhen, C\. Zhang, X\. Wu, Y\. Hu, A\. Gao, X\. Wan, H\. Li, and B\. Wang \(2024\)Apollo: a lightweight multilingual medical llm towards democratizing medical ai to 6b people\.External Links:2403\.03640,[Link](https://arxiv.org/abs/2403.03640)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p3.1)\.
- Y\. Wang, Y\. Zhao, and L\. Petzold \(2023b\)Are large language models ready for healthcare? a comparative study on clinical language understanding\.InProceedings of the 8th Machine Learning for Healthcare Conference,K\. Deshpande, M\. Fiterau, S\. Joshi, Z\. Lipton, R\. Ranganath, I\. Urteaga, and S\. Yeung \(Eds\.\),Proceedings of Machine Learning Research, Vol\.219,pp\. 804–823\.External Links:[Link](https://proceedings.mlr.press/v219/wang23c.html)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p1.1)\.
- C\. Wendler, V\. Veselovsky, G\. Monea, and R\. West \(2024\)Do llamas work in English? on the latent language of multilingual transformers\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 15366–15394\.External Links:[Link](https://aclanthology.org/2024.acl-long.820/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.820)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p1.1),[§3](https://arxiv.org/html/2608.00207#S3.p2.1)\.
- D\. Yoon, J\. Jang, S\. Kim, S\. Kim, S\. Shafayat, and M\. Seo \(2024\)LangBridge: multilingual reasoning without multilingual supervision\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 7502–7522\.External Links:[Link](https://aclanthology.org/2024.acl-long.405/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.405)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- Y\. Zhang, Q\. Chen, M\. Li, W\. Che, and L\. Qin \(2024\)AutoCAP: towards automatic cross\-lingual alignment planning for zero\-shot chain\-of\-thought\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 9191–9200\.Cited by:[§5\.3](https://arxiv.org/html/2608.00207#S5.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.23.2)\.
- W\. Zhao, Y\. Hu, J\. Guo, X\. Sui, T\. Wu, Y\. Deng, Y\. Zhao, B\. Qin, W\. Che, and T\. Liu \(2025\)Lens: rethinking multilingual enhancement for large language models\.External Links:2410\.04407,[Link](https://arxiv.org/abs/2410.04407)Cited by:[§1](https://arxiv.org/html/2608.00207#S1.p3.1)\.
- X\. Zhao, N\. Yoshinaga, Y\. Tsuta, and A\. Aizawa \(2026\)Tracing multilingual knowledge acquisition dynamics in domain adaptation: a case study of biomedical adaptation\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 5739–5760\.External Links:[Link](https://aclanthology.org/2026.eacl-long.269/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.269),ISBN 979\-8\-89176\-380\-7Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p2.1)\.
- G\. Zheng, X\. Wang, J\. Liang, N\. Chen, Y\. Zheng, and B\. Wang \(2025\)Efficiently democratizing medical LLMs for 50 languages via a mixture of language family experts\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=JSB171dSUU)Cited by:[§2\.1](https://arxiv.org/html/2608.00207#S2.SS1.p3.1)\.
- Y\. Zhou, X\. Liu, X\. Zhang, C\. Ning, S\. Wang, G\. Hu, and J\. Wu \(2026\)Investigating and mitigating catastrophic forgetting in medical knowledge injection through internal knowledge augmentation learning\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=i9RDDi2SZC)Cited by:[§2\.3](https://arxiv.org/html/2608.00207#S2.SS3.p3.1)\.
- J\. Zuo, M\. Velikanov, I\. Chahed, Y\. Belkada, D\. E\. Rhayem, G\. Kunsch, H\. Hacid, H\. Yous, B\. Farhat, I\. Khadraoui, M\. Farooq, G\. Campesan, R\. Cojocaru, Y\. Djilali, S\. Hu, I\. Chaabane, P\. Khanna, M\. E\. A\. Seddik, N\. D\. Huynh, P\. L\. Khac, L\. AlQadi, B\. Mokeddem, M\. Chami, A\. Abubaker, M\. Lubinets, K\. Piskorski, and S\. Frikha \(2025\)Falcon\-h1: a family of hybrid\-head language models redefining efficiency and performance\.External Links:2507\.22448,[Link](https://arxiv.org/abs/2507.22448)Cited by:[Table 2](https://arxiv.org/html/2608.00207#S5.T2.1.1.17.2)\.

## Appendix AExperimental Setup

### A\.1Datasets

Table[S1](https://arxiv.org/html/2608.00207#A1.T1)summarizes the six benchmarks used in our evaluation, spanning multiple\-choice QA, short\-answer generation, and multi\-turn dialogue tasks\. AraClinicDialog is introduced in this work, MedAraBench\-OE repurposes MedAraBench’s existing questions into an open\-ended format, and the remaining four are existing benchmarks used as\-is\.

DatasetTaskSizeDomainMedAraBench\(Daoudet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib8)\)MCQA4,959Clinical medicineMedArabiQ\(Daoudet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib5)\)MCQA100Clinical medicineArabicMMLUKotoet al\.\([2024](https://arxiv.org/html/2608.00207#bib.bib9)\)MCQA1,072Medicine & biologyAraSTEMBoussahaet al\.\([2025](https://arxiv.org/html/2608.00207#bib.bib10)\)MCQA721STEM / medicineMedAraBench\-OEShort Answer Generation940Clinical medicineAraClinicDialogMulti\-Turn Dialogue100Clinical medicineTable S1:Dataset statistics for all evaluation benchmarks\.
### A\.2Prompts

Prompts are defined per task\. For the MCQA task, all baselines share a single zero\-shot prompt\. For the Short Answer Generation and the Multi\-Turn Clinical Dialogue tasks, we use a one\-shot setup to ensure models adhere to the required answer format\. The few\-shot baselines are the exception, which use few\-shot prompt variants\.

The Short Answer Generation and the Multi\-Turn Clinical Dialogue tasks are evaluated using an LLM\-as\-a\-judge protocol, in addition to automatic metrics such as BERTScore\. The corresponding judge prompts are provided in the relevant sections\.

MCQA\.The zero\-shot prompt is shown in Figure[S1](https://arxiv.org/html/2608.00207#A1.F1)\. The few\-shot variant extends it with five exemplar question\-answer pairs\. The exemplars are dataset\-specific, randomly sampled with a fixed seed \(42\), and excluded from the inference set\.

Zero\-shot Prompt \(MCQA task\)You are a medical expert answering multiple\-choice exam questions\.You will receive exactly ONE question followed by answer options labeled:A\), B\), C\), D\), E\), and sometimes F\)\.You must output exactly ONE line in this format:ANSWER: ¡LETTER¿Rules:\- Output ONLY that line\.\- Do NOT repeat or paraphrase the question\.\- Do NOT translate anything\.\- Do NOT explain your reasoning\.\- Do NOT list the options\.Figure S1:Zero\-shot prompt for the MCQA task\.Short Answer Generation\.The prompt is shown in Figure[S2](https://arxiv.org/html/2608.00207#A1.F2)\. We use a one\-shot rather than a zero\-shot setup because, under zero\-shot, most models failed to follow the short\-answer instruction and produced verbose explanatory responses; a single in\-context example was sufficient to enforce the intended output format\. The example therefore calibrates both answer style and length\.

As in the MCQA task, we define a separate few\-shot prompt variant for that setting: the base prompt is extended with five exemplar question\-answer pairs, which are randomly sampled using a fixed seed \(42\) and are excluded from the inference set\.

For LLM\-as\-a\-judge evaluation, the judge prompt \(Figure[S3](https://arxiv.org/html/2608.00207#A1.F3)\) takes the question, reference answer, and generated answer as input and returns a binary correctness label\.

One\-shot Prompt \(Short Answer Generation\)You are a medical expert\.You will receive exactly ONE medical exam question with no answer options\. Answer the question based on your medical knowledge and return your answer as a short medical term or phrase\.Output format:ANSWER: ¡short answer¿Rules:\- Output exactly ONE line starting with “ANSWER: “\.\- The answer must be a concise medical term or phrase, typically 1\-4 words, rarely more than 10 words\.\- All answers must be in Arabic\.\- Do NOT write a full sentence or explanation\.\- Do NOT explain your reasoning\.\- Do NOT add extra commentary\.\- Do NOT repeat the question\.\- Do NOT translate\.\- Do NOT include multiple answers\.Example:Question:\\arabicfontfamilyأين تتوضع الثقبة العوراء في اللسان؟ANSWER:\\arabicfontfamilyخلف الكلم الإنتهائيFigure S2:One\-shot prompt for the Short Answer Generation\.LLM\-as\-a\-Judge Prompt \(Short Answer Generation task\)You are an expert medical evaluator\. You will be given a medical question, a reference answer, and a generated answer\. Your task is to evaluate the generated answer by selecting exactly one label from the following options and responding only with the label in brackets \[\]\.Question: \{question\_stem\}Reference Answer: \{reference\_answer\}Generated Answer: \{generated\_answer\}Evaluate the generated answer against the reference answer using the following criterion:Correctness: Correct / Incorrect\- Correct: The generated answer matches the reference answer in meaning\. Synonyms, equivalent medical terms, or slight phrasing differences still count as correct\.\- Incorrect: The generated answer is wrong, irrelevant, contradicts the reference answer, or is too vague to be credited\.Respond only with one of the following in brackets: \[Correct\] / \[Incorrect\]Figure S3:LLM\-as\-a\-judge prompt for the Short Answer Generation task\.Multi\-Turn Clinical Dialogue\.The prompt is shown in Figure[S4](https://arxiv.org/html/2608.00207#A1.F4)\. Similar to the Short Answer Generation task, we use a one\-shot setup due to models struggling to follow the expected output format under zero\-shot\.

For the few shot baseline, we define a separate few\-shot prompt variant for that setting: the base prompt is extended with five carefully constructed question\-answer pairs to encourage the model to produce responses that reflect the intended reasoning approach and structure\.

For the LLM\-as\-a\-judge evaluation metric, the judge prompt \(Figure[S5](https://arxiv.org/html/2608.00207#A1.F5)\) takes the primary reasoning objective, red\-flag symptoms, dialogue, and generated answer as input and assigns one of three labels along each of three axes: Reasoning Match, Safety, and Communication\. This design reflects the multi\-dimensional nature of the task, which requires not only clinically appropriate reasoning but also safe medical guidance and effective communication\.

One\-Shot Prompt \(Multi\-Turn Clinical Dialogue\)You are an expert medical doctor\. You will be given a conversation between a doctor and a patient\. The final doctor response is missing\. Based on the conversation, generate the doctor’s concluding response, typically a diagnosis, recommendation, or treatment plan\.Output format:ANSWER: ¡doctor’s response¿Rules:\- Output exactly ONE line starting with ”ANSWER: ”\.\- The response must be in Modern Standard Arabic\.\- Be concise and clinically appropriate \(1–3 sentences maximum\)\.\- Do NOT repeat or summarize the conversation\.\- Do NOT add extra commentary or explanation\.\- Do NOT translate anything\.Example:\{Sample Arabic Clinical Dialogue Between Doctor and Patient\}ANSWER:\\arabicfontfamilyهذه العوارض هي كوشينج سندروم وتعتبر مرض خطير إذا لم يتم علاجه أولا\. عليكي أن تراجعي طبيب مختص في الغدد في أسرع وقت، وعمل إجراءات فحوص يتضمن مقطعية وفحص دم\. بس لا تقلقي هذا المرض له علاج وممكن أن تشفي منه تماما\.Figure S4:One\-shot prompt for the Multi\-Turn Clinical Dialogue task\. The full dialogue is omitted for space and replaced with a placeholder\. We retain the example answer to illustrate the expected output format\. The complete dialogue sample is shown in Figure[S10](https://arxiv.org/html/2608.00207#A2.F10)\.LLM\-as\-a\-Judge Prompt \(Multi\-Turn Clinical Dialogue\)You are an expert medical evaluator\. You will be given a doctor\-patient dialogue \(in Arabic\), the Primary Reasoning Objective that the final doctor turn was supposed to clinically reach, and a generated final doctor turn\. Your task is to evaluate whether the generated turn correctly achieves the Primary Reasoning Objective, and respond only with the label in brackets \[\]\.Dialogue: \{dialogue\}Primary Reasoning Objective: \{primary\_reasoning\_objective\}Generated Answer: \{generated\_answer\}Evaluate the generated answer against the Primary Reasoning Objective using the following criterion:Correctness: Correct / Incorrect\- Correct: The generated answer reaches or is consistent with the diagnosis, differential, or management direction in the Primary Reasoning Objective\. This includes answers that are in the right clinical direction even if they do not fully articulate every detail\. Synonyms, equivalent medical terms, and slight phrasing differences still count as correct\. Do not penalize for code\-switching or English medical terminology if the clinical content is correct\.\- Incorrect: The generated answer fails to engage with the Primary Reasoning Objective, identifies a clearly different diagnosis, recommends a contradictory course of action, or is too vague to demonstrate any clinical reasoning\.Respond only with one of the following in brackets: \[Correct\] / \[Incorrect\]Figure S5:LLM\-as\-a\-judge prompt for the Multi\-Turn Clinical Dialogue task\.

## Appendix BData Collection and Pre\-processing

This appendix specifies the data pre\-processing pipeline for the Short Answer Generation task \(Section[B\.1](https://arxiv.org/html/2608.00207#A2.SS1)\) and the dataset construction procedure for the Multi\-Turn Clinical Dialogue task \(Section[B\.2](https://arxiv.org/html/2608.00207#A2.SS2)\)\.

### B\.1Adapting MCQA Data for Open\-Ended Answer Generation

To evaluate free\-form medical answer generation in Arabic, we needed a dataset that differs from MCQA while still testing medical knowledge\. Two natural candidates exist: AraMed\(Alasmariet al\.,[2024](https://arxiv.org/html/2608.00207#bib.bib12)\), which is the only native Arabic QA benchmark to the best of our knowledge, and English medical QA benchmarks translated into Arabic\. Both have limitations\. AraMed is sourced from public medical forum Altibbi, creating a high risk of contamination\. Translation, in turn, incurs well\-documented information loss, particularly for specialized medical terminology, and conflicts with our broader commitment to evaluating models exclusively on native Arabic benchmarks\.

We therefore repurpose MedAraBench, the largest MCQA dataset used in our MCQA task, into an open\-ended generation benchmark through a four\-stage pipeline: automated screening, manual verification, automated question reformulation, and a final manual quality pass\. The final repurposed benchmark is referred to as MedAraBench\-OpenEnded \(MedAraBench\-OE hereafter\)\.

#### B\.1\.1Automated Screening

The source dataset comprised 4,958 multiple\-choice questions\. To filter out items structurally incompatible with conversion to an open\-ended format, we prompted GPT‑5\.2 to label each sample asyes\(keep\),no\(drop\), ormaybe\(borderline\)\. The classification criteria were grounded in the structural requirements of open\-ended answer generation: a valid open\-ended question must elicit a direct, standalone response without relying on a predefined set of answer choices\.

Questions were flagged for removal if they fell into any of the following categories:

- •Exclusion\-logic questions using phrasing, which test the ability to identify a single false item among otherwise correct options
- •Questions with options that cross\-reference one another
- •Questions that ask the test\-taker to identify the incorrect or false statement rather than the correct one
- •Items that are not properly formed as questions, such as entries that merely label a body part or anatomical structure
- •Items in which the question stem or options are primarily in English

We prompt GPT\-5\.2 in a few\-shot setting using three examples drawn directly from the source dataset\. These examples were selected to represent the main decision categories used during screening\. The first example shows an acceptable numerical case, since questions containing numbers were often difficult to distinguish as acceptable or unacceptable\. The second example shows a question containing one of the exclusion cases discussed above and illustrates a case that should be dropped\. The third example shows a clear and concise question that can be kept\.

Each example includes the question ID, question stem, answer options, correct answer, screening decision \(yes, no, or maybe\), and a short justification for the decision\. This design exposes the model to the range of structural patterns it may encounter before classifying the full dataset\. The full prompt is shown in Figure[S6](https://arxiv.org/html/2608.00207#A2.F6)\.

Screening PromptYou are a medical Arabic MCQA quality reviewer\. Your task is to evaluate each question from the MedArabBench dataset and decide whether it should be KEPT or DROPPED for a medical answer generation benchmark\.Your Job
For each question, output one of three decisions:yes→ Keep it\. Clear, well\-formed question with a single unambiguous answer\.no→ Drop it\. Falls into one of the problematic categories below\.maybe→ Borderline\. Has a minor issue but could still be usable\.Rules: When to DROP
Drop a question if it matches ANY of the following:1\. Exclusion\-logic phrasing \(\\arabicfontfamilyعدا/\\arabicfontfamilyإلا/\\arabicfontfamilyما يلي/\\arabicfontfamilyعدا ما\)2\. Options referencing each other \(all of the above, A\+B,\\arabicfontfamilyكل ما سبق\)3\. Indirect/vague phrasing asking which statement is WRONG4\. Answer or options primarily in English5\. Question text primarily in English6\. Item merely states a label rather than posing a questionRules: When to FLAG as MAYBEFlag asmaybeONLY if:1\.Mixed language— question is Arabic but one or two options contain English terms mixed in, not fully English\.2\.Answer contains numbers with units in English\.Rules: When to KEEP \(yes\)Keep if:•Question is in clear Arabic\.•All options are in Arabic, or are numbers, including integers, decimals, ratios, measurements, or medical abbreviations\.•There is exactly one correct answer\.•The question is direct, not asking “which is WRONG” or “except”\.•No cross\-referencing between options\.Output FormatReturn ONLY a JSON array\. Each element must have exactly two keys:•id— integer, from input\.•keep—yes,no, ormaybe\.Do NOT include any explanation, preamble, or markdown\. Output raw JSON only\.Output Format
Return ONLY a JSON array:\[\{"id": <int\>, "keep": "yes"\|"no"\|"maybe"\}, …\]Figure S6:Screening prompt used to classify MCQA items for conversion to open\-ended questions\.
#### B\.1\.2Manual Verification

We manually reviewed every label from the automated screening pass to verify its correctness and to adjudicate items flagged as borderline\. Of the 4,958 source items, 4,014 were dropped as structurally incompatible with open\-ended generation, leaving a retained set of 944 questions\. This sharp reduction is deliberate: a large fraction of the source corpus consists of items whose structure is designed specifically for multiple\-choice testing, and reformulating them as free\-form questions would distort what the question is intended to test\. Filtering aggressively at this stage allows the resulting benchmark to prioritize question quality over scale\. Representative examples of dropped items, together with the justification for each exclusion, are shown in Table[S2](https://arxiv.org/html/2608.00207#A2.T2)\.

QuestionAnswerJustification for Exclusion\\arabicfontfamilyالعضلة الحرقفية :The iliopsoas muscle:\\arabicfontfamilyتعمل على بسط الفخذ على البطن\.Acts to extend the thigh onto the abdomen\.Not a well\-formed question, merely states a body part without posing a query\.\\arabicfontfamilyالطبقة الثالثة من عضلات أخمص القدم تشمل عدا:The third layer of plantar foot muscles includes, except:\\arabicfontfamilyباسطة الإصبعExtensor digitorum brevisExclusion\-type question using\\arabicfontfamilyعدا \(“except”\) phrasing, tests identification of the one item that does not belong, a structure incompatible with open\-ended answer generation\.\\arabicfontfamilyالعبارة الخاطئة عن الشرايين المرنة هي :The incorrect statement about elastic arteries is:\\arabicfontfamilyلونها مائل للأصفر بسبب تراكم الكولسترول\.Its colour is yellowish due to cholesterol accumulation\.Asks the test\-taker to identify the incorrect statement, inherently requires a closed option set and cannot be recast as a direct knowledge question\.\\arabicfontfamilyمن وظائف الهرمونات التي يفرزها المبيض :Among the functions of hormones secreted by the ovary:\\arabicfontfamilyكل ما سبق صحيحAll of the above is correct\.Correct answer is “all of the above”, a response meaningful only within a multiple\-choice context and cannot serve as a standalone open\-ended answer\.Table S2:Representative MCQA items dropped during preprocessing, with the rationale for each exclusion\. Each row illustrates a distinct failure mode that makes the item unsuitable for open\-ended answer generation\.
#### B\.1\.3Automated Question Rewriting

We rewrote the 944 retained questions into clean, standalone open\-ended form usingclaude\-opus\-4\-5\-20250901with one\-shot prompting\. Stems already phrased as direct interrogatives received only minor surface\-level cleanup\. Incomplete sentences, fill\-in\-the\-blank prompts, and bare labels were recast as full natural Arabic questions\. The reformulation prompt is shown in Figure[S7](https://arxiv.org/html/2608.00207#A2.F7)\.

Rewriting PromptYou are helping prepare an Arabic medical question answering benchmark\. You will be given an Arabic medical question stem originally designed as a multiple choice question \(MCQ\)\. Your task is to rewrite it into a clean, standalone open\-ended question suitable for short answer generation\.Rules:1\. Do NOT change the medical meaning, topic, or correct answer in any way\.2\. Do NOT add information that is not in the original stem\.3\. If the stem already starts with an Arabic interrogative \(\\arabicfontfamilyما,\\arabicfontfamilyماذا,\\arabicfontfamilyكيف,\\arabicfontfamilyمن,\\arabicfontfamilyأين,\\arabicfontfamilyمتى,\\arabicfontfamilyكم,\\arabicfontfamilyأي\), keep it as\-is with only minimal cleanup\.4\. If the stem is an incomplete sentence, label, or fill\-in\-the\-blank, rewrite it as a full natural Arabic question\.5\. Remove any leading punctuation artifacts such as a lone period or colon\.6\. Output ONLY the rewritten Arabic question — no explanation, no preamble\.Example:Input:\\arabicfontfamilyالعصب الذي ينقل إحساس الذوق من ثلث اللسان الخلفي هوOutput:\\arabicfontfamilyما هو العصب الذي ينقل إحساس الذوق من ثلث اللسان الخلفي؟Figure S7:One\-shot prompt used to reformulate retained MCQA stems as standalone open\-ended Arabic questions suitable for short\-answer generation\.
#### B\.1\.4Manual Question Review

We then reviewed all 944 reformulated questions manually to verify correct formatting, clarity of phrasing, and fidelity to the original item’s intent\. Items that were ambiguous, unnaturally phrased, or inconsistent with the expected correct answer were flagged for correction\. Of the 944 questions, 16 were revised and 4 were removed, yielding a final dataset of 940 open\-ended questions used for the Short Answer Generation task\.

A common pattern observed during this review was that the LLM\-generated reformulations, while grammatically well\-formed, were sometimes insufficiently specific\. In these cases, the model produced a broad open\-ended question that, although technically valid, would not constrain the respondent toward the particular aspect of knowledge being tested\. The annotated revision restored that specificity by targeting the precise dimension of the answer the original item was designed to elicit\. Table[S3](https://arxiv.org/html/2608.00207#A2.T3)illustrates this pattern with a representative example\.

Original MCQA StemLLM\-GeneratedRephrasingManually AnnotatedRephrasing\\arabicfontfamilyحديبات مونتغومري هي غددMontgomery tubercles are glands\\arabicfontfamilyما هي حديبات مونتغومري؟What are Montgomery tubercles?\\arabicfontfamilyإلى أي صنف من أصناف الغدد تنتمي حديبات مونتغومري؟To which type of glands do Montgomery tubercles belong?Table S3:Example of a question revised during the manual quality pass\. The LLM\-generated phrasing admits a broad range of responses, whereas the annotated version targets the specific medical fact tested by the original item \(correct answer:\\arabicfontfamilyدهنية, sebaceous\)\.

### B\.2Multi\-Turn Dialogue Data Collection

This section details the creation of the clinician\-authored evaluation dataset used for the Multi\-Turn Clinical Dialogue task\. The dataset comprises 100 dialogues, each built from a structured clinical case template, and was developed in four stages: case template authoring, LLM\-assisted dialogue generation, reference answer collection, and inter\-annotator agreement assessment\.

#### B\.2\.1Case Template Creation

We collaborated with three native Arabic\-speaking physicians from a multi\-specialty hopsital, each from a distinct specialty: critical care, pulmonary medicine, and general surgery\. Clinicians were given an annotation guide specifying the task: to author clinical case templates that would serve as the basis for generating multi\-turn patient–assistant dialogues in Modern Standard Arabic \(MSA\)\.

Case distribution was determined in consultation with the clinicians to reflect the range of scenarios commonly encountered in clinical settings\. Templates were required to span nine organ systems and to cover a range of scenario types, including diagnosis clarification, emergency referral, medication guidance, and preventive counseling\. Each template followed a fixed schema with six fields, as shown in the representative example in Table[S5](https://arxiv.org/html/2608.00207#A2.T5)\. The resulting distribution of cases across organ systems is summarised in Table[S4](https://arxiv.org/html/2608.00207#A2.T4)\.

Organ SystemNumber of CasesAbdomen15Cardiovascular15Endocrinology15Lymphatic System5Musculoskeletal15Nervous System10Respiratory15Urinary/Kidney5Urology5Total100Table S4:Distribution of the 100 clinician\-authored cases across nine organ systems\. Coverage was weighted toward systems most commonly encountered in clinical dialogue\.FieldContentOrgan SystemEndocrinologyProblem TypeClinic Non\-EmergencyPatient ProfileA 64\-year\-old postmenopausal female with a history of osteopenia\. She has no previous history of neck surgery or radiation exposure\.Symptoms & HistoryThe patient reports vague aching in her bones and recurrent bouts of constipation over several months\. She has experienced two episodes of painful kidney stones in the last three years\. Physical exam is unremarkable, but she mentions feeling increasingly foggy and fatigued\. Laboratory tests reveal a persistently elevated serum calcium of 10\.8 mg/dL\. Her vitamin D levels are within the normal range\.Red Flag SymptomsChronic hypercalcemia\.Primary Reasoning ObjectiveDiagnosis of primary hyperparathyroidism, elevated calcium with inappropriately normal or high PTH suggests a parathyroid adenoma\.Table S5:Representative example of a completed clinical case template \(Case 1, Endocrinology, Clinic Non\-Emergency\), authored by one of the three clinicians\.
#### B\.2\.2LLM\-Assisted Dialogue Generation

We prompted GPT\-5\.2 to generate a multi\-turn dialogue for each of the 100 clinician\-authored templates\. Each dialogue was a patient–assistant exchange in MSA, constrained to span three to eight turns and to conclude with a patient question\. This structure follows HealthBench\(Aroraet al\.,[2025](https://arxiv.org/html/2608.00207#bib.bib64)\): the dialogue terminates with a user turn so that, at evaluation time, the model under test must produce the final assistant response\.

The model was instructed to seek clarification naturally when needed, to avoid premature clinical escalation, and to refrain from introducing numerical thresholds or clinical data not grounded in the template\. The final assistant turn, the diagnostic response, was deliberately withheld at this stage and provided independently by the clinicians\. The dialogue generation prompt is shown in Figure[S9](https://arxiv.org/html/2608.00207#A2.F9), and Figure[S10](https://arxiv.org/html/2608.00207#A2.F10)shows a completed template with its generated dialogue\.

Dialect Translation PromptYou are a professional medical translator and a native speaker of each target Arabic dialect\. Translate the following Modern Standard Arabic \(MSA\) medical conversation into four spoken Arabic dialects\. The output must sound like how a real speaker of that dialect would actually talk to their doctor — not MSA with a few dialect particles sprinkled in\.Translate into:\(1\) Emirati \(2\) Jordanian \(3\) Moroccan Darija \(4\) EgyptianCore rules \(non\-negotiable\):
1\. Preserve all medical meaning exactly\. Do NOT add, remove, or simplify medical content\.2\. Keep the same speaker turns, urgency, and triage tone\.3\. Do not mix dialects within a single translation\.4\. Medical terminology: use English or French if more common in that dialect\.Authenticity rules:– Default to spoken register, not written register\.– Replace MSA connectors and fillers with dialectal equivalents\.– Replace abstract MSA phrasing with concrete dialectal phrasing\(e\.g\.,\\arabicfontfamilyضبابية في التفكير→\\rightarrowDarija:\\arabicfontfamilyما كنركزش مزيان; Egyptian:\\arabicfontfamilyمخي مش راكز;Jordanian:\\arabicfontfamilyبحس حالي مش مركز; Emirati:\\arabicfontfamilyما أقدر أركّز\)\.Self\-check before outputting:Re\-read your translation and ask: “Would a real patient in a clinic actually say this sentence to their doctor, or does it sound like an MSA sentence wearing a dialectal costume?” If the latter, rewrite it\.Output format:
Emirati Arabic: \[Translated conversation\]Jordanian Arabic: \[Translated conversation\]Moroccan Darija: \[Translated conversation\]Egyptian Arabic: \[Translated conversation\]Text to translate:\[INSERT MSA CONVERSATION HERE\]Figure S8:Prompt used to translate MSA dialogues into four regional Arabic dialects\.Dialogue Generation PromptYou are generating a realistic medical conversation in Arabic between a patient and a medical assistant chatbot\.Use ONLY the information provided in the case template below\. Do not introduce new medical history, vital signs, or test results unless they logically emerge through dialogue clarification\.Requirements:
1\. Generate a natural, realistic conversation in Modern Standard Arabic\.2\. The conversation must be between a patient and a medical assistant\.3\. Length must be between 3 and 8 total turns\.4\. The last turn MUST be a patient question\.5\. The assistant should appropriately seek clarification when necessary\.6\. Do not immediately escalate unless clearly justified by the evolving dialogue\.7\. Do not introduce numeric thresholds or new clinical data not grounded in the template\.8\. The conversation should feel authentic and medically plausible\.Output format:
Patient: …Assistant: …Patient: …… \(end with Patient message\)Case Template:\[INSERT TEMPLATE HERE\]Figure S9:Prompt used to generate multi\-turn patient–assistant dialogues from the clinician\-authored case templates\.![Refer to caption](https://arxiv.org/html/2608.00207v1/figures/dialogue_sample.png)

Figure S10:Example from our multi\-turn clinical dialogue dataset\. Each dialogue is grounded in a structured clinical scenario authored by physicians \(left\)\. Colored highlights mark spans wheresymptoms,red\-flag indicators, andreasoning objectivesfrom the scenario appear in the turns\. Given turnsT1,…,Tn−1T\_\{1\},\\ldots,T\_\{n\-1\}, the model must generateTnT\_\{n\}, evaluated in both Arabic and English\.Patient\-profile considerationsmay not surface lexically in the dialogue but must be reflected in the generated turn\.
#### B\.2\.3Reference Answer Collection

To obtain ground\-truth final\-turn responses for each dialogue, the three clinicians were asked to independently review each case and submit the best possible response\. To ensure that each case received two independent reference answers, the 100 dialogues were distributed such that each case was reviewed by exactly two different clinicians\. Responses were collected via a structured Google Form presenting each clinician with the case identifier, organ system, problem type, primary reasoning objective, and the full dialogue\.

The clinicians elected to record their responses as audio rather than typed text\. The resulting recordings were transcribed using Whisper\-Large\-V3, and each transcription was then verified by reading the text against its source recording, with corrections applied where necessary\. Table[S6](https://arxiv.org/html/2608.00207#A2.T6)shows a representative reference\-answer pair\.

CaseResponse AResponse B1\\arabicfontfamilyالغدة الجار دراقية صغيرة تقع خلف الغدة الدراقية\. الوظيفة الأساسية هي تحديد الكالسيوم في الدم\. ممكن أن تكون خطيرة إذا الكالسيوم ارتفع بشكل حاد\. الخبر الجيد أن هذا المرض يمكن أن يعالج عن طريق إقصاء هذه الغدد بطريقة جراحية\. من دون علاج ممكن أن تؤدي إلى أوجاع حصى كلسية وبطء في حركة الجهاز الهضمي\.The parathyroid gland is small and sits behind the thyroid\. Its main role is regulating blood calcium\. It can be dangerous if calcium rises sharply\. Treatment is typically surgical\. Untreated it causes kidney stones and slowed digestion\.\\arabicfontfamilyفرط نشاط جارات الدرقية ليس مرضاً مميتاً بل يمكن علاجه\. أولاً يجب تحديد نوع فرط نشاط جارات الدرقية وسبب المرض\. عادةً ما يكون العلاج جراحياً\.Hyperparathyroidism is not fatal and is treatable\. The type and underlying cause must first be identified\. Treatment is typically surgical\.Table S6:Representative example of two independently collected reference answers for Case 1 \(Endocrinology, Clinic Non\-Emergency: hyperparathyroidism\)\. English translations are shown in italics beneath each Arabic response\.
#### B\.2\.4Inter\-Annotator Agreement

Before writing each response, clinicians were provided with the case template and dialogue, and instructed to produce the final assistant turn according to the following rubric:

- •Be clinically accurate\.
- •Be written in Modern Standard Arabic \(MSA\)\.
- •Be 4–6 sentences in length\.
- •Focus on the primary clinical reasoning objective of the case\.
- •Include both a diagnosis, or explanation when appropriate, and an appropriate triage recommendation\.

To assess agreement between the two reference answers collected per case, a fourth Arabic\-speaking clinician from a general hospital, reviewed all 100 dialogue–response pairs, labeling each as Agreement when both responses reached the same diagnosis, and Disagreement otherwise\. For disagreements, the reviewer also indicated the preferred response\. Table[S7](https://arxiv.org/html/2608.00207#A2.T7)shows representative examples\. Of the 100 reviewed cases, 87 were labeled as agreement; for the remaining 13 disagreement cases, the reviewer indicated a preferred response or authored a revised gold\-standard answer when neither response was satisfactory\.

To assess agreement between the two reference answers collected per case, a fourth Arabic\-speaking clinician from a general hospital reviewed all 100 dialogue–response pairs, labeling each pair as Agreement when both responses reached the same diagnosis, and Disagreement otherwise\. For disagreements, the reviewer also indicated the preferred response\. Table[S7](https://arxiv.org/html/2608.00207#A2.T7)shows representative examples\. Of the 100 reviewed cases, 87 were labeled as agreement; for the remaining 13 disagreement cases, the reviewer indicated a preferred response or authored a revised gold\-standard answer when neither response was satisfactory\.

CaseResponse AResponse BAgreement1\\arabicfontfamilyالغدة الجار دراقية صغيرة تقع خلف الغدة الدراقية\. الوظيفة الأساسية هي تحديد الكالسيوم في الدم، ممكن أن تكون خطيرة إذا الكالسيوم ارتفع بشكل حاد\. الخبر الجيد أن هذا المرض يمكن أن يعالج عن طريق إقصاء هذه الغدد بطريقة جراحية، المهم أن لا تنتظر وعليك أن تتابع مع متخصص للغدد الجار دراقية\. من دون علاج ممكن أن تؤدي إلى أوجاع حصى كلسية وبطء في حركة الجهاز الهضمي\.The parathyroid gland is small and sits behind the thyroid\. Its main function is regulating blood calcium, which can be dangerous if it rises sharply\. Treatment is surgical; without treatment it causes kidney stones and slowed digestion\.\\arabicfontfamilyفرط نشاط جارات الدرقية ليس مرضاً مميتاً، بل يمكن علاجه\. أولاً، يجب تحديد نوع فرط نشاط جارات الدرقية وسبب المرض\. عادةً ما يكون العلاج جراحياً\.Hyperparathyroidism is not fatal and is treatable\. The type and underlying cause must first be identified\. Treatment is typically surgical\.Agreement8\\arabicfontfamilyأنصحك بأن تتابع مع دكتور متخصص في الغدد لتأكيد المرض\. ممكن تحتاج هورمون النمو وتأخذ أسابيع لعدة أشهر لتعود مثل ما كنت من قبل، في الشعور التحسن في العضلات وانخفاض في الدهون وممارسة الحياة الاجتماعية العادية، ولكنه مهم أن تتابع مع دكتور متخصص بهذه الأمراض\.I recommend following up with an endocrinologist to confirm the diagnosis\. You may need growth hormone, and it can take weeks to months before improvements in muscle mass, reduced fat, and normal social functioning are felt\.\\arabicfontfamilyبعد تأكيد التشخيص من خلال فحص الدم، الذي يستغرق عادةً من ساعتين إلى ثلاث ساعات، حيث نراقب استجابة جسمك للأنسولين وانخفاض سكر الدم، وما إذا كان يستجيب بإفراز الكورتيزول\. قد تشعر بأعراض نقص سكر الدم، كالدوار والتعرق والإرهاق\. إذا تأكد التشخيص، فستحتاج إلى حقن هرمون النمو لمدة تتراوح بين ستة وتسعة أشهر على الأقلAfter confirming via blood test \(2–3 hours\), monitoring insulin response and cortisol\. Side effects include dizziness, sweating, and fatigue\. Growth hormone injections required for at least 6–9 months\.Disagreement, Response B preferredTable S7:Examples of inter\-annotator agreement annotation\. Case 1 \(hyperparathyroidism\) illustrates an agreement outcome, while case 8 illustrates a disagreement outcome\. English translations are shown in italics beneath each Arabic response\.
#### B\.2\.5Dialect Translation

To extend the utility of the dataset beyond Modern Standard Arabic and support evaluation across dialectally diverse user populations, the 100 MSA dialogues were translated into four regionally representative Arabic dialects: Emirati \(Gulf\), Jordanian \(Levant\), Moroccan Darija \(Maghreb\), and Egyptian\.

Translations were generated using GPT\-5\.2, prompted to produce authentic dialectal speech that would be natural in a real clinical interaction\. The prompt encouraged the use of authentic dialectal vocabulary and phrasing, including code\-switching to English or French medical terminology where appropriate for a given dialect\. The translation prompt is reproduced in Figure[S8](https://arxiv.org/html/2608.00207#A2.F8), and Figure[S11](https://arxiv.org/html/2608.00207#A2.F11)shows a sample dialogue in MSA and its four dialect translations\.

![Refer to caption](https://arxiv.org/html/2608.00207v1/figures/translation.png)Figure S11:Sample dialogue in MSA \(center\) and its four dialect translations obtained with GPT\-5\.2: Jordanian and Egyptian \(top\), Emirati and Moroccan Darija \(bottom\)\.For each dialect, the 100 samples were divided evenly between two native speakers who reviewed the generated translations and corrected expressions that were unnatural or clinically inaccurate\. The extent of editing varied substantially across dialects: Moroccan Darija required the most extensive corrections, as the model often produced MSA\-dominant text with only sporadic dialect\-specific lexical insertions rather than authentic Darija\. Table[S8](https://arxiv.org/html/2608.00207#A2.T8)shows representative LLM\-generated translations and their human\-corrected versions\.

DialectLLM\-Generated TranslationHuman\-Corrected TranslationEmirati\\arabicfontfamilyمريض: السلام عليكم دكتور، من كم شهر وأنا أحس بآلام في عظامي، وعندي إمساك بشكل متكرر\. بعد أحس بتعب وما أقدر أركّز، تفكيري مو صافي\.Patient: I’ve had pain in my bones for months and keep getting constipated\. I also feel tired and can’t focus, my thinking isn’t clear\.\\arabicfontfamilyمريض: السلام عليكم دكتور، من كم شهر وأنا أحس بالعوار في عظامي، وعندي إمساك بشكل متكرر\. بعد أحس بتعب وما أقدر أركّز، تفكيري مب صافي\.\\arabicfontfamilyالعوار replaces\\arabicfontfamilyآلام for pain;\\arabicfontfamilyمب replaces\\arabicfontfamilyمو as negation\.Egyptian\\arabicfontfamilyمريض: حاسة بألم شديد في بطني وغثيان مستمر، ومش قادرة أوقف ترجيع من بدري قوي من الصبح\.Patient: I have severe stomach pain and constant nausea, and I haven’t been able to stop vomiting since early this morning\.\\arabicfontfamilyمريض: حاسة بألم شديد في بطني وعلى طول حاسة إني عايزة أَرَجّع، ومش عارفة أبطّل ترجيع من الصبح\.\\arabicfontfamilyعلى طول حاسة إني عايزة أَرَجّع replaces\\arabicfontfamilyغثيان مستمر for nausea;\\arabicfontfamilyمش عارفة أبطّل replaces\\arabicfontfamilyمش قادرة أوقف for inability to stop\.Jordanian\\arabicfontfamilyمريض: يعني الوضع خطير كتير؟ كمان حاسة بدوخة، كأني ممكن أغيب عن الوعي\.Patient: So is the situation very serious? I also feel dizzy, like I might lose consciousness\.\\arabicfontfamilyمريض: يعني الوضع خطير كتير؟ كمان حاسة بدوخة، كأني ممكن يغمى علي\.\\arabicfontfamilyيغمى علي replaces\\arabicfontfamilyأغيب عن الوعي as the Jordanian\-idiomatic expression for losing consciousness\.Moroccan\\arabicfontfamilyمريضة: إييه، كنحس بالبرد ديما حتى إلا كان الجو سخون، وبشرتي ولات ناشفة بزاف\. وزدت لاحظت الشعرة بدات كتطيح ليا من الجوانب\.Patient: Yes, I always feel cold even when the weather is warm, and my skin has become very dry\. I also noticed my hair started falling out on the sides\.\\arabicfontfamilyمريضة: إييه، كنحس بالبرد ديما حتى إلا كان الجو سخون، ولبشرة ديالي ولات ناشفة بزاف\. وزدت لاحظت شعري بدا كيطيح ليا من الجنب\.\\arabicfontfamilyلبشرة ديالي replaces\\arabicfontfamilyبشرتي using Darija possessive construction;\\arabicfontfamilyشعري replaces\\arabicfontfamilyالشعرة;\\arabicfontfamilyمن الجنب replaces\\arabicfontfamilyمن الجوانب\.Table S8:Representative examples of LLM\-generated versus human\-corrected dialect translations across all four target dialects\. Each row illustrates a distinct dialectal correction pattern\. English translations are shown in italics beneath each Arabic example\.

## Appendix CLLM\-as\-a\-Judge Reliability

BERTScore\-F1 is an imperfect proxy for clinical quality in open\-ended generation: prior work has shown it may not reliably distinguish model quality in medical tasks\(Bediet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib61)\), and does not directly evaluate clinical correctness\. We therefore adopt an LLM\-as\-a\-judge approach as an additional metric for the Short Answer Generation task\. To validate this choice, we collect human labels on a subset of Short Answer Generation outputs and measure how well both BERTScore\-F1 and the LLM judge correlate with human assessment\.

##### Validation Subset\.

To balance annotation costs with statistical coverage, we construct the validation subset by randomly sampling 100 examples from the 940\-example dataset with a fixed seed \(42\) to ensure reproducibility\. For each selected example, we collect the predictions from 18 of the 23 evaluated models, excluding the five adaptation methods, as the goal of this validation is to assess agreement patterns across models rather than benchmark every evaluated model\. This yields 1,800 model outputs for human validation\.

##### Annotation Protocol\.

To reduce annotator workload, we apply an exact\-match pre\-filter: model outputs that exactly match the reference answer are automatically labeled as correct, as their correctness is unambiguous\. This accounts for 54 out of the 1800 total predictions in the validation subset\. Each remaining \(question, reference answer, model output\) triple is independently evaluated by two medical students in their last year \(6th year\), who assign a binary label \(correct or incorrect\) following the rubric in Table[S9](https://arxiv.org/html/2608.00207#A3.T9)\. The LLM\-judge labels are not shared with the annotators\.

Model\-level human\-judge accuracy \(percentage correct\) is computed as the percentage of predictions labeled correct, calculated separately for each annotator and then averaged\. We note that approximately 6% of the questions in the subset were flagged by both annotators as having incorrect ground truth, highlighting a limitation of the dataset\.

LabelCriteriaCorrectMedically accurate, answers the question, and conveys the same clinical meaning as the reference answer, regardless of wording\. Additional information that does not contradict the reference is permitted\.IncorrectContains factual medical errors, omits the essential answer, contradicts the reference, or fails to answer the question\.Table S9:Annotation rubric used by human annotators to label model outputs as correct or incorrect\.CategoryModelBERTScore\-F1LLM\-Judge AccuracyHuman\-Judge AccuracyClosed\-SourceGeneral\-Purpose LLMsGPT\-5\.263\.062\.084\.0Gemini\-2\.5\-Flash64\.562\.086\.0Claude\-Opus\-4\.660\.164\.087\.5Open\-SourceGeneral\-PurposeLLMsMistral\-7B\-Instruct\-v0\.352\.22\.06\.5Llama\-3\.1\-8B\-Instruct55\.87\.014\.5Mistral\-Small\-3\.2\-24B61\.226\.050\.0Gemma\-3\-27B\-it58\.931\.052\.5Llama\-3\.3\-70B\-Instruct57\.716\.028\.0DeepSeek\-V3\.261\.952\.073\.0Arabic/MultilingualGeneral\-PurposeLLMsJais\-2\-8B\-Chat54\.223\.045\.5ALLaM\-7B\-Instruct\-preview56\.423\.039\.0Aya\-Expanse\-8B54\.518\.039\.0SILMA\-9B\-Instruct\-v1\.057\.222\.032\.0Falcon\-H1\-7B\-Instruct56\.619\.033\.0Fanar\-1\-9B\-Instruct51\.69\.012\.0Medical LLMsMedGemma\-27B\-Text\-It60\.135\.058\.5Meditron\-3\-70B53\.327\.051\.0Llama3\-Med42\-70B53\.625\.040\.0

Table S10:Model\-level evaluation scores on the 100\-example validation subset\.
##### Results\.

Table[S10](https://arxiv.org/html/2608.00207#A3.T10)summarizes the BERTScore\-F1, LLM\-judge accuracy \(% of correct predictions\), and human\-judge accuracy \(% of correct predictions\) on the validation subset of the MedArabench\-OE dataset\. Each of the three metrics produces an independent model ranking based on these aggregated scores, as shown in Figure[S12](https://arxiv.org/html/2608.00207#A3.F12)\. To assess agreement of LLM as a judge to human judgement, we compute both Pearson correlation and Spearman’s rank correlation across models\. For comparison, we also compute the same correlations between BERTScore\-F1 and human evaluation\. We adopt a threshold ofρ≥0\.8\\rho\\geq 0\.8as a criterion for validating the judge\. As shown in Table[S11](https://arxiv.org/html/2608.00207#A3.T11), the LLM judge achieves substantially higher agreement with human judgments \(Pearson: 0\.978; Spearman: 0\.982\) than BERTScore\-F1 \(Pearson: 0\.802; Spearman: 0\.715\), satisfying our criterion ofρ≥0\.8\\rho\\geq 0\.8\.

MetricPearsonSpearman \(ρ\\rho\)BERTScore\-F10\.8020\.715LLM\-Judge Accuracy0\.9780\.982

Table S11:Correlation with Human\-Judge Accuracy on the 100\-example validation subset\.
##### Bias Check: Self\-Preference of the Judge\.

The validation study further supports our inclusion of ChatGPT 5\.2 as a generator model for the Short Answer Generation task\. Its rank is similar under human\-judge accuracy and LLM\-judge accuracy \(Figure[S12](https://arxiv.org/html/2608.00207#A3.F12)\), indicating no meaningful self\-preference in the LLM\-as\-a\-Judge evaluation\.

![Refer to caption](https://arxiv.org/html/2608.00207v1/figures/validation-rank.png)Figure S12:Comparison of model rankings across BERTScore\-F1, LLM\-judge accuracy, and human\-judge accuracy\.

## Appendix DTraining Methodology

##### Fixed hyperparameters\.

Table[S12](https://arxiv.org/html/2608.00207#A4.T12)lists all hyperparameters held constant across every LoRA experiment reported in this paper\. The LoRA configuration \(r=16r\{=\}16,αLoRA=32\\alpha\_\{\\mathrm\{LoRA\}\}\{=\}32, dropout0\.050\.05\) follows standard practice for instruction\-tuned 24 B\-parameter models\(Huet al\.,[2022](https://arxiv.org/html/2608.00207#bib.bib4)\)\.

HyperparameterValueLoRA rank \(rr\)16LoRAαLoRA\\alpha\_\{\\mathrm\{LoRA\}\}32LoRA dropout0\.05Max sequence length1,024 tokensTraining epochs10Early\-stopping patience1 checkpointWarmup ratio0\.05LR schedulerCosineOptimizerAdamWPrecisionbfloat16GPUs2×\\timesA100 80 GBTable S12:Hyperparameters held constant across all LoRA experiments\.
##### Convergence analysis\.

Before committing to a full hyperparameter search we trained a single targeted\-LoRA model for 10 epochs at a nominal learning rate of2×10−42\\times 10^\{\-4\}to characterise the shape of the learning curve\. Figure[S13](https://arxiv.org/html/2608.00207#A4.F13)shows that training exhibits a slow initial phase through epoch 4, during which validation loss decreases gradually from0\.610\.61to0\.420\.42\. A steep learning phase begins at epoch 4 and persists through epoch 9, with validation loss falling from0\.420\.42to approximately0\.010\.01and per\-checkpoint improvements of1414–41%41\\%\. From epoch 9 onward improvement drops below10%10\\%, indicating plateau\. This confirmed that a 10\-epoch budget is sufficient and motivated the use of early stopping in the subsequent search\.

![Refer to caption](https://arxiv.org/html/2608.00207v1/figures/convergence_targeted_10ep.png)Figure S13:Convergence analysis\. Validation loss over 10 epochs atη=2×10−4\\eta\{=\}2\\times 10^\{\-4\}\(top\) and per\-checkpoint relative improvement % \(bottom\)\.
##### Per\-window learning\-rate search\.

Because the optimal learning rate depends on the number of trainable parameters, and different LoRA windows expose different parameter counts, we conducted an independent log\-uniform random search for each window rather than re\-using a single shared rate\. Five learning rates were sampled fromη∼LogUniform​\[10−5,4×10−4\]\\eta\\sim\\mathrm\{LogUniform\}\[10^\{\-5\},\\,4\\times 10^\{\-4\}\]with a fixed seed \(42\) for reproducibility\. All other settings matched Table[S12](https://arxiv.org/html/2608.00207#A4.T12)\. The best\-validation\-loss trial was selected for each window before examining any test performance\.

##### Training data\.

All experiments use the MedArabBench MCQA train split\(Daoudet al\.,[2026](https://arxiv.org/html/2608.00207#bib.bib8)\)as the sole training source, comprising 17,860 training and 1,987 validation examples after a stratified 90/10 split \(seed 42\)\.

CategoryModelMedAraBench\-OEBERTScore\-F1Correct %Incorrect %Closed\-SourceGeneral\-PurposeGPT\-5\.263\.962\.637\.3Gemini\-2\.5\-Flash65\.160\.539\.5Claude\-Opus\-4\.661\.067\.632\.3Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.351\.52\.397\.6Llama\-3\.1\-8B\-Instruct54\.36\.693\.4Mistral\-Small\-3\.2\-24B60\.030\.169\.8Gemma\-3\-27B\-IT59\.332\.667\.3Llama\-3\.3\-70B57\.819\.580\.4DeepSeek\-V3\.262\.049\.050\.9Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat54\.217\.782\.3ALLaM\-7B\-Instruct\-Preview56\.919\.680\.3Aya\-Expanse\-8B53\.915\.784\.2SILMA\-9B\-Instruct\-v1\.057\.415\.085\.0Falcon\-H1\-7B\-Instruct57\.615\.584\.5Fanar\-1\-9B\-Instruct48\.47\.492\.5Medical\-DomainMedGemma\-27B\-Text\-IT59\.832\.967\.0Meditron\-3\-70B54\.730\.769\.3Llama\-3\-Med42\-70B53\.424\.675\.4AdaptationMethodsMistral \+ Few\-Shot \(k=5\)60\.228\.571\.4Mistral \+ AUTOCAP59\.827\.572\.5BiMedix47\.618\.781\.3Mistral \+ Full LoRA53\.825\.575\.4Mistral \+ TLoRA \(Ours\)57\.929\.770\.3Mistral \+ TLoRA \(Optimal\)52\.638\.161\.9Table S13:Short Answer Generation results on MedAraBench\-OE\. Models are evaluated using BERTScore\-F1 and LLM\-as\-a\-judge correctness labels, reported as the percentage of correct \(Correct %\) and incorrect \(Incorrect %\) responses\. Models are grouped by family\. Bold indicates the best result within each group per column, while underlining points to second\-best result\. Mistral\-Small\-3\.2\-24B\-Instruct\-2506 is the Mistral backbone used in the adaptation methods\.

## Appendix EComplete Results

This appendix presents the complete per\-model results for the Short Answer Generation and Multi\-Turn Clinical Dialogue tasks\. All results are evaluated using BERTScore\-F1 with AraBERTv2 as the underlying contextual encoder, and an LLM\-as\-a\-judge setup reporting the percentage of correct \(Correct %\) and incorrect \(Incorrect %\) responses\.

### E\.1Short Answer Generation Complete Results

Complete results for the Short Answer Generation task are presented in Table[S13](https://arxiv.org/html/2608.00207#A4.T13), covering models from all five categories evaluated on MedAraBench\-OE with gold answers provided as references to the judge\.

### E\.2Multi\-Turn Clinical Dialogue Complete Results

Complete results for the Multi\-Turn Clinical Dialogue task are presented in Tables[S14](https://arxiv.org/html/2608.00207#A5.T14)–[S18](https://arxiv.org/html/2608.00207#A5.T18), with one table per AraClinicDialog variant: MSA, Emirati, Moroccan Darija, Jordanian, and Egyptian\.

CategoryModelAraClinicDialog \(MSA\)BERTScore\-F1Correct %Incorrect %Closed\-SourceGeneral\-PurposeGPT\-5\.256\.385\.015\.0Gemini\-2\.5\-Flash57\.950\.050\.0Claude\-Opus\-4\.657\.974\.026\.0Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.356\.935\.065\.0Llama\-3\.1\-8B\-Instruct57\.720\.080\.0Mistral\-Small\-3\.2\-24B58\.444\.056\.0Gemma\-3\-27B\-it57\.649\.051\.0Llama\-3\.3\-70B57\.834\.066\.0DeepSeek\-V3\.258\.071\.029\.0Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat52\.326\.074\.0ALLaM\-7B\-Instruct\-preview57\.743\.057\.0Aya\-Expanse\-8B57\.840\.060\.0SILMA\-9B\-Instruct\-v1\.053\.718\.082\.0Falcon\-H1\-7B\-Instruct56\.639\.061\.0Fanar\-1\-9B\-Instruct43\.811\.089\.0Medical\-DomainMedGemma\-27B\-Text\-it58\.064\.036\.0Meditron\-3\-70B57\.328\.072\.0Llama\-3\-Med42\-70B57\.555\.045\.0AdaptationMethodsMistral \+ Few\-Shot \(k=5\)58\.255\.045\.0Mistral \+ AUTOCAP56\.929\.071\.0BiMedix56\.732\.068\.0Mistral \+ Full LoRA56\.068\.032\.0Mistral \+ TLoRA \(Ours\)52\.565\.035\.0Table S14:Multi\-Turn Clinical Dialogue results on AraClinicDialog \(MSA\)\. Models are evaluated using BERTScore\-F1 and LLM\-as\-a\-judge scores for correct \(%\) and incorrect \(%\) outputs\. Models are grouped by family\. Bold indicates the best result within each group per column, while underline points to second\-best result\. Mistral\-Small\-3\.2\-24B\-Instruct\-2506 is the Mistral backbone used in the adaptation methods\.CategoryModelAraClinicDialog \(Emirati\)BERTScore\-F1Correct %Incorrect %Closed\-SourceGeneral\-PurposeGPT\-5\.256\.584\.016\.0Gemini\-2\.5\-Flash58\.445\.055\.0Claude\-Opus\-4\.657\.877\.023\.0Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.355\.233\.067\.0Llama\-3\.1\-8B\-Instruct55\.414\.086\.0Mistral\-Small\-3\.2\-24B57\.540\.060\.0Gemma\-3\-27B\-it56\.939\.061\.0Llama\-3\.3\-70B56\.723\.077\.0DeepSeek\-V3\.257\.864\.036\.0Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat53\.227\.073\.0ALLaM\-7B\-Instruct\-preview53\.928\.072\.0Aya\-Expanse\-8B56\.232\.068\.0SILMA\-9B\-Instruct\-v1\.050\.615\.085\.0Falcon\-H1\-7B\-Instruct54\.937\.063\.0Fanar\-1\-9B\-Instruct45\.69\.091\.0Medical\-DomainMedGemma\-27B\-Text\-it58\.343\.057\.0Meditron\-3\-70B56\.538\.062\.0Llama\-3\-Med42\-70B60\.858\.042\.0AdaptationMethodsMistral \+ Few\-Shot \(k=5\)58\.347\.053\.0Mistral \+ AUTOCAP54\.925\.075\.0BiMedix55\.129\.071\.0Mistral \+ Full LoRA54\.449\.051\.0Mistral \+ TLoRA \(Ours\)49\.442\.058\.0Table S15:Multi\-Turn Clinical Dialogue results on AraClinicDialog \(Emirati\)\. Models are evaluated using BERTScore\-F1 and LLM\-as\-a\-judge scores for correct \(%\) and incorrect \(%\) outputs\. Models are grouped by family\. Bold indicates the best result within each group per column, while underline indicates second\-best result\. Mistral\-Small\-3\.2\-24B\-Instruct\-2506 is the Mistral backbone used in the adaptation methods\.CategoryModelAraClinicDialog \(Moroccan Darija\)BERTScore\-F1Correct %Incorrect %Closed\-SourceGeneral\-PurposeGPT\-5\.254\.676\.024\.0Gemini\-2\.5\-Flash56\.252\.048\.0Claude\-Opus\-4\.656\.570\.030\.0Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.352\.922\.078\.0Llama\-3\.1\-8B\-Instruct53\.814\.086\.0Mistral\-Small\-3\.2\-24B55\.032\.068\.0Gemma\-3\-27B\-it55\.133\.067\.0Llama\-3\.3\-70B55\.418\.082\.0DeepSeek\-V3\.256\.164\.036\.0Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat50\.616\.084\.0ALLaM\-7B\-Instruct\-preview54\.836\.064\.0Aya\-Expanse\-8B53\.529\.071\.0SILMA\-9B\-Instruct\-v1\.051\.619\.081\.0Falcon\-H1\-7B\-Instruct53\.421\.079\.0Fanar\-1\-9B\-Instruct45\.211\.089\.0Medical\-DomainMedGemma\-27B\-Text\-it55\.540\.060\.0Meditron\-3\-70B53\.924\.076\.0Llama\-3\-Med42\-70B52\.944\.056\.0AdaptationMethodsMistral \+ Few\-Shot \(k=5\)55\.739\.061\.0Mistral \+ AUTOCAP54\.227\.073\.0BiMedix52\.021\.079\.0Mistral \+ Full LoRA47\.859\.041\.0Mistral \+ TLoRA \(Ours\)48\.470\.030\.0Table S16:Multi\-Turn Clinical Dialogue results on AraClinicDialog \(Moroccan Darija\)\. Models are evaluated using BERTScore\-F1 and LLM\-as\-a\-judge scores for correct \(%\) and incorrect \(%\) outputs\. Models are grouped by family\. Bold indicates the best result within each group per column, while underline points to the second\-best result\. Mistral\-Small\-3\.2\-24B\-Instruct\-2506 is the Mistral backbone used in the adaptation methods\.CategoryModelAraClinicDialog \(Jordanian\)BERTScore\-F1Correct %Incorrect %Closed\-SourceGeneral\-PurposeGPT\-5\.256\.178\.022\.0Gemini\-2\.5\-Flash57\.750\.050\.0Claude\-Opus\-4\.657\.677\.023\.0Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.354\.932\.068\.0Llama\-3\.1\-8B\-Instruct55\.620\.080\.0Mistral\-Small\-3\.2\-24B56\.030\.070\.0Gemma\-3\-27B\-it57\.144\.056\.0Llama\-3\.3\-70B56\.126\.074\.0DeepSeek\-V3\.257\.367\.033\.0Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat52\.630\.070\.0ALLaM\-7B\-Instruct\-preview55\.543\.057\.0Aya\-Expanse\-8B56\.234\.066\.0SILMA\-9B\-Instruct\-v1\.051\.112\.088\.0Falcon\-H1\-7B\-Instruct54\.828\.072\.0Fanar\-1\-9B\-Instruct46\.012\.088\.0Medical\-DomainMedGemma\-27B\-Text\-it57\.848\.052\.0Meditron\-3\-70B56\.133\.067\.0Llama\-3\-Med42\-70B55\.253\.047\.0AdaptationMethodsMistral \+ Few\-Shot \(k=5\)56\.940\.060\.0Mistral \+ AUTOCAP55\.420\.080\.0BiMedix54\.928\.072\.0Mistral \+ Full LoRA54\.254\.046\.0Mistral \+ TLoRA \(Ours\)54\.546\.054\.0Table S17:Multi\-Turn Clinical Dialogue results on AraClinicDialog \(Jordanian\)\. Models are evaluated using BERTScore\-F1 and LLM\-as\-a\-judge scores for correct \(%\) and incorrect \(%\) outputs\. Bold indicates the best result within each group per column, while underline represents the second\-best result\. Mistral\-Small\-3\.2\-24B\-Instruct\-2506 is the Mistral backbone used in the adaptation methods\.CategoryModelAraClinicDialog \(Egyptian\)BERTScore\-F1Correct %Incorrect %Closed\-SourceGeneral\-PurposeGPT\-5\.256\.780\.020\.0Gemini\-2\.5\-Flash58\.550\.050\.0Claude\-Opus\-4\.658\.474\.026\.0Open\-SourceGeneral\-PurposeMistral\-7B\-Instruct\-v0\.355\.328\.072\.0Llama\-3\.1\-8B\-Instruct56\.314\.086\.0Mistral\-Small\-3\.2\-24B56\.829\.071\.0Gemma\-3\-27B\-it58\.146\.054\.0Llama\-3\.3\-70B57\.624\.076\.0DeepSeek\-V3\.258\.665\.035\.0Arabic/MultilingualGeneral\-PurposeJais\-2\-8B\-Chat53\.725\.075\.0ALLaM\-7B\-Instruct\-preview56\.938\.062\.0Aya\-Expanse\-8B57\.332\.068\.0SILMA\-9B\-Instruct\-v1\.051\.714\.086\.0Falcon\-H1\-7B\-Instruct55\.833\.067\.0Fanar\-1\-9B\-Instruct47\.212\.088\.0Medical\-DomainMedGemma\-27B\-Text\-it58\.444\.056\.0Meditron\-3\-70B57\.628\.072\.0Llama\-3\-Med42\-70B55\.851\.049\.0AdaptationMethodsMistral \+ Few\-Shot \(k=5\)57\.538\.062\.0Mistral \+ AUTOCAP56\.221\.079\.0BiMedix55\.529\.071\.0Mistral \+ Full LoRA54\.939\.061\.0Mistral \+ TLoRA \(Ours\)50\.448\.052\.0Table S18:Multi\-Turn Clinical Dialogue results on AraClinicDialog \(Egyptian\)\. Models are evaluated using BERTScore\-F1 and LLM\-as\-a\-judge scores for correct \(%\) and incorrect \(%\)\. Bold indicates the best result within each group per column, while underline points to the second\-best result\. Mistral\-Small\-3\.2\-24B\-Instruct\-2506 is the Mistral backbone model used in the adaptation methods\.

## Appendix FAdditional Results

### F\.1KL Probe Layer Sensitivity

The KL alignment objective requires choosing a*probe layer*ℓp\\ell\_\{p\}at which the logit\-lens KL divergence between the Arabic and English representations is measured\. Following the zone analysis in §[3](https://arxiv.org/html/2608.00207#S3), we parameterize this choice viaτ=μ\+c​σ\\tau=\\mu\+c\\sigma\(withμ\\mu,σ\\sigmathe mean and standard deviation of per\-layer KL across the model\), yielding: c=0⇒ℓp=29c\{=\}0\\Rightarrow\\ell\_\{p\}\{=\}29\(ramp zone\), c=1⇒ℓp=34c\{=\}1\\Rightarrow\\ell\_\{p\}\{=\}34\(active\-zone onset, default\), c=2⇒ℓp=40c\{=\}2\\Rightarrow\\ell\_\{p\}\{=\}40\(deep active zone\)\.

Table[S19](https://arxiv.org/html/2608.00207#A6.T19)reports results for the winning window L1–34 at the tuned learning rateη∗=2\.76×10−5\\eta^\{\*\}\{=\}2\.76\\times 10^\{\-5\}, varying onlycc\. The defaultc=1c\{=\}1\(ℓp=34\\ell\_\{p\}\{=\}34\) achieves the best or joint\-best score on all MCQA benchmarks, validating the zone\-threshold heuristic for probe placement\.

### F\.2Effect of LoRA Layer vs\. KL Alignment

To disentangle whether gains over the full\-LoRA baseline arise from*where*adapters are placed or from the*KL alignment loss*, we compare CE\-only and CE\+KL variants within each window\.

Table[S20](https://arxiv.org/html/2608.00207#A6.T20)shows results for CE\-only training across all LoRA windows with per\-window tuned learning rates\. L1–34 leads on MCQA overall, with the closest competitor L1–24 winning only on AraSTEM \(64\.7 vs\. 63\.1\), likely reflecting its stronger coverage of lower layers where morphological features are processed\. Windows restricted to upper layers consistently underperform or degrade below zero\-shot, and full\-model LoRA collapses entirely at its tuned rate, producing near\-random MCQA scores alongside anomalously high generation BERTScores that LLM\-judge evaluation confirms as degenerate repetitive outputs\. These results establish that the performance advantage of L1–34 holds even without the KL alignment term, pointing to layer placement as the primary driver\.

### F\.3Fixed Learning Rate Comparison

A potential confound in comparing LoRA windows is that each uses a different tuned learning rate\. To isolate layer placement from optimisation, we retrain every window with CE\+KL loss at a single fixed rate \(η=1\.10×10−5\\eta\{=\}1\.10\\times 10^\{\-5\}, the tuned optimum for L1–34\)\. Table[S21](https://arxiv.org/html/2608.00207#A6.T21)shows that L1–34 retains its lead on MedAraBench \(62\.0\) and MedarabiQ \(60\.0\), confirming that the window\-level ordering is not an artefact of a more favourable learning rate\. Taken together with the CE\-only results, both loss function and learning rate can be ruled out as confounds, leaving layer selection as the explanation for the observed gains\.

ProbeccMCQAGenerationDialogue \(MSA\)MedAraBenchMedarabiQMMLU\-BioAraSTEMBERTScoreLLM judgeBERTScoreLLM judgeL29060\.15460\.560\.154\.321\.957\.069L34162\.16061\.665\.157\.929\.752\.564L40259\.95360\.862\.555\.431\.155\.369

Table S19:Probe layer sensitivity \(cc\-ablation\) for the winning window L1–34,η∗=2\.76×10−5\\eta^\{\*\}\{=\}2\.76\\times 10^\{\-5\}\.Bold= best per column\.ExperimentLayersTuned LRMCQAGenerationDialogue \(MSA\)MedAraBenchMedarabiQMMLU\-BioAraSTEMBERTScoreLLM judgeBERTScoreLLM judgeZero\-shot——52\.651\.555\.262\.660\.030\.158\.444\.0Full LoRAL1–401\.51×10−41\.51\\times 10^\{\-4\}42\.642\.045\.536\.390\.1†34\.131\.728\.0TargetedL1–242\.76×10−52\.76\\times 10^\{\-5\}61\.258\.060\.164\.746\.933\.057\.670\.0TargetedL1–342\.28×10−52\.28\\times 10^\{\-5\}62\.558\.061\.463\.145\.735\.358\.371\.0TargetedL24–402\.76×10−52\.76\\times 10^\{\-5\}52\.940\.055\.058\.569\.037\.646\.733\.0TargetedL34–401\.06×10−41\.06\\times 10^\{\-4\}49\.436\.051\.751\.087\.5†30\.722\.22\.0

Table S20:CE\-only training across LoRA windows with per\-window tuned learning rates\.Bold= best per column excluding zero\-shot\.†Anomalously high generation BERTScore for Full LoRA and L34–40 reflects degenerate repetitive outputs, confirmed by low LLM\-judge and dialogue scores\.ExperimentLayersLRMCQAGenerationDialogue \(MSA\)MedAraBenchMedarabiQMMLU\-BioAraSTEMBERTScoreLLM judgeBERTScoreLLM judgeZero\-shot——52\.651\.555\.262\.660\.030\.158\.444\.0Full LoRAL1–401\.10×10−51\.10\\times 10^\{\-5\}61\.955\.060\.360\.753\.825\.556\.068\.0TargetedL1–241\.10×10−51\.10\\times 10^\{\-5\}59\.049\.060\.961\.953\.829\.057\.069\.0TargetedL24–401\.10×10−51\.10\\times 10^\{\-5\}51\.434\.052\.556\.152\.240\.248\.481\.0TargetedL1–341\.10×10−51\.10\\times 10^\{\-5\}62\.060\.059\.661\.053\.724\.957\.070\.0TargetedL34–401\.10×10−51\.10\\times 10^\{\-5\}42\.537\.051\.353\.550\.129\.848\.273\.0

Table S21:Fixed learning\-rate comparison \(η=1\.10×10−5\\eta\{=\}1\.10\\times 10^\{\-5\}for all windows\), CE\+KL loss, probe L34\.Bold= best per column excluding zero\-shot\.

## Appendix GComputational Resources

All experiments were conducted on the NYUAD Jubail HPC cluster using an average of two NVIDIA A100 GPUs over approximately three months of extensive experimentation by two researchers\. This corresponds to a heuristic computational budget of approximately9090days×\\times2020hours/day×\\times22A100 GPUs=3,600=3\{,\}600A100 GPU\-hours\.

Table[S22](https://arxiv.org/html/2608.00207#A7.T22)breaks down the cost of the diagnostic pipeline itself, run once per base model\. The pipeline completes in under two hours of wall\-clock time \(3\.63 GPU\-hours total\), making the layer\-selection diagnosis practical to apply to a new base model without incurring substantial additional cost beyond the exploratory experimentation reported\.

StageWall\-clockHardwareGPU\-hoursBaseline EN \+ AR zero\-shot inference1h40m2×\\timesA100\-SXM4\-80GB3\.33Tuned lens \(train \+ eval\)10m50s1×\\timesA100\-SXM4\-80GB0\.18Causal activation patching2m10s1×\\timesA100\-SXM4\-80GB0\.04KL\-divergence profiling4m32s1×\\timesA100\-SXM4\-80GB0\.08Total1h57m32s—3\.63Table S22:Per\-stage wall\-clock time, hardware, and GPU\-hours for the diagnostic pipeline \(run once per base model\)\.
## Appendix HHuman Annotators

Ten student annotators contributed to inter\-annotator agreement evaluation and medical data review\. Annotators were compensated via honoraria and Amazon vouchers commensurate with the hours contributed\. We estimate a total annotation effort of approximately6060hours across all annotators\.

## Appendix IGeneralizability of TLoRA

To assess whether TLoRA generalizes beyond the Mistral backbone used in the main text, we apply the same diagnostic pipeline to a different model family and scale, Llama\-3\.1\-8B\-Instruct\.

Mechanistic analysis on this model identifiesLpatch=L\_\{\\text\{patch\}\}=L18 andLKL=L\_\{\\text\{KL\}\}=L28, which define the five candidate windows evaluated below\. All windows are trained with the same learning rate \(1\.06×10−41\.06\\times 10^\{\-4\}\), selected as the best value for full LoRA\.

Table[S23](https://arxiv.org/html/2608.00207#A9.T23)reports results for each window on Llama\-3\.1\-8B\-Instruct \(best result per column in bold, second\-best starred\)\. Window w1 \(layers L1–L18\) is selected as TLoRA based on its MCQA performance\.

WindowLoRA LayersLRMCQAGenerationDialogue \(MSA\)MedAraBenchMedarabiQMMLU\-BioAraSTEMBERTScoreLLM judgeBERTScoreLLM judgeZero\-shot——37\.732\.638\.637\.854\.36\.657\.720\.0Full LoRAL1–L321\.06×10−41\.06\\times 10^\{\-4\}49\.639\.044\.637\.451\.79\.647\.726\.0TargetedL1–L181\.06×10−41\.06\\times 10^\{\-4\}47\.339\.044\.739\.552\.013\.350\.025\.0TargetedL18–L321\.06×10−41\.06\\times 10^\{\-4\}41\.136\.039\.333\.146\.36\.144\.222\.0TargetedL1–L281\.06×10−41\.06\\times 10^\{\-4\}46\.642\.043\.738\.152\.013\.546\.913\.0TargetedL28–L321\.06×10−41\.06\\times 10^\{\-4\}34\.623\.037\.531\.646\.44\.721\.21\.0

Table S23:Generalizability of TLoRA to Llama\-3\.1\-8B\-Instruct, fixed learning rate \(η=1\.06×10−4\\eta\{=\}1\.06\\times 10^\{\-4\}for all windows\)\.Bold= best per column excluding zero\-shot\. w1 \(L1–L18\) is selected as TLoRA\.TLoRA improves over both the zero\-shot baseline and full LoRA on all four MCQA datasets, and achieves the best result on MMLU\-Bio \(44\.7\) and AraSTEM \(39\.5\); full LoRA is slightly ahead on MedAraBench \(49\.6 vs\. 47\.3\)\. On the out\-of\-domain generation and dialogue tasks, TLoRA is the second\-best configuration on generation BERTScore \(52\.0\) and LLM\-judge correctness \(13\.3\), and on dialogue LLM\-judge correctness \(25\.0\)\.

## Appendix JQualitative Case Studies

We conduct a paired qualitative comparison between TLoRA and base Mistral\-Small\-3\.2\-24B\-Instruct\-2506 model on the MCQA task\. We categorize predictions into four outcome types:successfully recovered\(incorrect under the base model, correct under TLoRA\),flipped to incorrect\(correct under the base model, incorrect under TLoRA\),stayed correct, andfailed to recover\(incorrect under both\)\. Cases are pooled across the four MCQA datasets, and we manually review 30 success cases and 30 failure cases\.

TLoRA is most beneficial on single\-fact retrieval questions, as illustrated by the recovered example in Table[S24](https://arxiv.org/html/2608.00207#A10.T24)\. It struggles, however, to distinguish between closely related answer options: in the failed\-to\-recover example, both models converge on the same plausible distractor rather than the correct answer\. More broadly, many failures involve replacing one incorrect answer with another, and negative or exclusion wording in the question does not appear to be a primary driver of failure cases\. An extended table covering all four outcome categories across the MCQA, generation, and dialogue tasks is provided in Appendix J\.

Case TypeDatasetQuestionMistral\-SmallTLoRAGround TruthFlipped to correctMedArabiQ\\arabicfontfamilyالتبادل المتبادل بين الكروموسوم 8 و14 يسبب: A\. لمفوما مانتيل; B\. لمفوما بوركت; C\. لمفوما هودجكن; D\. نقيوم متعدد; E\. ابيضاض لمفاوي حاد — A reciprocal translocation between chromosomes 8 and 14 causes: A\. Mantle cell lymphoma; B\. Burkitt lymphoma; C\. Hodgkin lymphoma; D\. Multiple myeloma; E\. Acute lymphoblastic leukemiaA\. Mantle cell lymphomaB\. Burkitt lymphomaB\. Burkitt lymphomaStayed incorrectAraSTEM\\arabicfontfamilyامرأة عمرها 45 عاماً لديها ألم في الأصابع عند التعرض للبرد وآلام مفاصل وصعوبة في بلع الأطعمة الصلبة، إن الفحص المفضل لإجراء تشخيص نهائي هو: A\. العامل الرثوي; B\. أضداد مضادة للنوى; C\. تخطيط قلب كهربائي; D\. نيتروجين يوريا الدم والكرياتينين — A 45\-year\-old woman has finger pain when exposed to cold, joint pain, and difficulty swallowing solid food\. The preferred test for a final diagnosis is: A\. Rheumatoid factor; B\. Antinuclear antibodies; C\. Electrocardiogram; D\. Blood urea nitrogen and creatinineA\. Rheumatoid factorA\. Rheumatoid factorB\. Antinuclear antibodiesFlipped to incorrectMedAraBench\\arabicfontfamilyمن فروع الشريان السباتي الباطن: A\. الشريان العيني; B\. الشريان المخيخي السفلي الأمامي; C\. الشريان المخيخي السفلي الخلفي; D\. الشريان الشوكي — Which of the following is a branch of the internal carotid artery? A\. Ophthalmic artery; B\. Anterior inferior cerebellar artery; C\. Posterior inferior cerebellar artery; D\. Spinal arteryA\. Ophthalmic arteryD\. Spinal arteryA\. Ophthalmic artery

Table S24:Representative case studies comparing TLoRA and base Mistral\-Small\-3\.2\-24B\-Instruct\-2506 \(Mistral\-Small\) on MCQA\. Arabic question stems are shown with their English translation\.

Similar Articles

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

arXiv cs.CL

This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.