Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
Summary
This survey identifies three critical gaps in explainable AI for Arabic NLP—method, task, and linguistic—and proposes a taxonomy and research agenda for linguistically grounded explanations.
View Cached Full Text
Cached at: 08/28/26, 09:21 AM
# Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap Source: [https://arxiv.org/html/2608.26144](https://arxiv.org/html/2608.26144) Salima Lamsiyah1,Ruslan Mitkov2 1University of Luxembourg, Luxembourg 2University of Alicante, Spain salima\.lamsiyah@uni\.lu ###### Abstract Explainable AI \(XAI\) is now a major theme in NLP; however, Arabic NLP remains under\-explained in three connected senses\. First, there is a*method gap*: Arabic XAI relies heavily on a small set of post\-hoc techniques such as LIME, SHAP, attention visualization, and saliency, while broader NLP XAI offers richer diagnostic, counterfactual, probing, rationale\-based, and human\-centered methods\. Second, there is a*task gap*: existing Arabic XAI work is concentrated in classification tasks, especially sentiment analysis, hate/offensive language detection, fake news, and spam, with weaker coverage of generation, retrieval, translation, summarization, structured prediction, and dialogue\. Third, there is a*linguistic gap*: many explanations identify influential tokens, but rarely explain Arabic\-specific phenomena such as morphology, clitics, dialectal variation, diglossia, orthographic ambiguity, diacritics, code\-switching, named entities, cultural references, or Classical and religious registers\. This critical structured survey synthesizes the reviewed literature on Arabic XAI across text, speech, and multimodal settings\. We argue that Arabic NLP does not only need explanations of model decisions; it needs explanations that are faithful to Arabic as a linguistic, cultural, and sociotechnical object\. We introduce a taxonomy of tasks, methods, linguistic units, varieties, goals, and evaluation practices, and propose a research agenda for linguistically grounded Arabic XAI\. Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap Salima Lamsiyah1, Ruslan Mitkov21University of Luxembourg, Luxembourg2University of Alicante, Spainsalima\.lamsiyah@uni\.lu ## 1Introduction Arabic NLP has advanced rapidly through Arabic\-specific transformer models, pretrained multilingual and monolingual encoders, and Arabic\-centric large language models \(LLMs\)\(Abdul\-Mageedet al\.,[2021](https://arxiv.org/html/2608.26144#bib.bib39); Inoueet al\.,[2021](https://arxiv.org/html/2608.26144#bib.bib40); Khaderet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib22); Senguptaet al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib41); Huanget al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib42)\)\. However, the same developments that improve performance also make system behavior harder to inspect, a concern widely discussed in explainable NLP and LLM explainability research\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19); Zhaoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib37)\)\. This opacity is not merely a generic concern inherited from NLP\. Arabic is morphologically rich, diglossic, dialectally diverse, and orthographically variable, with challenges involving clitics, segmentation, undiacritized text, dialectal variation, and register\-specific usage\(Habash,[2010](https://arxiv.org/html/2608.26144#bib.bib43); Darwishet al\.,[2021](https://arxiv.org/html/2608.26144#bib.bib44); Inoueet al\.,[2021](https://arxiv.org/html/2608.26144#bib.bib40)\)\. A model may rely on a clitic, subword fragment, dialect marker, missing diacritic, named entity boundary, Qur’anic expression, or culturally loaded phrase in ways that are difficult to interpret from surface tokens alone\(Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6); Mustafaet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib1); Sahyoun and Shehata,[2023](https://arxiv.org/html/2608.26144#bib.bib7); Younes,[2026](https://arxiv.org/html/2608.26144#bib.bib9)\)\. Explanations that are adequate for one language, register, or task may therefore be misleading when applied unchanged to Arabic\. General NLP surveys describe a broad XAI landscape, including local interpretation, rationales, probing, counterfactuals, saliency, faithfulness evaluation, and LLM explanation\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19); Hassanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib36); Zhaoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib37)\)\. Arabic\-specific survey work has reviewed Arabic sentiment analysis\(Shi and Agrawal,[2025](https://arxiv.org/html/2608.26144#bib.bib13)\)and explainability for Arabic sentiment analysis\(Alsehaimiet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib23)\)\. However, the provided Arabic XAI literature reveals a broader*explainability gap*: Arabic systems are increasingly explained, but the explanations often remain methodologically narrow, task\-concentrated, and weakly connected to Arabic linguistic and socio\-cultural structure\. We unpack this gap along three dimensions\. Themethod gapis that many Arabic studies use LIME, SHAP, attention visualization, or saliency\-style evidence as the main explanation device\(Abdelwahabet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib28); Awadallahet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib30); Mustafaet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib1); Aljrees,[2024](https://arxiv.org/html/2608.26144#bib.bib27); Sweidanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib24); Boukeet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib2)\)\. These methods are valuable, but they represent only a subset of XAI for NLP\. Thetask gapis that Arabic XAI is concentrated in classification: sentiment analysis, hate/offensive language, fake news, spam, reviews, news, authorship, and related social\-media tasks\(Atabuzzamanet al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib11); Alwateeret al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib29); Mahouachi,[2026](https://arxiv.org/html/2608.26144#bib.bib3); Ibrahim Aboulola and Umer,[2024](https://arxiv.org/html/2608.26144#bib.bib33); Alansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib32); Alshammariet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib25)\)\. Thelinguistic gapis that many explanations highlight tokens or features without explaining Arabic\-specific evidence: morphology, clitics, dialectal variation, diglossia, diacritics, code\-switching, orthographic ambiguity, named entities, cultural references, or Classical and religious registers\. A smaller set of studies points toward a richer agenda by probing Arabic transformers for morphology and dialectal information\(Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6)\), designing explainable metrics for dialectal automatic speech recognition \(ASR\)\(Sahyoun and Shehata,[2023](https://arxiv.org/html/2608.26144#bib.bib7)\), developing diagnostic visual analytics for Arabic named entity recognition \(NER\)\(Younes,[2026](https://arxiv.org/html/2608.26144#bib.bib9)\), and using knowledge\-graph diagnostics for Arabic machine reading comprehension \(MRC\) hallucinations\(AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10)\)\. The central argument of this survey is therefore stronger than the claim that Arabic NLP needs more XAI\. Arabic NLP does not only need explanations of model decisions; it needs explanations that are faithful to Arabic as a linguistic, cultural, and sociotechnical object\. A heatmap over words may be a useful start, but it does not by itself explain whether a model used a dialect cue, an inflectional pattern, a named entity, a religious reference, a moderation norm, or a spurious dataset artifact\. This paper makes four contributions, structured as follows\. First, it frames Arabic XAI through the method, task, and linguistic gaps\. Second, it proposes a four\-level account of what explanations should capture in Arabic NLP: prediction\-level, model\-level, linguistic, and socio\-cultural evidence\. Third, it organizes the surveyed literature into a critical taxonomy and a task\-method coverage table that distinguish current practice from what is missing\. Fourth, it proposes a concrete agenda for linguistically grounded, faithful, useful, and reproducible Arabic XAI\. ## 2What Should an Explanation Explain in Arabic NLP? The phrase “explainable Arabic NLP” does not refer to a single explanatory target but encompasses several distinct ones\. We propose a four\-level distinction, synthesized from the reviewed literature, with the aim of rendering evaluation criteria more precise and tractable\. #### Prediction\-level explanation\. At the most fundamental level, an explanation accounts for why a model assigns a specific label, returns a particular score, retrieves a given result, or generates a specific output\. Most Arabic XAI classification papers operate at this level, using local feature attribution to identify influential words or features for sentiment, fake news, spam, or hate\-speech decisions\(Abdelwahabet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib28); Awadallahet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib30); Aljrees,[2024](https://arxiv.org/html/2608.26144#bib.bib27); Ibrahim Aboulola and Umer,[2024](https://arxiv.org/html/2608.26144#bib.bib33); Boukeet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib2); Alwateeret al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib29)\)\. Prediction\-level explanations can support debugging and user trust, but they are not automatically faithful or linguistically meaningful\. #### Model\-level explanation\. A model\-level explanation addresses the internal representations, layers, features, or mechanisms that the model employs\. This level is especially relevant for Arabic transformers and multilingual encoders, where behavior may be distributed across layers and subword representations\.Abdelaliet al\.\([2022](https://arxiv.org/html/2608.26144#bib.bib6)\), for example, probe Arabic transformer models for morphology, syntax, and dialect identification\. Attention\-based sentiment architectures\(Berrimiet al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib16); Ghasemi and Momtazi,[2023](https://arxiv.org/html/2608.26144#bib.bib17); Sweidanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib24)\)and LLM adaptation studies\(Khaderet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib22)\)are also relevant, although attention or benchmark performance alone does not settle whether the explanation is faithful\. #### Linguistic explanation\. A linguistic explanation identifies which Arabic linguistic features – morphological, syntactic, semantic, dialectal, or orthographic – are operative in the model’s predictions\. This level is where the Arabic explainability gap is most visible\. For Arabic, the relevant unit may be a stem, clitic, lemma, diacritic, named entity, multiword expression, dialect phrase, or orthographic variant rather than a whitespace\-delimited token\. Work on dialectal ASR evaluation\(Sahyoun and Shehata,[2023](https://arxiv.org/html/2608.26144#bib.bib7)\), Arabic NER diagnostics\(Younes,[2026](https://arxiv.org/html/2608.26144#bib.bib9)\), Arabic readability\(Rabih,[2025](https://arxiv.org/html/2608.26144#bib.bib26)\), and transformer probing\(Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6)\)begins to connect explanations to linguistic structure, but this remains less common than token attribution\. #### Socio\-cultural explanation\. A socio\-cultural explanation addresses how a model’s interpretation is shaped by cultural context, moderation norms, religious or classical references, identity terms, user expectations, and the broader social meaning of language use\. This level is essential for harmful\-content moderation, propaganda, Qur’anic semantic search, and high\-stakes applications\. User\-centered Arabic hate\-speech explanation work\(Al\-Ansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib5)\), neuro\-symbolic offensive\-language detection\(Mahouachi,[2026](https://arxiv.org/html/2608.26144#bib.bib3)\), propagandistic meme explanation\(Kmainasiet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib8)\), Qur’anic semantic search\(Mustafaet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib1)\), and knowledge\-graph diagnostics for Qur’anic MRC\(AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10)\)illustrate why Arabic XAI cannot be reduced to visually plausible word highlights\. ## 3A Critical Taxonomy Table[1](https://arxiv.org/html/2608.26144#S3.T1)summarizes the principal dimensions of the explainability gap, distinguishing current practice from missing explanatory capacity and grounding each absence in an Arabic\-specific motivation\. The taxonomy demonstrates that the Arabic XAI gap cannot be reduced to a shortage of explanation papers alone\. Rather, the same pattern recurs across several axes: explanations are often attached to already trained systems as post\-hoc artifacts, while the Arabic\-specific object of explanation remains underspecified\. For example, LIME, SHAP, attention visualization, and feature\-importance methods can identify influential surface units in sentiment, fake\-news, spam, or hate\-speech classifiers\(Abdelwahabet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib28); Awadallahet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib30); Aljrees,[2024](https://arxiv.org/html/2608.26144#bib.bib27); Boukeet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib2)\), but they do not by themselves establish whether the model has used a clitic, a stem, a dialectal marker, a diacritic\-sensitive ambiguity, or a named\-entity boundary in a linguistically meaningful way\. In this sense, the method, task, and linguistic\-unit axes are tightly coupled: when the task is framed as label prediction and the method is framed as token attribution, the resulting explanation is likely to remain at the prediction level even when the underlying error is morphological, dialectal, semantic, or socio\-cultural\. The table also makes clear that Arabic XAI requires a stronger and more explicit alignment between explanation goals and evaluation practice\. Some reviewed work already moves beyond generic attribution through probing, visual analytics, explainable ASR metrics, knowledge\-graph diagnostics, neuro\-symbolic modeling, and user\-centered evaluation\(Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6); Sahyoun and Shehata,[2023](https://arxiv.org/html/2608.26144#bib.bib7); Younes,[2026](https://arxiv.org/html/2608.26144#bib.bib9); AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10); Mahouachi,[2026](https://arxiv.org/html/2608.26144#bib.bib3); Al\-Ansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib5)\)\. These studies are important because they make the explanatory target more explicit: representations, error types, entity boundaries, semantic triples, rules, or user informativeness\. However, they remain scattered rather than forming a shared protocol\. The central challenge, therefore, is not to replace post\-hoc methods wholesale, but to embed them in Arabic\-aware evaluation designs that test faithfulness, stability, usefulness, and fairness under the linguistic variation that Arabic NLP systems actually encounter\. Table 1:Critical taxonomy of the Arabic XAI explainability gap\. ## 4Task\-Method Coverage Table[2](https://arxiv.org/html/2608.26144#S4.T2)maps representative tasks to explanation methods and gaps\. The asymmetry is clear: sentiment analysis and harmful\-content classification form the strongest clusters, whereas generation, retrieval\-augmented systems, translation, summarization, parsing, dialogue, and structured linguistic analysis remain weakly covered in Arabic XAI\. The table also shows that method coverage is narrower than task coverage: many tasks rely on the same small family of post\-hoc explainers\. Table 2:Task\-method coverage and asymmetries in the surveyed Arabic XAI literature\.### 4\.1Sentiment and Opinion Mining: Useful but Over\-Represented Arabic sentiment analysis is the clearest default testbed for XAI\. Studies combine sentiment classifiers with LIME\-style local explanations for domain\-specific medical opinions\(Abdelwahabet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib28)\), multi\-dialect sentiment classification\(Awadallahet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib30)\), and noisy deep explainable sentiment modeling\(Atabuzzamanet al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib11)\)\. Attention\-based sentiment architectures and cross\-lingual sentiment models use attention or contextual representations as interpretive evidence\(Berrimiet al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib16); Ghasemi and Momtazi,[2023](https://arxiv.org/html/2608.26144#bib.bib17)\), while newer work adds SHAP, multi\-self\-attention, empirical comparison of explanation techniques, or platform\-level explanation interfaces\(Azzemet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib4); Chafiquiet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib34); Sweidanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib24)\)\. These works contribute an important baseline: they make Arabic classifier decisions inspectable and show that XAI can be integrated into practical sentiment pipelines\. However, much of this work remains prediction\-level\. It often explains which words influenced a label, not whether the model captured negation, dialectal intensification, sarcasm, stance, topic confounds, or morphology\. This distinction matters because a visually plausible explanation can be useful for demonstration while still failing faithfulness, stability, or Arabic\-awareness\. The sentiment scoping review byAlsehaimiet al\.\([2025](https://arxiv.org/html/2608.26144#bib.bib23)\)is therefore important as a consolidation of the subfield, but it also underscores the limits of sentiment\-centered Arabic XAI as a proxy for the whole language\. ### 4\.2Harmful Content: Explanations Need Social Context Hate speech, offensive language, and propaganda expose the socio\-cultural level of explanation\.Alwateeret al\.\([2025](https://arxiv.org/html/2608.26144#bib.bib29)\)study interpretable Arabic hate\-speech detection with LLMs;Al\-Ansariet al\.\([2026](https://arxiv.org/html/2608.26144#bib.bib5)\)evaluate perceived informativeness of XAI representations for Arabic hate\-speech detection with Arabic\-speaking users;Mahouachi \([2026](https://arxiv.org/html/2608.26144#bib.bib3)\)uses a hybrid multi\-view and neuro\-symbolic approach for Arabic offensive\-language detection; andKmainasiet al\.\([2025](https://arxiv.org/html/2608.26144#bib.bib8)\)study explanation\-enhanced detection of propagandistic Arabic memes\. Compared with sentiment work, this cluster more clearly motivates user\-centered and safety\-oriented explanations\. It also shows why token salience is insufficient\. A word can be harmful, quoted, reclaimed, dialect\-specific, target\-dependent, or embedded in a meme\. The contribution of this line of work is to connect explainability with moderation and social meaning\. Its limitation is that the field still lacks shared protocols for testing whether explanations remain useful and fair across dialects, regions, identity terms, targets, and cultural references\. In our synthesis, harmful\-content XAI is one of the strongest arguments for treating Arabic as a sociotechnical object rather than a string of tokens\. ### 4\.3Information Integrity: Attribution Is Not Evidence Arabic fake\-news and spam studies use explainability to inspect high\-performing classifiers\.Aljrees \([2024](https://arxiv.org/html/2608.26144#bib.bib27)\)combines LIME with an ELMo\-based tri\-ensemble fake\-news model;Ibrahim Aboulola and Umer \([2024](https://arxiv.org/html/2608.26144#bib.bib33)\)use large\-language\-feature embeddings, CNN\-LSTM ensembles, and explainable AI for Arabic fake\-news classification; andBoukeet al\.\([2026](https://arxiv.org/html/2608.26144#bib.bib2)\)integrates LIME and SHAP into an Arabic spam\-detection pipeline based on LightGBM\. These studies are useful for feature inspection and model debugging\. The critical issue is that attribution is not the same as evidence\. A highlighted token may reveal what the classifier used, but not whether the model identified deception, propaganda framing, source credibility, temporal inconsistency, or a culturally grounded claim\. For Arabic misinformation, explanations should ideally connect entities, claims, sources, and context\. The reviewed work moves toward transparency, but it rarely explains the information\-integrity phenomenon itself\. ### 4\.4Diagnostics Beyond Classification Work that most substantively engages with Arabic\-specific phenomena frequently falls outside the conventions of standard post\-hoc explanation\.Mustafaet al\.\([2024](https://arxiv.org/html/2608.26144#bib.bib1)\)use LIME and SHAP for Arabic transformer models in Qur’anic semantic search, where explanation must confront high\-register and religious text\.AlGhamdiet al\.\([2026](https://arxiv.org/html/2608.26144#bib.bib10)\)use knowledge\-graph triple diagnostics for Arabic MRC hallucinations, shifting from word salience to relational grounding\.Abdelaliet al\.\([2022](https://arxiv.org/html/2608.26144#bib.bib6)\)probes Arabic transformers for morphology, syntax, and dialect information\.Sahyoun and Shehata \([2023](https://arxiv.org/html/2608.26144#bib.bib7)\)propose an explainable ASR metric for dialectal Arabic, andYounes \([2026](https://arxiv.org/html/2608.26144#bib.bib9)\)develops visual analytics for Arabic NER evaluation\.Rabih \([2025](https://arxiv.org/html/2608.26144#bib.bib26)\)studies interpretable measures for Arabic readability\. This group shows what linguistically grounded Arabic XAI can look like\. Rather than asking only which token affected a label, these works ask where linguistic information is encoded, what kind of recognition error occurred, which entity boundaries or annotations are problematic, which readability features matter, or whether a generated answer is semantically grounded\. Their limitation is scale and standardization: they are promising case studies and frameworks, not yet a shared paradigm for Arabic XAI evaluation\. ### 4\.5Multimodal and LLM\-Adjacent Work Arabic XAI is also expanding beyond text\-only classification\. Arabic sign language recognition uses LIME or Grad\-CAM\-style visual explanations\(Baghdadiet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib31); Balatet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib12)\); Arabic captioning uses interpretable visual concept integration\(Elchafei and Fashwan,[2025](https://arxiv.org/html/2608.26144#bib.bib14)\); and meme detection connects multimodal classification with rationale\-oriented resources\(Kmainasiet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib8)\)\. LLM\-oriented Arabic work discusses adaptation, evaluation, and ethical deployment\(Khaderet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib22)\), while AraLingBench evaluates Arabic linguistic capabilities of LLMs\(Zbibet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib15)\)\. These works widen the scope of Arabic XAI, but they also make evaluation harder\. For multimodal tasks, explanation quality must be separated from visual grounding and Arabic generation quality\. For LLMs, explanation must be connected to factual grounding, linguistic competence, hallucination, and user expectations\. Within the reviewed literature, Arabic LLM explainability remains more of an emerging requirement than a mature methodology\. ## 5Critical Discussion: The existing Gaps ### 5\.1The Method Gap The reviewed literature shows a clear preference for methods that are easy to attach to existing classifiers: LIME, SHAP, attention visualization, saliency, or feature importance\. This is understandable because model\-agnostic methods are accessible and produce intuitive visual artifacts\. The problem is that these artifacts are often treated as explanations without enough evidence that they are faithful, stable, or aligned with Arabic linguistic units\. General NLP XAI surveys emphasize that local explanations, rationales, probing, counterfactuals, and faithfulness tests answer different questions\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19); Hassanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib36)\)\. In the Arabic corpus, richer methods appear in isolated forms, including probing\(Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6)\), visual analytics\(Younes,[2026](https://arxiv.org/html/2608.26144#bib.bib9)\), knowledge\-graph diagnostics\(AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10)\), user\-centered evaluation\(Al\-Ansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib5)\), and neuro\-symbolic modeling\(Mahouachi,[2026](https://arxiv.org/html/2608.26144#bib.bib3)\), but they have not yet become common practice\. ### 5\.2The Task Gap Arabic XAI is disproportionately shaped by classification\. Sentiment analysis, hate/offensive language, fake news, spam, app reviews, news classification, authorship, and mental health detection supply many of the examples and methods\(Abdelwahabet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib28); Atabuzzamanet al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib11); Alwateeret al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib29); Aljrees,[2024](https://arxiv.org/html/2608.26144#bib.bib27); Boukeet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib2); Alansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib32); Kumaret al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib18); Alshammariet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib25)\)\. Classification is a useful starting point, but it encourages explanations that identify label\-supporting tokens\. Modern Arabic NLP increasingly includes retrieval, generation, machine reading comprehension, captioning, LLM adaptation, and linguistic benchmarking\(Mustafaet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib1); AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10); Elchafei and Fashwan,[2025](https://arxiv.org/html/2608.26144#bib.bib14); Khaderet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib22); Zbibet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib15)\)\. These tasks require explanations of grounding, hallucination, alignment, fluency, factuality, entity tracking, and discourse, not only explanations of class labels\. ### 5\.3The Linguistic Gap We argue that the most important gap is linguistic\. Many Arabic XAI papers explain which input words influenced a prediction, but not what Arabic evidence the model used\. This is not a minor technical issue\. Arabic preprocessing choices, including normalization, segmentation, stop\-word removal, dialect filtering, subword tokenization, and diacritic handling, can change the unit being explained\. If an explanation highlights a whitespace token while the model operates over subwords, or if preprocessing removes clitics or diacritics, the explanation may be difficult to interpret linguistically\. Work on transformer probing, dialectal ASR, NER diagnostics, readability, and Qur’anic MRC demonstrates that explanations can be tied to morphology, dialect, recognition errors, entity boundaries, readability features, and semantic relations\(Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6); Sahyoun and Shehata,[2023](https://arxiv.org/html/2608.26144#bib.bib7); Younes,[2026](https://arxiv.org/html/2608.26144#bib.bib9); Rabih,[2025](https://arxiv.org/html/2608.26144#bib.bib26); AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10)\)\. This is the direction in which Arabic XAI should move\. ### 5\.4The Evaluation Gap Faithfulness, plausibility, stability, usefulness, and fairness are distinct evaluation targets\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19)\)\. Arabic XAI papers increasingly acknowledge these criteria, and some compare explanation techniques or evaluate user informativeness\(Azzemet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib4); Al\-Ansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib5)\)\. However, many studies still rely on predictive accuracy plus selected explanation examples\. For Arabic, this is especially weak because an explanation can change under spelling variation, normalization, segmentation, dialectal paraphrase, or diacritic removal while the meaning remains similar\. In high\-stakes settings–health sentiment\(Abdelwahabet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib28); Sweidanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib24)\), mental health detection\(Kumaret al\.,[2023](https://arxiv.org/html/2608.26144#bib.bib18)\), hate speech\(Alwateeret al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib29); Al\-Ansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib5)\), fake news\(Aljrees,[2024](https://arxiv.org/html/2608.26144#bib.bib27); Ibrahim Aboulola and Umer,[2024](https://arxiv.org/html/2608.26144#bib.bib33)\), Qur’anic search and MRC\(Mustafaet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib1); AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10)\), and Arabic sign language recognition\(Baghdadiet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib31); Balatet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib12)\)–explanations should be evaluated for decision support, error discovery, harm reduction, and stakeholder usefulness, not only visual plausibility\. ## 6Toward Linguistically Grounded Arabic XAI A linguistically grounded agenda for Arabic XAI should move beyond adding generic explanation methods to Arabic models\. It should instead redesign explanation targets, evaluation protocols, and reporting standards around the linguistic and sociotechnical properties of Arabic itself\. The following directions outline how future work can make explanations more faithful to model behavior, more stable under Arabic\-specific variation, more useful to Arabic\-speaking users and domain experts, and more reproducible across tasks, datasets, and varieties\. #### Arabic\-aware units of explanation\. Explanations should be evaluated at levels that are meaningful for Arabic: morphemes, clitics, stems, lemmas, diacritics, named entities, multiword expressions, dialectal phrases, and discourse cues\. Token\-level heatmaps are insufficient when model decisions depend on morphology or orthography\. Future papers should report the relationship between the model’s internal tokenization and the units shown to users\. #### Arabic\-preserving perturbation tests\. Arabic XAI needs perturbation tests that preserve meaning while varying surface form: spelling variation, normalization, clitic segmentation, dialectal paraphrases, and diacritic restoration or removal\. Such tests would help distinguish robust explanations from artifacts of preprocessing or tokenization\. This would extend the faithfulness and stability concerns emphasized in general NLP XAI\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19)\)to Arabic\-specific forms\. #### Cross\-variety explanation evaluation\. Benchmarks should evaluate explanations across MSA, regional dialects, Classical Arabic, social\-media Arabic, Arabizi, and code\-switched Arabic\. Work on dialectal ASR and Arabic transformer probing shows that variety can be studied systematically\(Sahyoun and Shehata,[2023](https://arxiv.org/html/2608.26144#bib.bib7); Abdelaliet al\.,[2022](https://arxiv.org/html/2608.26144#bib.bib6)\); Arabic XAI should make variety an explanation variable, not merely a dataset description\. #### Human rationale benchmarks with Arabic expertise\. Arabic XAI needs shared datasets with human rationales, expert annotations, and task\-specific explanation quality labels\. User\-centered work on Arabic hate\-speech explanation\(Al\-Ansariet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib5)\)and explanation\-enhanced multimodal resources\(Kmainasiet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib8)\)point to this direction\. Annotators should include Arabic speakers with dialectal competence and, where necessary, domain expertise in moderation, health, education, religious text, or accessibility\. #### Explanation evaluation for Arabic LLMs, RAG, and generation\. Future work should target Arabic LLMs, retrieval\-augmented generation, QA, summarization, machine translation, dialogue, and hallucination detection\. Arabic LLM adaptation and linguistic benchmarking are already emerging\(Khaderet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib22); Zbibet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib15)\), and knowledge\-graph diagnostics show one path for hallucination analysis\(AlGhamdiet al\.,[2026](https://arxiv.org/html/2608.26144#bib.bib10)\)\. Explanations for these systems should address factual grounding, retrieval evidence, morphology, syntax, discourse, and answer\-source alignment\. #### Faithfulness, plausibility, stability, usefulness, and fairness protocols\. Arabic XAI needs evaluation protocols that separate what the model actually used from what users find plausible\. Faithfulness tests, human plausibility judgments, stability under Arabic\-preserving perturbations, usefulness for stakeholders, and fairness across varieties and identity terms should be reported separately\. This distinction is especially important in moderation, health, education, religious\-domain applications, and information integrity\. #### Reproducibility standards\. Arabic XAI papers should report dataset variety, preprocessing, normalization, tokenization, segmentation, diacritic handling, model checkpoints, explanation parameters, evaluation metrics, and known validity limits\. Where possible, studies should release code, trained models or inference scripts, explanation outputs, rationale annotations, and diagnostic error analyses\. Without these details, it is difficult to know whether a method is better for Arabic or simply easier to visualize\. ## 7Conclusion Arabic XAI is an emerging and necessary research direction, but the field is still shaped by a clear explainability gap\. Current work has demonstrated the value of LIME, SHAP, attention visualization, probing, visual analytics, knowledge graphs, neuro\-symbolic methods, and user\-centered explanations\. However, the dominant pattern remains output\-level explanation for classification tasks, especially sentiment analysis, harmful\-content detection, fake news, and spam\. This leaves important parts of Arabic NLP under\-explained, including generation, retrieval, QA, summarization, translation, structured prediction, and LLM\-based applications\. The main message of this survey is that Arabic NLP does not only need explanations of model decisions; it needs explanations that are faithful to Arabic as a linguistic, cultural, and sociotechnical object\. Explanations should account for morphology, clitics, dialectal variation, diglossia, diacritics, orthographic ambiguity, named entities, cultural references, and Classical or religious registers\. By identifying the method gap, task gap, and linguistic gap, this survey aims to provide a clearer foundation for future work in Arabic XAI\. Advancing this agenda is important for building Arabic NLP systems that are not only accurate, but also transparent, trustworthy, fair, and useful in the diverse contexts where Arabic is used\. Future research should therefore move from explaining predictions to explaining Arabic itself through Arabic\-aware explanation units, cross\-variety evaluation, perturbation testing, human rationale benchmarks, LLM/RAG explanation protocols, and stronger reproducibility standards\. ## 8Survey Methodology and Scope This paper is a*critical structured survey*, not a fully systematic review\. We searched across major academic databases and indexes, including ACL Anthology, ACM Digital Library, IEEE Xplore, ScienceDirect, SpringerLink, MDPI, arXiv/preprint venues, and Google Scholar, using combinations of keywords such as Arabic NLP, explainable AI, interpretability, transparency, XAI, LIME, SHAP, attention, saliency, LLMs, sentiment analysis, hate speech, fake news, spam, ASR, NER, MRC, hallucination, readability, and multimodal Arabic NLP\. General XAI\-for\-NLP and XAI\-for\-LLMs surveys are used for background and terminology\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19); Hassanet al\.,[2025](https://arxiv.org/html/2608.26144#bib.bib36); Zhaoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib37)\)\. We included studies when they met two criteria: \(i\) the task, dataset, model, or output was Arabic or Arabic\-language adjacent; and \(ii\) the work explicitly addressed explanation, interpretability, transparency, diagnostic analysis, or human\-facing explanation\. We excluded papers that only reported Arabic NLP performance without an explicit explanatory component, and we treat attention\-based models cautiously because attention is not necessarily an explanation\(Zini and Awad,[2022](https://arxiv.org/html/2608.26144#bib.bib35); Luoet al\.,[2024](https://arxiv.org/html/2608.26144#bib.bib19)\)\. Each included work was annotated by task, model family, XAI method, Arabic variety or register, explanation goal, evaluation practice, and degree of linguistic grounding\. The relatively small number of directly relevant papers confirms the lack of dedicated work on XAI for Arabic NLP and motivates framing this paper as a critical survey of an emerging, underdeveloped area\. ## Limitations This survey provides a critical synthesis of recent work on explainability for Arabic NLP, but it does not claim to be an exhaustive systematic review of all Arabic NLP papers that may contain interpretable components\. Its scope is limited to studies that explicitly frame their contribution in terms of explainability, interpretability, transparency, or diagnostic analysis\. As a result, some relevant work on Arabic model analysis, evaluation, or error diagnosis may fall outside the discussion if it does not use XAI terminology\. A second limitation is that the reviewed literature is uneven across tasks and publication stages\. Some areas, such as sentiment analysis and harmful\-content detection, are represented more strongly than others, while emerging topics such as Arabic LLM explainability, retrieval\-augmented generation, hallucination analysis, and multimodal Arabic XAI remain relatively new\. Consequently, the taxonomy and agenda proposed here should be seen as a structured snapshot of a developing field rather than a fixed classification\. ## Use of AI Assistance In preparing this work, we used AI\-assisted tools for language editing, such as spell checking and stylistic revisions with Grammarly and ChatGPT\. The authors take full responsibility for the scientific content, analyses, and conclusions presented in this work\. ## References - Post\-hoc analysis of arabic transformer models\.InProceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP,pp\. 91–103\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.4.3.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.5.4.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.6.5.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p3.1),[§4\.4](https://arxiv.org/html/2608.26144#S4.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.6.5.5.1.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§5\.3](https://arxiv.org/html/2608.26144#S5.SS3.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px3.p1.1)\. - Y\. Abdelwahab, M\. Kholief, and A\. A\. H\. Sedky \(2022\)Justifying arabic text sentiment analysis using explainable ai \(xai\): lasik surgeries case study\.Information13\(11\),pp\. 536\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p2.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - M\. Abdul\-Mageed, A\. R\. Elmadany, and E\. M\. B\. Nagoudi \(2021\)ARBERT & MARBERT: deep bidirectional transformers for Arabic\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 7088–7105\.External Links:[Link](https://aclanthology.org/2021.acl-long.551/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.551)Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1)\. - N\. Al\-Ansari, D\. Al\-Thani, and M\. Bahameish \(2026\)Explaining in context: perceived informativeness of explainable artificial intelligence \(xai\) in arabic hate speech detection\.Computers in Human Behavior Reports,pp\. 101083\.Cited by:[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.6.5.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.7.6.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p3.1),[§4\.2](https://arxiv.org/html/2608.26144#S4.SS2.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.3.2.5.1.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px4.p1.1)\. - A\. Alansari, D\. Alomari, S\. Mahmood, and I\. Ahmad \(2026\)Multi\-label classification of arabic app reviews with data augmentation and explainable ai\.Arabian Journal for Science and Engineering51\(9\),pp\. 12487–12506\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.7.6.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1)\. - N\. A\. AlGhamdi, S\. Al\-Azani, K\. Nuamah, and A\. Bundy \(2026\)A knowledge graph based diagnostic framework for analyzing hallucinations in arabic machine reading comprehension\.InProceedings of the 2nd Workshop on NLP for Languages Using Arabic Script,pp\. 413–421\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.4.3.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.6.5.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p3.1),[§4\.4](https://arxiv.org/html/2608.26144#S4.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.5.4.5.1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.9.8.5.1.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§5\.3](https://arxiv.org/html/2608.26144#S5.SS3.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px5.p1.1)\. - T\. Aljrees \(2024\)Improving prediction of arabic fake news using elmo’s features\-based tri\-ensemble model and lime xai\.IEEE Access12,pp\. 63066–63076\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p2.1),[§4\.3](https://arxiv.org/html/2608.26144#S4.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.4.3.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - A\. Alsehaimi, A\. Babour, and D\. Alahmadi \(2025\)Toward transparent modeling: a scoping review of explainability for arabic sentiment analysis\.Applied Sciences15\(19\),pp\. 10659\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p2.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p2.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1)\. - B\. S\. Alshammari, K\. N\. Alfraidi, M\. O\. Aljudaiey, and W\. Abdelhalim \(2025\)Robustness and explainability in arabic authorship models: tackling the challenge of ai\-generated texts\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.7.6.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1)\. - M\. Alwateer, I\. Gad, M\. Elmarhomy, G\. Elmarhomy, H\. Hashim, M\. Almaliki, and E\. Atlam \(2025\)Interpretable arabic hate speech detection using large language model\.In2025 2nd International Conference on Advanced Innovations in Smart Cities,pp\. 1–8\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[§4\.2](https://arxiv.org/html/2608.26144#S4.SS2.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.3.2.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - M\. Atabuzzaman, M\. Shajalal, M\. B\. Baby, and A\. Boden \(2023\)Arabic sentiment analysis with noisy deep explainable model\.InProceedings of the 2023 7th International Conference on Natural Language Processing and Information Retrieval,pp\. 185–189\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1)\. - M\. S\. Awadallah, F\. de Arriba\-Perez, E\. Costa\-Montenegro, M\. Kholief, and N\. El\-Bendary \(2022\)Investigation of local interpretable model\-agnostic explanations \(lime\) framework with multi\-dialect arabic text sentiment classification\.In2022 32nd International Conference on Computer Theory and Applications,pp\. 116–121\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.5.4.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p2.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1)\. - Y\. C\. H\. Azzem, F\. Harrag, and L\. Bellatreche \(2024\)Exploring explainability in arabic language models: an empirical analysis of techniques\.Procedia Computer Science244,pp\. 212–219\.Cited by:[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.7.6.3.1.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - N\. A\. Baghdadi, Y\. M\. Abdulazeem, H\. ZainEldin, T\. A\. Farrag, M\. K\. Aljohani, A\. Malki, M\. Badawy, and M\. A\. Elhosseini \(2024\)Toward robust arabic sign language recognition via vision transformers and local interpretable model\-agnostic explanations integration\.Journal of Disability Research3\(8\),pp\. 20240092\.Cited by:[§4\.5](https://arxiv.org/html/2608.26144#S4.SS5.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.8.7.5.1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - M\. Balat, R\. Awaad, A\. B\. Zaky, and S\. A\. Aly \(2025\)Revolutionizing communication with deep learning and xai for enhanced arabic sign language recognition\.arXiv preprint arXiv:2501\.08169\.Cited by:[§4\.5](https://arxiv.org/html/2608.26144#S4.SS5.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.8.7.5.1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - M\. Berrimi, M\. Oussalah, A\. Moussaoui, and M\. Saidi \(2023\)Attention mechanism architecture for arabic sentiment analysis\.ACM Transactions on Asian and Low\-Resource Language Information Processing22\(4\),pp\. 1–26\.Cited by:[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1)\. - M\. A\. Bouke, O\. I\. Alramli, A\. A\. A\. Abdalhafid, and A\. Abdullah \(2026\)A novel lightgbm model for arabic spam detection integrated with xai for enhanced explainability\.Computers and Electrical Engineering133,pp\. 111032\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.7.6.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p2.1),[§4\.3](https://arxiv.org/html/2608.26144#S4.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.4.3.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1)\. - Y\. Chafiqui, H\. Anoun, L\. Aloiradi, and F\. E\. Nazih \(2024\)XS2A: a platform for explainable sentiment analysis in arabic comments\.InInternational Conference on Advances in Communication Technology and Computer Engineering,pp\. 309–320\.Cited by:[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1)\. - K\. Darwish, N\. Habash, M\. Abbas, H\. Al\-Khalifa, H\. T\. Al\-Natsheh, H\. Bouamor, K\. Bouzoubaa, V\. Cavalli\-Sforza, S\. R\. El\-Beltagy, W\. El\-Hajj,et al\.\(2021\)A panoramic survey of natural language processing in the arab world\.Communications of the ACM64\(4\),pp\. 72–81\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1)\. - P\. Elchafei and A\. Fashwan \(2025\)Multimodal arabic captioning with interpretable visual concept integration\.arXiv preprint arXiv:2510\.03295\.Cited by:[§4\.5](https://arxiv.org/html/2608.26144#S4.SS5.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.8.7.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1)\. - R\. Ghasemi and S\. Momtazi \(2023\)How a deep contextualized representation and attention mechanism justifies explainable cross\-lingual sentiment analysis\.ACM Transactions on Asian and Low\-Resource Language Information Processing22\(11\),pp\. 1–15\.Cited by:[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1)\. - N\. Y\. Habash \(2010\)Introduction to arabic natural language processing\.Morgan & Claypool Publishers\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1)\. - M\. M\. Hassan, A\. Nag, R\. Biswas, M\. S\. Ali, S\. Zaman, A\. K\. Bairagi, and C\. Kaushal \(2025\)Explainable artificial intelligence for natural language processing: a survey\.Data & Knowledge Engineering160,pp\. 102470\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§8](https://arxiv.org/html/2608.26144#S8.p1.1)\. - M\. M\. Hossain, M\. S\. Hossain, M\. Safran, S\. Alfarhood, M\. Alfarhood, and M\. F\. Mridha \(2024\)A hybrid attention\-based transformer model for arabic news classification using text embedding and deep learning\.IEEE Access12,pp\. 198046–198066\.Cited by:[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.7.6.5.1.1)\. - H\. Huang, F\. Yu, J\. Zhu, X\. Sun, H\. Cheng, S\. Dingjie, Z\. Chen, M\. Alharthi, B\. An, J\. He,et al\.\(2024\)AceGPT, localizing large language models in arabic\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 8139–8163\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1)\. - O\. Ibrahim Aboulola and M\. Umer \(2024\)Novel approach for arabic fake news classification using embedding from large language features with cnn\-lstm ensemble model and explainable ai\.Scientific Reports14\(1\),pp\. 30463\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[§4\.3](https://arxiv.org/html/2608.26144#S4.SS3.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.4.3.5.1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - G\. Inoue, B\. Alhafni, N\. Baimukan, H\. Bouamor, and N\. Habash \(2021\)The interplay of variant, size, and task type in arabic pre\-trained language models\.InProceedings of the sixth Arabic natural language processing workshop,pp\. 92–104\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1)\. - F\. Kboubi and A\. Habacha Chaibi \(2025\)Fine\-grained arabic dialect identification: investigating various approaches across multiple datasets\.ACM Transactions on Asian and Low\-Resource Language Information Processing24\(10\),pp\. 1–33\.Cited by:[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.5.4.3.1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.6.5.5.1.1)\. - K\. A\. Khader, M\. S\. Hussein, and A\. S\. Abu\-Issa \(2025\)Adapting large language models for arabic: comparative evaluation, fine\-tuning, and ethical deployment\.IEEE Access\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px2.p1.1),[§4\.5](https://arxiv.org/html/2608.26144#S4.SS5.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.9.8.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px5.p1.1)\. - M\. B\. Kmainasi, A\. Hasnat, M\. A\. Hasan, A\. E\. Shahroor, and F\. Alam \(2025\)MemeIntel: explainable detection of propagandistic and hateful memes\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 30263–30279\.Cited by:[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.5.4.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.7.6.3.1.1),[§4\.2](https://arxiv.org/html/2608.26144#S4.SS2.p1.1),[§4\.5](https://arxiv.org/html/2608.26144#S4.SS5.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.3.2.5.1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.8.7.5.1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px4.p1.1)\. - A\. Kumar, J\. Kumari, and J\. Pradhan \(2023\)Explainable deep learning for mental health detection from english and arabic social media posts\.ACM Transactions on Asian and Low\-Resource Language Information Processing\.Cited by:[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.3.2.3.1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.7.6.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - S\. Luo, H\. Ivison, S\. C\. Han, and J\. Poon \(2024\)Local interpretations for explainable natural language processing: a survey\.ACM Computing Surveys56\(9\),pp\. 1–36\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px2.p1.1),[§8](https://arxiv.org/html/2608.26144#S8.p1.1),[§8](https://arxiv.org/html/2608.26144#S8.p2.1)\. - R\. Mahouachi \(2026\)A hybrid multi\-view and neuro\-symbolic approach for arabic offensive language detection: a step towards safer online spaces\.Expert Systems with Applications,pp\. 131854\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.6.5.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p3.1),[§4\.2](https://arxiv.org/html/2608.26144#S4.SS2.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.3.2.5.1.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1)\. - A\. M\. Mustafa, S\. Nakhleh, R\. Irsheidat, and R\. Alruosan \(2024\)Interpreting arabic transformer models: a study on xai interpretability for qur’anic semantic\-search models\.Jordanian Journal of Computers and Information Technology10\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px4.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.4.3.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.5.4.3.1.1),[§4\.4](https://arxiv.org/html/2608.26144#S4.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.5.4.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - N\. Rabih \(2025\)Interpretable measures for arabic readability\.Cited by:[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.4.3.3.1.1),[§4\.4](https://arxiv.org/html/2608.26144#S4.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.6.5.5.1.1),[§5\.3](https://arxiv.org/html/2608.26144#S5.SS3.p1.1)\. - A\. Sahyoun and S\. Shehata \(2023\)AraDiaWER: an explainable metric for dialectical arabic asr\.InProceedings of the Second Workshop on NLP Applications to Field Linguistics,pp\. 64–73\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.4.3.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.5.4.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p3.1),[§4\.4](https://arxiv.org/html/2608.26144#S4.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.6.5.5.1.1),[§5\.3](https://arxiv.org/html/2608.26144#S5.SS3.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px3.p1.1)\. - N\. Sengupta, S\. K\. Sahu, B\. Jia, S\. Katipomu, H\. Li, F\. Koto, W\. Marshall, G\. Gosal, C\. Liu, Z\. Chen,et al\.\(2023\)Jais and jais\-chat: arabic\-centric foundation and instruction\-tuned open generative large language models\.arXiv preprint arXiv:2308\.16149\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1)\. - Z\. Shi and R\. Agrawal \(2025\)A comprehensive survey of contemporary arabic sentiment analysis: methods, challenges, and future directions\.Findings of the Association for Computational Linguistics: NAACL 2025,pp\. 3760–3772\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p2.1)\. - A\. H\. Sweidan, N\. El\-Bendary, S\. A\. Taie, A\. M\. Idrees, and E\. Elhariri \(2025\)Explainable deep learning for covid\-19 vaccine sentiment in arabic tweets using multi\-self\-attention bilstm with xlnet\.Big Data and Cognitive Computing9\(2\),pp\. 37\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.26144#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.2.1.5.1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1)\. - A\. M\. Younes \(2026\)DeformAR: a visual analytics framework for evaluation of arabic named entity recognition\.InProceedings of the 2nd Workshop on NLP for Languages Using Arabic Script,pp\. 253–275\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p3.1),[§2](https://arxiv.org/html/2608.26144#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.2.1.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.4.3.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.6.5.3.1.1),[Table 1](https://arxiv.org/html/2608.26144#S3.T1.1.7.6.3.1.1),[§3](https://arxiv.org/html/2608.26144#S3.p3.1),[§4\.4](https://arxiv.org/html/2608.26144#S4.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.6.5.5.1.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§5\.3](https://arxiv.org/html/2608.26144#S5.SS3.p1.1)\. - M\. B\. Zbib, H\. A\. A\. K\. Hammoud, A\. Mohanna, N\. Rizk, F\. Karnib, S\. Moukaled, and B\. Ghanem \(2026\)AraLingBench: a human\-annotated benchmark for evaluating arabic linguistic capabilities of large language models\.InProceedings of the 2nd Workshop on NLP for Languages Using Arabic Script,pp\. 385–393\.Cited by:[§4\.5](https://arxiv.org/html/2608.26144#S4.SS5.p1.1),[Table 2](https://arxiv.org/html/2608.26144#S4.T2.1.9.8.5.1.1),[§5\.2](https://arxiv.org/html/2608.26144#S5.SS2.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px5.p1.1)\. - H\. Zhao, H\. Chen, F\. Yang, N\. Liu, H\. Deng, H\. Cai, S\. Wang, D\. Yin, and M\. Du \(2024\)Explainability for large language models: a survey\.ACM Transactions on Intelligent Systems and Technology15\(2\),pp\. 1–38\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p2.1),[§8](https://arxiv.org/html/2608.26144#S8.p1.1)\. - J\. E\. Zini and M\. Awad \(2022\)On the explainability of natural language processing deep models\.ACM Computing Surveys55\(5\),pp\. 1–31\.Cited by:[§1](https://arxiv.org/html/2608.26144#S1.p1.1),[§1](https://arxiv.org/html/2608.26144#S1.p2.1),[§5\.1](https://arxiv.org/html/2608.26144#S5.SS1.p1.1),[§5\.4](https://arxiv.org/html/2608.26144#S5.SS4.p1.1),[§6](https://arxiv.org/html/2608.26144#S6.SS0.SSS0.Px2.p1.1),[§8](https://arxiv.org/html/2608.26144#S8.p1.1),[§8](https://arxiv.org/html/2608.26144#S8.p2.1)\.
Similar Articles
Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
A systematic review of XAI research in the context of the EU Right to Explanation, analyzing gaps between legal requirements and technical implementations across GDPR and the AI Act.
Evaluating Explainability in Safety-Critical ATR Systems: Limitations of Post-Hoc Methods and Paths Toward Robust XAI
This paper evaluates explainability methods in safety-critical Automatic Target Recognition (ATR) systems, highlighting the limitations of post-hoc techniques like saliency and attention maps. It proposes a taxonomy and assessment framework to address issues such as spurious explanations and instability, advocating for more robust, causally grounded XAI approaches.
Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids
This paper investigates whether high-quality Natural Language Explanations (NLEs) generated by LLMs from XAI outputs actually improve task performance, finding they do not aid accuracy but inflate confidence, revealing a quality-usefulness gap.
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
A systematic survey of Federated Explainable Artificial Intelligence (FedXAI), covering roles, architectures, evaluation practices, and open challenges. It presents a multi-axis taxonomy and discusses model-agnostic to interpretable-by-design approaches, highlighting gaps in standardization and privacy-aware evaluation.
Applied Explainability for Large Language Models: A Comparative Study
A comparative study evaluating three explainability techniques (Integrated Gradients, Attention Rollout, SHAP) on fine-tuned DistilBERT for sentiment classification, highlighting trade-offs between gradient-based, attention-based, and model-agnostic approaches for LLM interpretability.