Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems
Summary
A comprehensive overview of twenty years of Arabic NLP research, discussing lessons, failures, and open problems in the field.
View Cached Full Text
Cached at: 05/21/26, 06:35 AM
# Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems Source: [https://arxiv.org/abs/2605.20786](https://arxiv.org/abs/2605.20786) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
This survey identifies three critical gaps in explainable AI for Arabic NLP—method, task, and linguistic—and proposes a taxonomy and research agenda for linguistically grounded explanations.
Automated Scoring of Arabic Text Using Large Language Models: A Literature Review
A literature review examining LLM-based approaches for automatic scoring of Arabic text, covering short answer grading and essay scoring, with a proposed taxonomy and comparative analysis.
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
This paper investigates why LLMs underperform in Arabic medical tasks, showing via mechanistic analysis that knowledge exists internally but fails to surface, then proposes TLoRA, a targeted low-rank adaptation method that outperforms full-network LoRA on medical QA and introduces a new Arabic clinical dialogue benchmark.
North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings
This paper presents a conceptual framework for using NLP to address mental healthcare barriers in Algeria and other low-resource, multilingual settings, proposing a research and policy roadmap.
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.