arabic

Tag

Cards List
#arabic

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

arXiv cs.CL · 2026-08-04 Cached

This paper investigates why LLMs underperform in Arabic medical tasks, showing via mechanistic analysis that knowledge exists internally but fails to surface, then proposes TLoRA, a targeted low-rank adaptation method that outperforms full-network LoRA on medical QA and introduces a new Arabic clinical dialogue benchmark.

0 favorites 0 likes
#arabic

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

arXiv cs.CL · 2026-07-31 Cached

This paper introduces AHA-Memes, the first large-scale Arabic hateful meme benchmark with fine-grained multi-label annotations, covering 5K manually annotated and ~66K silver-labeled memes, and benchmarks various multimodal models for culturally grounded hate detection.

0 favorites 0 likes
#arabic

Constrained CTC Decoding for Efficient Diacritic Restoration

arXiv cs.CL · 2026-07-22 Cached

This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.

0 favorites 0 likes
#arabic

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

arXiv cs.CL · 2026-07-07 Cached

This paper investigates methods to steer Arabic LLMs toward dialect-specific generation by identifying sparse neuron populations and extracting dialect activation directions, enabling dialect control at inference time without fine-tuning.

0 favorites 0 likes
#arabic

I swapped the TTS in my voice agent and it cut the lag people actually feel more than anything else

Reddit r/AI_Agents · 2026-07-06

The author shares their experience swapping the TTS in their voice agent to a custom model (Banter 1) designed for bilingual Arabic-English conversations, which significantly reduced perceived lag.

0 favorites 0 likes
#arabic

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

Hugging Face Daily Papers · 2026-07-06 Cached

PAST-TIDE is a stance detection system for the StanceNakba Shared Task, using statement tuning with cloze-style masked language modeling, prototypical contrastive learning, and topic-conditional layer normalization for cross-topic Arabic stance detection, achieving macro-F1 scores of 0.75 and 0.74 on subtasks A and B.

0 favorites 0 likes
#arabic

Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study

arXiv cs.CL · 2026-07-02 Cached

Introduces a comprehensive hate speech dataset for Turkish and Arabic, and develops state-of-the-art BERT-based models for hate speech analysis including classification, intensity prediction, target identification, and span detection.

0 favorites 0 likes
#arabic

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

arXiv cs.CL · 2026-07-02 Cached

This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.

0 favorites 0 likes
#arabic

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

arXiv cs.CL · 2026-07-01 Cached

This paper presents a benchmark for Arabic-Russian scientific translation, including a hybrid parallel corpus of 27,000 sentence pairs and fine-tuned multilingual models (mT5, NLLB, Qwen) using LoRA. The best model achieves BLEU 23.15, and the work aims to lower language barriers for scientific knowledge exchange between Arabic and Russian researchers.

0 favorites 0 likes
#arabic

CohereLabs/cohere-transcribe-arabic-07-2026

Hugging Face Models Trending · 2026-06-18 Cached

Cohere releases an open-source 2B parameter Arabic ASR model optimized for Arabic dialect performance and Arabic-English code-switching, based on the Conformer encoder-decoder architecture.

0 favorites 0 likes
#arabic

When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates

arXiv cs.CL · 2026-06-12 Cached

This paper introduces SemCog Bench, a curated benchmark of 1,858 Arabic-Hebrew word pairs with sentence-level annotations, to evaluate LLMs' ability to distinguish true cognates from false friends and loanwords. Results show high accuracy on true cognates but sharp drops on false friends, highlighting a key limitation in cross-lingual semantic reasoning.

0 favorites 0 likes
#arabic

ArabiGEE: A Hierarchical Taxonomy for Arabic Grammatical Error Explanation

arXiv cs.CL · 2026-06-10 Cached

Introduces ArabiGEE, the first comprehensive Arabic grammatical error explanation taxonomy with a hierarchical structure spanning orthographic, morphological, syntactic, and lexical dimensions, comprising 27 error types, 140 correction types, and 324 explanations.

0 favorites 0 likes
#arabic

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

Hugging Face Daily Papers · 2026-06-04 Cached

BloomBench is a cognitively grounded bilingual (English-Arabic) multimodal benchmark for Vision-Language Models, systematically evaluating six cognitive levels based on Bloom's Taxonomy. Experiments reveal significant cognitive asymmetries and cross-lingual performance gaps in current models.

0 favorites 0 likes
#arabic

AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse

arXiv cs.CL · 2026-05-25 Cached

This paper introduces AraHopeCorpus, the first annotated dataset of hope speech in Arabic social media, collected from YouTube comments about the war on Gaza. It provides a detailed annotation framework and analysis, showing that hopeful language dominates crisis discourse.

0 favorites 0 likes
#arabic

Pattern-and-root inflectional morphology: the Arabic broken plural

arXiv cs.CL · 2026-05-22 Cached

Presents a novel pattern-and-root model for describing Arabic noun inflection, focusing on broken plurals, with a taxonomy of 160 classes and an encoding scheme applied to 3,200 entries, aiming to improve computational language resources.

0 favorites 0 likes
#arabic

ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination

arXiv cs.CL · 2026-05-22 Cached

ArabDiscrim is a decade-long lexical resource and corpus of 293K Arabic Facebook posts about racism and discrimination, with engagement signals, morphological regex families, and discrimination axes, supporting fairness-oriented Arabic NLP research.

0 favorites 0 likes
← Back to home

Submit Feedback