arabic

Tag

Cards List
#arabic

ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts

arXiv cs.CL ↗ · 2d ago Cached

ArGuard is a shared task focused on detecting harmful content in Arabic memes and LLM prompts, highlighting challenges in fine-grained classification and releasing datasets for further research.

0 favorites 0 likes
#arabic

MudawanSn: A Gold-Standard Wolof-Arabic Parallel Corpus for Machine Translation

arXiv cs.CL ↗ · 2026-09-17 Cached

This paper introduces MudawanSn, a gold-standard parallel corpus for Wolof-Arabic machine translation, and shows that fine-tuning on it yields substantial improvements in translation quality for both language directions.

0 favorites 0 likes
#arabic

YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

arXiv cs.CL ↗ · 2026-09-10 Cached

YallaMorph is a large-scale benchmark for evaluating controlled Arabic morphological generation in large language models, revealing significant challenges with cliticized and morphologically rare forms.

0 favorites 0 likes
#arabic

A Personal Computer For Children Of All Cultures

Lobsters Hottest ↗ · 2026-08-20 Cached

Ramsey Nasser's talk addresses the English-centric nature of programming languages and introduces an Arabic-based programming language to promote inclusivity in computing education.

0 favorites 0 likes
#arabic

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

arXiv cs.CL ↗ · 2026-08-20 Cached

This research investigates whether fine-tuning large language models on cultural data improves figurative language understanding and vice versa, finding that while poetry fine-tuning enhances idiom comprehension, cultural fine-tuning can reduce proverb accuracy, highlighting a non-straightforward relationship.

0 favorites 0 likes
#arabic

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

arXiv cs.CL ↗ · 2026-08-04 Cached

This paper investigates why LLMs underperform in Arabic medical tasks, showing via mechanistic analysis that knowledge exists internally but fails to surface, then proposes TLoRA, a targeted low-rank adaptation method that outperforms full-network LoRA on medical QA and introduces a new Arabic clinical dialogue benchmark.

0 favorites 0 likes
#arabic

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper introduces AHA-Memes, the first large-scale Arabic hateful meme benchmark with fine-grained multi-label annotations, covering 5K manually annotated and ~66K silver-labeled memes, and benchmarks various multimodal models for culturally grounded hate detection.

0 favorites 0 likes
#arabic

Constrained CTC Decoding for Efficient Diacritic Restoration

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.

0 favorites 0 likes
#arabic

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

arXiv cs.CL ↗ · 2026-07-07 Cached

This paper investigates methods to steer Arabic LLMs toward dialect-specific generation by identifying sparse neuron populations and extracting dialect activation directions, enabling dialect control at inference time without fine-tuning.

0 favorites 0 likes
#arabic

I swapped the TTS in my voice agent and it cut the lag people actually feel more than anything else

Reddit r/AI_Agents ↗ · 2026-07-06

The author shares their experience swapping the TTS in their voice agent to a custom model (Banter 1) designed for bilingual Arabic-English conversations, which significantly reduced perceived lag.

0 favorites 0 likes
#arabic

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

Hugging Face Daily Papers ↗ · 2026-07-06 Cached

PAST-TIDE is a stance detection system for the StanceNakba Shared Task, using statement tuning with cloze-style masked language modeling, prototypical contrastive learning, and topic-conditional layer normalization for cross-topic Arabic stance detection, achieving macro-F1 scores of 0.75 and 0.74 on subtasks A and B.

0 favorites 0 likes
#arabic

Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study

arXiv cs.CL ↗ · 2026-07-02 Cached

Introduces a comprehensive hate speech dataset for Turkish and Arabic, and develops state-of-the-art BERT-based models for hate speech analysis including classification, intensity prediction, target identification, and span detection.

0 favorites 0 likes
#arabic

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

arXiv cs.CL ↗ · 2026-07-02 Cached

This paper introduces a cross-evaluation framework for benchmarking LLMs on Arabic cultural and sociolinguistic knowledge, using human SME ground truth and automated judges. The authors contribute a dataset of prompt-rubric pairs for Egyptian and Iraqi Arabic, evaluating frontier LLMs and finding that cultural reasoning remains a primary failure mode for automated grading.

0 favorites 0 likes
#arabic

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

arXiv cs.CL ↗ · 2026-07-01 Cached

This paper presents a benchmark for Arabic-Russian scientific translation, including a hybrid parallel corpus of 27,000 sentence pairs and fine-tuned multilingual models (mT5, NLLB, Qwen) using LoRA. The best model achieves BLEU 23.15, and the work aims to lower language barriers for scientific knowledge exchange between Arabic and Russian researchers.

0 favorites 0 likes
#arabic

CohereLabs/cohere-transcribe-arabic-07-2026

Hugging Face Models Trending ↗ · 2026-06-18 Cached

Cohere releases an open-source 2B parameter Arabic ASR model optimized for Arabic dialect performance and Arabic-English code-switching, based on the Conformer encoder-decoder architecture.

0 favorites 0 likes
#arabic

When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper introduces SemCog Bench, a curated benchmark of 1,858 Arabic-Hebrew word pairs with sentence-level annotations, to evaluate LLMs' ability to distinguish true cognates from false friends and loanwords. Results show high accuracy on true cognates but sharp drops on false friends, highlighting a key limitation in cross-lingual semantic reasoning.

0 favorites 0 likes
#arabic

ArabiGEE: A Hierarchical Taxonomy for Arabic Grammatical Error Explanation

arXiv cs.CL ↗ · 2026-06-10 Cached

Introduces ArabiGEE, the first comprehensive Arabic grammatical error explanation taxonomy with a hierarchical structure spanning orthographic, morphological, syntactic, and lexical dimensions, comprising 27 error types, 140 correction types, and 324 explanations.

0 favorites 0 likes
#arabic

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

Hugging Face Daily Papers ↗ · 2026-06-04 Cached

BloomBench is a cognitively grounded bilingual (English-Arabic) multimodal benchmark for Vision-Language Models, systematically evaluating six cognitive levels based on Bloom's Taxonomy. Experiments reveal significant cognitive asymmetries and cross-lingual performance gaps in current models.

0 favorites 0 likes
#arabic

AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse

arXiv cs.CL ↗ · 2026-05-25 Cached

This paper introduces AraHopeCorpus, the first annotated dataset of hope speech in Arabic social media, collected from YouTube comments about the war on Gaza. It provides a detailed annotation framework and analysis, showing that hopeful language dominates crisis discourse.

0 favorites 0 likes
#arabic

Pattern-and-root inflectional morphology: the Arabic broken plural

arXiv cs.CL ↗ · 2026-05-22 Cached

Presents a novel pattern-and-root model for describing Arabic noun inflection, focusing on broken plurals, with a taxonomy of 160 classes and an encoding scheme applied to 3,200 entries, aiming to improve computational language resources.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback