Tag
AraGenre 2026 is a shared task for hierarchical, definition-guided Arabic genre classification aimed at improving annotated data in low-resource languages, with results showing strong broad-genre recognition but gaps in fine-grained classification.
This paper introduces MudawanSn, a gold-standard parallel corpus for Wolof-Arabic machine translation, and shows that fine-tuning on it yields substantial improvements in translation quality for both language directions.
NepKANUN is an AI-powered legal assistant for Nepali using RAG and a fine-tuned LLaMA 3.2 3B model to provide accurate answers to legal queries, validated by expert reviews.
This paper proposes an end-to-end sequence-to-sequence approach for contextual Tamil spelling and grammar correction, using progressively fine-tuned mT5 and mBART models on synthetic data, achieving 69.3% exact-match accuracy on a diagnostic set.
This paper studies supervised fine-tuning and reinforcement learning for reasoning in low-resource languages, revealing that accuracy benchmarks are noisy while SFT builds language-specific reasoning and RL fixes format and leakage issues.
This paper investigates fine-tuning MoE models to reason in Greek, revealing that accuracy metrics are noisy, while supervised fine-tuning and reinforcement learning improve language-specific reasoning and fix behavioral defects, with proposed evaluation instruments.
BengaliMCQ is a structure-aware RAG framework using graph neural networks to model hierarchical document structures in Bengali textbooks, enabling automatic generation and answer prediction of academic multiple-choice questions with improved performance over baseline methods.
Introduces BavGround, a benchmark for evaluating LLMs' regional cultural grounding and dialect competence in Bavarian across English, German, and Bavarian, finding that models struggle with dialectal and localized cultural knowledge.
A two-dialect finite-state morphological analyzer for the Dungan language is presented, with a multi-genre evaluation measuring inflection, ambiguity, and lexical coverage.
Mwando is a virtual educational assistant leveraging AI to preserve and teach the Comorian language shiKomori, utilizing a multi-agent architecture with vector search, knowledge graph, and web fallback.
Introduces PatiGonit22K, an expanded Bengali mathematical word problem dataset with 22,441 problems, including complex multi-operation problems, to advance mathematical reasoning research for low-resource languages.
This paper introduces KyrgyzLLM-Bench, a benchmark suite for evaluating large language models in the Kyrgyz language, comprising both natively authored and translated datasets, and provides a systematic evaluation of 26 models.
Introduces BaFCo, a benchmark dataset for Bangla form comprehension focusing on Document Layout Analysis (DLA) and Key Information Extraction (KIE). It includes 200 multi-page complex Bangladeshi government forms with fine-grained annotations across 26 entity types and evaluates multiple MLLMs, revealing limitations in understanding complex Bangla forms.
This paper introduces BanglaMemeEvidence, a multimodal dataset of 2,917 Bengali memes annotated for explanatory evidence detection, and proposes BengaliMemeEvidenceNet, a hybrid framework achieving an F1 score of 0.74.
This paper investigates using text-to-speech (TTS) to generate synthetic training data for spoken question answering in Luxembourgish, a low-resource language, and evaluates multi-source TTS configurations with a parameter-efficient SLAM-style architecture.
Presents an open-source ASR system for assessing children's reading in Bambara, including field data collection, benchmark construction, model adaptation, and classroom validation, achieving significant word error rate reduction.
This paper introduces a Bangla event detection benchmark with noisy text (ASR, orthographic corruption) and evaluates encoder-only and decoder-only LLMs, finding decoder models more robust to noise.
Riazi-8B is an Urdu large language model fine-tuned for mathematical reasoning, achieving improved performance on MGSM-Urdu through continued pre-training and supervised fine-tuning on Urdu Chain-of-Thought data.
This paper introduces the first public multimodal dataset of 100 Turkish scam and benign phone calls, evaluating seven LLMs under raw audio, ASR transcripts, and human-corrected transcripts. Results show transcript-based inputs outperform direct audio, highlighting the need for inclusive AI safety research in low-resource languages.
This paper presents an end-to-end hybrid framework for rumour detection in low-resource Algerian dialect social media content, achieving an F1-score of 0.84 by combining transformer embeddings with a classical classifier.