low-resource-language

Tag

Cards List
#low-resource-language

AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task

arXiv cs.CL ↗ · 6d ago Cached

AraGenre 2026 is a shared task for hierarchical, definition-guided Arabic genre classification aimed at improving annotated data in low-resource languages, with results showing strong broad-genre recognition but gaps in fine-grained classification.

0 favorites 0 likes
#low-resource-language

MudawanSn: A Gold-Standard Wolof-Arabic Parallel Corpus for Machine Translation

arXiv cs.CL ↗ · 2026-09-17 Cached

This paper introduces MudawanSn, a gold-standard parallel corpus for Wolof-Arabic machine translation, and shows that fine-tuning on it yields substantial improvements in translation quality for both language directions.

0 favorites 0 likes
#low-resource-language

NepKANUN: A RAG-Based Nepali Legal Assistant

arXiv cs.CL ↗ · 2026-09-16 Cached

NepKANUN is an AI-powered legal assistant for Nepali using RAG and a fine-tuned LLaMA 3.2 3B model to provide accurate answers to legal queries, validated by expert reviews.

0 favorites 0 likes
#low-resource-language

Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper proposes an end-to-end sequence-to-sequence approach for contextual Tamil spelling and grammar correction, using progressively fine-tuned mT5 and mBART models on synthetic data, achieving 69.3% exact-match accuracy on a diagnostic set.

0 favorites 0 likes
#low-resource-language

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper studies supervised fine-tuning and reinforcement learning for reasoning in low-resource languages, revealing that accuracy benchmarks are noisy while SFT builds language-specific reasoning and RL fixes format and leakage issues.

0 favorites 0 likes
#low-resource-language

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

Hugging Face Daily Papers ↗ · 2026-08-18 Cached

This paper investigates fine-tuning MoE models to reason in Greek, revealing that accuracy metrics are noisy, while supervised fine-tuning and reinforcement learning improve language-specific reasoning and fix behavioral defects, with proposed evaluation instruments.

0 favorites 0 likes
#low-resource-language

BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language

arXiv cs.CL ↗ · 2026-08-18 Cached

BengaliMCQ is a structure-aware RAG framework using graph neural networks to model hierarchical document structures in Bengali textbooks, enabling automatic generation and answer prediction of academic multiple-choice questions with improved performance over baseline methods.

0 favorites 0 likes
#low-resource-language

BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

arXiv cs.CL ↗ · 2026-08-14 Cached

Introduces BavGround, a benchmark for evaluating LLMs' regional cultural grounding and dialect competence in Bavarian across English, German, and Bavarian, finding that models struggle with dialectal and localized cultural knowledge.

0 favorites 0 likes
#low-resource-language

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

arXiv cs.CL ↗ · 2026-08-03 Cached

A two-dialect finite-state morphological analyzer for the Dungan language is presented, with a multi-genre evaluation measuring inflection, ambiguity, and lexical coverage.

0 favorites 0 likes
#low-resource-language

Mwando: Leveraging AI to Preserve and Teach shiKomori

arXiv cs.CL ↗ · 2026-07-28 Cached

Mwando is a virtual educational assistant leveraging AI to preserve and teach the Comorian language shiKomori, utilizing a multi-agent architecture with vector search, knowledge graph, and web fallback.

0 favorites 0 likes
#low-resource-language

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

arXiv cs.CL ↗ · 2026-07-28 Cached

Introduces PatiGonit22K, an expanded Bengali mathematical word problem dataset with 22,441 problems, including complex multi-operation problems, to advance mathematical reasoning research for low-resource languages.

0 favorites 0 likes
#low-resource-language

KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper introduces KyrgyzLLM-Bench, a benchmark suite for evaluating large language models in the Kyrgyz language, comprising both natively authored and translated datasets, and provides a systematic evaluation of 26 models.

0 favorites 0 likes
#low-resource-language

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

arXiv cs.CL ↗ · 2026-07-08 Cached

Introduces BaFCo, a benchmark dataset for Bangla form comprehension focusing on Document Layout Analysis (DLA) and Key Information Extraction (KIE). It includes 200 multi-page complex Bangladeshi government forms with fine-grained annotations across 26 entity types and evaluates multiple MLLMs, revealing limitations in understanding complex Bangla forms.

0 favorites 0 likes
#low-resource-language

BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes

arXiv cs.CL ↗ · 2026-07-07 Cached

This paper introduces BanglaMemeEvidence, a multimodal dataset of 2,917 Bengali memes annotated for explanatory evidence detection, and proposes BengaliMemeEvidenceNet, a hybrid framework achieving an F1 score of 0.74.

0 favorites 0 likes
#low-resource-language

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

arXiv cs.CL ↗ · 2026-07-07 Cached

This paper investigates using text-to-speech (TTS) to generate synthetic training data for spoken question answering in Luxembourgish, a low-resource language, and evaluates multi-source TTS configurations with a parameter-efficient SLAM-style architecture.

0 favorites 0 likes
#low-resource-language

Building an ASR Solution for Training and Assessing Children's Reading

arXiv cs.CL ↗ · 2026-07-01 Cached

Presents an open-source ASR system for assessing children's reading in Bambara, including field data collection, benchmark construction, model adaptation, and classroom validation, achieving significant word error rate reduction.

0 favorites 0 likes
#low-resource-language

Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text

arXiv cs.CL ↗ · 2026-07-01 Cached

This paper introduces a Bangla event detection benchmark with noisy text (ASR, orthographic corruption) and evaluates encoder-only and decoder-only LLMs, finding decoder models more robust to noise.

0 favorites 0 likes
#low-resource-language

Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning

arXiv cs.CL ↗ · 2026-06-25 Cached

Riazi-8B is an Urdu large language model fine-tuned for mathematical reasoning, achieving improved performance on MGSM-Urdu through continued pre-training and supervised fine-tuning on Urdu Chain-of-Thought data.

0 favorites 0 likes
#low-resource-language

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper introduces the first public multimodal dataset of 100 Turkish scam and benign phone calls, evaluating seven LLMs under raw audio, ASR transcripts, and human-corrected transcripts. Results show transcript-based inputs outperform direct audio, highlighting the need for inclusive AI safety research in low-resource languages.

0 favorites 0 likes
#low-resource-language

An End-to-End Hybrid Framework for Rumour Detection in Low-Resources Algerian Dialect

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper presents an end-to-end hybrid framework for rumour detection in low-resource Algerian dialect social media content, achieving an F1-score of 0.84 by combining transformer embeddings with a classical classifier.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback