low-resource-language

Tag

Cards List
#low-resource-language

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

arXiv cs.CL · 2026-08-03 Cached

A two-dialect finite-state morphological analyzer for the Dungan language is presented, with a multi-genre evaluation measuring inflection, ambiguity, and lexical coverage.

0 favorites 0 likes
#low-resource-language

Mwando: Leveraging AI to Preserve and Teach shiKomori

arXiv cs.CL · 2026-07-28 Cached

Mwando is a virtual educational assistant leveraging AI to preserve and teach the Comorian language shiKomori, utilizing a multi-agent architecture with vector search, knowledge graph, and web fallback.

0 favorites 0 likes
#low-resource-language

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

arXiv cs.CL · 2026-07-28 Cached

Introduces PatiGonit22K, an expanded Bengali mathematical word problem dataset with 22,441 problems, including complex multi-operation problems, to advance mathematical reasoning research for low-resource languages.

0 favorites 0 likes
#low-resource-language

KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

arXiv cs.CL · 2026-07-21 Cached

This paper introduces KyrgyzLLM-Bench, a benchmark suite for evaluating large language models in the Kyrgyz language, comprising both natively authored and translated datasets, and provides a systematic evaluation of 26 models.

0 favorites 0 likes
#low-resource-language

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

arXiv cs.CL · 2026-07-08 Cached

Introduces BaFCo, a benchmark dataset for Bangla form comprehension focusing on Document Layout Analysis (DLA) and Key Information Extraction (KIE). It includes 200 multi-page complex Bangladeshi government forms with fine-grained annotations across 26 entity types and evaluates multiple MLLMs, revealing limitations in understanding complex Bangla forms.

0 favorites 0 likes
#low-resource-language

BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes

arXiv cs.CL · 2026-07-07 Cached

This paper introduces BanglaMemeEvidence, a multimodal dataset of 2,917 Bengali memes annotated for explanatory evidence detection, and proposes BengaliMemeEvidenceNet, a hybrid framework achieving an F1 score of 0.74.

0 favorites 0 likes
#low-resource-language

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

arXiv cs.CL · 2026-07-07 Cached

This paper investigates using text-to-speech (TTS) to generate synthetic training data for spoken question answering in Luxembourgish, a low-resource language, and evaluates multi-source TTS configurations with a parameter-efficient SLAM-style architecture.

0 favorites 0 likes
#low-resource-language

Building an ASR Solution for Training and Assessing Children's Reading

arXiv cs.CL · 2026-07-01 Cached

Presents an open-source ASR system for assessing children's reading in Bambara, including field data collection, benchmark construction, model adaptation, and classroom validation, achieving significant word error rate reduction.

0 favorites 0 likes
#low-resource-language

Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text

arXiv cs.CL · 2026-07-01 Cached

This paper introduces a Bangla event detection benchmark with noisy text (ASR, orthographic corruption) and evaluates encoder-only and decoder-only LLMs, finding decoder models more robust to noise.

0 favorites 0 likes
#low-resource-language

Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning

arXiv cs.CL · 2026-06-25 Cached

Riazi-8B is an Urdu large language model fine-tuned for mathematical reasoning, achieving improved performance on MGSM-Urdu through continued pre-training and supervised fine-tuning on Urdu Chain-of-Thought data.

0 favorites 0 likes
#low-resource-language

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

arXiv cs.CL · 2026-06-24 Cached

This paper introduces the first public multimodal dataset of 100 Turkish scam and benign phone calls, evaluating seven LLMs under raw audio, ASR transcripts, and human-corrected transcripts. Results show transcript-based inputs outperform direct audio, highlighting the need for inclusive AI safety research in low-resource languages.

0 favorites 0 likes
#low-resource-language

An End-to-End Hybrid Framework for Rumour Detection in Low-Resources Algerian Dialect

arXiv cs.CL · 2026-06-12 Cached

This paper presents an end-to-end hybrid framework for rumour detection in low-resource Algerian dialect social media content, achieving an F1-score of 0.84 by combining transformer embeddings with a classical classifier.

0 favorites 0 likes
#low-resource-language

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

arXiv cs.CL · 2026-06-08 Cached

UrduMMLU is a new benchmark of 26,431 multiple-choice questions across 26 subjects for evaluating LLMs on Urdu language understanding, sourced from native educational materials. Evaluation of 30 LLMs reveals Gemini-3.5-Flash performs best, while open-source models and region-specific subjects pose significant challenges.

0 favorites 0 likes
#low-resource-language

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

arXiv cs.CL · 2026-05-22 Cached

This paper evaluates four text chunking strategies for Retrieval-Augmented Generation on Khmer agricultural documents, finding that character-based Recursive chunking with 300 characters yields the best retrieval and relevance performance.

0 favorites 0 likes
#low-resource-language

Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

arXiv cs.CL · 2026-05-15 Cached

This paper proposes using reinforcement learning with semantic rewards (via GRPO) to expand LLMs to low-resource languages without the typical alignment tax of catastrophic forgetting, showing improved semantic quality and transferability over supervised fine-tuning.

0 favorites 0 likes
#low-resource-language

Towards High-Quality Machine Translation for Kokborok: A Low-Resource Tibeto-Burman Language of Northeast India

arXiv cs.CL · 2026-04-23 Cached

Researchers develop KokborokMT, a neural MT system for the low-resource Kokborok language, achieving BLEU scores of 17.30 en→trp and 38.56 trp→en by fine-tuning NLLB-200 on a 36k-sentence parallel corpus.

0 favorites 0 likes
#low-resource-language

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models

arXiv cs.CL · 2026-04-20 Cached

VLegal-Bench is a cognitively grounded benchmark for evaluating large language models on Vietnamese legal reasoning tasks, containing 10,450 expert-annotated samples designed to address the gap in legal benchmarks for civil law systems. The benchmark assesses multiple levels of legal understanding through question answering, multi-step reasoning, and scenario-based problem solving, providing a replicable framework for evaluating LLMs in non-English, codified legal contexts.

0 favorites 0 likes
← Back to home

Submit Feedback