low-resource

Tag

Cards List
#low-resource

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

arXiv cs.CL · 2d ago Cached

This paper presents a method to build a compact fixed-voice Thai TTS system using synthetic speech from a larger model, evaluating its performance and introducing an 82M-parameter model for on-device deployment.

0 favorites 0 likes
#low-resource

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

arXiv cs.CL · 3d ago Cached

This paper introduces SpeakPay and a Nepali financial speech dataset, showing that LoRA fine-tuning of Whisper reduces Word Error Rate by 67.2% and improves transaction success rates for low-resource language accessibility.

0 favorites 0 likes
#low-resource

Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training

arXiv cs.LG · 4d ago Cached

The paper introduces TPGC, a dual-prior prompt initialization method for multi-task graph pre-training that combines task and structural priors to improve alignment and transferability, achieving superior performance in few-shot scenarios.

0 favorites 0 likes
#low-resource

Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study

arXiv cs.CL · 5d ago Cached

This paper proposes a cross-lingual romanization ecosystem for Sinitic languages, develops specific schemes for Mandarin and Cantonese, and shows improved performance in speech-to-romanization tasks compared to baseline methods.

0 favorites 0 likes
#low-resource

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems

Hugging Face Daily Papers · 2026-08-28 Cached

This paper presents an empirical study on bootstrapping conversational recommender systems using synthetic data generated from non-conversational signals, demonstrating that it outperforms zero-shot and scarce real-data methods in low-resource settings.

0 favorites 0 likes
#low-resource

L\"etzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

arXiv cs.CL · 2026-08-25 Cached

LëtzCross is a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, comparing text-only and multimodal retrievers in low-resource settings.

0 favorites 0 likes
#low-resource

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

arXiv cs.CL · 2026-08-21 Cached

This study develops an Automatic Speech Recognition system for Mizo, a low-resource language, by fine-tuning Whisper and SraVaani 1.0 models, achieving a morphology-aware WER of 7.22% with Whisper-large-v3.

0 favorites 0 likes
#low-resource

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

arXiv cs.CL · 2026-08-20 Cached

The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.

0 favorites 0 likes
#low-resource

HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA

arXiv cs.CL · 2026-08-14 Cached

This paper proposes HybridRAG-BN, a retrieval-augmented framework for Bangla knowledge-base question answering that combines hybrid retrieval, Gemma-based generation, and LoRA fine-tuned verification, achieving first place with F1 scores of 0.71654 and 0.72912.

0 favorites 0 likes
#low-resource

Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects

arXiv cs.CL · 2026-08-13 Cached

This paper presents a unified poly-dialectal neural machine translation system for 12 Bangla regional dialects, introducing the largest multi-dialect parallel corpus to date and achieving state-of-the-art BLEU scores with a fine-tuned BanglaT5 model using DoRA.

0 favorites 0 likes
#low-resource

When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation

arXiv cs.CL · 2026-08-13 Cached

This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.

0 favorites 0 likes
#low-resource

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

arXiv cs.CL · 2026-08-12 Cached

This paper argues that single-run evaluations in low-resource ASR are unreliable and demonstrates with a new multi-seed Garhwali ASR benchmark that many reported gains vanish under seed-level testing, while standard CTC with w2v-BERT 2.0 remains the most robust approach.

0 favorites 0 likes
#low-resource

North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings

arXiv cs.CL · 2026-08-11 Cached

This paper presents a conceptual framework for using NLP to address mental healthcare barriers in Algeria and other low-resource, multilingual settings, proposing a research and policy roadmap.

0 favorites 0 likes
#low-resource

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

arXiv cs.CL · 2026-08-11 Cached

DialectS2S is an end-to-end speech dialogue model for low-resource Chinese dialects, introducing a scalable data synthesis pipeline and a two-stage post-training strategy with self-aligned speech supervision. Experiments show improvements in dialect consistency, response quality, and intelligibility, with fully open-sourced models, data, and code.

0 favorites 0 likes
#low-resource

Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

arXiv cs.CL · 2026-08-11 Cached

This paper trains five GPT-2-style models from scratch to compare dedicated monolingual models for Tamil, Telugu, Kannada, and Malayalam against a joint multilingual model, finding monolingual models outperform mGPT on sentiment classification and NER with more efficient tokenizers.

0 favorites 0 likes
#low-resource

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

arXiv cs.CL · 2026-08-07 Cached

This paper introduces MameLoshnLM, the first open-source 8B-parameter Yiddish language model, along with the Oytser pretraining corpus and Kashes evaluation benchmark. It demonstrates that continued pretraining on high-quality Yiddish data outperforms general multilingual models, highlighting the value of dedicated low-resource language modeling.

0 favorites 0 likes
#low-resource

Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders

arXiv cs.CL · 2026-08-06 Cached

This paper introduces MSRT, a framework with a resource-aware Mixture of Speech Encoders (MoSE) to overcome the curse of multilinguality in many-to-many speech-to-text translation. The 4B-parameter model achieves state-of-the-art results across 45 languages, particularly improving low-resource speech translation with only 10 hours of paired data per language.

0 favorites 0 likes
#low-resource

HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?

arXiv cs.CL · 2026-08-05 Cached

This paper proposes HomoEnsNER, a homogeneous ensemble of five GujaratiBERT models for Gujarati named entity recognition, and shows it outperforms heterogeneous alternatives that rely on architectural diversity, achieving state-of-the-art F1 on the Naamapadam test split.

0 favorites 0 likes
#low-resource

Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

arXiv cs.CL · 2026-08-04 Cached

Introduces OSCD, a post-training algorithm to improve native multilingual chain-of-thought reasoning in low-resource Southeast Asian languages, achieving up to 3.2x improvements on math benchmarks.

0 favorites 0 likes
#low-resource

Cross-Lingual Transfer for Machine Translation in Turkic Languages

arXiv cs.CL · 2026-08-03 Cached

This paper studies cross-lingual transfer for machine translation among five Turkic languages using pairwise transfer matrices with mT5, finding that transfer is strongest between closely related pairs and that Latinization helps in script-mismatched settings.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback