Tag
This paper presents a method to build a compact fixed-voice Thai TTS system using synthetic speech from a larger model, evaluating its performance and introducing an 82M-parameter model for on-device deployment.
This paper introduces SpeakPay and a Nepali financial speech dataset, showing that LoRA fine-tuning of Whisper reduces Word Error Rate by 67.2% and improves transaction success rates for low-resource language accessibility.
The paper introduces TPGC, a dual-prior prompt initialization method for multi-task graph pre-training that combines task and structural priors to improve alignment and transferability, achieving superior performance in few-shot scenarios.
This paper proposes a cross-lingual romanization ecosystem for Sinitic languages, develops specific schemes for Mandarin and Cantonese, and shows improved performance in speech-to-romanization tasks compared to baseline methods.
This paper presents an empirical study on bootstrapping conversational recommender systems using synthetic data generated from non-conversational signals, demonstrating that it outperforms zero-shot and scarce real-data methods in low-resource settings.
LëtzCross is a benchmark for cross-lingual page-level retrieval over Luxembourgish PDF documents, comparing text-only and multimodal retrievers in low-resource settings.
This study develops an Automatic Speech Recognition system for Mizo, a low-resource language, by fine-tuning Whisper and SraVaani 1.0 models, achieving a morphology-aware WER of 7.22% with Whisper-large-v3.
The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.
This paper proposes HybridRAG-BN, a retrieval-augmented framework for Bangla knowledge-base question answering that combines hybrid retrieval, Gemma-based generation, and LoRA fine-tuned verification, achieving first place with F1 scores of 0.71654 and 0.72912.
This paper presents a unified poly-dialectal neural machine translation system for 12 Bangla regional dialects, introducing the largest multi-dialect parallel corpus to date and achieving state-of-the-art BLEU scores with a fine-tuned BanglaT5 model using DoRA.
This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.
This paper argues that single-run evaluations in low-resource ASR are unreliable and demonstrates with a new multi-seed Garhwali ASR benchmark that many reported gains vanish under seed-level testing, while standard CTC with w2v-BERT 2.0 remains the most robust approach.
This paper presents a conceptual framework for using NLP to address mental healthcare barriers in Algeria and other low-resource, multilingual settings, proposing a research and policy roadmap.
DialectS2S is an end-to-end speech dialogue model for low-resource Chinese dialects, introducing a scalable data synthesis pipeline and a two-stage post-training strategy with self-aligned speech supervision. Experiments show improvements in dialect consistency, response quality, and intelligibility, with fully open-sourced models, data, and code.
This paper trains five GPT-2-style models from scratch to compare dedicated monolingual models for Tamil, Telugu, Kannada, and Malayalam against a joint multilingual model, finding monolingual models outperform mGPT on sentiment classification and NER with more efficient tokenizers.
This paper introduces MameLoshnLM, the first open-source 8B-parameter Yiddish language model, along with the Oytser pretraining corpus and Kashes evaluation benchmark. It demonstrates that continued pretraining on high-quality Yiddish data outperforms general multilingual models, highlighting the value of dedicated low-resource language modeling.
This paper introduces MSRT, a framework with a resource-aware Mixture of Speech Encoders (MoSE) to overcome the curse of multilinguality in many-to-many speech-to-text translation. The 4B-parameter model achieves state-of-the-art results across 45 languages, particularly improving low-resource speech translation with only 10 hours of paired data per language.
This paper proposes HomoEnsNER, a homogeneous ensemble of five GujaratiBERT models for Gujarati named entity recognition, and shows it outperforms heterogeneous alternatives that rely on architectural diversity, achieving state-of-the-art F1 on the Naamapadam test split.
Introduces OSCD, a post-training algorithm to improve native multilingual chain-of-thought reasoning in low-resource Southeast Asian languages, achieving up to 3.2x improvements on math benchmarks.
This paper studies cross-lingual transfer for machine translation among five Turkic languages using pairwise transfer matrices with mT5, finding that transfer is strongest between closely related pairs and that Latinization helps in script-mismatched settings.