low-resource

Tag

Cards List
#low-resource

From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition

arXiv cs.CL · 2026-07-08 Cached

This research investigates cross-lingual transfer learning from Sinhala to Dhivehi for automatic speech recognition, achieving significant improvements in word error rate compared to Dhivehi-only baselines.

0 favorites 0 likes
#low-resource

CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script

arXiv cs.CL · 2026-07-08 Cached

This paper proposes CoPiT, a cognitively motivated pivot-based translation pipeline for digraphic Mongolian that routes translation through the better-resourced Cyrillic script to improve translation from the low-resource Traditional script, achieving significant BLEU and COMET gains and releasing a new multi-script parallel dataset.

0 favorites 0 likes
#low-resource

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

arXiv cs.CL · 2026-07-07 Cached

This paper proposes a multimodal framework that jointly improves Automatic Speech Recognition (ASR) and Dialect Identification (DID) for Indian languages, using a Bottleneck Encoder and RoBERTa with a gating mechanism. Evaluated on eight languages with 33 dialects, it achieves 81.63% DID accuracy and reduces CER/WER to 4.65%/17.73%.

0 favorites 0 likes
#low-resource

Small AI Models Gain Traction In places with unreliable networks

Hacker News Top · 2026-07-06 Cached

Small AI models are proving valuable in regions with unreliable networks, enabling life-saving applications like counterfeit drug detection and disease identification in crops without needing constant internet connectivity.

0 favorites 0 likes
#low-resource

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

Hugging Face Daily Papers · 2026-07-06 Cached

PAST-TIDE is a stance detection system for the StanceNakba Shared Task, using statement tuning with cloze-style masked language modeling, prototypical contrastive learning, and topic-conditional layer normalization for cross-topic Arabic stance detection, achieving macro-F1 scores of 0.75 and 0.74 on subtasks A and B.

0 favorites 0 likes
#low-resource

I built an open, from-scratch MT pipeline + parallel corpus for Tunisian Darija (Arabizi) early baseline, and I'm growing it into a curated community corpus [P]

Reddit r/MachineLearning · 2026-07-05

An 18-year-old Tunisian student introduces an open-source machine translation pipeline and parallel corpus for Tunisian Darija in Arabizi script, built from scratch with a small 15.6M-parameter Transformer and an honest baseline BLEU of 3.89, and calls for contributors to ethically expand the corpus.

0 favorites 0 likes
#low-resource

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

arXiv cs.CL · 2026-07-03 Cached

This paper analyzes the use of LLM-as-a-Judge in multilingual and low-resource settings, finding inconsistent evaluation outcomes and overtrust in LLM judgments, and provides recommendations for better practices.

0 favorites 0 likes
#low-resource

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

arXiv cs.CL · 2026-07-03 Cached

SPARCLE is a speaker-aware grapheme representation model that uses contrastive learning to align grapheme embeddings with acoustic representations, improving text-to-speech quality especially in low-resource settings.

0 favorites 0 likes
#low-resource

Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian

arXiv cs.CL · 2026-07-01 Cached

This paper investigates cross-lingual relation extraction for Romanian by translating the SemEval-2010 Task 8 benchmark and evaluating Gemma 4 under zero-shot, few-shot, and QLoRA fine-tuning, comparing with smaller encoder baselines.

0 favorites 0 likes
#low-resource

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

arXiv cs.CL · 2026-07-01 Cached

This paper proposes a tone-conditioned curriculum learning framework for low-resource Bantu speech recognition, combining hybrid difficulty scoring, gated adapters, and staged curriculum training. Evaluations on six Southern Bantu languages show that W2V-BERT outperforms Whisper on Nguni languages while Whisper performs better on Sotho-Tswana languages.

0 favorites 0 likes
#low-resource

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

arXiv cs.CL · 2026-06-30 Cached

This paper introduces sinhala-ocr-lk-acts-1010, the first publicly available real-world page-level dataset for Sinhala OCR, and fine-tunes three vision language models (DeepSeek-OCR V1, DeepSeek-OCR V2, LightOnOCR-2-1B) using QLoRA. LightOnOCR-2-1B achieves a CER of 1.05%, outperforming both open-source and commercial OCR models, and maintains consistent performance across degraded documents from different time periods.

0 favorites 0 likes
#low-resource

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

arXiv cs.CL · 2026-06-30 Cached

This paper investigates the distributional gap between synthetic and real speech in LLM-based ASR systems, identifies where the LLM separates them, and proposes using layer-selection and RIR augmentation to match real-data baselines with less real data.

0 favorites 0 likes
#low-resource

DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums

arXiv cs.CL · 2026-06-29 Cached

This paper proposes DysLexLens, a low-resource LLM framework for analyzing dyslexic learners' experiences with AI tools using online forum data, featuring dictionary-driven filtering, knowledge-graph reasoning, and evaluation metrics.

0 favorites 0 likes
#low-resource

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

arXiv cs.CL · 2026-06-26 Cached

This paper investigates LoRA fine-tuning of the VoxCPM2 TTS model to improve quality for low-resource languages like Khmer, while showing no gain for Korean which the base model already handles well. The adapter yields significant MOS improvement for Khmer with minimal parameter training.

0 favorites 0 likes
#low-resource

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

arXiv cs.CL · 2026-06-26 Cached

This paper presents NEST-V1, a proof-of-concept multimodal framework for generating emotion-conditioned Nepali Sign Language avatars from spoken input, achieving 81.1% ASR accuracy and 79.21% emotion recognition accuracy on a dataset of 600 audio samples from 50 speakers.

0 favorites 0 likes
#low-resource

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

arXiv cs.CL · 2026-06-25 Cached

This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.

0 favorites 0 likes
#low-resource

Neural Machine Translation for Low-Resource Tangkhul--English

arXiv cs.CL · 2026-06-25 Cached

Presents a neural machine translation system for the severely under-resourced Tangkhul–English language pair, achieving strong BLEU, chrF++, BERTScore, and COMET scores using fine-tuned ByT5-large and mT5-small models.

0 favorites 0 likes
#low-resource

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

arXiv cs.CL · 2026-06-18 Cached

The paper proposes a novel framework (CDDTLDA) using transfer learning and data augmentation to improve Chinese dialects discrimination under low-resource conditions, achieving state-of-the-art results on two benchmark corpora.

0 favorites 0 likes
#low-resource

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

arXiv cs.CL · 2026-06-17 Cached

This study evaluates bilingual fine-tuning with language identification tokens for improving ASR in low-resource languages across nine diverse language pairs, finding that high LID accuracy is beneficial and that providing the LID token at inference can boost performance when LID accuracy is low.

0 favorites 0 likes
#low-resource

Distilling Examples into Task Instructions: Enhanced In-Context Learning for Real-World B2B Conversations

arXiv cs.CL · 2026-06-16 Cached

This paper introduces the Call Playbook dataset for classifying real-world B2B conversations and proposes methods to distill examples into compact, interpretable task instructions, achieving 99% token reduction and up to 7% AUC improvement over traditional in-context learning.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback