underrepresented-languages

Tag

Cards List
#underrepresented-languages

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Hugging Face Daily Papers · 2026-07-11 Cached

This paper introduces GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio with a HuBERT-style objective, addressing data scarcity for Central Asian languages (Kazakh, Kyrgyz, Uzbek). It employs cluster-level data balancing and domain-aware sampling to outperform strong baselines like Whisper Large v3 on target languages.

0 favorites 0 likes
#underrepresented-languages

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices

Papers with Code Trending · 2025-09-02 Cached

This paper presents Flavors of Moonshine, a suite of tiny specialized ASR models for edge devices. The authors show that monolingual models trained on a balanced mix of human-labeled, pseudo-labeled, and synthetic data outperform larger multilingual models like Whisper, achieving state-of-the-art error rates for small models and enabling on-device ASR for underrepresented languages.

0 favorites 0 likes
← Back to home

Submit Feedback