data-balancing

Tag

Cards List
#data-balancing

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Hugging Face Daily Papers · 2026-07-11 Cached

This paper introduces GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio with a HuBERT-style objective, addressing data scarcity for Central Asian languages (Kazakh, Kyrgyz, Uzbek). It employs cluster-level data balancing and domain-aware sampling to outperform strong baselines like Whisper Large v3 on target languages.

0 favorites 0 likes
← Back to home

Submit Feedback