Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
Summary
Introduces Vividh-ASR, a complexity-tiered benchmark for Hindi and Malayalam ASR, identifies studio-bias in fine-tuning, and proposes R-MFT to improve spontaneous speech performance efficiently.
View Cached Full Text
Cached at: 05/14/26, 08:17 AM
Paper page - Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
Source: https://huggingface.co/papers/2605.13087
Abstract
Research identifies studio-bias in multilingual ASR fine-tuning and proposes R-MFT method to improve spontaneous speech performance while maintaining efficiency.
Fine-tuningmultilingual ASR models likeWhisperforlow-resource languagesoften improves read speech but degrades spontaneous audio performance, a phenomenon we termstudio-bias. To diagnose this mismatch, we introduceVividh-ASR, a complexity-stratified benchmark for Hindi and Malayalam across four tiers: studio, broadcast, spontaneous, and synthetic noise. Through a controlled study of learning-rate timing and curriculum ordering, we find that early large parameter updates improve global WER by 12 absolute points, while a hard-to-easy curriculum adds gains for spontaneous speech. These findings motivate reverse multi-stagefine-tuning(R-MFT), a training recipe that enables a parameter-efficient 244MWhispermodel to match or exceed conventionally fine-tuned 769M counterparts. Representational analysis viaCKAandSVDreveals effective schedules concentrate adaptation in the decoder, preserving the pre-trained encoder’s acoustic geometry. We release the benchmark and models.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2605\.13087
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.13087 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.13087 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.13087 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
Researchers introduce Voice of India, a 536-hour closed benchmark of unscripted telephonic conversations across 15 Indian languages and 139 regional clusters, exposing geographic and demographic ASR performance disparities.
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.
SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR
SCRIBE is a diagnostic evaluation framework for automatic speech recognition that provides categorical error decomposition for Indic languages, releasing benchmarks and open-weight rich transcription models for Hindi, Malayalam, and Kannada.
Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion
This paper proposes a multimodal framework that jointly improves Automatic Speech Recognition (ASR) and Dialect Identification (DID) for Indian languages, using a Bottleneck Encoder and RoBERTa with a gating mechanism. Evaluated on eight languages with 33 dialects, it achieves 81.63% DID accuracy and reduces CER/WER to 4.65%/17.73%.
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages
This paper audits multilingual clinical ASR systems on psychiatric interviews in Indian languages and proposes SamaVaani, a unified debiasing technique to improve performance and fairness across demographic groups.