SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

arXiv cs.CL Papers

Summary

This paper audits multilingual clinical ASR systems on psychiatric interviews in Indian languages and proposes SamaVaani, a unified debiasing technique to improve performance and fairness across demographic groups.

arXiv:2606.26901v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, we first conduct the systematic audit of ASR performance on real-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state-of-the-art models including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini. Our results reveal substantial variability across models and languages, with some systems performing competitively in Indian English but failing in regional speech. We further fine-tune two of the best performing opensource models, i.e., Gemma3n and OmniLingual, using various methods. With this, we uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings, which are further mitigated by fairness-aware fine-tuning. To this end, we propose SamaVaani, a unified debiasing technique that simultaneously improves ASR performance and improves fairness across demographic groups.
Original Article
View Cached Full Text

Cached at: 06/26/26, 05:20 AM

# Auditing and Debiasing Multilingual Clinical ASR for Indian Languages
Source: [https://arxiv.org/html/2606.26901](https://arxiv.org/html/2606.26901)
\\fontspec\_if\_language:nTF

ENG\\addfontfeatureLanguage=English

Subham Kumar†Prakrithi Shivaprakash‡Abhishek Manoharan‡Astut Kurariya‡ Diptadhi Mukherjee\*Prabhat Chand‡Pratima Murthy‡ Koustav Rudra†Lekhansh Shukla‡Animesh Mukherjee† †IIT, Kharagpur,‡NIMHANS, Bangalore,\*LGBRIMH, Tezpur \{kumarshubham209, prakrithishivaprakash, 12\.abhishek\.m, astutnamo, diptadhimukherjee\}@gmail\.com prabhat@vknnimhans\.in, \{pratimamurthy, krudra5, drlekhansh, animeshm\}@gmail\.com

###### Abstract

Automatic Speech Recognition \(ASR\) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown\. In this study, we first conduct the systematic audit of ASR performance on real\-world psychiatric interview data spanning Kannada, Hindi and Indian English, comparing eight state\-of\-the\-art models includingIndicWhisper,WhisperLargeV3,Sarvam,GoogleS2T,Gemma3n,OmniLingual,Vaani, andGemini\. Our results reveal substantial variability across models and languages, with some systems performing competitively in Indian English but failing in regional speech\. We further fine\-tune two of the best performing open\-source models, i\.e\.,Gemma3nandOmniLingual, using various methods\. With this, we uncover systematic performance gaps tied to speaker role and gender, raising concerns about equitable deployment in clinical settings, which are further mitigated by fairness\-aware fine\-tuning\. To this end, we proposeSamaVaani, a unified debiasing technique that simultaneously improves ASR performance and improves fairness across demographic groups\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1Introduction

The field of psychiatry is highly dependent on language, with a detailed psychiatric interview the primary diagnostic toolsommers2015clinicalrather than laboratory or radiological investigations\. Subsequently, verbatim transcripts of these interviews are widely used for clinical diagnosis, academic training, qualitative research, and, more recently, for the development of AI systems in psychiatric tasksso2024aligning\. However, producing accurate transcripts remains a major bottleneck: manual transcription is strenuous, time\-intensive \(5\-8 hours for every hour of audio\) and prone to human errormays2019measuring\. ASR systems provide a scalable alternative albeit transcription errorsbasma2011errorin psychiatric settings can significantly change clinical interpretationciampelli\_combining\_2023;shikino\_clinical\_2023\. Modern ASR systems, including proprietary platforms \(GoogleS2Tgoogle\_speech\_to\_text, Microsoft Azureazure\_speech\_to\_text, Amazon Transcribeamazon\_transcribe\) and open\-source models likeWhisperradford2022whisperhave improved transcription quality\. Although these systems work well for standard English \(American and British\) and in controlled environments, performance deteriorates substantially for non\-standard English, conversational and accented Englishrussell\_what\_2024, code\-mixing and code\-switching, which is common in multilingual contexts like Indiasitaram2020surveycodeswitchedspeechlanguage;koenecke2020racial\. In addition, Indian English and Indian regional languages remain underrepresented in global training corpora, leading to lower accuracy and greater variabilityjaved\_svarah\_2023;rai\_deep\_2024\. These difficulties are amplified with clinically distinctive speech patterns in psychiatric interviews\. For instance, low tone, hesitations, and long pauses are common in depressionalpert2001;yang2012;durao2025; fast and loud speech in maniakaczmarekmajer2024;anmella\_automated\_2024;diflorio2021; stammering and repeating in anxietysilber\_appropriate\_2018;harrigan1994disfluency;teferra2022; and disordered agrammatical speech containing neologisms \(made\-up words to which the patient attaches special meaning\) in schizophreniacovington2005;stokes2023;hinzen2015\. Beyond this, recordings are often made under acoustically difficult conditions \(in wards with noise from ceiling fans, hospital instruments, or ambulance sirens\)baroudi2026doctor\. Another big concern is the fairness and accuracy of the transcription\. These problems may be magnified in psychiatric interviews, as doctors and patients are very different in terms of education, socio\-economic background, conversational role, and speaking stylekoenecke2020racial;tatman\_gender\_2017\. While recent advances in multilingual ASR from open\-source models such asWhisperradford2022whisper,IndicWhisperindicwhisper2024,OmniLingualomnilingualasrteam2025omnilingualasropensourcemultilingual,Vaanivaani\_whisper\_hi\_2025and commercial models likeSarvamsarvam2025have improved support for multilingual and regional speech in India, existing evaluations rely mainly on general\-purpose datasets and do not capture the complexities of real\-world psychiatric interviews\. Our work addresses this gap by systematically analyzing and further improving the performance of ASR on multilingual psychiatric interactions\. Our contributions and findings: In this work, we present the first systematic audit of ASR systems on real\-world multilingual psychiatric interviews in the Indian context\. Leveraging a novel dataset spanning Kannada, Hindi, and Indian English, we evaluate eight state\-of\-the\-art ASR models across languages, speaker roles, and demographic groups, and introduce a comprehensive fairness analysis grounded in WER and fine\-grained error patterns\. We find substantial variability in performance across models and languages, with consistently higher error rates for low\-resource languages such as Kannada, as well as systematic disparities across speaker roles and gender\. Therefore, we proposeSamaVaani, a simple yet effective fairness\-aware fine\-tuning framework combining contrastive learning and CTC alignment, which significantly improves both transcription accuracy \(up to∼50%\\sim 50\\%WER reduction\) and fairness across demographic groups\. Together, our study highlights critical limitations of current ASR systems in clinical settings and provides actionable pathways toward more equitable and robust multilingual ASR deployment in healthcare\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2Related work

Research on ASR in Indian multilingual psychiatric settings spansthreeinterconnected areas: \(i\) ASR for psychiatric interviews, \(ii\) Bias and speaker\-level disparities in ASR, and \(iii\) ASR for Indian languages and accents\. We highlight the challenges and gaps in each area that motivate the current study\. ASR for psychiatric interviews: A growing set of studies has examined the use of ASR to transcribe or analyze psychiatric interviews\.ciampelli\_combining\_2023andjust\_moving\_2025evaluated ASR in schizophrenia patients, comparing them with healthy controls using interviews conducted in Dutch and German respectively\. Several studiesmolina2025automatic;hwang2025evaluatingin Spanish and French psychiatric interviews in patients with psychosis has found high WER performance\. Lastly,seyedi2023usingstudied ASR in psychiatric interviews in American English in patients with depression vs\. controls with past history of depression, and found no difference in WER in either group\. Bias and speaker\-level disparities in ASR: Prior works has identified gender differences in YouTube ASRtatman\_gender\_2017, higher error rates for African\-American speakers in commercial ASR systemskoenecke2020racial, and accent\-based inequity among non\-native German speakers in clinical contextsjust\_moving\_2025\. Biases related to age, gender, accents and low\-resource languages also affect the performance of ASR\(feng2024towards\)\. However, these studies are not designed for multilingual psychiatric interviews, where conversational roles are inherently unequal with clinicians typically producing longer, more structured speech and patients providing shorter and more uncertain responses\. ASR for Indian languages and accents: Recent work has focused on ASR challenges for Indian English and regional Indian languages\.Svarahhad significantly higher WER for English with Indian accent than for LibriSpeechjaved2023svarah\.rai\_deep\_2024observed substantial disparities in gender, region, and speech rate by analyzing 8,740 hours of NPTEL’s Indian lecture in English\. Considering Indic languages the large datasetsIndicSUPERBandIndicVoicesfurther highlight the linguistic, morphological, and prosodic diversity of Indian languages which pose challenges to ASRjaved\_indicsuperb\_2022;javed\_indicvoices\_2024\. However, they do not contain clinical conversations and analysis of error types\. In this regard, the Eka Medical ASR Evaluation dataset\(eka\_care\_eka\_2024\)offers valuable Indian accents and drug vocabulary with over 3,900 recordings but is limited to brief, static conversations\. On the other hand, theDISPLACE\-Mdataset contains 55 hours of annotated conversational speech in the healthcare domain\. However, this dataset lacks interviews of psychiatric cases\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3Dataset

CharacteristicsOverallEnglishHindiKannadapp\-value\#N = 202∗N = 54∗N = 78∗N = 70∗Duration30\.936\.525\.327\.2<<0\.001\(minutes\)\(6\.1, 45\.6\)\(28\.7, 51\.1\)\(4\.5, 36\.5\)\(3\.6, 45\.2\)Total words4756\.55636\.53877\.53637\.0<<0\.001\(1042, 6111\.5\)\(4331, 6216\.8\)\(845\.8, 7135\)\(526, 4873\)Unique words1152\.01269\.5885\.51450\.5<<0\.001\(407\.5, 1412\.2\)\(1028\.8, 1404\.2\)\(316\.2, 1185\.5\)\(248\.8, 1759\.8\)Moving average0\.640\.610\.640\.69<<0\.001Type\-token ratio\(0\.61, 0\.67\)\(0\.59, 0\.62\)\(0\.62, 0\.66\)\(0\.67, 0\.71\)\(window = 100\)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 1:Dataset summary\. \(∗\) indicates median \(Q1, Q3\) and \(\#\) indicates Kruskal\-Wallis rank sum test\.The data for this study comes from a tertiary teaching hospital dedicated to the treatment of psychiatric and neurological conditions\. The hospital provides free inpatient and outpatient treatment for economically disadvantaged patients, and thus, a majority of beneficiaries are from such a background\. We collected 202 audio recordings of patient and doctor/therapist interactions\. These recordings were collected using Android mobile phones in mp3 format\. While an attempt was made to make recordings in a quiet environment, no special arrangements were made for this\. Therefore, the data represent a real\-world setting where recordings are made in busy wards and outpatient department rooms\. The language, duration and lexical diversity of the dataset are summarised in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1](https://arxiv.org/html/2606.26901#S3.T1)and the speaker profiles are detailed in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2](https://arxiv.org/html/2606.26901#S3.T2)\.

CharacteristicsOverallEnglishHindiKannadapp\-value\#N = 202∗N = 54∗N = 78∗N = 70∗Patient’s genderF51 \(25\.2%\)2 \(3\.7%\)18 \(23\.1%\)31 \(44\.3%\)<<0\.001M151 \(74\.8%\)52 \(96\.3%\)60 \(76\.9%\)39 \(55\.7%\)<<0\.001Doctor’s genderF104 \(51\.5%\)54 \(100%\)30 \(38\.5%\)20 \(28\.6%\)<<0\.001M98 \(48\.5%\)0 \(0%\)48 \(61\.5%\)50 \(71\.4%\)<<0\.001Patient’s education level<<Graduate132 \(65\.3%\)18 \(33\.3%\)58 \(74\.4%\)56 \(80\.0%\)<<0\.001\>=\>=Graduate70 \(34\.7%\)36 \(66\.7%\)20 \(25\.6%\)14 \(20\.0%\)<<0\.001\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 2:Summary of speaker profiles\. \(∗\) Median \(Q1, Q3\), \(\#\) Kruskal\-Wallis rank sum test\.This dataset contains speech from 130 unique speakers, including 7 doctors/therapists and 123 patients\. All conversations are between two individuals, making 202 unique doctor\-patient/therapist\-patient pairs\. Preprocessing: All recordings were listened to by two psychiatrists to ensure they were intelligible\. It was ensured that the recordings did not contain the name of the patient or any numerical identifier like phone number, etc\. However, we did not exclude segments that contained names of places, dates, etc\., as we wish to evaluate if such named entities \(Section[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English7](https://arxiv.org/html/2606.26901#S7)\) lead to more errors in ASR\. Annotation: As part of earlier research, transcripts in native languages were available for 103 of these recordings\. For the remaining transcripts \(99\), two psychiatrists transcribed the recordings\. The annotation guidelines for transcribing speech into text are given in the Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishC](https://arxiv.org/html/2606.26901#A3)with examples for each of the three languages\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4Methodological details

In this section, we first note the base ASR models that we use to perform the audit\. We also discuss different fine\-tuning approaches to improve the WER and the fairness of the base models\. Finally, we introduce the debiasing algorithm used to develop theSamaVaaniframework\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.1ASR models

Base models: We evaluate a total ofeightASR models, as mentioned earlier\. These includeIndicWhisperindicwhisper2024,WhisperLargeV3openai\_whisper\_large\_v3,Sarvamsarvam2025,GoogleS2Tgoogle\_speech\_to\_text,Gemma3ngemmateam2025gemma3technicalreport,OmniLingualomnilingualasrteam2025omnilingualasropensourcemultilingual,Vaanivaani\_whisper\_hi\_2025, andGeminigemini25pro\_2025\. Of these,GoogleS2T,Sarvam, andGeminiare proprietary models inferenced through APIs, while others are open\-source\. Recall that we have the audio files in an interview format where each of them has exactly two speakers, i\.e\., the patient and the doctor\. We generate transcripts for each of these long\-form audio files\. Couple of these ASR models \(Sarvam’s Saarika\-2\.5 andGemini\) can generate transcripts with long\-form audio as input while the other models \(IndicWhisper,WhisperLargeV3,Vaani,Gemma3nandGoogleS2T\) can only transcribe audio in 30\-second chunks\.

Fine\-tuned models: One of the straightforward ways to improve the overall WER and fairness across demographic groups is fine\-tuning the base models\. For this we choose two of the best performing open\-source ASR models –Gemma3n\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishhttps://huggingface\.co/google/gemma\-3n\-E4B\-it](https://huggingface.co/google/gemma-3n-E4B-it)andOmniLingualthat have the best scores across groups \(see Section[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6](https://arxiv.org/html/2606.26901#S6)\)\. We perform two types of fine\-tuning as follows\. FTStd\.\{\}^\{\\textsc\{Std\.\}\}: This refers to the standard LoRA fine\-tuning using part of our dataset for training\. FTPS\{\}^\{\\textsc\{PS\}\}: Here we double the dataset by augmenting the pitch of the audio files\. The hypothesis is that this augmentation of synthetic data would result in stronger fine\-tuning, and therefore better WER and fairness\. For our experiments, we have usedPitchShift\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishhttps://docs\.pytorch\.org/audio/main/generated/torchaudio\.transforms\.PitchShift\.html](https://docs.pytorch.org/audio/main/generated/torchaudio.transforms.PitchShift.html)to augment our original audio by randomly selecting semitones in the range of \[−5,\+5\-5,\+5\]\.

![Refer to caption](https://arxiv.org/html/2606.26901v1/figures/pipeline.png)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 1:Architecture for proposed fine\-tuning pipeline illustrating LoRA, contrastive and CTC head\. \(→\\to\) indicates fine\-tuning flow while \(⇢\\dashrightarrow\) shows inference on test set\.
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.2TheSamaVaaniarchitecture

While straightforward fine\-tuning can potentially improve the WER and the fairness of the models they do not specifically target the issue of debiasing\. Recall that the transformer based ASR models follow the modern encoder decoder paradigm where the encoder learns the audio representation and the decoder follows the next word prediction task to generate transcription\. Our approach integrates a contrastive learning and a CTC \(connectionist temporal classification\)graves2006ctchead in the decoding stage while fine\-tuning\. We only train the LoRA adapters to adapt for Indic languages, freezing the transformer layers as shown in Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1](https://arxiv.org/html/2606.26901#S4.F1)\. The entire architecture and the associated loss functions are discussed below\. Contrastive learning: There are several studies that shows the effectiveness of improving fairness when using a contrastive learning approachkoudounas2024contrastive;shen2021contrastivelearningfairrepresentations;9746929\_contrastive;ye\-etal\-2022\-contrastive\. As in the case of FTPS\{\}^\{\\textsc\{PS\}\}setup, here also we have used thePitchShiftalgorithm to augment our original audio by randomly selecting semitones in the range of \[−5,\+5\-5,\+5\]\. Thus, in a batch ofNNaudio samples, we have one original sample and its corresponding pitch\-shifted sample constituting a positive pair, and the rest of theN−1N\-1audio samples are naturally considered as negative pairs\. This forces the model to learn the same utterance regardless of the pitch \(a key point of distinction between demographic groups\)\. Further, this doubles the training data and makes the model robust to the high phonetic variance found in languages like Kannada\. For an anchor representationzaz\_\{a\}and its pitch\-shifted positive pairzpz\_\{p\}, withN−1N\-1remaining negative original samples, the contrastive loss for each sample is formulated as:

ℒC​L=−log⁡esim\(​za,zp​\)τesim\(​za,zp​\)τ​\+​Σk=1N−1​esim\(​za,zn,k​\)τ\\mathcal\{L\}\_\{CL\}=\-\\log\\frac\{e^\{\\frac\{\\text\{sim\(\}z\_\{a\},z\_\{p\}\\text\{\)\}\}\{\\tau\}\}\}\{e^\{\\frac\{\\text\{sim\(\}z\_\{a\},z\_\{p\}\\text\{\)\}\}\{\\tau\}\}\\text\{ \+ \}\\Sigma\_\{k=1\}^\{N\-1\}e^\{\\frac\{\\text\{sim\(\}z\_\{a\},z\_\{n,k\}\\text\{\)\}\}\{\\tau\}\}\}\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\(1\)wherezn,kz\_\{n,k\}are the otherN−1N\-1audio samples in the batch which are considered as negative pairs andτ\\tauis the temperature that dictates the sharpness of the probability distribution, ensuring the model is heavily penalized for any phonetic overlap between the anchor and the negative samples\. The contrastive loss implemented is a variation of the NT\-Xent \(normalized temperature\-scaled cross entropy\) lossntxentlosswhich is the standard objective for self\-supervised frameworks like SimCLRsimclr\. Here, we only create the pitch augmentation for the anchor file\. CTC head: In this stage, the logits from last hidden state \(HH\) of transformer layers is projected to the vocabulary space \(VV\)\. By projecting the last hidden states through the CTC head, the model receives a secondary signal that rewards correct character\-level sequencing\. This head is initialized from scratch and is fully trainable \(WC​T​CW\_\{CTC\}=H×VH\\times V\) with CTC loss function represented as,

ℒC​T​C=−log⁡P​\(y∣x\)\\mathcal\{L\}\_\{CTC\}=\-\\log P\\left\(y\\mid x\\right\)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\(2\)wherexx= \(x0,x1,…,xTx\_\{0\},x\_\{1\},\\dots,x\_\{T\}\) denotes the input sequence of audio features,TTis the time steps,yy= \(y0,y1,…,ymy\_\{0\},y\_\{1\},\.\.\.,y\_\{m\}\) is the target label sequence text \(transcription\), andP​\(y∣x\)P\\left\(y\\mid x\\right\)is the total probability of correct transcription over various alignment paths\. The CTC loss enables alignment\-free training by allowing the model to learn mappings between input frames and output sequences without explicit frame\-level labels\. Overall loss function: The weighted sum of the standard cross entropy \(CE\) loss, the contrastive loss \(CL\) and the CTC loss constitutes the final loss as follows\.

ℒt​o​t​a​l=α×ℒC​E​\+​β×ℒC​L​\+​γ×ℒC​T​C\\mathcal\{L\}\_\{total\}=\\mathcal\{\\alpha\\times L\}\_\{CE\}\\text\{ \+ \}\\mathcal\{\\beta\\times L\}\_\{CL\}\\text\{ \+ \}\\mathcal\{\\gamma\\times L\}\_\{CTC\}\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\(3\)whereα,β\\alpha,\\beta, andγ\\gammaare the weights for each of the loss components\. These weights are optimized usingOptuna\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishhttps://optuna\.org/](https://optuna.org/)over 20 trials\. The best values ofα,β\\alpha,\\beta, andγ\\gammaare selected based on the WER score on the validation set\. LoRA adaptation: We train all the attention and MLP modules of the transformer based on the above loss function in eq\. \([\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3](https://arxiv.org/html/2606.26901#S4.E3)\)\. This allows the model to repurpose the internal representations for Indic languages with less compute\. In addition, it contributes toward the causal cross entropy loss for next word prediction in the speech to text task\. Finally, it ensures that the model retains the ability to generate contextually and grammatically correct sentences in text from the speech in a multilingual setting\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5Experimental setup

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.1Evaluation metric

Performance metric: The main indicator used to assess the accuracy of ASR transcription is the word error rate \(WER\)\. Since WER is the most widely used metric for assessing ASRs and has been utilized by several researchers in the literature, we have chosen it as the evaluation metric\. The WER is mathematically defined as the ratio of the sum of substitutions \(SS\), deletions \(DD\), and insertions \(II\) to the total number of words \(N\) in the reference transcript\.

𝒲​ℰ​ℛ%=S​\+​D​\+​IN×100\\mathcal\{WER\\%\}=\\frac\{S\\text\{ \+ \}D\\text\{ \+ \}I\}\{N\}\\times 100\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\(4\)First, we normalize both the ground truth and the generated transcripts by preprocessing the text\. This preprocessing includes lowercasing, removing punctuation, and standardizing numbers to reduce superficial mismatches\. Second, we use the JiWER Python library for the calculation of𝒲​ℰ​ℛ\\mathcal\{WER\}\. Fairness metric: The fairness score \(ℱ​𝒮\\mathcal\{FS\}\) combines a couple of components, namely, the𝒲​ℰ​ℛ\\mathcal\{WER\}gap and the average𝒲​ℰ​ℛ\\mathcal\{WER\}among groups \(say two groups –𝒢1\\mathcal\{G\}\_\{1\}and𝒢2\\mathcal\{G\}\_\{2\}\)\.

ℱ​𝒮=−δ×𝒲​ℰ​ℛa​v​g−θ×𝒲​ℰ​ℛg​a​p;δ,θ\>=0\\mathcal\{FS\}=\-\\delta\\times\\mathcal\{WER\}\_\{avg\}\-\\theta\\times\\mathcal\{WER\}\_\{gap\};\\delta,\\theta\>=0\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English\(5\)where𝒲​ℰ​ℛa​v​g\\mathcal\{WER\}\_\{avg\}and𝒲​ℰ​ℛg​a​p\\mathcal\{WER\}\_\{gap\}are the average𝒲​ℰ​ℛ\\mathcal\{WER\}of the two groups \(𝒢1\\mathcal\{G\}\_\{1\},𝒢2\\mathcal\{G\}\_\{2\}\) and the absolute difference in𝒲​ℰ​ℛ\\mathcal\{WER\}between the groups \(𝒢1\\mathcal\{G\}\_\{1\},𝒢2\\mathcal\{G\}\_\{2\}\) respectively\. A lower𝒲​ℰ​ℛg​a​p\\mathcal\{WER\}\_\{gap\}suggests better fairness across groups\. In addition, a higher value ofℱ​𝒮\\mathcal\{FS\}\(ranging from \-∞\\inftyto 0\) indicates overall balanced performance across groups\. In this study, we setδ=θ\\delta=\\theta= 0\.5 to assign equal importance to both units\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.2Implementation details

We employedGemma3nandOmniLingualmultimodal transformer\-based models as the backbone for all of our experiments\. We implement LoRA by injecting trainable rank 8 decomposition matrices into the query \(qp​r​o​j\{q\}\_\{proj\}\), key \(kp​r​o​j\{k\}\_\{proj\}\), value \(vp​r​o​j\{v\}\_\{proj\}\), output \(op​r​o​j\{o\}\_\{proj\}\) and MLP projection layers \(g​a​t​ep​r​o​j\{gate\}\_\{proj\},u​pp​r​o​j\{up\}\_\{proj\}, andd​o​w​np​r​o​j\{down\}\_\{proj\}\)\. The experimental configuration and the different hyperparameters are noted in Appendix[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishA](https://arxiv.org/html/2606.26901#A1)\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6Results

ModelsWER \(%\)\(S,D,IS,D,I\) %Sarvam\\cellcolorlangeng 34\.33\\cellcolorlangeng \(8\.94 ,16\.40, 36\.87\)\\cellcolorlanghin 39\.03\\cellcolorlanghin \(14\.47, 6\.76, 39\.15\)\\cellcolorlangkan 54\.37\\cellcolorlangkan \(27\.87, 17\.68, 39\.03\)GoogleS2T\\cellcolorlangeng 74\.60\\cellcolorlangeng \(36\.13, 40\.22, 9\.71\)\\cellcolorlanghin 85\.55\\cellcolorlanghin \(10\.09, 75\.60, 0\.13\)\\cellcolorlangkan 94\.90\\cellcolorlangkan \(11\.34, 84\.01, 0\.005\)Gemini\\cellcolorlangeng14\.15\\cellcolorlangeng \(5\.28, 18\.58, 4\.94\)\\cellcolorlanghin18\.52\\cellcolorlanghin \(11\.32, 4\.35, 8\.74\)\\cellcolorlangkan35\.01\\cellcolorlangkan \(22\.71, 16\.86, 7\.53\)WhisperLargeV3\\cellcolorlangeng 46\.76\\cellcolorlangeng \(9\.48, 19\.81, 21\.76\)\\cellcolorlanghin 71\.68\\cellcolorlanghin \(26\.23, 39\.17, 9\.21\)\\cellcolorlangkan 98\.55\\cellcolorlangkan \(22\.22, 76\.51, 0\.0014\)IndicWhisper\\cellcolorlangeng \-\\cellcolorlangeng \-\\cellcolorlanghin 70\.3\\cellcolorlanghin \(23\.44, 45\.58, 4\.27\)\\cellcolorlangkan 97\.05\\cellcolorlangkan \(29\.03, 68\.21, 0\.02\)Vaani\\cellcolorlangeng \-\\cellcolorlangeng \-\\cellcolorlanghin 44\.42\\cellcolorlanghin \(21\.54, 15\.60, 6\.29\)\\cellcolorlangkan 77\.21\\cellcolorlangkan \(5\.56, 24\.92, 1\.73\)Gemma3n\\cellcolorlangeng40\.22\\cellcolorlangeng \(14\.63, 17\.64, 19\.54\)\\cellcolorlanghin 48\.14\\cellcolorlanghin \(19\.88, 6\.45, 26\.48\)\\cellcolorlangkan 90\.90\\cellcolorlangkan \(48\.57, 16\.67, 30\.47\)OmniLingual\\cellcolorlangeng 58\.64\\cellcolorlangeng \(23\.40, 31\.20, 4\.00\)\\cellcolorlanghin43\.55\\cellcolorlanghin \(23\.76, 12\.20, 6\.53\)\\cellcolorlangkan75\.35\\cellcolorlangkan \(41\.08, 32\.19, 2\.06\)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 3:Overall performance of ASR models acrossEnglish,Hindi,Kannada\.SS,DD,II\(%\) denote median insertion, deletion, and substitution components of WER\. Best results among closed and open\-source models areboldandunderlinedrespectively\.This section is divided into two major parts\. In the first part we audit the performance of the eight ASR models across three languages – Hindi, Kannada and Indian English\. Next, we present the performance of the different fine\-tuning methods as well as the debiasing algorithmSamaVaaniproposed by us\. Audit outcomes: Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3](https://arxiv.org/html/2606.26901#S6.T3)reports the WER and its constituent parts, substitution \(SS\), deletion \(DD\), and insertion \(II\) in percentages, to compare ASR models in all the three languages\. As we observe from the table,Geminiachieves the best WER scores \(English: 14\.15%, Hindi: 18\.52%, Kannada: 35\.01%\), demonstrating better multilingual and generalization abilities\. Among the three languages, Kannada poses relatively higher challenge to the ASR models possibly due to the lack of enough pre\-training data\. Further, models likeWhisperLargeV3,GoogleS2T, andGemma3nshow a very high WER for Kannada because their architectures and training corpora are not deeply optimised for rich Indian phonetic diversity and retroflex phonemes in Dravidian languages\. Fine\-tuning and debiasing results: For fine\-tuning, we chooseGemma3nandOmniLingualas both have the lowest𝒲​ℰ​ℛ\\mathcal\{WER\}averaged over three languages\. We split the data into train, development \(dev\) and test folds with language as a stratification variable\. Train set had 83\.16 hours of audio \(English 30\.34, Hindi 27\.56, Kannada 25\.26\), dev set had 9\.64 hours \(English 3\.52, Hindi 3\.02, Kannada 3\.10\) and test set had 10\.22 hours \(English 3\.80, Hindi 2\.96 and Kannada 3\.46 hours\)\.

MetricBaseFTStd\.\{\}^\{\\textsc\{Std\.\}\}FTPS\{\}^\{\\textsc\{PS\}\}FTCL\{\}^\{\\textsc\{CL\}\}FTCTC\{\}^\{\\textsc\{CTC\}\}SamaVaaniOverall WER \(↓\\downarrow\)\\cellcolorgemmaval70\.47\\cellcoloromnival65\.85\\cellcolorgemmaval47\.62\\cellcoloromnival47\.41\\cellcolorgemmaval41\.14\\cellcoloromnival45\.83\\cellcolorgemmaval39\.08\\cellcoloromnival40\.7\\cellcolorgemmaval37\.92\\cellcoloromnival38\.12\\cellcolorgemmaval35\.19\\cellcoloromnival35\.29English WER \(↓\\downarrow\)\\cellcolorgemmaval57\.92\\cellcoloromnival47\.02\\cellcolorgemmaval25\.60\\cellcoloromnival26\.02\\cellcolorgemmaval24\.13\\cellcoloromnival24\.13\\cellcolorgemmaval23\.17\\cellcoloromnival24\.00\\cellcolorgemmaval22\.22\\cellcoloromnival22\.44\\cellcolorgemmaval20\.43\\cellcoloromnival20\.54Hindi WER \(↓\\downarrow\)\\cellcolorgemmaval58\.33\\cellcoloromnival63\.44\\cellcolorgemmaval44\.34\\cellcoloromnival43\.95\\cellcolorgemmaval38\.80\\cellcoloromnival42\.31\\cellcolorgemmaval37\.25\\cellcoloromnival38\.23\\cellcolorgemmaval36\.36\\cellcoloromnival36\.00\\cellcolorgemmaval33\.08\\cellcoloromnival32\.96Kannada WER \(↓\\downarrow\)\\cellcolorgemmaval84\.48\\cellcoloromnival90\.36\\cellcolorgemmaval75\.00\\cellcoloromnival75\.00\\cellcolorgemmaval61\.11\\cellcoloromnival72\.36\\cellcolorgemmaval58\.65\\cellcoloromnival60\.94\\cellcolorgemmaval57\.82\\cellcoloromnival56\.41\\cellcolorgemmaval52\.84\\cellcoloromnival52\.88Male \(M\) \(↓\\downarrow\)\\cellcolorgemmaval69\.23\\cellcoloromnival76\.55\\cellcolorgemmaval53\.33\\cellcoloromnival53\.19\\cellcolorgemmaval44\.23\\cellcoloromnival50\.00\\cellcolorgemmaval41\.77\\cellcoloromnival42\.31\\cellcolorgemmaval40\.71\\cellcoloromnival39\.76\\cellcolorgemmaval37\.25\\cellcoloromnival37\.25Female \(F\) \(↓\\downarrow\)\\cellcolorgemmaval71\.43\\cellcoloromnival74\.62\\cellcolorgemmaval48\.53\\cellcoloromnival48\.37\\cellcolorgemmaval42\.04\\cellcoloromnival42\.86\\cellcolorgemmaval39\.41\\cellcoloromnival42\.31\\cellcolorgemmaval38\.33\\cellcoloromnival39\.08\\cellcolorgemmaval34\.83\\cellcoloromnival35\.48ℱ​𝒮\\mathcal\{FS\}: M vs F \(↑\\uparrow\)\\cellcolorgemmaval\-36\.27\\cellcoloromnival\-43\.76\\cellcolorgemmaval\-27\.87\\cellcoloromnival\-27\.80\\cellcolorgemmaval\-22\.66\\cellcoloromnival\-26\.79\\cellcolorgemmaval\-21\.48\\cellcoloromnival\-22\.46\\cellcolorgemmaval\-20\.95\\cellcoloromnival\-20\.05\\cellcolorgemmaval\-19\.23\\cellcoloromnival\-19\.07Patient \(P\) \(↓\\downarrow\)\\cellcolorgemmaval76\.47\\cellcoloromnival84\.44\\cellcolorgemmaval53\.85\\cellcoloromnival53\.85\\cellcolorgemmaval47\.62\\cellcoloromnival49\.04\\cellcolorgemmaval45\.45\\cellcoloromnival47\.49\\cellcolorgemmaval44\.71\\cellcoloromnival43\.40\\cellcolorgemmaval41\.67\\cellcoloromnival41\.18Doctor \(D\) \(↓\\downarrow\)\\cellcolorgemmaval63\.16\\cellcoloromnival66\.92\\cellcolorgemmaval50\.00\\cellcoloromnival50\.00\\cellcolorgemmaval41\.54\\cellcoloromnival46\.18\\cellcolorgemmaval39\.24\\cellcoloromnival41\.46\\cellcolorgemmaval37\.60\\cellcoloromnival38\.03\\cellcolorgemmaval34\.40\\cellcoloromnival34\.69ℱ​𝒮\\mathcal\{FS\}: P vs D \(↑\\uparrow\)\\cellcolorgemmaval\-41\.56\\cellcoloromnival\-51\.60\\cellcolorgemmaval\-27\.88\\cellcoloromnival\-27\.88\\cellcolorgemmaval\-25\.33\\cellcoloromnival\-26\.48\\cellcolorgemmaval\-24\.28\\cellcoloromnival\-25\.25\\cellcolorgemmaval\-24\.13\\cellcoloromnival\-23\.04\\cellcolorgemmaval\-22\.65\\cellcoloromnival\-22\.21\>\>=Graduate \(↓\\downarrow\)\\cellcolorgemmaval63\.64\\cellcoloromnival75\.23\\cellcolorgemmaval28\.45\\cellcoloromnival28\.28\\cellcolorgemmaval26\.66\\cellcoloromnival28\.45\\cellcolorgemmaval25\.16\\cellcoloromnival27\.27\\cellcolorgemmaval25\.16\\cellcoloromnival26\.43\\cellcolorgemmaval22\.28\\cellcoloromnival22\.42<<Graduate \(↓\\downarrow\)\\cellcolorgemmaval80\.00\\cellcoloromnival83\.44\\cellcolorgemmaval70\.43\\cellcoloromnival70\.47\\cellcolorgemmaval61\.53\\cellcoloromnival64\.80\\cellcolorgemmaval59\.17\\cellcoloromnival62\.17\\cellcolorgemmaval58\.82\\cellcoloromnival57\.14\\cellcolorgemmaval53\.64\\cellcoloromnival53\.39ℱ​𝒮\\mathcal\{FS\}:\>\>=G vs<<G \(↑\\uparrow\)\\cellcolorgemmaval\-44\.09\\cellcoloromnival\-48\.07\\cellcolorgemmaval\-45\.71\\cellcoloromnival\-39\.81\\cellcolorgemmaval\-39\.49\\cellcoloromnival\-38\.90\\cellcolorgemmaval\-38\.09\\cellcoloromnival\-36\.81\\cellcolorgemmaval\-37\.83\\cellcoloromnival\-36\.25\\cellcolorgemmaval\-34\.66\\cellcoloromnival\-34\.44\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 4:Comparison of WER scores ofGemma3nandOmniLingualacross different fine\-tuning techniques and different demographic attributes\. Lower WER \(↓\\downarrow\) and higherℱ​𝒮\\mathcal\{FS\}\(↑\\uparrow\) indicates better performance\. FTCL\{\}^\{\\textsc\{CL\}\}: An ablation ofSamaVaaniwhere only the contrastive loss is used\. FTCTC\{\}^\{\\textsc\{CTC\}\}: An ablation ofSamaVaaniwhere only the CTC loss is used\.Boldindicates the best score obtained from one model\.Main results: We present the results from the different fine\-tuned models in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2606.26901#S6.T4)\. The results are organized under different categories including overall𝒲​ℰ​ℛ\\mathcal\{WER\}, language\-wise𝒲​ℰ​ℛ\\mathcal\{WER\}and various demographic\-wise𝒲​ℰ​ℛ\\mathcal\{WER\}\. Further we also present the fairness score \(ℱ​𝒮\\mathcal\{FS\}\) for each demographic group\. The key observations are as follows\.

1. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\.We observe thatSamaVaaniresults in a reduction of 50% in overall𝒲​ℰ​ℛ\\mathcal\{WER\}when compared to theBasepre\-trained model\. The overall𝒲​ℰ​ℛ\\mathcal\{WER\}ofSamaVaaniis also substantially better than standard fine\-tuning setups FTStd\.\{\}^\{\\textsc\{Std\.\}\}and FTPS\{\}^\{\\textsc\{PS\}\}\.
2. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.Across all the three languagesSamaVaaniis remarkably better than theBasemodel as well as FTStd\.\{\}^\{\\textsc\{Std\.\}\}and FTPS\{\}^\{\\textsc\{PS\}\}\. Among the three languages, evenSamaVaanistruggles the most with Kannada, like all the other models\.
3. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.For all the demographic groupsSamaVaaniis not only better in terms of the group\-wise𝒲​ℰ​ℛ\\mathcal\{WER\}but also in terms ofℱ​𝒮\\mathcal\{FS\}when compared toBase, FTStd\.\{\}^\{\\textsc\{Std\.\}\}and FTPS\{\}^\{\\textsc\{PS\}\}\.

Ablation study: A natural question in the design ofSamaVaaniregards the necessity of both the CL and CTC heads\. In order to check whether any one of them is as good as the combination, we present in columns 4 and 5 of Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2606.26901#S6.T4)the results from two additional fine\-tuning setups – \(i\) FTCL\{\}^\{\\textsc\{CL\}\}: an ablation ofSamaVaaniwhere only the contrastive loss is used and \(ii\) FTCTC\{\}^\{\\textsc\{CTC\}\}: an ablation ofSamaVaaniwhere only the CTC loss is used\. We observe that though these models outperform the standard fine\-tuning setups FTStd\.\{\}^\{\\textsc\{Std\.\}\}and FTPS\{\}^\{\\textsc\{PS\}\}in terms of both𝒲​ℰ​ℛ\\mathcal\{WER\}andℱ​𝒮\\mathcal\{FS\}, they are not as good asSamaVaani\. This quantitatively justifies the benefit of combining the two loss terms\. In the next section, we present some qualitative insights into the advantages of each of these components\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English7Discussion

TypeTranscriptGround truthfriends, relatives\. Yeah, if he values…\\dotsif he values that conflicted person a very high level in in his own mind…\\dotsit is uh, good decision that to worry…\\dotsfor him…\\dotsthat he didn’t come\. If he is a normal person who had a conflict, he won’t mind that\. He won’t mind that\.Basefriends,celebrities\. Yeah, ifif if if if…\\dots\(repeated 210 times more\)FTStd\.\{\}^\{\\textsc\{Std\.\}\}friends,celebrities\. Yeah, ifif if if if…\\dots\(repeated 212 times more\)FTCL\{\}^\{\\textsc\{CL\}\}friends,celebrities\. Yeah, ifif if if if…\\dots\(repeated 212 times more\)FTCTC\{\}^\{\\textsc\{CTC\}\}friends,celebrities\. Yeah, ifif if if you valueifif you valuethat conflicted person a very high level in in his own mind, it is a good decision that to worry for him that he didn’t come\. If he is a normal person who had a conflict, he won’t mind that\. He won’t mind that\.SamaVaanifriendscigarettes\. Yeah, if you value ifif you valuethat conflicted person a very high level in in his own mind, it isagood decision that to worry for him that he didn’t come\. If he is a normal person who had a conflict, he won’t mind that\. He won’t mind that\.\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 5:Qualitative comparison of transcriptions across models ofGemma3n\. Here the speaker is a male patient\.In this section, we discuss representative qualitative advantages of the CL and CTC loss terms \(forGemma3n\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4Similar observations hold forOmniLingual\.\) and finally discuss some error cases\. We posit that while the CTC head provides essential character alignment from speech to text, the contrastive learning objective acts as a phonetic regularizer\. Role of contrastive learning: Recall that we usePitchShiftto obtain a pitch\-shifted variant of the anchor audio\. With this, the objective of contrastive learning is to recognize the semantic equivalence between the original audio and its pitch augmented variant to ignore acoustic noise and focus on phonetic content\. By restricting the augmentation to a single anchor\-positive pair \(zi,zi\+z\_\{i\},z\_\{i\}^\{\+\}\) within a pool ofN−1N\-1negative samples \(zkz\_\{k\}\), the framework creates an asymmetric learning signal that is highly effective for phonetically dense languages\. We set the temperature \(τ\\tau\) at 0\.05 to sharpen the distribution, forcing the model to be very certain about the representations of samples in the latent space\.ℱ​𝒮\\mathcal\{FS\}improve consistently on all demographic attributes by 13\-41% compared to the base model and 15\-22% compared to the standard fine\-tuned model\. Role of CTC head: Standard LLMs use auto regressive decoding, which are prone to hallucinations with same words repeating sequentially, for example,ifif if if…\(get caught in a word repeating indefinitely\)\. The CTC heads enforces monotonic alignment\. While the generative head ensures the sentence makes sense, the CTC head acts as a “sanity check” to ensure every word/token corresponds to an actual speech in the file\. This balance significantly reduces instances of “word\-skipping” or adding “filler” that wasn’t in the original audio\. In the raw audio signal, a single phoneme \(like ‘s’ in “speech”\) lasts for many frames\. Without a special mechanism, a model might predict the letter ‘s’ many times in a row\. CTC solves this using a unique decoding rule involving blank tokens \(ϕ\\phi\)\. The CTC decoding algorithm follows two simple but powerful rules to convert a long sequence of frame\-by\-frame predictions into a clean word as follows – \(i\)collapse identical consecutive tokens: if the model predicts ‘aaaa\-bbbb\-cccc’ due to slow speech, CTC collapses them into ‘abc’; \(ii\)blank as a separator: to actually output two of the same letter \(like the ‘ll’ in ‘hello’\), the model must predict a blank token between them \(e\.g\., h\-e\-l\-ϕ\\phi\-l\-o\)\. Error analysis: Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5](https://arxiv.org/html/2606.26901#S7.T5)shows qualitative examples as to how auto\-regressive models can get stuck in a loop where they start repeating the same word sequentially\. Fine\-tuning with only contrastive learning also fails to get out of the repeating loop\. On the other hand, incorporating a CTC head on top LoRA is able to break out of this loop and generates better text\. Furthermore,SamaVaanigenerates the best transcripts compared to ground truth\. Not only it can get out of the indefinite loop, it also removes the same repeating words with a few inconsistencies\. Lastly, the error in the transcripts generated bySamaVaaniis essentially divided into the following three categories\.

1. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\.Inverse text normalization: The generation should appear in text as spoken word\-by\-word\. For instance, “So, no matter what time I sleep,8:00o’clock is when I have to wake up\.” Here, the time should appear in words‘eight’as it is spoken and not as it is represented\.
2. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.Named entity errors: There are several instances where the model fails to correct text for named entities, especially for organization and drug names\. For example, the organization name‘NIMHANS’in ground truth is being substituted by terms like‘neeman’s’or‘2 months’\. Similarily, the drug name‘benzodiazipines’is being transcribed as‘benzodiazepam pains’, suggesting a lack of acoustic robustness, where the model struggles to align the correct token sequences\.
3. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.Deletion errors: There are a few missing phonemes while transcribing speech to text\. Forcing monotonic alignment with the CTC head prevents the model from skipping any word/token\. It forces the model to attempt a phonetic transcription of every sound\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English8Conclusion

In this study, we highlighted the disparities in ASR models, particularly in a clinical psychiatric interview setting between a patient and a doctor\. We comprehensively audited eight state\-of\-the\-art models across three linguistically diverse languages and reported multiple disparities\. Next, we proposeSamaVaanithat leads to simultaneous improvement of the overall𝒲​ℰ​ℛ\\mathcal\{WER\}as well as the demographic fairness\. The key uniqueness lies in combining the two loss functions based on contrastive learning and CTC that serve complementary roles\.

While our approach is better than the other fine\-tuning setups, there is still a lot of scope for improvement, especially for languages like Kannada\. The values suggest that it is important to build customised ASR models for a clinical psychiatric interview setting, pre\-trained from scratch\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English9Limitations

While this study poses a solid foundation towards multilingual ASR in the psychiatric interview setting, there are a couple of limitations\.First, the experimental setup depends on LoRA fine\-tuning with rank of 8 \(low\)\. This setup was chosen due to our hardware constraints\. With a high\-end infrastructure, a full supervised fine\-tuning may further improve the transcription accuracy and fairness\. Andsecond, our evaluation is limited only to Indian English, Hindi and Kannada psychiatric interviews\. India contains substantial linguistic diversity, and ASR behaviour may differ considerably across other Indic languages, dialects, and code\-mixed settings\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English10Ethical consideration

This study was approved by the Institute Ethics Committee\. The speech data were collected from a tertiary teaching hospital with a specialised addiction treatment centre offering 24\-hour emergency services dedicated to the treatment of psychiatric and neurological conditions, along with inpatient and outpatient services\. Written informed consent was obtained from patients to audio\-record psychiatric interviews\. Although deidentified, the data cannot be made public as it contains highly sensitive personal health information, which can compromise patient privacy and confidentiality\.

## References

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix AImplementation details & hyperparameters

We implement LoRA by injecting trainable rank 8 decomposition matrices into the query \(qp​r​o​j\{q\}\_\{proj\}\), key \(kp​r​o​j\{k\}\_\{proj\}\), value \(vp​r​o​j\{v\}\_\{proj\}\), output \(op​r​o​j\{o\}\_\{proj\}\) and MLP projection layers \(g​a​t​ep​r​o​j\{gate\}\_\{proj\},u​pp​r​o​j\{up\}\_\{proj\}, andd​o​w​np​r​o​j\{down\}\_\{proj\}\) on the base ASR models\. The experimental configuration and the different hyperparameters are noted in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6](https://arxiv.org/html/2606.26901#A1.T6)\.

ConfigurationValueLoRA configurationRank \(rr\)8Scaling factor \(αLoRA\\alpha\_\{\\text\{LoRA\}\}\)16Dropout0\.09Target modulesqproj,kproj,vprojq\_\{\\text\{proj\}\},k\_\{\\text\{proj\}\},v\_\{\\text\{proj\}\},oproj,g​a​t​eprojo\_\{\\text\{proj\}\},gate\_\{\\text\{proj\}\},u​pproj,d​o​w​nprojup\_\{\\text\{proj\}\},down\_\{\\text\{proj\}\}Base model quantization4\-bitOptimizationOptimizerAdamW \(8\-bit\)Base learning rate5×10−55\\times 10^\{\-5\}Learning rate scheduleCosineLearning rate warmup period10% \(for 3 epochs\)Max sequence length1024 tokensTemperature \(τ\\tau\)0\.05Training infrastructurePer\-device batch size1Gradient accumulation steps8Effective global batch size32GPUs4×\\timesNVIDIA A6000\(48 GB each\)Validation & early stoppingValidation metricMedian WEREarly stopping patience3 evaluation steps\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 6:Common experimental configuration and hyperparameters\.The loss coefficients \(α,β\\alpha,\\beta,γ\\gamma\) for each of the three components in the total loss function for both the fine\-tuned models,Gemma3nandOmniLingualare optimized throughOptuna\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englishhttps://optuna\.org/](https://optuna.org/)over 20 trials\. The exact coefficient values forGemma3nandOmniLingualare \(0\.4135, 0\.2186, 0\.3679\) and \(0\.4385, 0\.2480, 0\.3135\) respectively\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix BAcoustic features and fairness analysis of dataset

In this section, we analyze the acoustic features of the dataset used in this stdy\. We evaluate several key features, which indicates how good the voice quality of the data is\. Performance evaluations indicate that males, patients, and Kannada speakers are structurally disadvantaged subgroups which are disproportionately affected by higher Word Error Rates \(𝒲​ℰ​ℛ\\mathcal\{WER\}\)\. Importantly, the performance gap stems from acoustic differences, not data size\. Although males represent 65% of the dataset, their speech leads to worse WERs because of a naturally lower fundamental frequency \(r=0\.84r=0\.84\) and a lower voice quality, indicated by lower Harmonics\-to\-Noise Ratio \(HNR\) and higher amplitude instability \(shimmer\)\. Likewise, the speech of patients is a reflection of clinical realities such as psychomotor symptoms or emotional distress, which manifest acoustically as lower pitch and degraded voice quality\. We depict the differences in pitch and voice quality of the speech data among gender and speaker role respectively from Figures[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2](https://arxiv.org/html/2606.26901#A2.F2)to[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5](https://arxiv.org/html/2606.26901#A2.F5)\. In addition, the illustrate the speech intelligibility analysis in Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6](https://arxiv.org/html/2606.26901#A2.F6)\.

![Refer to caption](https://arxiv.org/html/2606.26901v1/figures/pitchbyrole.png)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 2:Pitch comparison based on the role of the speaker, i\.e, doctor or patient\.![Refer to caption](https://arxiv.org/html/2606.26901v1/figures/vcbyrole.png)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 3:Voice quality comparison based on the role of the speaker, i\.e, doctor or patient\.![Refer to caption](https://arxiv.org/html/2606.26901v1/figures/pitchbygender.png)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 4:Pitch comparison based on the gender of the speaker, i\.e, male or female\.![Refer to caption](https://arxiv.org/html/2606.26901v1/figures/vcbygender.png)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 5:Voice quality comparison based on the gender of the speaker, i\.e, male or female\.![Refer to caption](https://arxiv.org/html/2606.26901v1/figures/stoi.png)\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 6:Speech intelligibility report of the speech data used in this study\.
## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishAppendix CAnnotation guidelines

In this section, we detail the exact transcription instructions given to annotators with examples for each language\.

Transcription instructions to annotators: Your task is to carefully listen to the provided audio file and create a\.txtor\.srtfile\. About the recordings: ∙\\bulletThese are audio\-recorded clinical interviews between doctors and patients\. ∙\\bulletAccuracy of transcription is very important for clinical and research purposes\. ∙\\bulletPlease follow the following guidance material to ensure accurate transcriptions\. Carefully read the instructions below on how to do the transcriptions and their examples for each of the three language, namely, English, Hindi, and Kannada \(in order\)\. Speaker diarisation: Clearly distinguish between the speakers\. Use the following labels at the beginning of each new turn of speech\. ∙\\bulletFor English→\\to\[Doctor:\] \[Patient:\] ∙\\bulletFor Hindi→\\to\[\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiडॉक्टर:\] \[\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमरीज़:\] ∙\\bulletFor Kannada→\\to\[\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaವೈದ್ಯ:\] \[\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaರೋಗಿ:\] Use of punctuation: Allowed punctuations are full\-stop, question mark, comma, ellipsis, em\-dash and exclamation marks\. PunctuationGuidanceFull Stop \(\.\)Use for the end of a sentence\.Comma \(,\)Use for short pauses and separating multiple items or phrases\.Ellipsis \(…\)Use for long pauses\.Em\-dash \(–\)Use for interruptions or cut\-offs\. English example, Doctor: So what brings you h– Hindi example,\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiचिकित्सक: तुम यहाँ क्यों – Kannada example,\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಡಾಕ್ಟರ್: ನೀವು ಯಾಕೆ ಇಲ್ಲಿಗೆ –\-Question Mark \(?\)Use for direct questions\.Exclamation \(\!\)Use a single mark to indicate strong emotions\. English example, It just feels so bad I can’t even tell you\! Hindi example,\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमैं आपको बता नहीं सकता कि मैं कितना परेशान हूँ\! Kannada example,\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನಂಗೆ ಎಷ್ಟು ಬೇಜಾರ್ ಆಗುತ್ತೆ ಅಂದ್ರೆ ಹೇಳೋಕ್ಕೆ ಆಗಲ್ಲ\!Strict verbatim transcription: Preserve all dysfluencies like pauses, filler words, stammers etc\. Do not correct if there are grammatical errors in the speech itself\. See table below\. GuidanceEnglish examplesHindi examplesKannada examplesCapture filler words and phrasesUm, Yeah, Hmm, Mmmm, etc\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiअच्छा, हाँ, हम्म, हा, आ, etc\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಉಮ್, ಅಹ್, ಹಮ್ಮ್, ಹಾ, ಆ, etcInclude all incomplete phrases as they are saidI… I yesterday… no, the day before, I had gone there\.\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमैं…मैं कल…नहीं, परसों वहाँ गया था।\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನಾನು… ನಾನು ನಿನ್ನೆ… ಇಲ್ಲ, ಮೊನ್ನೆ ಅಲ್ಲಿಗೆ ಹೋಗಿದ್ದೆInclude stutters and stammersI\-I\-I get scared\.\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमैं\-मैं\-मैं डर जाता हूं—\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನ\-ನ\-ನನಗೆ ಭಯ ಆಗುತ್ತೆUse ellipsis \(three dots\) for long pauses or incomplete sentencesI don’t know… I feel very scared\.\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमुझे नहीं पता… मुझे बहुत डर लग रहा है—\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನನಗೆ ಗೊತ್ತಿಲ್ಲ… ತುಂಬಾ ಭಯ ಆಗುತ್ತೆTranscribe as spoken, do not change informal to formal style \(applies for languages like Kannada where spoken/colloquial and written styles are very different\)\. Do not correct grammar or polish the output in any manner\.I am afraid→\\toI am feeling scared\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमुझे डर है→\\toमुझे डर लग रहा है\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನಂಗೆ ಭಯ ಆಗ್ತಿದೆ→\\toನನಗೆ ಭಯ ಆಗುತ್ತಿದೆThere can be frequent language mixing in these audios\. You can expect English, Hindi, Telugu, Tamil and Malayalam\. Write these words in their native script as of speech\.main samajh gaya, gottaytu\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमैं समझ गया, गोथायतू\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಠೀಕ್ ಹೈ, ನಂಗೆ ಅರ್ಥ ಆಯಿತು \[theek hai, nange artha aytu\]If there are numbers, type them out in words and not as numerals\.Incorrect: I took 10 tablets that day\. Correct: I took ten tablets that day\.Incorrect:\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमैंने 10 गोलियाँ लीं / मैंने १० गोलियाँ लीं। Correct:\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमैंने दस गोलियाँ ले लींIncorrect:\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನಾನು ಅವತ್ತು 10 ಮಾತ್ರೆಗಳು ತಗೊಂಡೆ / ನಾನು ಅವತ್ತು ೧೦ ಮಾತ್ರೆಗಳು ತಗೊಂಡೆ Correct:\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನಾನು ಅವತ್ತು ಹತ್ತು ಮಾತ್ರೆಗಳು ತಗೊಂಡೆ

Non\-speech occurrences: Five relevant non\-speech sounds must be noted\. Use square brackets \[…\\dots\] for these annotations\. For example, when the discussion is about depression, sounds of the speaker crying becomes important\. See the indicators below\. Do not use any other indicators\. TypeEnglish examplesHindi examplesKannada examplesThree Special tokens are allowed for emotional expressions\. They must be in square brackets and are the only ones allowed\.\[laughs\], \[cries\], and \[shouts\]\[\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiहंसने की आवाज़\], \[\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiरोने की आवाज़\] and \[\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiचिल्लाना\]\[\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaನಗುವಿನ ಸದ್ದು\], \[\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಅಳುವಿನ ಸದ್ದು\] and \[\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಕೂಗುತ್ತಾ\]One special token is allowed for unclear speech\[unclear\]\[\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiअस्पष्ट\]\[\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಅಸ್ಪಷ್ಟ\]One special token is allowed for other noises\[noise\]\. Use this for all non\-human background sounds like phone ringing or buzzing, ambulance siren, car honks, furniture scraping, other mechanical sounds\.\[00:03:16\] Patient: I had come then, but \[noise\] could not meet him\.\[00:03:16\]\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमरीज़: मैं तब आया था, लेकिन \[noise\]\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमिला नहीं\[00:03:16\]\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaರೋಗಿ: ಆವಾಗ್ಲೇ ಬಂದಿದ್ದೆ, ಆದ್ರೆ \[noise\]\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಸಿಗ್ಲಿಲ್ಲ\.

Timestamping and General Formatting: Use uniform font size and line spacing\. The document must be in \.txt or \.srt format\. The transcript must be timestamped in \[HH:MM:SS\] format \(hours: minutes: seconds\)\. See the example below: English examplesHindi examplesKannada examples\[00:00:05\] Doctor: Please come in, have a seat, and tell me, how are you? \[00:00:18\] Patient: Hello, Doctor\. I haven’t been feeling well for the past one or two weeks and I feel sad all the time\. \[00:01:05\] Patient: Yes, Doctor\. I used to enjoy talking to my friends and gardening, but now I feel like just sitting alone all the time\.\[00:00:05\]\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiडॉक्टर: आइए, बैठिए और बताइए कि आप कैसे हैं? \[00:00:18\]\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमरीज़: नमस्कार डॉक्टर साहब\. पिछले एक\-दो सप्ताह से मेरी तबीयत ठीक नहीं है और हमेशा उदास रहता हूं\. \[00:01:05\]\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमरीज़: हाँ डॉक्टर, पहले मुझे अपने दोस्तों से बात करना और बागवानी करना अच्छा लगता था, लेकिन अब हर समय अकेले बैठे रहने का मन करता है\.\[00:00:05\]\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaವೈದ್ಯರು: ಬನ್ನಿ, ಕೂತ್ಕೊಳ್ಳಿ\. ಹೇಳಿ, ಈವಾಗ ಹೇಗಿದ್ದೀರ? \[00:00:18\]\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaರೋಗಿ: ನಮಸ್ಕಾರ ಡಾಕ್ಟರ್\. ಒಂದು ಎರಡು ವಾರದಿಂದ ಏನೋ ಸರಿ ಇಲ್ಲ\. ಯಾವಾಗಲೂ ಬೇಜಾರು… ಒಂಥರಾ ಅನ್ಸುತ್ತೆ\. \[00:01:05\]\\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaರೋಗಿ: ಹೌದು ಡಾಕ್ಟರ್\. ಫ್ರೆಂಡ್ಸ್ ಜೊತೆ ಮಾತಾಡೋದು, ಗಿಡ ನೋಡಿಕೊಳ್ಳೋದು ಅಂದ್ರೆ ಇಷ್ಟ\. ಆದ್ರೆ ಈಗ ಏನೂ ಮಾಡೋಕೆ ಇಷ್ಟ ಆಗಲ್ಲ\. ಸುಮ್ಮನೆ ಕೂತಿರ್ತೀನಿ ಅಷ್ಟೇ\. A speaker’s entire turn must be kept in one line\. Do not press ENTER \(start a new line\) in the middle of their speech even if it contains multiple sentences\. Use line breaks between separate speakers\. See example below\. English examplesHindi examplesKannada examplesCorrect Doctor: When was the last time you slept well? Patient: I can’t remember\. It’s been many months\. Incorrect Doctor: When was the last time you slept well? Patient: I can’t remember\. It’s been many months\.Correct \\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiचिकित्सक: पिछली बार आपको अच्छी नींद कब आई थी? \\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमरीज़: मुझे याद नहीं\. लगभग एक महीना हो गया\. Incorrect \\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiचिकित्सक: पिछली बार आपको कब अच्छी नींद आई थी? \\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiमरीज़: मुझे याद नहीं\. \\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiलगभग एक महीना हो गया\.Correct \\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaವೈದ್ಯರು: ನೀವು ಕೊನೆಯದಾಗಿ ಯಾವಾಗ ಚೆನ್ನಾಗಿ ನಿದ್ದೆ ಮಾಡಿದ್ದೀರಿ? \\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaರೋಗಿ: ನನಗೆ ನೆನಪಿಲ್ಲ\. ಸುಮಾರು ಒಂದು ತಿಂಗಳಾಯ್ತು\. Incorrect \\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaವೈದ್ಯರು: ನೀವು ಕೊನೆಯದಾಗಿ ಯಾವಾಗ ಚೆನ್ನಾಗಿ ನಿದ್ದೆ ಮಾಡಿದ್ದೀರಿ? \\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaರೋಗಿ: ನನಗೆ ನೆನಪಿಲ್ಲ\. \\fontspec\_if\_script:nTFknda\\addfontfeatureScript=Kannada\\fontspec\_if\_language:nTFKAN\\addfontfeatureLanguage=Kannadaಸುಮಾರು ಒಂದು ತಿಂಗಳಾಯ್ತು\.

Similar Articles

Towards a Phonology-Informed Evaluation of Multilingual TTS

arXiv cs.CL

This paper proposes a classifier-based framework to audit multilingual TTS systems for phonological faithfulness, using Assamese ATR vowel harmony as a case study. It reveals that Meta's MMS TTS frequently misproduces advanced tongue root vowels, a bias absent in human speech.