Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

arXiv cs.CL Papers

Summary

This paper evaluates forced alignment for Hindi-English code-mixed speech using the Montreal Forced Aligner, demonstrating that bootstrapping strategies and code-mixed training data achieve a tenfold improvement in alignment accuracy over monolingual alternatives.

arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We address 2 problems: (1) free variation involving native vs non-native pairs and (2) phonemic boundary detection for mid-utterance English words. Bootstrapping strategies substantially outperform unmodified lexicons. Acoustic models trained on sentence-level code-mixed data achieve a mean error of 4.15ms, ie. ten times lower than monolingual Hindi (38.18ms) or isolated English (37.58ms) alternatives. Principled lexicon design and code-mixed training data are both essential for reliable alignment of bilingual speech.
Original Article
View Cached Full Text

Cached at: 07/29/26, 09:55 AM

# Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
Source: [https://arxiv.org/html/2607.25581](https://arxiv.org/html/2607.25581)
Pandey Gogoi Tang\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English

PamirKevin1Karya, India 2Department of English Language and Linguistics, Institute of English and American Studies, Faculty of Arts and Humanities, Heinrich Heine University Düsseldorf, Germany 3Department of Linguistics, University of Florida, United States of America[ayushi@karya\.in, pamir\.gogoi@karya\.in, kevin\.tang@hhu\.de](https://arxiv.org/html/2607.25581v1/mailto:[email protected],%[email protected],%[email protected])

###### Abstract

Code\-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation\. We evaluate forced alignment of Hindi\-English code\-mixed speech using the Montreal Forced Aligner\. We address 2 problems: \(1\) free variation involving native vs non\-native pairs and \(2\) phonemic boundary detection for mid\-utterance English words\. Bootstrapping strategies substantially outperform unmodified lexicons\. Acoustic models trained on sentence\-level code\-mixed data achieve a mean error of 4\.15ms, ie\. ten times lower than monolingual Hindi \(38\.18ms\) or isolated English \(37\.58ms\) alternatives\. Principled lexicon design and code\-mixed training data are both essential for reliable alignment of bilingual speech\.

###### keywords:

code\-mixing, code\-switching, forced alignment, pronunciation variation, speech recognition

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1Introduction

Code\-mixing \(word\-level insertions of another language in a sentential frame\) frequently occurs in quotidian speech of bilingual and multilingual communities\. English, as one of the official languages of India, is recognized as a widespread medium of education, and enjoys a superior diglossic position within the country\[parasher1981indian\]\. Therefore, speakers of Indian languages exhibit routine switches between their regional mother tongues \(L1\) and English\.

The phenomenon of code\-mixing has been actively addressed in the full stack of speech technologies for Indian languages: from dataset curation\[pandey2018phonetically,nayak2022l3cube,senthamizhselvi2025building\], speech recognition\[palivela2025code,bhogale2026towards\], to reasoning in language models, and in audio generation\[gourav2025code,murthy2025building\]for voice interfaces\. However, the foundational task of phonemic alignment in code\-mixing has been largely overlooked for Indian languages\. This means that the scope and understanding of computational resources required for phonological, linguistic analysis of code\-mixing is significantly lacking\.

Phonemic alignment, which produces time\-aligned phoneme boundaries, is a critical component of large\-scale phonological analysis of languages\. Forced aligners\[mcauliffe17\_interspeech,gorman2011prosodylab,bain2023whisperx\]automate this process, and enable phoneticians to document low\-resource languages\[chodroff2025comparing,wang2023evaulating,Tang\_Bennett2019\_bootstrapping\_FA\], model speaker variation and complement their lab\-produced results with real\-world data\[chodroff2014burst\]\. Forced alignment of code\-mixed speech offers a unique opportunity in studying the advanced phonemic inventory of bilingual speakers\[amazouz2019exploring,pandey2020understanding\]and examine their results in real\-world, often low\-resource settings\[ahn\-etal\-2025\-automatic\]\. Therefore, this paper conducts structured experiments for forced\-alignment on code\-mixed language, Hindi and English\. It identifies, and addresses two problems that arise in such developments\.

The first problem is of a phonological nature\.RQ1:How can pronunciation variation in code\-mixed speech be automatically modeled within a forced\-alignment framework?Bilingual production in Hindi frequently exhibits free variation in segments shared across languages\. For example, the English wordphonemay be realized as either \[phon\] or \[fon\]\. This is different from speaker variation of the same target phoneme, because the target sometimes shiftswithinthe speaker itself\. The problem is further exacerbated when the script does not clearly specify the target phoneme\. We address this issue in Experiment I, where we identify that the exact site of such variability exists in the \[\]∼\\sim\[z\] variation\. This issue has been addressed through bootstrapping methods\.

A second problem in code\-mixed forced alignment is of the technical nature\.RQ2:What acoustic training data \(code\-mixed, monolingual Hindi, or in\-corpus English\) is most suitable for aligning mid\-utterance English words in bilingual speech?Accurate phonemic boundary detection for mid\-utterance English words depends on the composition of the acoustic training data\. Acoustic models trained exclusively on monolingual Hindi or monolingual English may not adequately reflect bilingual production patterns\. It therefore remains unclear whether sentence\-level code\-mixed training data provides advantages over monolingual alternatives for forced alignment\[ahn\-etal\-2025\-automatic,ahn2025investigating\]\.

RQ1 addresses the problem of modeling bilingual segmental variability within the pronunciation lexicon, moving beyond monolingual phoneme assumptions to better reflect attested production patterns\. RQ2 evaluates how the composition of acoustic training data influences alignment quality in mixed\-language contexts, particularly for embedded English items\. To address these questions, we conduct two experiments: Experiment 1 focuses on lexicon design for bilingual variability \(RQ1\), while Experiment 2 examines the effect of acoustic training data on alignment performance \(RQ2\)\. Together, these experiments provide a principled evaluation of lexicon design and acoustic model selection for forced alignment in code\-mixed speech\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2Sources of variation: a problem for forced alignment

Two types of variations are discussed here: the phonological and the orthographic\. The phonological variation is an example of variation in the pronunciation of code\-mixed speech, where the acoustic realization of the target phoneme is not fixed\. The orthographic variation, on the other hand, is an example of variation in the written form\. Even though the PBCM corpus comes from newspapers, some orthographic inconsistencies persist\. A detailed explanation of both forms is written below:

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.1Segmental free variation

#### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.1\.1\[ph\]∼\\sim\[f\]

Hindi’s native voiceless bilabial aspirate \[ph\] alternates with the voiceless labiodental fricative \[f\]\. Both sounds involve turbulent labiodental airflow, aspiration in the case of \[ph\], frication in the case of \[f\], which likely drives their perceptual overlap in bilingual speech\. For example, the English wordphonemay be realized as either \[phon\] or \[fon\], and the Hindi wordsafal\(‘successful’\) as either \[saph@l\] or \[saf@l\]\. There is a tendency in monolingual speakers to use the native phoneme ph\], whereas a bilingual speaker may either alternate correctly, or default to the non\-native phoneme \[f\]\.

#### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.1\.2\[\]∼\\sim\[z\]

In the case of loan words that contain the voiced alveolar fricative \[z\] \(e\.g, zebra, zero\), the Hindi’s native voiced palatal affricate \[\] alternates with the voiced alveolar fricative \[z\]\. For example,zebramay be realized as either \[i:br@\] or \[zi:br@\]\.

There is a tendency in monolingual speakers to use the native phoneme, whereas a bilingual speaker can be expected to choose the correct alternative\.\. It is important to note here, that a similar pattern exists between loanwords from Urdu \(e\.g, zulfein \(hair\), fizool \(waste\)\), where monolingual speakers default to their native counterparts described above\. However, here we focus only on the code\-mixing between English\-Hindi\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.2Orthographic variation

Devanagari script employs a subscript diacritic, thenuqta\(\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindi\), to distinguish borrowed phonemes from their native counterparts: \(i\)\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiज \[\] vs\.\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiज़ \[z\] and \(ii\)\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiफ \[ph\] vs\.\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiफ़ \[f\]\.

In written work, the nuqta is frequently omitted, neutralizing the orthographic contrast within each pair\. This introducessystematic ambiguityin grapheme\-to\-phoneme \(G2P\) conversion\. The problem is further compounded by a cultural, native\-speaker awareness of this omission\. This means thatbilingual speakerswho know the diacritic is frequently omitted maycompensate: when a \[z\]\-word \(likezebra\) is written with a bare\\fontspec\_if\_script:nTFdeva\\addfontfeatureScript=Devanagari\\fontspec\_if\_language:nTFHIN\\addfontfeatureLanguage=Hindiज, they either infer that the nuqta was erroneously dropped and correctly produce the \[z\]\. And when a \-word \(likejuice\) is written without the nuqta, they remain faithful to the orthograph and produce a \. This yieldsbidirectional variation, in which both members of each pair surface in environments where only one was intended\. Because of this inconsistencies, a one\-to\-one mapping between the grapheme to phoneme is rendered invalid\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3Data

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.1Acoustic data

The Phonetically Balanced Code Mixed \(PBCM\) corpus\[pandey2018phonetically\]consists of 6,941\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1more phonetically balanced sentences were added after the original publication, which reports 6,126 utterancesphonetically balanced read\-speech utterances recorded at IIIT\-Hyderabad\. Textual prompts for this corpora were sampled from selected sections \(LifeStyle, Technology, Sports\) of a leading national newspaper, Dainik Bhaskar\. The dataset was recorded by 113 speakers \(58 male, and 55 female\), whose L1 was Hindi and who were educated in English medium schools\. For the purpose of analysis, Hindi and English tags were manually assigned to each word\-type in the corpus\.

\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 1:Distributions of Hindi\-English words and phonemes\.HindiEnglishword \(types\)4,7903,754word \(tokens\)54,96118,839phoneme \(types\)7352phoneme \(tokens\)194,67297,137
### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.2Lexical Refinement and Phone\-set Standardization

To address the inherent limitations of off\-the\-shelf G2P systems for Hindi\-English code\-mixed data, we implemented a multi\-stage refinement pipeline to generate a high\-fidelity lexicon:

1. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\.Bilingual Phoneme Mapping:We harmonized the disparate phonetic outputs by mapping British English \(\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=Englisheng\-uk\) G2P results onto a Hindi\-proximate phonetic space\. This ensured that English words in both Roman and Devanagri scripts shared a consistent acoustic representation\.
2. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.Syllabic De\-noising:We manually addressed systematic G2P errors, specifically the over\-insertion of schwas caused by incorrect syllabification rules in the\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglisheSpeakmodel\.
3. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.Nasal Disambiguation:To resolve the confusion between nasalized vowels and nasal consonants, we implemented phonologically conditioned nasal insertion \(e\.g\., inserting /m/ before bilabials and /n/ before dentals\)\.
4. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.Phonological Bootstrapping:For ambiguous cases where orthography was under\-specified \(e\.g\., /ph/ vs /f/ and // vs /z/\), we utilized bootstrapping to map voiced/aspirated segments to more robust voiceless proxies, significantly reducing alignment drift\. More details are presented in Section 4\.1\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.3Gold\-standard pronunciation and alignment data

For Experiment 1, 100 words for each of the free\-variation categories were hand\-annotated\. Similarly for Experiment 2, 10% of code\-mixed English words were selected for creating gold\-standard annotations\. In other words, 6 utterances out of the 62 utterances spoken by each speaker were randomly selected\. Code\-mixed words found in those utterances were first hand\-annotated by one fluent Hindi speaker \(second author\), and cross\-checked by another native Hindi speaker \(first author\)\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4Methods

This section gives a detailed description of the experimental procedure for both the experiments\. In the first experiment, we describe the handling of the pronunciation variants in the /ph\-f/ and /\-z/ context\. In the second experiment, we describe the approach towards selecting the most suitable acoustic model for forced\-alignment of English words\. The code for these experiments is available\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2https://github\.com/Ayushi113/mfa\-hindi\-code\-mixed\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.1Experiment 1

Montreal Forced Aligner \(MFA\) v1\.0\(a Kaldi\-based GMM\-HMM system using fMLLR\) was employed throughout the experiments\. Although the model is relatively less modern, its sufficiency of has been attested in recent works on Forced Alignment\[rousso2024tradition\]\. This architecture was tested with five bootstrapping configurations to identify the optimal proxy phones for two pairs of free variation:

1. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1\.Model 1 \(No Mapping\):This serves as the control model where the lexicon remains unedited\. Forced alignment is performed using raw G2P output, treating \[ph, f, , z\] as four distinct acoustic units\.
2. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2\.Model 2 \(Majority Baseline\):This model maps all instances of a variation to the single most frequent realization observed in the specific speaker’s data, testing if a speaker\-dependent dominant phone can override orthographic errors\.
3. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.Model 3 \(Dominant Baseline Mapping\):This version targets the primary native Hindi counterparts\. It maps the aspirated \[ph\] to \[p\] \(unaspirated\) or the voiced affricate \[\] to \[c\] \(voiceless\), aiming to align the speech to the most structurally similar native phone\.
4. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.Model 4 \(Fricative/Sibilant Mapping\):This model maps variants to their closest fricative or sibilant substitutes \(e\.g\., \[f\] or \[z\] to \[s\]\), testing if the aligner performs better when targets are reduced to a generic \[\+strident\] feature\.
5. \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.Model 5 \(Maximum Bootstrapping\):The most aggressive reduction strategy, where both variants in a pair are mapped to two different proxy phones \(e\.g\., to c and z to s\)\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.2Experiment 2

Three acoustic models were trained and were used to force\-align the hand\-annotated code\-mixed English words from the corpus as described in Section[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3\.3](https://arxiv.org/html/2607.25581#S3.SS3)\.

- \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English•Full dataset:All 6,941 sentences were used to train and align the sentence\-level in the corpus\. Acoustic\-phonetic transcriptions for code\-mixed English words were generated within the utterance\. This experiment aimed to generate phonemic boundaries for code\-mixed English words using naturally occurring Hindi\-English sentence level data from the corpus\.
- \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English•Hindi monolingual chunks:Using the TextGridTools\[buschmeier2013textgridtools\]package, we separated the audio data and TextGrids into contiguous Hindi phrase\-level and English word\-level chunks\. English words with Hindi inflections \(for example: “amerik\-i”\) were excluded from the analysis\. This sub\-experiment aimed to evaluate the accuracy of monolingual acoustic data on aligning code\-mixed English words from the same corpus\.
- \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English•English words: Code\-mixed English words extracted in the previous step were used to generate phoneme level alignment\. This sub\-experiment aimed to examine the efficacy of word\-level English data on aligning itself, with similar training and testing environments\.

![Refer to caption](https://arxiv.org/html/2607.25581v1/x1.png)

![Refer to caption](https://arxiv.org/html/2607.25581v1/x2.png)

![Refer to caption](https://arxiv.org/html/2607.25581v1/x3.png)

\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishFigure 1:Absolute difference between midpoints of phonemes in three models: Code\-mixed \(left\), Hindi \(middle\) and English \(right\)

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5Results and analysis

In this section, we present the results obtained from the two experiments\. The best performing model in Experiment 1 was chosen on the basis of the highest obtained F\-score\. For Experiment 2, the best training environment was chosen on the basis of absolute error in the midpoint values, when measured against the gold\-standard annotations\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.1Experiment 1

#### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.1\.1Words with the \[ph\]∼\\sim\[f\] variation

\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 2:Performance of five lexicon mapping strategies for resolving ph\-f variation against a gold\-standard transcript\.Table 2 describes the diagonal accuracies of the confusion matrix between the gold\-standard phoneme labels, and those predicted by different variations of the MFA\. The results in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English2](https://arxiv.org/html/2607.25581#S5.T2)reveal a discrepancy between orthographic representation and acoustic reality in Hindi\-English code\-mixed speech\. While the No mapping column shows an F\-score of only 0\.72, the gold\-standard data reveals that the speakers in this demographic almost unanimously produced the fricative /f/, despite the “nukta\-less” Hindi script suggesting the aspirated stop /ph/\. This mismatch causes the MFA to trigger frequent alignment errors when forced to use the\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglisheSpeakdefault\. Interestingly, the data suggests that for these specific code\-mixed terms, the expected “free variation” is absent in actual production; the speakers in the PBCM corpus have stabilized on the /f/ variant\.

As seen in the ph→ p column, simply removing aspiration provides a massive boost to an F\-score of 0\.97\. This works because reverting to a plain labio\-dental or bilabial stop creates a sufficiently distinct acoustic category from the fricative /f/, allowing the aligner to better distinguish the segments\. In contrast, the f → s mapping fails , as the sibilant /s/ is acoustically too distant from the target\. Given that the Majority baseline hits a perfect 1\.0, the most robust recommendation from this data is that for code\-mixed Hindi\-English corpora, a direct lexicon mapping to /f/, or at minimum, de\-aspirating to /p/, is necessary to correct for the orthographic “nukta” deficit\.

#### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.1\.2Words with the \[\]∼\\sim\[z\] variation

\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 3:Performance of five lexicon mapping strategies for resolving \-z variation against a gold\-standard transcript\.The results in Table[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English3](https://arxiv.org/html/2607.25581#S5.T3)for the \-z variation reflect a more complex alignment task than the previous set, as the ground truth contains both the voiced affricate and the voiced sibilant\. The No mapping baseline fails significantly with an F\-score of 0\.16, indicating that the default G2P’s singular output is incompatible with the speakers’ actual productions\. Even the Majority baseline, which defaults to the most frequent variant in the gold standard, only achieves an F\-score of 0\.56, highlighting that a single\-label approach cannot sufficiently capture the variation present in words like “Rajput” or “Resistance”\.

As shown in the table, providing the aligner with multiple pronunciation variants, allowing a “competition” between labels, yields a substantial performance increase\. While the individual mappings → c and z → s show incremental gains, the maximum bootstrapping approach \( → c and z → s\) proves most effective with an F\-score of 0\.74 and a Precision of 0\.84\. These results suggest that by mapping both voiced targets to their voiceless counterparts in the lexicon, the MFA can more reliably identify the phoneme boundaries and recover the intended labels from the acoustic signal\.

### \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English5\.2Experiment 2

The results of Experiment 1 gave us a variant\-free lexicon, where different pronunciation variants were encoded as different lexical items\. This design reduced ambiguity and enabled a clearer investigation of RQ2 by preventing free\-variation errors from propagating to the subsequent experiment\. As discussed in Section[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4\.2](https://arxiv.org/html/2607.25581#S4.SS2), there were three different types of acoustic models for forced alignment\. For evaluation of the obtained TextGrids, we had 4 datasets: a\) the hand\-annotated boundaries, and forced\-aligned boundaries from b\) the code\-mixed sentence level data, c\) the phrase\-level monolingual Hindi chunks, and d\) the word\-level English only chunks\. For every phoneme, the timestamp of themidpointof the left and right boundaries was extracted, for each of the forced\-alignment conditions\. Then, the midpoint was compared against the gold\-standard midpoint, and absolute errors were computed\. Figure[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English1](https://arxiv.org/html/2607.25581#S4.F1)displays the comparative distribution of the absolute errors for each of the forced\-aligned conditions\. We observe that the code\-mixed sentence level alignment outperform the two monolingual conditions\.

The mean absolute error by the code\-mixed sentence level model was 4\.15 ms, which is 10 times lower than the monolingual Hindi \(38\.18 ms\) and English models \(37\.58 ms\)\.

\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=EnglishTable 4:Comparison of tolerance \(in msec\) of the three modelsTable[\\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English4](https://arxiv.org/html/2607.25581#S5.T4)displays the coverage of phonemes under each training environment, for different levels of error tolerance\. For example, when trained with code\-mixed sentences, 87\.06% phonemes were found to comparable with the gold\-standard annotations, with an error of <10 ms\. As can clearly be seen \(in a row\-wise/model comparison\), the model trained on code\-mixed sentences shows the highest percentage of phonemes in the lower error \(<10 ms\) range\. A closer look reveals that when monolingual Hindi chunks are compared with word\-level English as training, the former outperforms the latter only at low tolerance \(<10/20\), but then the pattern reverses from 30ms onwards\. This indicates that while on average, many tokens of Hindi models were better aligned \(e\.g\. 7% of the phones were better aligned at 10% tolerance\), it nonetheless contains more tokens that are particularly poorly aligned \(\>30ms errors\)\. In other words, the word\-level English model has on average 1\-2% fewer tokens with errors of \>30ms, than the monolingual Hindi model\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English6Discussion & conclusion

In this paper, we addressed the problem of free variation, and orthographic inconsistencies in code\-mixed Hindi\-English as a challenge for forced alignment\. We found that bootstrapping techniques, used commonly in low\-resource languages are useful for addressing such variation\. Next, we identified that the use of an acoustic model trained on code\-mixed sentences is most suitable for accurate alignment of the test corpus in code\-mixed speech\. An important observation is that code\-mixed data, even at the word\-level is more useful than a larger monolingual corpus of Hindi\.

Our work is among the first to present a structured analysis of computational tools for a phonological analysis of code\-mixing in Indian languages, and extends previous research\[pandey2020understanding,ahn\-etal\-2025\-automatic,amazouz2019exploring\]\. We have shown that speakers compensate for the inconsistency in orthography caused by the missing diacritic \(”nukta”\)\. In the ph\-f variation, most speakers of our dataset defaulted to the /f/, and did not use the aspirated stop\. The stop is more characteristic of a vernacular expression, and it is possible that in read speech, its manifestation is excluded\. However, the words present a unique challenge where speaker compensation does not lead to a single majority variant\. In words where the script omits the nukta, speakers exercise their linguistic competence to produce both and z depending on the intended lexical target\. This necessitates the use of maximum bootstrapping \( → c and z → s\), which provides the aligner with the phonetic flexibility to recover both variants where a simple lexicon correction cannot\. This confirms that while these speakers possess stable phonemic categories, the “invisible” nukta in the script triggers a split in production\. This problem is addressed in this paper through bootstrapping\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English7Limitation

At the time the experiments were conducted, an older version of Montreal Forced Aligner \(MFA\) was used\. Repeating the experiments with a recent version of MFA \(e\.g\., v3\.x\) may provide additional insights\. In Experiment 1, pronunciation probabilities could be explicitly defined and conditioned on speaker demographics \(e\.g\., education level\), and pretrained G2P models for Hindi, Indian English, and British English could be incorporated\. In Experiment 2, future work could experiment with the pretrained English acoustic model \(v3\.1\.0\), trained on multiple English varieties, including Indian and British English\.

## \\fontspec\_if\_language:nTFENG\\addfontfeatureLanguage=English8Generative AI Use Disclosure

Generative AI was used only for editing and polishing manuscripts, and not for producing a significant part of the manuscript\.

## References

Similar Articles

Towards a Phonology-Informed Evaluation of Multilingual TTS

arXiv cs.CL

This paper proposes a classifier-based framework to audit multilingual TTS systems for phonological faithfulness, using Assamese ATR vowel harmony as a case study. It reveals that Meta's MMS TTS frequently misproduces advanced tongue root vowels, a bias absent in human speech.