基于文本约束声学重评分的无需训练发音转录
摘要
本文提出一种无需训练的发音转录流程,通过集成词法资源和声学模型,实现低错误率和高效率,性能优于包括商用多模态LLMs在内的基线方法。
arXiv:2609.30924v1 Announce Type: new
Abstract: Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.
查看缓存全文
缓存时间: 2026/09/28 09:45
# Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
Source: [https://arxiv.org/html/2609.30924](https://arxiv.org/html/2609.30924)
\\CJKencfamily
UTF8ipxm\\CJK@envStartUTF8
## Training\-Free Pronunciation Transcription via Text\-Constrained Acoustic RescoringThanks:This work was done during Hikaru Asano’s internship at Sakana AI\.
###### Abstract
Accurate and efficient pronunciation transcription is essential for preparing text\-to\-speech training data at scale\. Existing approaches have different limitations: grapheme\-to\-pronunciation \(G2P\) and speech\-to\-pronunciation \(S2P\) methods each capture only partial information, using only text or only speech, while speech\-and\-text\-to\-pronunciation \(ST2P\) methods use both but require costly pronunciation\-annotated data\. To address this problem, we propose a training\-free ST2P pipeline that integrates both lexical and acoustic information at inference time\. Lexical resources and G2P tools generate text\-constrained candidates, and a left\-to\-right greedy search selects the best one using whole\-sequence negative log\-likelihoods from frozen pretrained S2P models\. On three Japanese corpora, our method reduces Character Error Rate \(CER\) from 0\.60–1\.40% \(text\-only baseline\) to 0\.04–0\.17% with reference transcripts, and 0\.64–1\.58% with ASR transcripts\. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs\. Our greedy search method is 3–3\.5×\\timesfaster than beam search at similar CER, and the cascade is 2×\\timesfaster than direct decoding ensuring the efficiency and accuracy\. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods\.
###### Index Terms:
pronunciation transcription, training\-free inference, CTC rescoring, grapheme\-to\-phoneme conversion
††address:1The University of Tokyo2Sakana AI## 1Introduction
Pronunciation transcription is a technique that recovers pronunciation symbols from speech signals or transcribed texts\[[1](https://arxiv.org/html/2609.30924#bib.bib1),[2](https://arxiv.org/html/2609.30924#bib.bib2)\]\. It is widely used in diverse applications including training data preparation for text\-to\-speech \(TTS\) models, language education, and study of low\-resourced languages\[[3](https://arxiv.org/html/2609.30924#bib.bib3),[4](https://arxiv.org/html/2609.30924#bib.bib4)\]\. In particular, keeping pronunciation transcription inexpensive while maintaining low error rates is crucial for preparing TTS training data at scale\[[5](https://arxiv.org/html/2609.30924#bib.bib5),[6](https://arxiv.org/html/2609.30924#bib.bib6)\]\.
Figure 1:Processing cost versus pronunciation error rates on JVS\-dev\. The horizontal axis is the real\-time factor \(log scale; lower is faster\) and the vertical axis is character error rate \(CER\) \(log scale; lower is better\)\. The proposed methods \(ours\) achieves the lowest CER at low computational cost\.Two conventional approaches address this task using different inputs\. Grapheme\-to\-pronunciation \(G2P\) conversion\[[7](https://arxiv.org/html/2609.30924#bib.bib7),[8](https://arxiv.org/html/2609.30924#bib.bib8)\]predicts a pronunciation sequenceyyfrom an orthographic transcriptww\. Speech\-to\-pronunciation \(S2P\) conversion\[[9](https://arxiv.org/html/2609.30924#bib.bib9),[10](https://arxiv.org/html/2609.30924#bib.bib10)\]instead predicts pronunciation from a speech signalxx\. For the application of TTS data preparation, it is common that both graphemes and speech signals are available for pronunciation prediction\. However, both G2P and S2P approaches are insufficient because they discard information from the other clue for pronunciation transcription\.
Speech\-and\-text\-to\-pronunciation \(ST2P\) models address this limitation by jointly using speech and text to predict pronunciation, corresponding to the mapping\(x,w\)→y\(x,w\)\\to y\. Related approaches incorporate text through additional text embeddings in an S2P model\[[11](https://arxiv.org/html/2609.30924#bib.bib11)\]or introduce textual supervision through multitask training that combines speech\-to\-text and S2P objectives\[[12](https://arxiv.org/html/2609.30924#bib.bib12)\]\. Despite their effectiveness, they require pronunciation\-annotated training data, which is typically more costly than ordinary text transcription and limits scalability\[[13](https://arxiv.org/html/2609.30924#bib.bib13)\]\.
In this paper, we instead explore a “training\-free” ST2P pipeline that leverages existing lexical resources and G2P tools, together with open pretrained S2P models\. The key idea is to generate text\-constrained pronunciation candidates forwwfrom lexical resources and G2P tools, and to select the best one by greedy search using the acoustic scores of frozen S2P models onxx\(Fig\.[2](https://arxiv.org/html/2609.30924#S1.F2)\)\. Both clues are thus integrated purely at inference time: the pipeline requires neither audio–text–pronunciation\(x,w,y\)\(x,w,y\)triplets nor any model training, and is therefore inexpensive while achieving low transcription error rates\.
Our contributions are: \(i\) We propose a training\-free ST2P pipeline that combines lexical candidates with frozen S2P models \(Sec\.[3](https://arxiv.org/html/2609.30924#S3)\)\. \(ii\) On three Japanese corpora, our method achieves lower character error rates \(CER\) than all evaluated baselines, using both reference and ASR transcripts, while maintaining fast inference as shown in Fig\.[1](https://arxiv.org/html/2609.30924#S1.F1)\(Sec\.[4\.2](https://arxiv.org/html/2609.30924#S4.SS2),[4\.3](https://arxiv.org/html/2609.30924#S4.SS3)\)\. \(iii\) We also demonstrate improvements on Spanish, French, and preliminary English, outperforming four open MLLMs and text\-only specialists with the best configuration for each language \(Sec\.[4\.4](https://arxiv.org/html/2609.30924#S4.SS4)\)\.
Figure 2:Training\-free pronunciation transcription\. \(a\) Dictionaries, G2P, and rules generate readings for transcript spans\. Complete pronunciationsy\(𝐫\)y\(\\mathbf\{r\}\)are converted to model\-specific tokens and evaluated by a frozen S2P model given speechxx, combining lexical constraints and acoustic evidence through negative log\-likelihood \(NLL\)\. \(b\) Left\-to\-right greedy search for “read the record” \(P=0P=0; illustrative NLLs\)\. Each node represents a complete pronunciation with earlier choices fixed and later spans at their defaults\. The lowest\-NLL candidate \(green\) is committed; gray nodes are rejected, and dashed outlines mark span defaults\. The final node is the search output; cascade rescoring and the subsequent margin check are not shown\.
## 2Task and training\-free setting
Task\.Given speechxxand its transcriptww, we estimate the pronunciation sequenceyyactually realized inxx, rather than its canonical reading\. The transcript may come from reference text or ASR output; the latter enables an automatic pipeline\. The pronunciation alphabet matches the language and available resources \(e\.g\., kana for Japanese, IPA otherwise\)\.
## 3Method
Fig\.[2](https://arxiv.org/html/2609.30924#S1.F2)shows the training\-free ST2P pipeline\. Candidate pronunciationsyyfor transcriptwware generated from lexical resources and scored by frozen S2P models to yield a single acoustic score integrating acoustic and lexical cues \(Fig\.[2](https://arxiv.org/html/2609.30924#S1.F2)\(a\)\)\. The best combination of pronunciation candidates is found by a left\-to\-right greedy search minimizing negative log\-likelihood \(NLL\) under these models given the speechxx\(Fig\.[2](https://arxiv.org/html/2609.30924#S1.F2)\(b\)\)\.
### 3\.1Candidates and frozen acoustic scores
As in Fig\.[2](https://arxiv.org/html/2609.30924#S1.F2)\(a\), the transcriptwwis segmented intoNNspans, e\.g\., by morphological analysis\. For each spannn, lexical resources such as dictionaries, G2P, or rules provide a finite setℛn\\mathcal\{R\}\_\{n\}of candidate readings\. One of them, the default readingrn0∈ℛnr\_\{n\}^\{0\}\\in\\mathcal\{R\}\_\{n\}, is the most probable reading predicted from the text alone\. Choosing one readingrn∈ℛnr\_\{n\}\\in\\mathcal\{R\}\_\{n\}for every span gives an assignment𝐫=\(r1,…,rN\)∈ℛ=∏nℛn\\mathbf\{r\}=\(r\_\{1\},\\ldots,r\_\{N\}\)\\in\\mathcal\{R\}=\\prod\_\{n\}\\mathcal\{R\}\_\{n\}, whose complete pronunciation is the concatenation
y\(𝐫\)=r1r2⋯rN\.y\(\\mathbf\{r\}\)=r\_\{1\}\\,r\_\{2\}\\cdots r\_\{N\}\.\(1\)Instead of scoring each span separately, we evaluatey\(𝐫\)y\(\\mathbf\{r\}\)as a whole on the audioxx, jointly judging all readings in context, and seek
𝐫^=argmin𝐫∈ℛJ\(𝐫,x\),J\(𝐫,x\)=S\(y\(𝐫\),x\)\+P\(𝐫\),\\hat\{\\mathbf\{r\}\}=\\operatorname\*\{arg\\,min\}\_\{\\mathbf\{r\}\\in\\mathcal\{R\}\}J\(\\mathbf\{r\},x\),\\quad J\(\\mathbf\{r\},x\)=S\(y\(\\mathbf\{r\}\),x\)\+P\(\\mathbf\{r\}\),\(2\)whereSSis the acoustic score, and the optional penaltyPPfavors more reliable candidates in ambiguous cases\. The acoustic score is the NLL of the candidate under a frozen S2P model,
S\(y,x\)=−logp\(z\(y\)∣x\),S\(y,x\)=\-\\log p\(z\(y\)\\mid x\),\(3\)wherezzrendersyyinto the alphabet and tokens of the model\.
### 3\.2Greedy search and cascade
As\|ℛ\|\|\\mathcal\{R\}\|grows exponentially withNN, we approximate the objective in Eq\. \([2](https://arxiv.org/html/2609.30924#S3.E2)\) by a left\-to\-right greedy search \(Fig\.[2](https://arxiv.org/html/2609.30924#S1.F2)\(b\)\), assuming each span’s acoustic evidence is mostly local\. Let𝐫\[n←r\]\\mathbf\{r\}\[n\\\!\\leftarrow\\\!r\]denote𝐫\\mathbf\{r\}with itsnn\-th span replaced byrr\. Starting from𝐫^\(0\)=𝐫0=\(r10,…,rN0\)\\hat\{\\mathbf\{r\}\}^\{\(0\)\}=\\mathbf\{r\}^\{0\}=\(r\_\{1\}^\{0\},\\ldots,r\_\{N\}^\{0\}\), we visit each span once\. At spannn, we select the best candidate and commit it:
rn′=argminr∈ℛnJ\(𝐫^\(n−1\)\[n←r\],x\),𝐫^\(n\)=𝐫^\(n−1\)\[n←rn′\],r\_\{n\}^\{\\prime\}=\\operatorname\*\{arg\\min\}\_\{r\\in\\mathcal\{R\}\_\{n\}\}J\(\\hat\{\\mathbf\{r\}\}^\{\(n\-1\)\}\[n\\\!\\leftarrow\\\!r\],x\),\\;\\;\\hat\{\\mathbf\{r\}\}^\{\(n\)\}=\\hat\{\\mathbf\{r\}\}^\{\(n\-1\)\}\[n\\\!\\leftarrow\\\!r\_\{n\}^\{\\prime\}\],\(4\)so that𝐫^\(n\)=\(r1′,…,rn′,rn\+10,…,rN0\)\\hat\{\\mathbf\{r\}\}^\{\(n\)\}=\(r\_\{1\}^\{\\prime\},\\ldots,r\_\{n\}^\{\\prime\},r\_\{n\+1\}^\{0\},\\ldots,r\_\{N\}^\{0\}\)\. Thus, the\|ℛn\|\|\\mathcal\{R\}\_\{n\}\|candidates are scored as complete sequences, with earlier choices fixed and later spans at their defaults\. The final assignment is𝐫^=𝐫^\(N\)\\hat\{\\mathbf\{r\}\}=\\hat\{\\mathbf\{r\}\}^\{\(N\)\}\.
Second scorer and cascade\.For added robustness, we can combine two frozen S2P models using the fused scoreS=L1\+λL2S=L\_\{1\}\+\\lambda L\_\{2\}, whereLiL\_\{i\}is the NLL from modeliiandλ≥0\\lambda\\geq 0controls the second model’s weight\. To reduce computation, we use a cascade: we first run the greedy search withS=L1S=L\_\{1\}\. Ifλ\>0\\lambda\>0and this pass changes any default reading, we rerun the search from𝐫0\\mathbf\{r\}^\{0\}with the fused score; otherwise, we retain the first\-pass result\. We take the last pass’s output as𝐫^\\hat\{\\mathbf\{r\}\}and use its acoustic scoring functionSSin the margin check below\.
### 3\.3Margin check
The search may still leave the default on weak evidence\. We therefore initialize𝐫~=𝐫^\\tilde\{\\mathbf\{r\}\}=\\hat\{\\mathbf\{r\}\}and revisit the changed spans from left to right\. With a hyperparameterτ≥0\\tau\\geq 0, at each spanjjwe keep the current reading only if its score advantage over the default meets a length\-scaled threshold:
S\(y\(𝐫~\[j←rj0\]\),x\)−S\(y\(𝐫~\),x\)≥τℓj,S\\big\(y\(\\tilde\{\\mathbf\{r\}\}\[j\\\!\\leftarrow\\\!r\_\{j\}^\{0\}\]\),x\\big\)\-S\\big\(y\(\\tilde\{\\mathbf\{r\}\}\),x\\big\)\\geq\\tau\\,\\ell\_\{j\},\(5\)whereℓj=max\(1,\|r~j\|,\|rj0\|\)\\ell\_\{j\}=\\max\(1,\|\\tilde\{r\}\_\{j\}\|,\|r\_\{j\}^\{0\}\|\), with\|r\|\|r\|counting symbols in the normalized pronunciation alphabet; thus,τ\\tauis the required score gain per symbol\. If this condition is not met, the span is reverted,𝐫~←𝐫~\[j←rj0\]\\tilde\{\\mathbf\{r\}\}\\leftarrow\\tilde\{\\mathbf\{r\}\}\[j\\\!\\leftarrow\\\!r\_\{j\}^\{0\}\], before the next span is checked\. Settingτ=0\\tau=0disables this check\.
## 4Experiments
### 4\.1Japanese resources and protocol
Japanese is a challenging language for acoustic disambiguation: one orthographic string can have several valid readings \(ipxm明日:*asu*,*ashita*, or*myōnichi*\), and the goal is to select the reading actually spoken in the recording\.
Reading candidates\.For reading candidate generation, we use Japanese morphological analyzers and G2P tools such as Sudachi\[[14](https://arxiv.org/html/2609.30924#bib.bib14)\], Open JTalk111[https://open\-jtalk\.sp\.nitech\.ac\.jp/](https://open-jtalk.sp.nitech.ac.jp/), MeCab/UniDic\[[15](https://arxiv.org/html/2609.30924#bib.bib15)\], and rule\-based variants\. Readings found only by isolated\-span analysis are less reliable, so the penaltyPPin Eq\. \([2](https://arxiv.org/html/2609.30924#S3.E2)\) is set to0\.03mnctc0\.03\\,m\\,n\_\{\\mathrm\{ctc\}\}, wheremmcounts spans switched from the default to such a reading andnctcn\_\{\\mathrm\{ctc\}\}is the utterance length in token count of the main scorerL1L\_\{1\}\.
Scorers\.We use a 320M\-parameter wav2vec2\.0 kana CTC model\[[16](https://arxiv.org/html/2609.30924#bib.bib16)\]as the first scorer and an 810M\-parameter autoregressive kana\-whisper\[[6](https://arxiv.org/html/2609.30924#bib.bib6)\]as the second\. Kana\-whisper allows multiple spellings for a pronunciation \(e\.g\., long vowels, ipxmは/ipxmわ\), soL2L\_\{2\}is the negative logarithm of the teacher\-forced likelihoods summed over up to 16 spellings, caching cross\-attention per recording\.
Hyperparameters\.We use second scorer weightλ=1\\lambda=1\(from\{0\.5,1,2\}\\\{0\.5,1,2\\\}\) and margin check thresholdτ=0\.6\\tau=0\.6\(see Eq\.[5](https://arxiv.org/html/2609.30924#S3.E5)\), with the margin check using the finalSSandℓj\\ell\_\{j\}in normalized kana\. For calibration, we use 1,000 utterances each from JSUT and JVS\-par, transferred to JVS\-dev; the spelling limit is set on JVS\-dev\. Reading candidate and prompt settings also use these calibration subsets, which are excluded from evaluation\.
Data and transcripts\.We use JVS\-dev \(600 nonparallel utterances, 20 speakers, manual kana\)\[[17](https://arxiv.org/html/2609.30924#bib.bib17)\], JSUT BASIC5000 \(5,000 utterances, one speaker, verified labels\)\[[18](https://arxiv.org/html/2609.30924#bib.bib18)\], and JVS\-par \(9,997 utterances, 100 speakers, alignment\-derived silver labels\)222[https://github\.com/r9y9/jvs\_r9y9](https://github.com/r9y9/jvs_r9y9), evaluating all systems on 512 common utterances per corpus\. Transcripts are corpus text \(reference\) or Whisper large\-v3 output\[[19](https://arxiv.org/html/2609.30924#bib.bib19)\]\(ASR\)\.
Baselines and references\(Tab\.[1](https://arxiv.org/html/2609.30924#S4.T1)\)\. We compare five categories of baselines:*\(A\) Text\-only G2P:*Open JTalk, Sudachi \+ Open JTalk, and Sudachi \+ Open JTalk with spoken\-form rules \(our default, used as the control for acoustic selection\)\.*\(B\) Audio\-only decoding:*greedy Kana CTC and kana\-whisper with the same checkpoints as our scorers\.*\(C\) Same\-model text conditioning:*kana\-whisper with a transcript prompt, and Kana CTC with dictionary\-constrained greedy decoding inspired by\[[11](https://arxiv.org/html/2609.30924#bib.bib11)\], with and without a long\-vowel lattice\.*\(D\) Released ST2P annotator:*Furigana Whisper333[https://huggingface\.co/Parakeet\-Inc/furigana\_whisper\_small\_jsut](https://huggingface.co/Parakeet-Inc/furigana_whisper_small_jsut), a public audio\-and\-text kana model, with its published decoding and with the same dictionary constraint over our candidates \(not a reproduction of the unavailable model of\[[11](https://arxiv.org/html/2609.30924#bib.bib11)\]\); JSUT results are omitted because it was trained on JSUT\.*\(E\) Multimodal LLMs \(MLLMs\):*Qwen3\-Omni\-30B\-A3B, Qwen2\.5\-Omni\-7B, Gemma\-3n\-E4B, and Phi\-4\-multimodal\[[20](https://arxiv.org/html/2609.30924#bib.bib20),[21](https://arxiv.org/html/2609.30924#bib.bib21),[22](https://arxiv.org/html/2609.30924#bib.bib22)\]\. We report Gemini 2\.5/3/3\.5/3\.6 Flash\[[23](https://arxiv.org/html/2609.30924#bib.bib23)\]as commercial references, which are not highlighted for best scores due to their high labeling cost\. All MLLMs are given the same audio, transcripts, three fixed examples, prompt tuned by validation, and one output repair\.
Metric\.We report corpus\-level character error rate \(CER, %\) after normalization \(katakana, punctuation, long vowels, fixed mappings\) to ignore notational differences\. CER is computed on normalized kana, or on silver labels for JVS\-par\.
Table 1:Japanese CER \(%,↓\\downarrow\) on 512 common utterances per corpus\. dev/par: JVS\-dev/JVS\-par\.†\\dagger: dictionary\-constrained decoding inspired by\[[11](https://arxiv.org/html/2609.30924#bib.bib11)\]\. Furigana Whisper’s JSUT cells are omitted \(training data\)\. Gray: commercial references\. Bold: column minima excluding commercial models\.
### 4\.2Japanese results
Main results \(Tab\.[1](https://arxiv.org/html/2609.30924#S4.T1)\)\.Our cascade outperforms all baselines in both the reference\-text and ASR\-text settings, including ST2P Furigana Whisper with dictionary constraint, which requires pronunciation\-labeled speech; ours uses only frozen components\. It also substantially improves over greedy Kana CTC decoding \(5\.38/5\.61/13\.02 to 0\.154/0\.171/0\.042\) and the text\-only default \(0\.85/1\.40/0\.60\), showing that combining S2P scores and G2P candidates is more effective than either alone\. The cascade also outperforms CTC\-only and AR\-only rescoring on all corpora\.
Ablations \(Tab\.[2](https://arxiv.org/html/2609.30924#S4.T2)\)\.We validate each component by removing the margin check \(w/o gate,τ=0\\tau=0\), removing the cascade \(w/o cascade, fusing every utterance\), using beam search with widthB=5B=5\(w/ beam\), or replacing the input audio with unrelated or silent audio \(other audio, silent audio\)\. For reference, we also report a greedy oracle that selects candidates with minimum edit distance to the reference reading, an approximate oracle within the candidate set\.
Tab\.[2](https://arxiv.org/html/2609.30924#S4.T2)shows that removing the gate or cascade raises CER by up to 0\.022 and 0\.016, indicating that both components are effective\. Replacing audio with other recordings or silence also increases CER to 12\.4–16\.2%, confirming that acoustic evidence is necessary\. Finally, beam search performs similarly to greedy search, supporting that each span’s acoustic evidence is mostly local\. We thus prefer greedy search for its lower cost \(Sec\.[4\.3](https://arxiv.org/html/2609.30924#S4.SS3)\)\.
The cascade is only 0\.107/0\.050/0\.042 points above the oracle, recovering 87–96% of the oracle’s CER reduction over text\-only\.
Table 2:Japanese ablations \(CER %, reference text, 512 utterances\)\. Proposed: cascade, gate \(τ=0\.6\\tau=0\.6\), greedy search\. w/o gate:τ=0\\tau=0; w/o cascade: fusion every utterance; w/ beam:B=5B=5\. Other/silent audio: use unrelated or silent audio, keeping text/candidates\.
### 4\.3Cost and accuracy \(Fig\.[1](https://arxiv.org/html/2609.30924#S1.F1)\)
We measured real\-time factor \(RTF\) for all audio\-based systems on 512 JVS\-dev utterances with reference transcripts, using a single H100 GPU \(batch size one\)\.
Fig\.[1](https://arxiv.org/html/2609.30924#S1.F1)plots RTF against CER: the proposed methods achieve the lowest CER while remaining comparatively inexpensive\. Greedy search is 3–3\.5×\\timesfaster than beam search at comparable CER \(Tab\.[2](https://arxiv.org/html/2609.30924#S4.T2)\)\. The cascade method is also twice as fast as kana\-whisper decoding because it runs kana\-whisper only on selected spans and scores G2P candidates with a single forward pass instead of stepwise decoding\.
### 4\.4Multilingual evaluation
We test transfer beyond Japanese with reference transcripts, IPA candidates, and language\-specific resources\.
Candidates and scorers\.Spanish and French candidates come from CharsiuG2P\[[7](https://arxiv.org/html/2609.30924#bib.bib7)\]with eSpeak NG and Epitran as fallbacks; English uses the MFAenglish\_us\_arpadictionary\[[24](https://arxiv.org/html/2609.30924#bib.bib24)\]\. Each pool is extended to up to 16 candidates with additional G2P outputs and local substitution or deletion of phonetically similar segments to cover unlisted variants\. PhoneticXeus\[[9](https://arxiv.org/html/2609.30924#bib.bib9)\]is the main scorer, and fusion adds the POWSM CTC head\[[10](https://arxiv.org/html/2609.30924#bib.bib10)\], both scored via Eq\.[3](https://arxiv.org/html/2609.30924#S3.E3)\.
We reuse the Japanese settings \(Sec\.[4\.1](https://arxiv.org/html/2609.30924#S4.SS1)\): greedy search, cascade withλ=1\\lambda=1, and the margin check \(Sec\.[3\.3](https://arxiv.org/html/2609.30924#S3.SS3)\) withℓj\\ell\_\{j\}counted in normalized IPA segments\. The two main changes are:P=0P=0, since these G2P and dictionary candidates contain no isolated\-span readings that the Japanese penalty targets; andτ\\tauis set per language \(2\.0 for Spanish, 1\.5 for French, 0 for English\) according to each development set to minimize PFER\.
Data\.We evaluate Spanish DIMEx100\[[25](https://arxiv.org/html/2609.30924#bib.bib25)\], French Rhapsodie\[[26](https://arxiv.org/html/2609.30924#bib.bib26)\], and English Buckeye\[[27](https://arxiv.org/html/2609.30924#bib.bib27)\], sampling 200 utterances per language that contain at least one variable candidate span \(192 for French after removing utterances without audio\)\. Spanish and French use six speaker groups each for evaluation and configuration; English uses four groups each, all drawn from the corpus’s designated development speakers, so its results are preliminary\.
Baselines and references\(Tab\.[3](https://arxiv.org/html/2609.30924#S4.T3)\)\.*\(A\) Text\-only specialist:*the default reading from the language’s dictionary or G2P\.*\(B\) Open MLLMs:*models from Sec\.[4\.1](https://arxiv.org/html/2609.30924#S4.SS1), evaluated alongside commercial references with the same audio, text, and prompting protocol\.
Metric\.We report phonological feature error rate,PFER=100∑iwidF\(yi,y^i\)/∑iwi\|yi\|\\mathrm\{PFER\}=100\\sum\\nolimits\_\{i\}w\_\{i\}\\,d\_\{F\}\(y\_\{i\},\\hat\{y\}\_\{i\}\)/\\sum\\nolimits\_\{i\}w\_\{i\}\\,\|y\_\{i\}\|, wherewiw\_\{i\}is an inverse\-sampling weight \(eligible\-to\-sampled ratio\), anddFd\_\{F\}is an edit distance using PanPhon feature distances\[[28](https://arxiv.org/html/2609.30924#bib.bib28)\]for substitutions and unit costs for insertions and deletions\. Fixed IPA mappings align conventions within each language\. PFER scores are not directly comparable across languages\.
Table 3:PFER \(%,↓\\downarrow\) on shared subsets with candidate variation\. \(A\): reference text only; others: text and audio\. Gray: commercial references\. Bold: best \(non\-commercial\)\.Results \(Tab\.[3](https://arxiv.org/html/2609.30924#S4.T3)\)\.With only a change of lexical resources and scorers, and no training, the pipeline improves on the language\-specific specialist in all three languages \(2\.47→\\to2\.38 in Spanish with PhoneticXeus\-only; 4\.39→\\to4\.11 and 13\.76→\\to13\.21 in French and English with the cascade\)\. Both configurations outperform every open MLLM\. The best configuration is within 0\.4 points of the top Gemini model in Spanish and English\. Gains depend on resource and scorer matching, so choosing resources is key\.
## 5Conclusion
We proposed a training\-free ST2P pipeline based on open models and lexical resources\. For Japanese, it is faster and outperforms all baselines\. The method also works for Spanish, French, and English\.
## References
- \[1\]A Waibel, T Hanazawa, G Hinton, K Shikano, and K J Lang,“Phoneme recognition using time\-delay neural networks,”IEEE Transactions on Acoustics, Speech, and Signal Processing, 1989\.
- \[2\]Alex Graves, Abdel\-Rahman Mohamed, and Geoffrey Hinton,“Speech recognition with deep recurrent neural networks,”in2013 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\), 2013\.
- \[3\]John Kominek and Alan W\. Black,“The CMU Arctic speech databases,”in5th ISCA Workshop on Speech Synthesis \(SSW 5\), 2004\.
- \[4\]Horacio Franco, Leonardo Neumeyer, María Ramos, and Harry Bratt,“Automatic detection of phone\-level mispronunciation for language learning,”in6th European Conference on Speech Communication and Technology, 1999\.
- \[5\]Jaehyeon Kim, Jungil Kong, and Juhee Son,“Conditional variational autoencoder with adversarial learning for end\-to\-end text\-to\-speech,”inProceedings of the 38th International Conference on Machine Learning \(ICML\), 2021\.
- \[6\]Lianbo Liu, Shiao Zhu, Kai Washizaki, Reo Yoneyama, Haesung Jeon, Mengjie Zhao, Yusuke Fujita, Hao Shi, Nao Yoshida, Yuan Gao, Roman Koshkin, Yukiya Hono, and Yui Sudo,“Sarashina2\.2\-TTS: Tackling kanji polyphony in japanese speech generation via data scaling and targeted data synthesis,”arXiv \[cs\.SD\], 2026\.
- \[7\]Jian Zhu, Cong Zhang, and David Jurgens,“ByT5 model for massively multilingual grapheme\-to\-phoneme conversion,”inProceedings of Interspeech, 2022\.
- \[8\]Tomoki Koriyama,“Benchmarking large language models for grapheme\-to\-phoneme conversion: A japanese case study,”inProceedings of Interspeech, 2026\.
- \[9\]Shikhar Bharadwaj, Chin\-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, and David R\. Mortensen,“An empirical recipe for universal phone recognition,”inProceedings of Interspeech, 2026\.
- \[10\]Chin\-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David R Mortensen, and Shinji Watanabe,“POWSM: A phonetic open whisper\-style speech foundation model,”inProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), 2026\.
- \[11\]Hien Ohnaka, Yuma Shirahata, Byeongseon Park, and Ryuichi Yamamoto,“Grapheme\-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning,”inProceedings of Interspeech, 2025, pp\. 2525–2529\.
- \[12\]Yotaro Kubo, Richard Sproat, Chihiro Taguchi, and Llion Jones,“Building tailored speech recognizers for Japanese speaking assessment,”inProceedings of Interspeech, 2026\.
- \[13\]John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, David S Pallett, and Nancy L Dahlgren,“DARPA TIMIT::acoustic\-phonetic continuous speech corpus CD\-ROM, NIST speech disc 1\-1\.1,” Jan\. 1993\.
- \[14\]Kazuma Takaoka, Sorami Hisamoto, Noriko Kawahara, Miho Sakamoto, Yoshitaka Uchida, and Yuji Matsumoto,“Sudachi: a japanese tokenizer for business,”inProceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC\), 2018\.
- \[15\]Taku Kudo, Kaoru Yamamoto, and Yuji Matsumoto,“Applying conditional random fields to japanese morphological analysis,”inProceedings of the 2004 Conference on Empirical Methods in Natural Language Processing \(EMNLP\), 2004\.
- \[16\]Sakasegawa,“hiragana\-asr: Lightweight japanese asr with dual ctc,” 2026\.
- \[17\]Shinnosuke Takamichi, Kentaro Mitsui, Yuki Saito, Tomoki Koriyama, Naoko Tanji, and Hiroshi Saruwatari,“JVS corpus: free japanese multi\-speaker voice corpus,”arXiv \[cs\.SD\], 2019\.
- \[18\]Ryosuke Sonobe, Shinnosuke Takamichi, and Hiroshi Saruwatari,“JSUT corpus: free large\-scale japanese speech corpus for end\-to\-end speech synthesis,”arXiv \[cs\.CL\], 2017\.
- \[19\]Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever,“Robust speech recognition via large\-scale weak supervision,”inProceedings of the 40th International Conference on Machine Learning \(ICML\), 2023\.
- \[20\]Qwen Team,“Qwen3\-omni technical report,”arXiv \[cs\.CL\], 2025\.
- \[21\]Qwen Team,“Qwen2\.5\-omni technical report,”arXiv \[cs\.CL\], 2025\.
- \[22\]Microsoft,“Phi\-4\-mini technical report: Compact yet powerful multimodal language models via mixture\-of\-LoRAs,”arXiv \[cs\.CL\], 2025\.
- \[23\]Gemini Team, Google,“Gemini 2\.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,”arXiv \[cs\.CL\], 2025\.
- \[24\]Kyle Gorman, Jonathan Howell, and Michael Wagner,“Prosodylab\-aligner: A tool for forced alignment of laboratory speech,”Canadian Acoustics, vol\. 39, no\. 3, pp\. 192–193, 2011\.
- \[25\]Luis A Pineda, Luis Villaseñor Pineda, Javier Cuétara, Hayde Castellanos, and Ivonne López,“DIMEx100: A new phonetic and speech corpus for mexican spanish,”inAdvances in Artificial Intelligence – IBERAMIA 2004, 2004\.
- \[26\]Anne Lacheret, Sylvain Kahane, Julie Beliao, Anne Dister, Kim Gerdes, Jean\-Philippe Goldman, Nicolas Obin, Paola Pietrandrea, and Atanas Tchobanov,“Rhapsodie: a prosodic\-syntactic treebank for spoken french,”inProceedings of the Ninth International Conference on Language Resources and Evaluation \(LREC\), 2014\.
- \[27\]M\.A\. Pitt, L\. Dilley, K\. Johnson, S\. Kiesling, W\. Raymond, E\. Hume, and E\. Fosler\-Lussier,“Buckeye Corpus of Conversational Speech \(2nd release\)\.,” 2007\.
- \[28\]David R Mortensen, Patrick Littell, Akash Bharadwaj, Kartik Goyal, Chris Dyer, and Lori Levin,“PanPhon: A resource for mapping IPA segments to articulatory feature vectors,”inProceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, 2016\.
\\CJK@envEnd相似文章
视觉语音识别中基于课程学习的音素到文本重构噪声适应
本文提出渐进式错误课程训练(PECT),通过逐步适应现实中的音素预测错误,提高视觉语音识别中音素到文本重构的鲁棒性,在LRS2和LRS3基准测试上实现了词错误率的降低。
转录儿童语音:ASR性能与获取可靠的正字法转写
这篇论文评估了九种ASR模型(Whisper、Parakeet、Wav2Vec2)在荷兰语儿童语音数据集JASMIN和DART上的表现,发现微调后的Whisper-medium取得了最佳性能(在JASMIN上WER为5.54%,在DART上为70.37%)。它还提出了一种选择方法,能够以高精度自动识别发音正确的录音片段,从而减少人工验证的需求。
对齐就是一切:面向通用音频-语言模型的无指令训练
本文介绍了一种无指令的纯对齐方法,用于构建大型音频-语言模型,该方法通过冻结LLM和音频编码器,仅在自生成数据上训练一个轻量级投影器,实现了与传统多阶段流程相比更少数据下的竞争性性能。
基于梯度的语音到文本对齐方法,适用于任何ASR模型:从CTC到语音大语言模型
本文介绍了一种基于梯度的语音到文本对齐方法,适用于任何可微分的ASR模型,包括CTC、transducer、基于注意力的编码器-解码器以及语音大语言模型,无需训练或模型修改。
通过LLM增强的音频-文本对齐实现零样本呼吸声音分类
本文提出了一种零样本呼吸声音分类框架,通过LLM合成的报告将音频编码器与医学术语对齐,在临床诊断任务中优于CLAP和Qwen2-Audio等模型。