标签
本文提出了一种方法,使用更大模型的合成语音构建紧凑的固定语音泰语TTS系统,评估其性能,并介绍了一个用于设备端部署的82M参数模型。
This paper studies a failure mode in ASR-roundtrip evaluation for Chinese news TTS, showing that fluency errors in context-dependent reading decisions can be masked by ASR surface recovery. A targeted audit across TTS and ASR systems quantifies these false negatives and proposes a human-audited protocol.
本文提出了一种基于分类器的框架,用于审计多语言TTS系统的音系忠实性,以阿萨姆语ATR元音和谐为案例研究。结果显示,Meta的MMS TTS频繁错误生成舌根前伸元音,而这种偏差在人类语音中不存在。