Tag
This paper proposes a cross-lingual romanization ecosystem for Sinitic languages, develops specific schemes for Mandarin and Cantonese, and shows improved performance in speech-to-romanization tasks compared to baseline methods.
This paper explores how intonation and lexical tone interact in Mandarin varieties, using boundary phenomena to examine f0 cues for sentence-level functions like questions versus statements, and discusses theoretical models.
This study investigates whether contextualized embeddings from a large language model predict spoken word duration and pitch contours for Mandarin monosyllabic words, demonstrating above-chance prediction and the ability to back-transform normalized f0 contours to ms time scale.
This paper introduces CLeaD, a supervised contrastive alignment framework for cross-lingual depression detection from speech using WavLM embeddings. It reveals that previous results were inflated due to speaker identity leakage and achieves modest improvements on Mandarin speakers.
A research paper proposing a new metric and stress-aware system for evaluating and preserving lexical stress in English-to-Chinese speech-to-speech translation, demonstrating significant improvements over existing approaches while maintaining translation quality.
This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.