Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study
Summary
This paper proposes a cross-lingual romanization ecosystem for Sinitic languages, develops specific schemes for Mandarin and Cantonese, and shows improved performance in speech-to-romanization tasks compared to baseline methods.
View Cached Full Text
Cached at: 09/01/26, 12:18 PM
# Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study Source: [https://arxiv.org/abs/2608.29170](https://arxiv.org/abs/2608.29170) [View PDF](https://arxiv.org/pdf/2608.29170) > Abstract:This paper proposes the Sinitic Romanization Ecosystem, a cross\-lingual Sinitic romanization design framework with supporting digital infrastructure and a community\-driven open\-source workflow\. The design framework addresses the lack of systematic cross\-lingual romanization alignment among Sinitic languages through four design principles: phonetic correspondence for representing similar sounds with similar romanized symbols, historical\-phonological correspondence for aligning cognate romanization strings, one\-phoneme\-one\-symbol, and basic Latin\-letter use, with a balancing consideration recognizing trade\-offs among these principles\. For the main paired case study, we devel\-op CantRomZJ1 and MandRomZJ1, Cantonese and Manda\-rin romanization schemes following the design framework, respectively\. We also develop schemes for several other Sinitic languages, including Meixian Hakka, Shanghai Wu, and Nanjing Jianghuai Mandarin, following the same de\-sign framework\. To bring the romanization schemes into practical use, we develop open\-source infrastructure for structured romanization storage, conversion, parsing, dic\-tionary construction, and input\-method generation\. Finally, we evaluate the design framework through speech\-to\-romanization experiments based on Meta's Massively Mul\-tilingual Speech \(MMS\) fine\-tuning\. Compared with the Pinyin\+Jyutping baseline, our Man\-dRomZJ1\+CantRomZJ1 condition reduces Cantonese WER and CER by 7\.80% and 10\.61%, respectively\. These results suggest that cross\-lingual romanization alignment can improve transfer in low\-resource Sinitic speech technology\. ## Submission history From: Zijie Zhang \[[view email](https://arxiv.org/show-email/1b7e3707/2608.29170)\] **\[v1\]**Sat, 29 Aug 2026 09:35:14 UTC \(3,867 KB\)
Similar Articles
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
DialectS2S is an end-to-end speech dialogue model for low-resource Chinese dialects, introducing a scalable data synthesis pipeline and a two-stage post-training strategy with self-aligned speech supervision. Experiments show improvements in dialect consistency, response quality, and intelligibility, with fully open-sourced models, data, and code.
Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation
A research paper proposing a new metric and stress-aware system for evaluating and preserving lexical stress in English-to-Chinese speech-to-speech translation, demonstrating significant improvements over existing approaches while maintaining translation quality.
Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation
This paper evaluates whether Cantonese-adapted language models better predict Cantonese reading using eye-tracking data, comparing models like CKIP GPT-2 and CantoneseLLM-7B. Results indicate that more extensive Cantonese-specific training improves predictive fit, though performance varies by information-theoretic measure.
Speech-Driven End-to-End Language Discrimination towards Chinese Dialects
This paper investigates speech-driven features for fine-grained discrimination among Chinese dialects, using an end-to-end model that combines MFCC-based features with word-level embeddings via a CNN, outperforming text-driven methods.
Efficiently Adapting Spoken Language Models for the Singaporean Context
This paper presents a strategy to adapt an open-source spoken language model to the Singaporean Home Team context using LoRA fine-tuning, a surrogate text-QA dataset, and a multi-task objective, achieving competitive performance across five speech tasks in Singapore's four official languages.