Exploring the Capability Boundaries of LLMs in Mastering Chinese Chouxiang Language
Summary
This paper introduces Mouse, a specialized benchmark for evaluating LLMs on Chinese Chouxiang Language tasks across six NLP domains, revealing that current state-of-the-art models have significant limitations with this subcultural internet language despite performing well on contextual understanding tasks.
View Cached Full Text
Cached at: 04/20/26, 08:29 AM
# Exploring the Capability Boundaries of LLMs in Mastering Chinese Chouxiang Language Source: https://arxiv.org/html/2604.15841 Dianqing Lin,Tian Lan11footnotemark:1,Jiali Zhu11footnotemark:1,Jiang Li,Wei Chen Aruukhan,Xu Liu,Xiangdong Su,Hongxu Hou,Guanglai Gao College of Computer Science, Inner Mongolia University, China lindian7ing@163\.com, velikayascarlet@gmail\.com umaru4fun@gmail\.com, cshhx@imu\.edu\.cn ###### Abstract Warning: This paper contains content that may be offensive or harmful While large language models (LLMs) have achieved remarkable success in general language tasks, their performance on Chouxiang Language, a representative subcultural language in the Chinese internet context, remains largely unexplored. In this paper, we introduce Mouse, a specialized benchmark designed to evaluate the capabilities of LLMs on NLP tasks involving Chouxiang Language across six tasks. Experimental results show that current state-of-the-art (SOTA) LLMs exhibit clear limitations on multiple tasks, while performing well on tasks that involve contextual semantic understanding. In addition, we further discuss the reasons behind the generally low performance of SOTA LLMs on Chouxiang Language, examine whether the LLM-as-a-judge approach adopted for translation tasks aligns with human judgments and values, and analyze the key factors that influence Chouxiang translation. Our study aims to promote further research in the NLP community on multicultural integration and the dynamics of evolving internet languages. Our code and data are publicly available at https://github.com/csdq777/Mouse. ![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/graph/title-icon.png)Exploring the Capability Boundaries of LLMs in Mastering Chinese Chouxiang Language Dianqing Lin††thanks:Equal contribution, Tian Lan11footnotemark:1, Jiali Zhu11footnotemark:1, Jiang Li, Wei ChenAruukhan,Xu Liu,Xiangdong Su,Hongxu Hou††thanks:Corresponding Author,Guanglai GaoCollege of Computer Science, Inner Mongolia University, Chinalindian7ing@163\.com, velikayascarlet@gmail\.comumaru4fun@gmail\.com, cshhx@imu\.edu\.cn Refer to captionFigure 1:Overall structure of the proposed Mouse benchmark. ## 1 Introduction With the widespread use of social media, internet language and memes have become an integral part of digital platforms and everyday communication(Kostadinovska-Stojchevska and Shalevska, 2018; Vlasos et al., 2024). In the Chinese internet context, Chouxiang Language represents a distinctive linguistic variant. Originating around 2015, it initially served as a mechanism to express negative sentiments and evade censorship. Consequently, the term historically carried negative connotations. However, it has evolved significantly over the past decade. Driven by the widespread popularity of Chouxiang Culture, a vast amount of non-offensive content has emerged. Chouxiang Language has thus become a neutral and highly symbolic subcultural code. Characterized by its specific expressive forms, it is now widely adopted by Chinese youth and online communities. A more detailed description of Chouxiang Culture is given in Appendix A. Chouxiang Language is usually formed by transforming sentences originally composed entirely of Chinese characters into expressions that combine text, emojis, and metaphorical elements, mainly through homophonic substitution, visual symbol analogy, and literal semantic translation. For example, in the expression "宁可真是个小![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)" (You're such a clever person). The character (宁) functions as a homophonic substitute for "You" (你); the emoji "![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)" metaphorically implies "cleverness" through the visual association of a brain; and "![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)" retains the literal semantics of "ghost (鬼)." Although this mode of expression significantly deviates from Standard Chinese in both form and semantics, thereby creating a non-standard semantic space, it maintains high intelligibility within communities that share the same subcultural context. Despite the widespread influence of Chouxiang Language as a representative internet language within the Chinese internet and society, a systematic analysis of this phenomenon remains absent in the existing natural language processing (NLP) community. Particularly, given the remarkable performance of Large Language Models (LLMs) across various NLP tasks in recent years(Brown et al., 2020; Achiami et al., 2023; Liu et al., 2024), an interesting question arises: What are the capabilities of LLMs in mastering Chouxiang Language? We consider this problem important for three reasons: First, from the perspective of computational social science and culture, existing LLMs and benchmarks exhibit a pronounced Western-centric bias, predominantly reflecting Western mainstream values(Cao et al., 2023; Naous et al., 2024; DURMUS et al., 2024; Singh et al., 2025). Since language is the carrier of cultural essence(Wang et al., 2024a; Zhang et al., 2024; Wang et al., 2025, 2026), exploring Chouxiang Language, a typical non-Western subcultural linguistic variant, is essential. It not only fills the gap in multicultural research for LLMs but is also crucial for understanding linguistic practices within complex cultural contexts. Second, existing studies focusing on Chinese internet language often confine such linguistic phenomena to negative pragmatic dimensions, such as toxic language detection and perturbed language detection(Xiao et al., 2024; Wu et al., 2025a; Bai et al., 2025; Guo et al., 2025). This focus overlooks the neutral and even positive functions that have emerged during the long-term evolution of Chouxiang Language. These non-negative semantic spaces remain largely underexplored. Finally, although prior studies have made impressive progress in the study of Chinese memes and Chinese buzzwords(Xie et al., 2025; Huang et al., 2025), these are only a subset of Chouxiang Language. Given that Chouxiang Language possesses more complex semantic structures and linguistic features, this paper aims to bridge this research gap. We strive to construct a more comprehensive analytical framework of Chouxiang Language for the NLP community, thereby fostering a deeper understanding of such online linguistic phenomena. To bridge this gap, we introduce Mouse, a benchmark designed to evaluate LLMs' proficiency in Chouxiang Language across six tasks. Our results show that while these LLMs demonstrate some understanding of contextual information, they have difficulty handling other aspects. In addition, we conducted a detailed analysis, hoping that our study can contribute to the development of the NLP community focused on subcultural languages. In summary, the main contributions of this paper are as follows: - **Subculture Formalization**: We introduce Chouxiang Language, a unique internet subcultural language, to the NLP community. - **Evaluation Benchmark**: We propose Mouse, the first LLM evaluation benchmark tailored for Chouxiang Language. Comprising six NLP tasks, aiming to evaluate LLMs' processing of this subcultural language. - **Experimental Analysis**: We conduct extensive experiments on SOTA LLMs. Furthermore, we analyze the potential factors underlying their performance and offer insights for future research. ## 2 Preliminaries ComponentOriginal TextDerivational LogicStandard ChineseEnglish ReferenceHomophonic主包 (zhǔ bāo)Near-homophone substitution主播 (zhǔ bō)Streamer91安![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)上![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)→\\rightarrow牌 (pái)→\\rightarrow排 (pái)91安排上Arrange you with 91Visual彳亍口巴Structural decomposition of characters行吧 (xíng ba)That's OK我扬了你![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)灰Iconographic metaphor (![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)→\\rightarrow骨)我扬了你骨灰Scatter your ashes¿e ua m i u onh si uInverted Pinyin你说你妈呢What the hell are you talking aboutSemantic踩到![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)皮Direct symbolic literalism踩到香蕉皮Step on a banana peel滚粗克 (gǔn cū kè)Dialectal register transformation滚出去 (gǔn chū qù)Get out Table 1: Representative examples across three representational components of Chouxiang Language. ### 2.1 The Definition of Chouxiang Language Chouxiang Language is a distinctive variant of Chinese internet language. It serves as a concrete manifestation of Chinese online subculture. Its core mechanism integrates diverse elements, such as special characters, homophones, Pinyin acronyms, dialects, emojis, Chinese radical combinations, and internet memes(Chen, 2021). Characterized by its implicit nature where meanings are felt rather than explicitly stated, Chouxiang Language functions as a subcultural mode of communication that emphasizes the conveyance of emotion over literal information. ### 2.2 Taxonomy To systematically analyze the complexity of Chouxiang Language and clarify its underlying logic, we categorize it into two dimensions: representational components and intents. This fine-grained taxonomy provides the theoretical foundation for our subsequent evaluation. By jointly modeling linguistic structure and pragmatic function, the taxonomy enables a more comprehensive evaluation of model capabilities. #### 2.2.1 The Representational Component of Chouxiang Language Prior studies(Chen, 2021) primarily categorized Chouxiang Language based on its origins, dividing it into symbols, homophones, dialects, and memes. Although these classifications documented early linguistic phenomena, they exhibit significant feature overlap and fail to capture recent, more deconstructive practices. Consequently, we propose a systematic classification of representational components from the perspective of symbolic representation(Shelestiuk, 2003). We categorize these components into three core dimensions: homophonic, visual, and semantic. Within this framework, a single sentence may simultaneously exhibit characteristics from multiple dimensions. The examples across three representational components can be found in Table 1. ##### Homophonic Component This dimension exploits the phonological redundancy of the Chinese language. Users construct Chouxiang expressions through homophonic substitution using Chinese characters, alphanumeric symbols, or multi-stage "image–name–homophone" mapping chains. This process maps the target vocabulary to characters with similar or identical pronunciations. ##### Visual Component Leveraging the ideographic nature of Chinese characters and the pictographic properties of emojis, this component exploits visual analogy through geometric structures, radicals, and other iconic imagery and emoji. It manifests through three mechanisms: (1) Character Decomposition, which fragments glyphs into constituent radicals to increase textual discreteness; (2) Visual Metaphor, where characters and emojis undergo semantic extension based on intuitive visual associations; and (3) Geometric Transformation, involving inverted or deformed typography to disguise sensitive content. ##### Semantic Component This dimension focuses on meaning-level mapping. It includes (1) Symbolic Literalism, which uses the direct or socially shared meanings of emojis, and (2) Dialectal Borrowing, which draws on regional pronunciation or writing variants to add humor or shift style while preserving the core meaning. #### 2.2.2 The Intent of Chouxiang Language In contemporary social media, Chouxiang Language serves not merely as a marker of identity but also functions as a vehicle for diverse communicative intents, akin to natural language. These communicative acts include, but are not limited to: comments on specific events (e.g., sarcasm or praise), direct emotional expressions (e.g., venting anger or helplessness), basic factual statements, and subculturally characteristic humor and memes. Furthermore, within specific contexts, it exhibits action-oriented directives or functions as a tool for implicit sexual reference. As Chouxiang Language enters broader use, analysis must move beyond surface-level symbols and consider its role in social interaction and behavioral intent. Consequently, we categorize these intents into eight distinct classes: Comment (e.g., complaint, praise), Emotional Expression, General Statement, Sexualized Reference, Making a Joke & Memes, Group Identity, Urging, and Others. AttributeExample (ZH)Example (EN)Original Text小![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)汁你8要命![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)N/AReference小伙子你不要命了Young man, are you out of your mind?Representational ComponentHomophonicCommentIntentComments (criticisms, praises, etc.)Toxicity00 Table 2: Chinese and English examples for each attribute in a CXEI. The conversion process is as follows: ![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)→\\rightarrow火 (huǒ)→\\rightarrow伙 (huǒ); 汁 (zhī)→\\rightarrow子 (zi); 8→\\rightarrow八 (bā)→\\rightarrow不 (bù); and ![[Uncaptioned image]](https://arxiv.org/html/2604.15841v1/all-twemojis.pdf)→\\rightarrow辣椒 (là jiāo)→\\rightarrow辣 (là)→\\rightarrow了 (le). ### 2.3 Chouxiang Language Evaluation Instance Drawing inspiration from McBE(Lan et al., 2025b), we integrate the Chouxiang Language Evaluation Instance (CXEI) into Mouse, which is a structured evaluation concept. As the core unit of our benchmark, CXEI enables a detailed assessment of model performance in processing Chouxiang Language. Mouse comprises a total of 1,099 CXEIs. Each CXEI is characterized by the following attributes: ##### Original Text The raw text in Chouxiang Language, typically composed of a mixture of emojis, Chinese characters, Latin letters, punctuation marks and other characters. ##### Reference The corresponding text consisting exclusively of Chinese, serving as a translation. ##### Representational Component These are categorized into three types: Homophonic, Visual, and Semantic. ##### Intent The categories include Comment (e.g., Complaint, Praise), Emotional Expression, General Statement, Sexualized Reference, Humor & Memes, Urging, Group Identity, and Others. ##### Toxicity A binary label indicating whether the text contains toxic content (labeled as 1 for toxic, and 0 otherwise). An example of a CXEI is presented in Table 2.
Similar Articles
LLMs for automatic annotation of Mandarin narrative transcripts
This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
DLawBench is a new benchmark for evaluating large language models in multi-turn legal consultation, covering Chinese and US law with four client types. Experiments show significant room for improvement, with the best model achieving only 0.562 on legal reasoning.
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
CulturALL introduces a 2,610-sample benchmark across 14 languages and 51 regions to evaluate LLMs on real-world, culturally grounded tasks; top model scores only 44.48%, highlighting large room for improvement.
The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
GaoYao introduces a 182k-sample benchmark across 26 languages and 51 regions to systematically evaluate LLMs’ multilingual and multicultural capabilities, revealing large geographical performance gaps.
Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand
This paper introduces the Cross-Lingual Comprehension Gap (CLCG) metric to measure how LLM response quality degrades when content is presented in non-English languages. Across 18 languages and multiple models, it finds a significant performance drop, especially for low-resource languages, questioning the assumption of English-centric capability transfer.