Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics
Summary
This study uses word embeddings to analyze reduplicative constructions in Mandarin Chinese, revealing semantic and pragmatic differentiation and validating distributional semantics for linguistic investigation.
View Cached Full Text
Cached at: 09/16/26, 08:54 AM
# Socio-emotional profiling through distributional semantics
Source: [https://arxiv.org/html/2609.16860](https://arxiv.org/html/2609.16860)
\[ BoldFont = texgyretermes\-bold\.otf, ItalicFont = texgyretermes\-italic\.otf, BoldItalicFont = texgyretermes\-bolditalic\.otf \]\\setCJKmainfont\[ BoldFont=FandolSong\-Bold\.otf, ItalicFont=FandolKai\-Regular\.otf \]FandolSong\-Regular\.otf\\setCJKsansfont\[ BoldFont=FandolHei\-Bold\.otf \]FandolHei\-Regular\.otf\\setCJKmonofontFandolFang\-Regular\.otf
## Reduplicative constructions in Mandarin: Socio\-emotional profiling through distributional semantics
Yu\-Hsiang TsengAffiliation:University of TübingenR\. Harald BaayenAffiliation:University of Tübingen
###### Abstract
Mandarin Chinese has two productive reduplicative constructions that repeat either two\-character base words or their constituents \(e\.g\., 健健康康 ‘in good health’, 讨论讨论 ‘discuss a bit’\)\. Their varied meanings have been described as realizing plurality, valence coloring, sound symbolism and pragmatic functions\. The aim of this study is twofold\. A first goal is to clarify whether it is possible to come to a more precise understanding of the variegated semantics of Mandarin reduplication by using word embeddings from distributional semantics\. A second goal is to explore how useful embeddings are for understanding the details of a semantically complex word\-formation process\. We show that the embedding space recovers the semantic and grammatical properties of reduplications previously identified in the literature, validating Tencent embeddings for morphological investigation\. Semantic profiling revealed that reduplicative constructions are often strongly represented on multiple dimensions\. The two patterns exhibit clear semantic and pragmatic differentiation in distributional space\. Procrustes analysis clarified that the overall organization of the base\-word space is largely preserved in the reduplication space, with local mismatches highlighting regions of discourse\-pragmatic reorganization\. Taken together, these results show that high\-dimensional word embeddings can recover established linguistic generalizations, and capture the semantic versatility of Mandarin reduplication and constructional transparency\.
Keywords:Mandarin reduplicative construction; distributional semantics; semantic vector; semantic profiling; semantic shift; semantic transparency
## 1Introduction
This study presents a quantitative investigation of two four\-syllable reduplicative constructional patterns in Mandarin Chinese which are derived from disyllabic bases\. Using word embeddings, we investigate how these reduplicative constructions are distributed in semantic space, what semantic changes distinguish them from their base words, and to what extent the reduplicative constructions remain semantically transparent with respect to their bases\. Our study builds on previous research in Mandarin linguistics on reduplication\([Chao, 1968](https://arxiv.org/html/2609.16860#bib.bib1);[Li and Thompson, 1981](https://arxiv.org/html/2609.16860#bib.bib2);[Zhu, 1982](https://arxiv.org/html/2609.16860#bib.bib3);[Hua, 2003](https://arxiv.org/html/2609.16860#bib.bib4);[Wang, 2023](https://arxiv.org/html/2609.16860#bib.bib5);[Lu and Müller, 2026](https://arxiv.org/html/2609.16860#bib.bib6)\), as well as on previous quantitative work on word and construction formation using distributional semantics\([Marelli and Baroni, 2015](https://arxiv.org/html/2609.16860#bib.bib7);[Perek and Hilpert, 2017](https://arxiv.org/html/2609.16860#bib.bib8);[Yang and Baayen, 2026](https://arxiv.org/html/2609.16860#bib.bib9)\), research on semantic shifts in inflectional and derivational morphology\([Nikolaev et al\., 2022](https://arxiv.org/html/2609.16860#bib.bib10);[Stupak and Baayen, 2022](https://arxiv.org/html/2609.16860#bib.bib11);[Shafaei\-Bajestan et al\., 2024](https://arxiv.org/html/2609.16860#bib.bib12)\), studies of semantic transparency in Mandarin lexicons\([Shen and Baayen, 2022a](https://arxiv.org/html/2609.16860#bib.bib13);[Shen and Baayen, 2022b](https://arxiv.org/html/2609.16860#bib.bib14)\), and work comparing lexical semantic spaces using Procrustes analysis\([Yang and Baayen, 2025](https://arxiv.org/html/2609.16860#bib.bib15)\)\.
The two reduplicative constructions investigated in the present study repeat either the individual constituents \(henceforth the AABB construction\) or the whole word \(henceforth the ABAB construction\)\. For example, the base 健康jian4\-kang1‘health, healthy’ gives rise to the AABB construction 健健康康jian4\-jian4\-kang1\-kang1‘in good health’\. For 讨论讨论 \(tao3\-lun4\-tao3\-lun4‘discuss a bit’\), by contrast, the disyllabic base 讨论 \(tao3\-lun4‘discuss’\) is reduplicated as a whole\.
In previous standard grammatical descriptions, the reduplicative constructions have been associated with a bewilderingly wide range of semantic and grammatical properties\. Compared with their disyllabic bases, reduplicative constructions have been described as expressing meanings such as plurality, degree intensification, affective coloring, or the pragmatics of persuasive force\([Chao, 1968](https://arxiv.org/html/2609.16860#bib.bib1);[Hua, 2003](https://arxiv.org/html/2609.16860#bib.bib4)\)\. Furthermore, the AABB and ABAB patterns show different tendencies: AABB reduplications have been associated with plurality of entities in nominal uses, iteration of events in verbal uses\([Zhang, 2015](https://arxiv.org/html/2609.16860#bib.bib16)\), and intensification of degree in adjectival uses\([Zhu, 2003](https://arxiv.org/html/2609.16860#bib.bib17);[Melloni and Basciano, 2018](https://arxiv.org/html/2609.16860#bib.bib18)\)\. ABAB constructions, on the other hand, especially in verbal uses, are commonly associated with tentative or suggestive interpretations\([Lu and Müller, 2026](https://arxiv.org/html/2609.16860#bib.bib6)\)\.
In the present study, we investigate to what extent the semantic and pragmatic properties attributed to Mandarin reduplication can be recovered in distributional semantic space\. In other words, we ask whether high\-dimensional word embeddings capture linguistically interpretable aspects of the meanings of Mandarin reduplicative constructions, whether AABB and ABAB show different distributions in semantic space, what changes in meaning arise when going from a disyllabic lexical base to its derived reduplicative forms, and to what extent the reduplicative constructions remain semantically transparent with respect to their base words, by comparing the semantic spaces of the base\-word and reduplicative constructions\.
To address these questions, we extract all the Mandarin reduplicative constructions from the Center for Chinese Linguistics Corpus\(CCL Corpus;[Zhan et al\., 2019](https://arxiv.org/html/2609.16860#bib.bib19)\), represent their meanings and those of their disyllabic bases using 200\-dimensional Tencent embeddings\([Song et al\., 2018](https://arxiv.org/html/2609.16860#bib.bib20)\), and use the Mandarin non\-reduplicative words investigated by[Yang and Baayen \(2025\)](https://arxiv.org/html/2609.16860#bib.bib15)as a reference baseline\. We first examine how reduplicative constructions are distributed in semantic space and profile these constructions in terms of the linguistically semantic and grammatical properties\. We then represent the change from a base word to its derived reduplicative construction as a shift vector and investigate what these semantic shifts are\. Finally, we compare the base\-word and reduplication spaces using Procrustes analysis in order to assess their global correspondence and identify local departures from semantic transparency\.
In the remainder of this study, we address the following research questions:
1. 1\.Do Mandarin reduplicative constructions form a distinct and internally structured domain in the broader Mandarin semantic space?
2. 2\.How are reduplicative constructions profiled along linguistically grammatical and semantic properties, and to what extent do the AABB and ABAB patterns differ in terms of these properties?
3. 3\.What semantic changes arise when going from a disyllabic base word to its derived reduplicative construction?
4. 4\.To what extent are reduplicative constructions semantically transparent with respect to their base words?
The remainder of this paper is structured as follows\. Section[2](https://arxiv.org/html/2609.16860#S2)introduces the dataset and the semantic representations used in the study\. Section[3](https://arxiv.org/html/2609.16860#S3)examines the clustering structure of reduplicative constructions in semantic space\. Section[4](https://arxiv.org/html/2609.16860#S4)characterizes these clusters further in terms of semantic profiling, including lexical category, valence, self\-relatedness, and sound\-related meaning\. Section[5](https://arxiv.org/html/2609.16860#S5)investigates the semantic shifts between reduplicative constructions and their identifiable base words\. Section[6](https://arxiv.org/html/2609.16860#S6)compares the semantic structures of the base\-word space and the reduplication space by means of Procrustes analysis\. Section[7](https://arxiv.org/html/2609.16860#S7)concludes with a general discussion\.
## 2Data
From the corpus compiled by the Center for Chinese Linguistics \(CCL;[Zhan et al\., 2019](https://arxiv.org/html/2609.16860#bib.bib19)\), we extracted all four\-character reduplications with the AABB and ABAB patterns\. Of the 1,010 reduplications, 520 instantiated the AABB pattern and 490 the ABAB pattern\. We collected the corresponding embeddings for these reduplications from the Tencent repository\([Song et al\., 2018](https://arxiv.org/html/2609.16860#bib.bib20)\)\. These embeddings make use of the algorithms underlying word2vec\([Mikolov et al\., 2013a](https://arxiv.org/html/2609.16860#bib.bib22)\)\.
We then identified the corresponding base word for each reduplication\. In most cases, this base word is the disyllabic form AB underlying the reduplicative construction\. For example, the base word of 健健康康jian4\-jian4\-kang1\-kang1‘in good health’ is 健康jian4\-kang1‘healthy’, and the base word of 讨论讨论tao3\-lun4\-tao3\-lun4‘discuss a bit’ is 讨论tao3\-lun4‘discuss’\.
However, not every reduplication has a corresponding base word\. For instance, 盆盆罐罐pen2\-pen2\-guan4\-guan4, ‘pots and pans’, refers to kitchen utensils, but the disyllabic form 盆罐pen2\-guan4is not an existing word\. For such cases, no base word embedding is available\. Some reduplications have base words that are linguistically implausible\. For example, for 林林总总lin2\-lin2\-zong3\-zong3, meaning ‘various and miscellaneous’, 林总lin2\-zong3is attested, but its meaning is unrelated: ‘Manager Lin’\. These forms were not included in the set of base words\. The resulting dataset contained 955 reduplication\-base pairs involving 916 distinct base words\.
It should be mentioned that some reduplications share the same base word while differing in reduplicative construction\. For example, the base word 热闹re4\-nao‘lively’ corresponds to both the AABB form 热热闹闹re4\-re4\-nao\-nao‘very lively’ and the ABAB form 热闹热闹re4\-nao\-re4\-nao‘liven up’\. In total, 39 base words in the dataset give rise to both possible constructions\.
## 3Clusters of reduplicative constructions
Before examining the semantic distribution of reduplicative constructions, we first asked whether they occupy identifiable regions in a broader Mandarin lexical semantic space\. To this end, we combined the 1,010 reduplicative constructions in our dataset with the Mandarin reference word set analyzed by[Yang and Baayen \(2025\)](https://arxiv.org/html/2609.16860#bib.bib15), which contains 2,173 non\-reduplicated Mandarin words from 21 semantic categories\.
Figure 1:Two\-dimensional t\-SNE plot situating reduplications in a broader lexical semantic space together with Mandarin reference words\. Reduplications occupy several identifiable regions rather than being randomly scattered among the reference words\.As shown in Figure[1](https://arxiv.org/html/2609.16860#S3.F1), reduplicative constructions are not randomly scattered among the reference words, but occupy several identifiable regions of the t\-SNE plane\. To further assess this separation, we used linear discriminant analysis to test whether the embeddings distinguish reduplicative constructions from the reference words\. Under leave\-one\-out cross\-validation, the LDA achieved an accuracy of 98\.7%, substantially above the majority\-class baseline of 68\.3%\. These results prove that reduplicative constructions are distributionally distinguishable from non\-reduplicated Mandarin words in the broader lexical semantic space\.
To address our first research question, which asks whether reduplication constructions form distinct clusters in distributional semantic space, we applied k\-means clustering\([MacQueen, 1967](https://arxiv.org/html/2609.16860#bib.bib23)\)to the 200\-dimensional Tencent embeddings of all reduplications in our dataset\. This allowed us to examine whether semantically similar Mandarin reduplications tend to cluster together rather than being randomly dispersed across the vector space\. We evaluated several values ofkkand selectedk=10k=10as the final clustering solution\. This choice was supported by a subsequent linear discriminant analysis\([Venables and Ripley, 2002](https://arxiv.org/html/2609.16860#bib.bib24)\)with leave\-one\-out cross\-validation, which achieved an accuracy of 90\.50%, see Table[1](https://arxiv.org/html/2609.16860#S3.T1)\(majority baseline: 16\.24%\)\. Table[2](https://arxiv.org/html/2609.16860#S3.T2)presents two characteristic examples for each cluster\.


Figure 2:Two\- and three\-dimensional t\-SNE maps for Mandarin reduplicative constructions, with color highlighting by the clusters identified by the k\-means algorithm\. An interactive version is available in the anonymized supplementary materials\.The left panel of Figure[2](https://arxiv.org/html/2609.16860#S3.F2)presents the 10 groups in a two\-dimensional plane obtained with t\-distributed stochastic neighbor embedding \(t\-SNE;[van der Maaten and Hinton, 2008](https://arxiv.org/html/2609.16860#bib.bib25)\), a dimensionality\-reduction technique for visualization\. The t\-SNE map shows that the 10 groups are remarkably well distinguishable even after reducing 200\-dimensional embeddings onto a two\-dimensional space\. The right panel of Figure[2](https://arxiv.org/html/2609.16860#S3.F2)presents the results of a three\-dimensional t\-SNE clustering\. Again, we find that the formations of the 10 clusters form good clusters\. These exploratory plots indicate that Mandarin reduplicative constructions are not randomly distributed in the embedding space, but instead form remarkably differentiated clusters\.
Table 1:Misclassification table for LDA prediction of k\-means clusters under leave\-one\-out cross\-validation\.Table 2:Examples for each k\-means cluster\.Table[2](https://arxiv.org/html/2609.16860#S3.T2)shows two examples for each cluster\. Some clusters appear to be relatively specialized\. Cluster 2, for instance, is dominated by ABAB forms\. The words in this cluster carry attenuative or delimitative verbal meanings, as illustrated by 讨论讨论tao3\-lun4\-tao3\-lun4, ‘discuss a bit’, derived from 讨论tao3\-lun4, ‘discuss’\. Cluster 3 comprises mostly AABB formations, including both verbal and nominal formations\. These formations bring opposites together, as in 进进出出jin4\-jin4\-chu1\-chu1, ‘pass in and out’, with 进出jin4\-chu1, ‘pass in and out’, as the base verb with approximately the same meaning\.
Other clusters are more mixed\. Cluster 6, for example, includes both AABB and ABAB forms\. These reduplications are associated with quantitative increase, as in 许许多多xu3\-xu3\-duo1\-duo1, ‘a great many’, from 许多xu3\-duo1, ‘many’, and intensified degree, as in 很好很好hen3\-hao3\-hen3\-hao3, ‘very good’, from 很好hen3\-hao3, ‘good’\.
Clusters 4 and 6 mainly comprise adjectives with intensified meanings, and cluster 5 contains many adverbial expressions\. Both verbs and adjectives are found in cluster 8\. Cluster 9 brings together various onomatopoeic formations, but also comprises verbs such as 摇摇晃晃yao2\-yao2\-huang4\-huang4‘to be swaying and wobbling’\. Clusters 1 and 10 comprise formations expressing incremental and completive semantics\. Cluster 7 consists entirely of ABAB forms and brings together expressions used in interpersonal exchange such as politeness routines, congratulations, and reassurance\. Examples include 客气客气ke4\-qi4\-ke4\-qi4, ‘you are too polite’, and 恭喜恭喜gong1\-xi3\-gong1\-xi3, ‘many congratulations’\. These expressions typically are independent utterances by themselves\.
The cluster analysis shows that there is considerable structure in the semantic space of reduplicated formations\. However, the cluster analysis remains silent as to which properties underlie the observed clusters\. In what follows, we therefore complement our informal characterization of the 10 clusters with a series of analyses that probe the factors that co\-determine the meanings of the different reduplicative formations, and that give rise to the clusters detected by the k\-means algorithm\. We first examine the role of construction type \(AABB/ABAB\) and lexical category\. We then turn to a series of non\-categorical factors relating to valence, ego\-relatedness, and sound\-symbolism\.
## 4Semantic profiling
The previous section showed that reduplicative constructions are not randomly scattered in semantic space, and instead show considerable structure\. This leads to the question of which grammatical and semantic properties underlie this structure\. In this section, we profile the clusters along a set of linguistic properties in order to determine which properties characterize individual clusters and to what extent they differentiate the two reduplicative constructions\. The linguistic properties that we consider range from relatively straightforward ones such as lexical category, to more gradient abstract cognitive dimensions, such as valence, self\-relatedness, and sound\-relatedness\.
### 4\.1Methods
#### 4\.1\.1Profiling method
Our profiling procedure adapts the category\-defining vector \(CDV\) approach introduced by[Westbury and Hollis \(2019\)](https://arxiv.org/html/2609.16860#bib.bib26);[Westbury et al\. \(2015\)](https://arxiv.org/html/2609.16860#bib.bib27);[Westbury and Wurm \(2022\)](https://arxiv.org/html/2609.16860#bib.bib28)\. The basic idea is that a semantic or grammatical property can be represented in semantic space by the averaged embeddings of a set of prototypical anchor words, i\.e\., the centroid of these embeddings\. The degree to which a given word instantiates that property can be estimated from its position relative to the centroid\. For a dimensionkkwith anchor setAkA\_\{k\}, the CDV is defined as
𝐜k=1\|Ak\|∑a∈Ak𝐯a\\mathbf\{c\}\_\{k\}=\\frac\{1\}\{\|A\_\{k\}\|\}\\sum\_\{a\\in A\_\{k\}\}\\mathbf\{v\}\_\{a\}whereAkA\_\{k\}is the anchor set for propertykk, and𝐯a\\mathbf\{v\}\_\{a\}is the embedding vector of anchor wordaa\.
For a unipolar property such as self\-relatedness, the profile score of a reduplicative construction is defined as the cosine similarity between its embedding vector and the corresponding centroid \(CDV\):
s\(r,k\)=cos\(𝐯r,𝐜k\)s\(r,k\)=\\cos\(\\mathbf\{v\}\_\{r\},\\mathbf\{c\}\_\{k\}\)where𝐯r\\mathbf\{v\}\_\{r\}is the embedding vector of reduplicative constructionrr\. Higher values indicate a stronger association with the relevant property\.
For bipolar properties, such as pairs of opposite emotions, we first constructed a semantic axis from the difference between two opposing centroids and then projected each reduplicative construction onto that axis:
𝐚i,j=𝐜i−𝐜j,p\(r,i,j\)=𝐯r⋅𝐚i,j‖𝐚i,j‖\\mathbf\{a\}\_\{i,j\}=\\mathbf\{c\}\_\{i\}\-\\mathbf\{c\}\_\{j\},\\qquad p\(r;i,j\)=\\mathbf\{v\}\_\{r\}\\cdot\\frac\{\\mathbf\{a\}\_\{i,j\}\}\{\\\|\\mathbf\{a\}\_\{i,j\}\\\|\}Positive and negative values indicate closer alignment with one or the other pole of the corresponding property\.
#### 4\.1\.2Anchor\-word selection
For each profiling property, we selected 100 high\-frequency anchor words that served as prototypical representatives of the semantic pole or category of interest\. Emotion\-related anchor words were selected with reference to the Chinese Emotional Lexicon Ontology\([Xu et al\., 2008](https://arxiv.org/html/2609.16860#bib.bib29)\)\. Lexical category information was determined with reference toA Frequency Dictionary of Mandarin Chinese\([Xiao et al\., 2009](https://arxiv.org/html/2609.16860#bib.bib30)\), which provides frequency\-based lexical entries together with part\-of\-speech indexing\. For selecting anchors, we ordered candidate words by decreasing frequency, and selected the most frequent candidates that most closely and most clearly represented a given prototypical pole\. Furthermore, only those anchor words were included for which an embedding was available in the Tencent resource\.
### 4\.2Construction: AABB vs\. ABAB
The two construction types show markedly different distributions across the clusters, as shown in Table[3](https://arxiv.org/html/2609.16860#S4.T3)\. While some clusters are strongly associated with AABB, others are dominated by ABAB\. For example, Cluster 5 contains only AABB items, whereas Cluster 7 contains only ABAB items\. Clusters 2, 6, and 8 also show a strong bias toward one construction type\. These patterns suggest that construction type is not randomly distributed across the semantic space\.
Table 3:Distribution of AABB and ABAB patterns across the 10 clusters\.If the two patterns encode partly distinct constructional meanings, construction types should be predictable from the embeddings\. To address this issue, we trained a linear discriminant analysis \(LDA\) classifier using the embedding vectors as predictors and construction type \(AABB vs\. ABAB\) as the response variable\. Model performance was evaluated using leave\-one\-out cross\-validation\. The classifier achieved an accuracy of 96\.44%, substantially above the majority\-class baseline of 51\.49%, with 506 of 520 AABB items and 468 of 490 ABAB items correctly identified\.
The LDA analysis shows that the two reduplication patterns are much more separable in the semantic space than expected on the basis of the counts of construction types within clusters presented in Table[3](https://arxiv.org/html/2609.16860#S4.T3)\. One cluster comprises only AABB constructions \(cluster 5\) and one cluster \(cluster 7\) has only ABAB constructions\. The other clusters mix the two constructions to different degrees\. This suggests that the clusters are not only determined by construction pattern, but by many other linguistic predictors as well\. One obvious candidate is lexical category\.
### 4\.3Lexical category
Table[4](https://arxiv.org/html/2609.16860#S4.T4)presents the counts of verbs, nouns, adjectives, and adverbs, broken down by cluster\. Nouns don’t show any clear preference for clusters, verbs favor clusters 2 and 7, adjectives favor clusters 4, 5 and 8, and adverbs favor clusters 9 and 10\. This distributional differentiation is supported by a linear discriminant analysis \(LDA\) using embeddings to predict lexical category\. Under leave\-one\-out cross\-validation, the model achieved an accuracy of 84\.06%, well above the majority\-class baseline of 45\.05%, correctly classifying 405 of 455 adjectives, 154 of 199 adverbs, 22 of 33 nouns, and 268 of 323 verbs\.
However, this discrete way of assigning lexical category is rather coarse, as in Mandarin Chinese, a word can be used across many different lexical categories\. For example, ‘明明白白’ming2\-ming2\-bai2\-bai2‘clear; clearly’ can be used in a more adjectival way, as in ‘大家都明明白白’da4\-jia1 dou1 ming2\-ming2\-bai2\-bai2‘everyone understands it clearly’, or in a more adverbial way, as in ‘他曾明明白白地吩咐过’ta1 ceng2 ming2\-ming2\-bai2\-bai2 de fen1\-fu4 guo4‘he had once instructed \[someone\] very clearly’, depending on context\.111Both examples are drawn from the CCL corpus used in this study\.
Therefore, we constructed a Noun\-Verb axis and an Adjective\-Adverb axis\. We then projected all reduplicative constructions onto these two continuous lexical axes, to address whether particular regions of the reduplicative semantic space show stronger noun\-like, verb\-like, adjective\-like, or adverb\-like tendencies\.
Table 4:Distribution of lexical categories across the 10 clusters\.Figure[3](https://arxiv.org/html/2609.16860#S4.F3)shows that, as expected, lexical categories are not randomly distributed across the semantic space\. In the left panel, Clusters 10, 5, and 1 show relatively stronger noun\-like tendencies\. Illustrative examples from these clusters are 方方面面fang1\-fang1\-mian4\-mian4‘all aspects’ in Cluster 10, 山山水水shan1\-shan1\-shui3\-shui3‘hills and rivers’ in Cluster 1, and 条条框框tiao2\-tiao2\-kuang4\-kuang4‘rules and regulations’ in Cluster 5\. Clusters 6 and 2 are relatively more verb\-like, with some of the darkest blue datapoints corresponding to items such as 选择选择xuan3\-ze2\-xuan3\-ze2‘choose a bit’ in Cluster 6 and 请教请教qing3\-jiao4\-qing3\-jiao4‘consult briefly’ in Cluster 2\.
A similar distribution can be observed on the Adjective\-Adverb axis, shown in the right panel of Figure[3](https://arxiv.org/html/2609.16860#S4.F3)\. The left part of the semantic space, especially Clusters 5 and 6, is more adjective\-like, with more red dots such as 轰轰烈烈hong1\-hong1\-lie4\-lie4‘vigorous and dynamic’ in Cluster 5 and 很贵很贵hen3\-gui4\-hen3\-gui4‘very expensive indeed’ in Cluster 6\. By contrast, adverb\-like tendencies are found mainly in Cluster 3 and, to a lesser extent, in Cluster 6, represented by deeper blue datapoints such as 前前后后qian2\-qian2\-hou4\-hou4‘back and forth’ in Cluster 3 and 许久许久xu3\-jiu3\-xu3\-jiu3‘for a very long time’ in Cluster 6\.


Figure 3:Three\-dimensional t\-SNE maps of Mandarin reduplicative constructions, with color highlighting for lexical category along the Noun\-Verb axis \(left panel\) and the Adjective\-Adverb axis \(right panel\)\. Darker red dots indicate more noun\-like and adjective\-like items, whereas deeper blue dots indicate more verb\-like and more adverb\-like items\. An interactive version is available in the anonymized supplementary materials\.We also found that lexical categories are related to reduplicative pattern\. More verb\-like tendencies are mainly associated with ABAB constructions, whereas more noun\-like tendencies are mainly associated with AABB constructions\. Adjectival and adverbial tendencies are attested in both patterns, but both are more common in AABB constructions and less frequent in ABAB constructions\.
To assess the extent to which our gradient measures for lexical categories contribute to cluster structure, we used the Noun\-Verb and Adjective\-Adverb axis scores to predict cluster, using linear discriminant analysis \(LDA\)\. Under leave\-one\-out cross\-validation, the model achieved an accuracy of 32\.97% \(majority\-class baseline: 16\.24%\)\. These results suggest that lexical category contributes to the organization of the cluster structure, but does not fully account for it\. Some clusters, especially Clusters 2 and 8, were identified with moderate success, whereas others, such as Clusters 3, 4, and 6, were not recoverable from the lexical\-category scores alone \(see Table[5](https://arxiv.org/html/2609.16860#S4.T5)\)\. This in turn raises the possibility that part of the remaining structure is related to more abstract semantic properties\. In what follows, we therefore extend the profiling analysis to dimensions such as valence, self\-relatedness, and sound\-related meaning\.
Table 5:Accuracies by cluster for semantic\-space profiling across four categories\.
### 4\.4Valence
For valence profiling, we selected six basic emotions\([Ekman, 1992](https://arxiv.org/html/2609.16860#bib.bib31)\)and organized them along three contrastive axes: Happy\-Sad, Anger\-Fear and Surprise\-Disgust\. We then projected all reduplicative constructions onto these three axes, resulting in three scores for each formation\. The three panels of Figure[4](https://arxiv.org/html/2609.16860#S4.F4)present the 3\-D t\-SNE space, with color highlighting for the three valence contrasts\. Higher scores for happiness, anger, and surprise are presented with darker red colors, whereas more negative scores, indicating closer proximity to sadness, fear, and disgust, are represented by deeper shades of blue\. Figure[4](https://arxiv.org/html/2609.16860#S4.F4)shows that different regions of the semantic space are associated with different aspects of emotion\.



Figure 4:Three\-dimensional t\-SNE maps of Mandarin reduplicative constructions, with color highlighting valence along the Happy\-Sad axis \(left panel\), Anger\-Fear axis \(middle panel\), and Surprise\-Disgust axis \(right panel\)\. Darker red dots indicate higher degrees of happiness, anger and surprise, whereas darker blue dots indicate higher degrees of sadness, fear and disgust\. An interactive version is available in the anonymized supplementary materials\.In the left panel, reduplications that score high on the happy\-sad axis tend to cluster in the left\-central part of the semantic space, especially in Clusters 5, 2, and 7, with reduplications such as 开开心心kai1\-kai1\-xin1\-xin1‘very happy’ in Cluster 5, 高兴高兴gao1\-xing4\-gao1\-xing4‘cheer up a bit’ in Cluster 2, and 祝贺祝贺zhu4\-he4\-zhu4\-he4‘many congratulations’ in Cluster 7 \(for cluster identifiers, compare Figure[2](https://arxiv.org/html/2609.16860#S3.F2)\)\. Reduplications that score low on the happy\-sad axis are more common in the middle and right parts of the semantic space, especially in Clusters 8, 6, and 9\. Among the more extreme formations we find 悲悲切切bei1\-bei1\-qie4\-qie4‘deeply sorrowful’ in Cluster 8, 很痛很痛hen3\-tong4\-hen3\-tong4‘very painful’ in Cluster 6, and 滴滴答答di1\-di1\-da1\-da1‘dripping continuously’ in Cluster 9\.
In the middle panel, the reduplications that score high on the anger\-fear axis are more common on the left side of the semantic space, especially in Clusters 8 and 2, with items such as 教训教训jiao4\-xun4\-jiao4\-xun4‘teach someone a lesson’ and 骂骂咧咧ma4\-ma4\-lie1\-lie1‘to swear while talking’\. Reduplications that score low on the anger\-fear axis, by contrast, are more common on the right side of the t\-SNE space, including 紧紧张张jin3\-jin3\-zhang1\-zhang1‘in a highly strained state’ in Cluster 8, 迷迷茫茫mi2\-mi2\-mang2\-mang2‘all lost and bewildered’ in Cluster 1, and 颠颠簸簸dian1\-dian1\-bo3\-bo3‘bumping and jolting all the way’ in Cluster 3\.
The third panel shows the results for the scores on the surprise\-disgust axis\. Words with higher scores are more common in the right\-hand part of the semantic space, and include formations such as 飘飘渺渺piao1\-piao1\-miao3\-miao3‘floating and indistinct’ \(Cluster 1\), 扑通扑通pu1\-tong1\-pu1\-tong1‘pounding rapidly’ \(Cluster 9\), and 恭喜恭喜gong1\-xi3\-gong1\-xi3‘many congratulations’ \(Cluster 7\)\. Why 飘飘渺渺 scores so high on this axis is unclear to us\. Reduplications with lower scores, indicating closer alignment with the disgust pole of the axis, are concentrated on the left side of the semantic space, especially in Cluster 8\. Among them we have 讨厌讨厌tao3\-yan4\-tao3\-yan4‘very annoying’ and 邋邋遢遢la1\-la1\-ta1\-ta1‘sloppy and untidy’\.
To assess whether the scores on these three emotional axes predict cluster membership, we again used LDA\. Under leave\-one\-out cross\-validation, the model achieved an accuracy of 27\.03% \(majority\-class baseline: 16\.24%\)\. This indicates that these emotional axes are somewhat related to cluster structure, but can only provide a partial explanation\. Some clusters, especially Clusters 1, 2, and 5, were recovered with moderate success, whereas others, such as Clusters 3, 4, and 10, were not recoverable from the emotion scores alone \(cf\. Table[5](https://arxiv.org/html/2609.16860#S4.T5)\)\.
### 4\.5Self\-relatedness
To investigate self\-relatedness, we calculated the cosine similarity of the reduplications with the centroid of the anchor words for the “self”\. Figure[5](https://arxiv.org/html/2609.16860#S4.F5)highlights words with higher scores for the self with darker shades of red\. Cluster 7 shows particularly strong scores for the ‘self’, such as 好说好说hao3\-shuo1\-hao3\-shuo1‘no problem’, 罢了罢了ba4\-le5\-ba4\-le5‘let it go’, and 多谢多谢duo1\-xie4\-duo1\-xie4‘many thanks’\. These constructions are often used independently in interaction to express the speaker’s own stance or attitude toward a situation\. Examples of formations that are orthogonal to the ‘self’ are 开开停停 \(Cluster 3\)kai1\-kai1\-ting2\-ting2‘stop and go’ and 角角落落 \(Cluster 1\)jiao3\-jiao3\-luo4\-luo4‘all the corners’ and 甜甜美美 \(Cluster 5\)tian2\-tian2\-mei3\-mei3‘sweet and pleasant’\.
Figure 5:Three\-dimensional t\-SNE map of Mandarin reduplicative constructions, with color highlighting the cosine similarity with the centroid of self\-related anchor words\. Deeper shades of red represent higher cosine similarities, deeper shades of blue near orthogonality\. An interactive version is available in the anonymized supplementary materials\.We also used ego scores to predict cluster membership with a linear discriminant analysis \(LDA\)\. Under leave\-one\-out cross\-validation, the model achieved an accuracy of 28\.12% \(majority\-class baseline: 16\.24%\)\. The effect is especially visible for Cluster 7, which also contains 19 of the 20 reduplications with the highest ego scores, but 6 out of 10 clusters are not recoverable at all from ego scores alone \(cf\. Table[5](https://arxiv.org/html/2609.16860#S4.T5)\)\.
### 4\.6Sound\-related meaning
Manual inspection suggests that Cluster 9 is the cluster most strongly associated with onomatopoeic meaning in our data\. To examine more systematically what kinds of sounds are encoded by reduplications, we drew on a classification of everyday sound sources and distinguished three broad categories: vibrating solids, gases, and liquids\([Gaver, 1993](https://arxiv.org/html/2609.16860#bib.bib32)\)\. Since many reduplications in our data also reflect sounds made by humans, we added a fourth category, human sounds\.




Figure 6:Three\-dimensional t\-SNE maps of Mandarin reduplicative constructions, with color highlighting the cosine similarity with the centroid of anchor words of Gases \(top left panel\), Liquids \(top right panel\), Vibrating Solids \(bottom left panel\), and Humans \(bottom right panel\) categories\. Deeper shades of red represent higher cosine similarities, deeper shades of blue near orthogonality\. An interactive version is available in the anonymized supplementary materials\.Figure[6](https://arxiv.org/html/2609.16860#S4.F6)shows that three sound\-source categories, gases, liquids, and vibrating solids, display broadly similar spatial tendencies, with darker shades of red concentrated especially in the upper\-right part of the semantic space, where Cluster 9, Cluster 1, and Cluster 8 are located\. Among them, Cluster 9 is especially associated with the gases category, with reduplications such as 轰隆轰隆hong1\-long1\-hong1\-long1‘continuous rumble’ and 一股一股yi4\-gu3\-yi4\-gu3‘in waves’, typically said of hot air\. Examples of formations that are orthogonal to the gases are 圆圆满满 \(Cluster 5\)yuan2\-yuan2\-man3\-man3‘fully satisfactory’ and 活络活络 \(Cluster 2\)huo2\-luo4\-huo2\-luo4‘loosen up a bit’\.
Liquid\-related formations are predominant in Clusters 9 and 1, and include formations such 哗啦哗啦hua1\-la1\-hua1\-la1‘a splashing sound’ in Cluster 9 and 淅淅沥沥xi1\-xi1\-li4\-li4‘dripping lightly’ in Cluster 1\. Examples of reduplications that are orthogonal to the category of liquids are found in Cluster 2 and Cluster 5, such as 宣传宣传xuan1\-chuan2\-xuan1\-chuan2‘give wide publicity’ and 体体面面ti3\-ti3\-mian4\-mian4‘very decent’\.
Cluster 9 also contains formations relating to vibrating solids, with items such as 叮叮当当ding1\-ding1\-dang1\-dang1‘jingling of bells’ and 咔嚓咔嚓ka1\-cha1\-ka1\-cha1‘cracking or snapping sounds’\. Examples of reduplications that are orthogonal to the vibrating solids category are found in Cluster 10 and Cluster 2, including examples such as 方方面面fang1\-fang1\-mian4\-mian4‘all aspects’ and 帮助帮助bang1\-zhu4\-bang1\-zhu4‘give some help’\.
The human sound category shows a more extensive distribution than the three physical sound\-source categories\. In addition to the upper\-right region, the middle and left parts of the semantic space also show shades of dark red, especially in Cluster 7 and Cluster 8\. Cluster 7 is mainly associated with vocal expressions, with items such as 哎呀哎呀ai1\-ya1\-ai1\-ya1‘repeated exclamations of surprise or concern’, compare to “oh dear, oh dear” in English\. By contrast, Cluster 8 is more closely associated with sounds or sound\-related effects of human action, as in 哆哆嗦嗦duo1\-duo1\-suo1\-suo1‘trembling and shivering’\. Examples of reduplications that are orthogonal to the human\-sound category are found in Cluster 10 and Cluster 6, such as 事事物物shi4\-shi4\-wu4\-wu4‘all things’ and 选择选择xuan3\-ze2\-xuan3\-ze2‘do some choosing’\.
We also used the Gases, Liquids, Vibrating Solids, and Humans scores to predict clusters using LDA\. Under leave\-one\-out cross\-validation, the model achieved an accuracy of 49\.11% \(majority\-class baseline: 16\.24%\)\. In particular, Cluster 7, Cluster 8, and Cluster 9 were recovered with moderate to high accuracy, as shown in Table[5](https://arxiv.org/html/2609.16860#S4.T5)\.
### 4\.7Predictability of clusters from semantic profiling
The preceding sections profiled the semantic space starting from four categories: lexical category, valence, self\-relatedness, and sound\-related meaning\. We next examine how well the semantic space can be explained when all semantic profiling categories are considered together\. Therefore, we used all semantic categories, combined with construction type \(AABB/ABAB\) in an LDA analysis, asking the LDA to predict cluster identity\. Under leave\-one\-out cross\-validation, the model achieved an accuracy of 71\.78%, which is substantially above the majority\-class baseline of 16\.24% \(see Table[6](https://arxiv.org/html/2609.16860#S4.T6)\)\.
Table 6:Misclassification table for LDA of k\-means clusters from four semantic categories, using leave\-one\-out cross\-validation\.Several clusters were predicted with high accuracy, including Cluster 2 \(0\.88\), Cluster 7 \(0\.87\), and Cluster 1 \(0\.71\), although others, especially Cluster 3 and Cluster 4, are not well recoverable\.
In summary, reduplications show considerable clustering in semantic space\. A k\-means clustering suggested 10 clusters that are well supported by an LDA analysis predicting cluster from embeddings\. We have shown that cluster identity can be predicted to a considerable extent \(and far above majority baselines\) from a range of linguistic features: reduplication construction type, word category, valency, self\-relevance, and type of sound\. Nevertheless, the individual clusters are not fully defined by these linguistic features, from which we conclude that there is further systematic variation in the semantic space that resists reduction to the linguistic variables we investigated\.
In the next section, we turn to the question of to what extent reduplicative meanings are related to the meanings of their base words and, if so, how this relation can be characterized\.
## 5Shift vectors
In distributional\-semantic research on morphology, shift vectors have been used to represent the change in meaning from a base form to a morphologically more complex form\(see, e\.g\.,[Drozd et al\., 2016](https://arxiv.org/html/2609.16860#bib.bib33);[Boleda, 2020](https://arxiv.org/html/2609.16860#bib.bib34);[Mikolov et al\., 2013b](https://arxiv.org/html/2609.16860#bib.bib21)\)\. This approach has proved useful for a range of phenomena, including English noun plurals\([Shafaei\-Bajestan et al\., 2022](https://arxiv.org/html/2609.16860#bib.bib35)\), Finnish noun inflection\([Nikolaev et al\., 2022](https://arxiv.org/html/2609.16860#bib.bib10)\), Russian nominal paradigms\([Chuang et al\., 2022](https://arxiv.org/html/2609.16860#bib.bib36)\), German particle verbs and derivation\([Stupak and Baayen, 2022](https://arxiv.org/html/2609.16860#bib.bib11)\), and Mandarin suffixation\([Shen and Baayen, 2022b](https://arxiv.org/html/2609.16860#bib.bib14)\)\. Across these studies, shift vectors have been used to ask whether a change in morphological form \(e\.g\., realized with suffixation\) goes hand in hand with a change in embedding space\. If a morphological meaning change is systematic, words sharing the same form structure are expected to show similar displacement in semantic space\. In the present study, we apply this approach to Mandarin reduplication, analyzing how reduplicative meanings relate to the meanings of their base words\. Of particular interest is whether the semantic change from base words to reduplicative constructions is systematically different across the different kinds of reduplicative constructions\.
Let𝒗r\\bm\{v\}\_\{r\}denote the vector of a reduplication and𝒗b\\bm\{v\}\_\{b\}the vector of its base word\. The semantic shift𝒔r\\bm\{s\}\_\{r\}from base to reduplication is defined as
𝒔r=𝒗r−𝒗b\.\\bm\{s\}\_\{r\}=\\bm\{v\}\_\{r\}\-\\bm\{v\}\_\{b\}\.The shift vector𝒔r\\bm\{s\}\_\{r\}captures both the direction and the extent of semantic change from the base words to the corresponding reduplicative constructions\. For example, for the reduplication 健健康康jian4\-jian4\-kang1\-kang1‘healthy and well’ and its base word 健康jian4\-kang1‘healthy’, the shift vector is
𝒔健健康康=健健康康→−健康→\.\\bm\{s\}\_\{\\text\{健健康康\}\}=\\overrightarrow\{\\text\{健健康康\}\}\-\\overrightarrow\{\\text\{健康\}\}\.Here,𝒔健健康康\\bm\{s\}\_\{\\text\{健健康康\}\}captures the semantic change from the base word 健康jian4\-kang1to the reduplicated form 健健康康jian4\-jian4\-kang1\-kang1\.
### 5\.1Semantic transparency: base word similarity in shift space
Focusing on the reduplicative constructions for which a base word exists \(955 out of the total 1,010 reduplications\), we projected the shift vectors onto a two\-dimensional t\-SNE space\. The resulting t\-SNE map is shown in Figure[7](https://arxiv.org/html/2609.16860#S5.F7)\. Data points are colored according to the clusters identified by the k\-means clustering \(see section[3](https://arxiv.org/html/2609.16860#S3)\)\. This makes it possible to ask to what extent the original clusters of reduplications are also visible in the shift space\. If the original clusters remain reasonably well separated, this suggests that the semantic change from base to reduplication is relatively consistent within clusters\. If, by contrast, the clusters become more intermixed, this points to more variable or less transparent semantic transformations from base word to reduplication\.
Figure 7:Two\-dimensional t\-SNE map of shift vectors from base words to reduplicative constructions, colored by the clusters of reduplicative constructions identified by the k\-means algorithm\.Compared with the semantic space of the reduplications themselves, shown in the left panel of Figure[2](https://arxiv.org/html/2609.16860#S3.F2), it is clear that considerable structure remains visible in the shift space, although not equally clearly for all clusters\. Cluster 2 \(orange, bottom center\) remains relatively compact\. This cluster is dominated by verbal reduplicative constructions with strong eventive semantics, and the corresponding semantic shifts serve to shorten, soften, and render more tentative the event meaning of the base word\. For example, the base word 琢磨zuo2\-mo2means ‘to figure out’, whereas the reduplicated form 琢磨琢磨zuo2\-mo0\-zuo2\-mo0,‘think it over a bit’, expresses a lighter and more tentative version of the same process\. Similarly, 讨论tao3\-lun4means ‘to discuss’, while 讨论讨论tao3\-lun4\-tao3\-lun4means ‘have a brief discussion’\. Words that are used across different word categories, such as 热闹re4\-nao0, which can be interpreted as an adjective \(‘lively’\), a verb \(‘liven up’\), or a noun \(‘scene of bustle’\), appear in cluster 2 with the ABAB form 热闹热闹re4\-nao0\-re4\-nao0, ‘to liven things up a bit’\. By contrast, the AABB form 热热闹闹re4\-re4\-nao0\-nao0, ‘very lively’, is associated with the adjectival interpretation of the base word, and belongs to Cluster 8\.
Clusters 4 \(purple\) and 5 \(pink\), in the outer upper left quadrant, are also relatively well preserved in the shift space\. The formations in Cluster 4 change the meanings of their base words by increasing their sensory perceptibility\. This can be seen in 高高大大gao1\-gao1\-da4\-da4‘burly and imposing’, cf\. 高大gao1\-da4‘tall and big’, where reduplication makes the size more visually salient\. A similar effect is found in the tactile domain\. Both 冰凉冰凉bing1\-liang2\-bing1\-liang2‘gelid’ and 冰冰凉凉bing1\-bing1\-liang2\-liang2‘\(pleasantly\) ice cool’, cf\. 冰凉bing1\-liang2‘ice\-cold’, construe the base property as a more immediate bodily sensation\. The same tendency is also visible in the domain of color, as in 通红通红tong1\-hong2\-tong1\-hong2‘red through and through’, cf\. 通红tong1\-hong2‘bright red’\.
The shift in Cluster 5 involves emphatic reinforcement of the meaning of the base word\. This can be seen in 扎扎实实zha1\-zha1\-shi2\-shi2‘in a down\-to\-earth manner’, cf\. 扎实zha1\-shi2‘solid’, where reduplication reinforces the sense of firmness and reliability\. Similarly, 清清白白qing1\-qing1\-bai2\-bai2‘completely blameless’, cf\. 清白qing1\-bai2‘innocent’, presents the state as complete and unequivocal, rather than simply stronger in degree\. In 快快乐乐kuai4\-kuai4\-le4\-le4‘happily; happy and carefree’, cf\. 快乐kuai4\-le4‘happy; happiness’, the reduplicated form foregrounds a sustained affective state\.
This semantic shift is often accompanied by a stronger tendency toward adverbial use, as reflected in the occurrence before the adverbial particle 地de, where the reduplicated form modifies a following verb phrase, as in 清清白白地生活qing1\-qing1\-bai2\-bai2 de sheng1\-huo2\-zhe‘live blamelessly’, 扎扎实实地工作zha1\-zha1\-shi2\-shi2 de gong1\-zuo4‘do a solid job’, 快快乐乐地上学kuai4\-kuai4\-le4\-le4 de shang4\-xue2‘go to school happily’\. The proximity of Clusters 4 and 5 in the shift space reflects that both clusters involve reinforcement of the base meaning: the reduplicative shift in Cluster 4 emphasizes the intensity of the sensory perception expressed by the base word, whereas the shift in Cluster 5 emphasizes the desirability or reliability of the — typically positive — meaning of the base word\.
At the same time, some clusters, such as Cluster 3 \(brown\), Cluster 8 \(light yellow\), and Cluster 9 \(dark yellow\), become more diffuse in the shift space than in the original reduplication space\. Among them, Cluster 3 shows the greatest dispersion, spreading across all four quadrants of the shift space\.
In the first quadrant, the reduplicative constructions of Cluster 3 are clustered around the repeated actions encoded in the base words, like 颠颠簸簸dian1\-dian1\-bo3\-bo3‘bump along the way’ cf\. 颠簸dian1\-bo3‘bump’, 沸沸腾腾fei4\-fei4\-teng2\-teng2,‘bubbling and seething’, cf\. 沸腾fei4\-teng2‘boil’\. 分分合合fen1\-fen1\-he2\-he2‘repeated separations and reunions’, cf\.分合fen1\-he2‘separation and reunion’\.
By contrast, the reduplicative constructions of Cluster 3 in the fourth quadrant are organized around the quantification of the entities denoted by the base words, such as 家家户户jia1\-jia1\-hu4\-hu4‘every family’, cf\. 家户jia1\-hu4‘household’, 里里外外li3\-li3\-wai4\-wai4‘all aspects; completely’, cf\. 里外li3\-wai4‘inside and outside’ and 老老少少lao3\-lao3\-shao4\-shao4‘people of all ages’, cf\. 老少lao3\-shao4‘old and young’\.
In the second quadrant, the reduplicative constructions of Cluster 3 are organized around vertical iteration, as exemplified by 沉沉浮浮chen2\-chen2\-fu2\-fu2‘sinking and floating repeatedly, floating life’, cf\. 沉浮chen2\-fu2‘sink and float’, 起起伏伏qi3\-qi3\-fu2\-fu2‘repeated ups and downs’, cf\. 起伏qi3\-fu2‘up and down’ and 坎坎坷坷kan3\-kan3\-ke3\-ke3‘repeatedly bumpy’, cf\. 坎坷kan3\-ke3‘bumpy’\.
The reduplicative constructions of Cluster 3 in the third quadrant comprise formations such as 进进出出jin4\-jin4\-chu1\-chu1‘in and out frequently’ \(cf\. 进出jin4\-chu1‘in and out’\) and 来来去去lai2\-lai2\-qu4\-qu4‘comes and goes repeatedly’ \(cf\. 来去lai2\-qu4‘come and go’\), both of which express repeated coming and going, with the former more prevalent with human subjects \(68\.9% according to CCL database\), as in 孩子们进进出出hai2\-zi\-men jin4\-jin4\-chu1\-chu1‘the children run in and out’, and the latter more readily extended to abstract subjects \(35\.3%222The examples and percentage were retrieved and calculated from the Chinese–English bilingual section of the CCL Corpus Search System\) and more global situations, such as the coming and going of generations or the ebb and flow of people between countries, as in 苦与乐来来去去ku3 yu3 le4 lai2\-lai2\-qu4\-qu4‘suffering and happiness come and go’\.
Formations such as 前前后后qian2\-qian2\-hou4\-hou4‘throughout, completely’ \(cf\. 前后qian2\-hou4‘front and back’\) express a vivid sense of step\-by\-step thoroughness and comprehensiveness, similar to the formations in the fourth quadrant\. For instance, 前前后后、里里外外叙述个遍qian2\-qian2\-hou4\-hou4, li3\-li3\-wai4\-wai4 xu4\-shu4 ge bian4literally translates as ‘front\-front\-back\-back, inside\-inside\-outside\-outside narrate CLASSIFIER once\-through’, meaning ‘Tell me, covering every detail, inside and out’\.
Other clusters overlap more strongly with neighboring regions than they did in the original reduplication space\. For example, Cluster 1 and Cluster 10 appear to overlap more strongly in the shift space\. Items in both clusters share the same constructional pattern, 一yi1‘one’ \+ classifier \+ 一yi1‘one’ \+ classifier, which expresses a unit\-by\-unit meaning\. Examples include 一年一年yi4\-nian2\-yi4\-nian2‘year after year’, cf\. 一年yi4\-nian2‘one year’, and 一行一行yi4\-hang2\-yi4\-hang2‘line by line’, cf\. 一行yi4\-hang2‘one line’, in Cluster 10; and 一片一片yi2\-pian4\-yi2\-pian4‘piece by piece’, cf\. 一片yi2\-pian4‘one piece’, and 一桶一桶yi4\-tong3\-yi4\-tong3‘bucket by bucket’, cf\. 一桶yi4\-tong3‘one bucket’, in Cluster 1\.
The comparison between the reduplication space and the shift space shows that some clusters remain relatively compact in both spaces, suggesting that these reduplicative constructions are similar not only in the meanings of the reduplicative forms themselves, but also in the semantic shifts from base word to reduplicative form \(e\.g\., Clusters 2, 4, and 5\)\. Other clusters become more dispersed in the shift space than in the reduplication space, indicating that a single cluster may involve several different types of shift change from base to reduplication \(e\.g\., Clusters 3, 8, and 9\)\. Conversely, some clusters that are relatively distant in the reduplication space become closer or more overlapping in the shift space, suggesting that they may instantiate similar types of semantic change despite belonging to different clusters in the reduplication space \(e\.g\., Clusters 1 and 10\)\.
To examine the structure of the shift space in more detail, the next section investigates what types of semantic shift from base words to reduplicative forms can be identified\.
### 5\.2Types of base\-to\-reduplication shifts
The preceding analysis showed that some reduplication clusters remain relatively compact in the shift space, whereas others become more diffuse or overlap strongly with neighboring regions, suggesting that the semantic shifts from base words to reduplicative constructions may themselves be structured into several types\. In this section, we therefore examine the structure of the shift space in order to identify shift types, and further ask whether these shift types are associated differently with the two reduplicative patterns, AABB and ABAB\.
To identify shift types, we applied k\-means clustering to the 955 shift vectors for reduplicative constructions whose corresponding base words exist\. Different values ofkkin the range \[2, 15\] yielded very similar results, with LDA cross\-validation accuracies ranging from 0\.97 fork=2k=2to 0\.84 fork=14k=14\. For our exploration, we opted fork=10k=10, the number of clusters previously observed for the semantic space of the reduplications\. Under leave\-one\-out cross\-validation, LDA classification accuracy for these 10 clusters reached 87\.33%, substantially above the majority\-class baseline of 16\.65%\.
For visualization, the shift vectors were projected from the shift space onto a two\-dimensional t\-SNE map, shown in Figure[8](https://arxiv.org/html/2609.16860#S5.F8)\. Each panel highlights one cluster in blue, with the remaining data points shown in grey\. Circle\-shaped and cross\-shaped data points represent AABB and ABAB constructions respectively\. Red crosses mark cluster centroids\. Most clusters are reasonably well separated, but cluster 6 is of low quality and overlaps largely with cluster 7\. On the other hand, some clusters \(e\.g\., 3, 5, 8, 9, and 10\) are clearly distinct in the t\-SNE projection\. Furthermore, some clusters have a much wider scatter than others \(compare, e\.g\., 1 and 2 with 3 and 8\)\. Although the clustering is not perfect, it is useful as a guide to the shift space\.
Figure 8:Two\-dimensional t\-SNE map of shift vectors for Mandarin reduplicative constructions\. Each panel highlights one k\-means cluster in the shift\-vector space, with the remaining data points shown in grey\. Circle\-shaped and cross\-shaped data points indicate AABB and ABAB forms respectively\. Red crosses denote cluster centroids\.Figure[8](https://arxiv.org/html/2609.16860#S5.F8)shows that some shift clusters are dominated by ABAB, such as Clusters 3, 8, and 9, and that others are dominated by AABB, such as Cluster 5\. A majority of the clusters contain both patterns \(Clusters 1, 2, 4, 6, 7, 10\)\. Nevertheless, LDA leave\-one\-out classification accuracy for AABB and ABAB reached 92\.88%, far above the majority\-class baseline \(50\.89%\)\. This high classification accuracy is consistent with the hypothesis that the two constructions realize different types of semantic shifts\. The high classification accuracy differs from the clusters in the t\-SNE map\. This is likely due to the t\-SNE that necessarily highlights the major ‘latent’ dimensions of variation in the shift space, whereas the LDA can take into account the full information in this space\.
The semantic shifts within the individual clusters differ in their degree of coherence\. Some clusters, such as Clusters 1, 3, and 5, have similar semantic shifts among their members and therefore receive descriptive labels in Figure[8](https://arxiv.org/html/2609.16860#S5.F8)\. Other clusters contain members with several semantic shifts and therefore do not receive a single cluster\-level label\. For these clusters, we will describe their internal composition and the types of the semantic shifts\.
Cluster 5, labeledpositive affect, consists entirely of AABB forms \(n=82\), and the semantic shifts in this cluster tend to construe the meaning of the base words in a more positive way\. For instance, 安静an1\-jing4‘quiet’ is relatively neutral as a base word, denoting the absence of noise, as in the imperative 安静\!an1\-jing4‘Be quiet\!’ or in predicative uses such as 那地方很安静na4 di4\-fang hen3 an1\-jing4‘That place is quiet’\. By contrast, 安安静静an1\-an1\-jing4\-jing4profiles quietness as something maintained over the course of an event, often with the implication that this is the appropriate or expected way\. In expressions such as 安安静静地睡觉an1\-an1\-jing4\-jing4 de shui4\-jiao4‘sleep peacefully’, the reduplication expresses a clearer evaluative stance of the speaker than the base word 安静an1\-jing4‘quiet’\. A similar tendency can also be observed with base words expressing negative affect \(n=6\)\. For instance, the base word 孤单gu1\-dan1means ‘lonely’, as in 孤单的老人 ‘a lonely elderly person’, and straightforwardly denotes a specific negative emotional state\. By contrast, 孤孤单单gu1\-gu1\-dan1\-dan1means ‘alone’, as in 留下我一个人孤孤单单的liu2\-xia4 wo3 yi1\-ge4 ren2 gu1\-gu1\-dan1\-dan1 de‘leaving me alone’\.
When the base word is already positive, reduplication reinforces this positive evaluation and often shifts the reduplicative constructions from the adjective toward adverbial use\. This can be seen in the contrast between 开心kai1\-xin1‘happy’ and 开开心心kai1\-kai1\-xin1\-xin1‘happily’\. Concordance results333The concordance check was based on the first 1,000 lines inspected for each form in the CCL corpus[https://corpus\.pku\.edu\.cn/](https://corpus.pku.edu.cn/), after restricting the search to texts from the 2000s to 2020s\.show that the most frequent immediate left collocate of 开心kai1\-xin1‘happy’ is 很hen3‘very’ \(12%\), like in 很开心hen3 kai1\-xin1‘very happy’, while the most frequent immediate right collocate is 的de\(13%\), an attributive marker used in nominal modification, like 开心的回忆kai1\-xin1 de hui2\-yi4‘happy memory’\. By contrast, 开开心心 is followed by the adverbial marker 地designificantly more often than 开心 \(25\.0% vs\. 9\.0%,p<0\.01p<0\.01, proportions test\), as in 开开心心地过大年kai1\-kai1\-xin1\-xin1 de guo4 da4\-nian2‘celebrate the New Year happily’\. It is also followed by the event verb 过guo4‘spend \(time\)’ more often than 开心 \(10\.0% vs\. 0\.0%,p<0\.01p<0\.01, proportions test\), further suggesting a stronger preference for event\-modifying, adverbial use\.
Cluster 1 contains both AABB \(n=38\) and ABAB \(n=78\) forms, with the shift vectors amplifying the base meaning along scalar dimensions such as degree, quantity, and spatial coverage\. A small number of AABB and ABAB forms \(n=5\) in this cluster also share the same base word\. Among the AABB forms, this tendency can be illustrated by 上下shang4\-xia4‘up and down’ and 上上下下shang4\-shang4\-xia4\-xia4‘from top to bottom; throughout’\. Whereas 上下shang4\-xia4encodes a basic vertical contrast, 上上下下shang4\-shang4\-xia4\-xia4extends this contrast into more exhaustive spatial coverage\. According to the concordance results, 上上下下shang4\-shang4\-xia4\-xia4is followed by 都dou1‘all’ more often than 上下shang4\-xia4\(18\.0% vs\. 1\.0%,p<0\.001p<0\.001, proportions test\), as in 上上下下都非常团结shang4\-shang4\-xia4\-xia4 dou1 fei1\-chang2 tuan2\-jie2‘people at all levels are very united’\. By comparison, 上下shang4\-xia4‘up and down’ is followed by 都dou1‘all’, in only 1\.8%, and these cases typically occur with an overt collective expression headed by 全quan2‘whole’, as in 全国上下都quan2\-guo2 shang4\-xia4 dou1‘the whole country all’\.
The base forms of the ABAB members in this cluster often begin with degree modifiers such as 太tai4‘too’, 很hen3‘very’, and 好hao3‘quite’, as in 太多tai4\-duo1‘too many’, 很冷hen3\-leng3‘very cold’, and 好远hao3\-yuan3‘quite far’\. Across thetai4\-forms in this cluster, the intensifier 实在shi2\-zai4‘indeed’ occurs before the reduplicated forms more often than before the corresponding base forms \(4\.9% vs\. 3\.1%,p<0\.05p<0\.05, proportions test\), suggesting a shift toward a more emphatic scalar construal\.
Some base words in this cluster give rise to both reduplicated patterns\. For example, 许多xu3\-duo1‘many’ gives rise to both the AABB form 许许多多xu3\-xu3\-duo1\-duo1and the ABAB form 许多许多xu3\-duo1\-xu3\-duo1\. The base form already encodes plurality, but the reduplicated forms strengthen this meaning\. This is reflected in the fact that 许多许多 is preceded by 还有hai2\-you3‘there are still’ more often than 许多 \(13\.2% vs\. 2\.4%,p<0\.001p<0\.001, proportions test\), while 许许多多 shows the same tendency, though more weakly \(4\.5% vs\. 2\.4%,p<0\.05p<0\.05, proportions test\)\. The reduplicated forms also occur more often with nouns such as 故事gu4\-shi4‘story’ and 事例shi4\-li4‘case’ in the immediate right context: 2\.0% for 许多许多 and 1\.3% for 许许多多, compared with 0\.2% for 许多 \(p<0\.001p<0\.001andp<0\.01p<0\.01, respectively, proportions test\), yielding readings such as ‘many, many stories’ and ‘numerous cases’\. A similar change can be observed for 永远yong3\-yuan3‘forever’\. The base form already encodes temporal persistence, while the reduplicated form 永远永远yong3\-yuan3\-yong3\-yuan3reinforces this temporal meaning\. In the concordance sample, 永远永远 is followed by 爱ai4‘love’ much more often than 永远 \(7\.1% vs\. 0\.2%,p<0\.001p<0\.001, proportions test\), suggesting that reduplication extends the temporal meaning of the base toward a stronger construal of enduring commitment; cf\. also 永远永远忘不了yong3\-yuan3\-yong3\-yuan3 wang4 bu4 liao3‘never, ever forget’\.
Cluster 3 is highly homogeneous in formal terms, consisting entirely of ABAB forms \(n = 130\)\. The semantic shifts in this cluster tend to recast the base event as more delimited, more tentative, and more interactionally softened\. This can be seen in the contrast between 琢磨zuo2\-mo2‘figure out; ponder’ and 琢磨琢磨zuo2\-mo0\-zuo2\-mo0‘think it over a bit’\. Whereas the base verb 琢磨zuo2\-mo2can denote sustained or serious mental engagement, the reduplicated form more often presents the event as a bounded episode of consideration\. Concordance results show that 再zai4‘again; further’ occurs immediately before 琢磨琢磨 significantly more often than before 琢磨 \(7\.0% vs\. 0\.4%,p<0\.001p<0\.001, proportions test\), as in 再琢磨琢磨zai4 zuo2\-mo0\-zuo2\-mo0‘think it over a bit more’\. Similarly, 好好hao3\-hao3‘carefully’ also occurs immediately before 琢磨琢磨 more often than before 琢磨 \(11\.1% vs\. 1\.2%,p<0\.001p<0\.001, proportions test\), as in 好好琢磨琢磨hao3\-hao3 zuo2\-mo0\-zuo2\-mo0‘think it over carefully’\. Together, these patterns suggest that the reduplicated form is more readily used in contexts of bounded, exploratory reconsideration\.
Some clusters do not support a single cluster\-level label for their shift vectors \(Clusters 4, 6, and 7\)\. Cluster 6, for example, is internally differentiated and contains several subtypes of semantic shifts\. One subtype is configurational depiction, as in 高低gao1\-di1‘high and low’ and 高高低低gao1\-gao1\-di1\-di1‘uneven’\. The base form 高低gao1\-di1‘high and low’ expresses a simple contrast in height, and as a noun means ‘height’\. By contrast, 高高低低gao1\-gao1\-di1\-di1‘uneven’ presents contrasts in height as distributed in space\. The concordance results show that 高高低低 is followed by 的de, the attributive marker, in 46\.2% of instances, compared to 6\.7% for 高低, indicating a strong preference for use as an adjective\. 高高低低 also occurs more often in descriptions of scenes, with words such as 路lu4‘road’, 山shan1‘mountain’, 树shu4‘tree’, 屋wu1‘house’, and 建筑jian4\-zhu4‘building’ \(50\.2% vs\. 20\.8%,p<0\.01p<0\.01, proportions test\)\. A second subtype involves onomatopoeia \(note, however, that a majority of onomatopoeia are found in cluster 10\)\. For instance, the base word 叮咚ding1\-dong1‘tinkle’ has as reduplication 叮叮咚咚ding1\-ding1\-dong1\-dong1‘tinkling repeatedly’, which typically occurs in rhythmic contexts with words such as 音乐yin1\-yue4‘music’, 节奏jie2\-zou4‘rhythm’, 琴qin2‘instrument’, 鼓点gu3\-dian3‘drumbeat’, and 乐音yue4\-yin1‘musical sound’ \(21\.5% vs\. 12\.4%,p<0\.01p<0\.01, proportions test\)\. The semantic shift recasts a sound label as the depiction of a repeated sound\. The last subtype shifts the base event toward a habitual construal\. The base form 唱跳chang4\-tiao4‘sing and dance’ simply denotes a combination of two activities, whereas 唱唱跳跳chang4\-chang4\-tiao4\-tiao4‘singing and dancing’ presents this activity as a repeated habitual event\. In the inspected CCL concordance samples, compared to 唱跳, 唱唱跳跳 occurs more often in texts discussing cultural activities, cf\. 文艺活动无非是唱唱跳跳,玩玩乐乐wen2\-yi4 huo2\-dong4 wu2\-fei1 shi4 chang4\-chang4\-tiao4\-tiao4, wan2\-wan2\-le4\-le4 de‘literary and artistic activities are nothing more than singing, dancing and having fun’, and 过年过节唱唱跳跳guo4\-nian2\-guo4\-jie2 chang4\-chang4\-tiao4\-tiao4‘singing and dancing during the holidays and festivals \(14\.9% vs\. 5\.7%,p<0\.01p<0\.01, proportions test\)\. 唱唱跳跳 is also more often associated with temporal expressions such as 整天zheng3\-tian1‘all day’ and 成天cheng2\-tian1‘all day long’ \(8\.1% vs\. 0\.6%,p<0\.01p<0\.01, proportions test\)\. The analysis of shift\-vectors clarifies how base words are related to their corresponding reduplicative forms\. However, for 5\.4% of the reduplicative constructions, no corresponding base word was included in the analysis\. The next section therefore turns from individual base\-reduplication pairs to the comparison of the structures of the base\-word space and the reduplication space by means of Procrustes analysis\.
## 6Procrustes analysis
The analyses presented thus far have shown that reduplicative constructions show considerable clustering in both the semantic space and the shift space\. The calculation of shift vectors requires that there is a base word that can be matched to the reduplication\. However, for some 5\.4% of the reduplications, no base word exists\. To include these words in our analyses, we made use of Procrustes analysis\.
Procrustes analysis is a statistical method originally developed for comparing shapes, such as the shapes of leaves\. We use it to compare the embeddings of base words with the embeddings of the reduplications\. The basic idea is that if sets of points in a high\-dimensional space have a similar structure, then after differences in location, scale, and orientation are removed, a rotation should be sufficient to bring one set of points into close correspondence with the other set of points\. In linguistic studies, Procrustes analysis has been used to compare semantic spaces across languages\([Yang and Baayen, 2025](https://arxiv.org/html/2609.16860#bib.bib15);[Mohiuddin and Joty, 2020](https://arxiv.org/html/2609.16860#bib.bib37)\), and to align word representations across different developmental stages\([Jorge\-Botana et al\., 2018](https://arxiv.org/html/2609.16860#bib.bib38)\)\. In the present study, we use procrustes analysis to examine whether the semantic space of base words and the semantic space of reduplicative constructions can be brought into close correspondence\. If the reduplication space can be aligned with the base\-word space with only minor distortion, this will indicate that reduplications largely preserve the semantic organization of their base words\.
We calculated the Procrustes alignment based on the centroids of the 10 clusters given by the k\-means algorithm for the base words, and the centroids of the 10 clusters that we obtained for the reduplications, also using the k\-means algorithm\. However, because the two spaces were clustered independently, their numerical cluster labels are not directly comparable\. For example, cluster 1 in the base\-word space does not necessarily correspond to cluster 1 in the reduplication space\. We therefore computed the pairwise cosine similarities between all base\-word cluster centroids and all reduplication cluster centroids, and used the Hungarian algorithm\([Kuhn, 1955](https://arxiv.org/html/2609.16860#bib.bib39)\)to select the one\-to\-one matching that maximized the total similarity between paired centroids:
max∑i=1Kπcos\(𝐛i,𝐫π\(i\)\)\\max\_\{\\pi\}\\sum\_\{i=1\}^\{K\}\\cos\\left\(\\mathbf\{b\}\_\{i\},\\mathbf\{r\}\_\{\\pi\(i\)\}\\right\)where𝐛i\\mathbf\{b\}\_\{i\}is the centroid of clusteriiin the base\-word space,𝐫j\\mathbf\{r\}\_\{j\}is the centroid of clusterjjin the reduplication space, andπ\\piis a permutation of cluster indices\. The Hungarian algorithm therefore finds the one\-to\-one assignment of reduplication clusters to base\-word clusters that maximizes the overall cosine similarity between matched centroid pairs\. Operationally, because the Hungarian algorithm solves a cost\-minimization problem, cosine similarities were converted into costs by subtracting each similarity value from the maximum similarity in the matrix\. Minimizing this cost is equivalent to maximizing the total cosine similarity between matched centroids\.
An asymmetric Procrustes alignment was then carried out on the resulting paired cluster centroids\. We used an asymmetric alignment because our goal was to inspect the reduplication space in relation to the base\-word space\. The base\-word centroid configuration was treated as the reference configuration, and the matched reduplication centroid configuration was transformed to best fit it\. The Procrustes transformation was therefore estimated from 10 matched pairs of 200\-dimensional centroid vectors, and was subsequently applied to all reduplication vectors to obtain Procrustes\-rotated reduplication vectors in the base\-word semantic space\.
The analysis was carried out using theprocrustesandprotestfunctions from theveganpackage\([Oksanen et al\., 2022](https://arxiv.org/html/2609.16860#bib.bib40)\), with significance assessed by 999 permutations\. The Procrustes analysis yielded a high correlation between the two centroid configurations,r=0\.9336r=0\.9336, with a residual sum of squares ofm122=0\.1284m\_\{12\}^\{2\}=0\.1284\. The fit was significant under permutation testing,p=0\.001p=0\.001, indicating that the base\-word and reduplication centroid configurations share a highly similar global geometry\.
The resulting Procrustes transformation was then applied to all reduplication vectors, yielding Procrustes\-transformed reduplication vectors in the base\-word semantic space\. These aligned reduplication vectors were visualized together with the base\-word vectors using t\-SNE\. The left panel of Figure[9](https://arxiv.org/html/2609.16860#S6.F9)shows the aligned semantic space\. Circles indicate base words, triangles indicate reduplicative constructions, and colors indicate the cluster correspondences established by Hungarian matching\. Clusters in which base\-word and reduplication points overlap closely provide evidence for high\-quality local alignment, whereas clusters that remain more clearly separated indicate greater divergence between the two spaces\.
The middle panel presents the cluster\-level Procrustes residuals for the matched centroid pairs\. These residuals quantify how far each rotated reduplication centroid remains from its corresponding base\-word centroid after alignment, with the smaller residuals indicating the higher degree of semantic continuity \(or semantic transparency\) between the base words and the reduplicative constructions, like clusters 3, 5, 7, 2 and 9, while larger residuals indicating stronger semantic reorganization, like clusters 6, 8, 1, 10 and 4\.
Figure 9:Procrustes analysis of the base word and reduplicative construction spaces\. Left panel: two\-dimensional t\-SNE map of base\-word vectors and Procrustes\-rotated reduplication vectors\. Circles indicate base words and triangles indicate reduplicative constructions\. Colors indicate the cluster correspondences by Hungarian matching\. Middle panel: cluster\-level Procrustes residuals for the matched centroid pairs\. Each vertical line represents the residual for one matched base\-reduplication cluster pair after Procrustes alignment\. Larger residuals indicate clusters for which the rotated reduplication centroid remains farther away from the corresponding base\-word centroid, like Clusters 6, 8, 1, 10 and 4\. Right panel: regression plot relating Hungarian matching values to Procrustes residuals\. Each point represents one matched cluster after Hungarian matching, with the number indicating the cluster label\. Cluster 2 \(Cook’s distance\>2\>2\) was excluded from this regression plot\. The negative association indicates that clusters with higher Hungarian matching values tend to have smaller Procrustes residuals after alignment \(r=−0\.72r=\-0\.72,p=\.028p=\.028\)\.As shown in the middle panel of Figure[9](https://arxiv.org/html/2609.16860#S6.F9), Cluster 3 has the smallest Procrustes residual among the ten matched cluster pairs \(0\.3230\.323\), whereas Cluster 6 has the largest residual \(0\.6300\.630\)\. A smaller residual indicates that the rotated reduplication centroid can be brought into closer correspondence with its matched base\-word centroid, while a larger residual indicates that a greater discrepancy remains after alignment\. The contrast between clusters 3 and 6 is consistent with the results of the Hungarian matching: The matched centroid pair for Cluster 3 has a cosine similarity of0\.870\.87, which is higher than the corresponding value of0\.800\.80for Cluster 6\.
In fact, the Procrustes residuals are negatively correlated with the Hungarian correlation score \(r=−0\.72,t\(7\)=−2\.766,p=0\.028r=\-0\.72,t\(7\)=\-2\.766,p=0\.028, after removal of one outlier \(Cook’s distance\>2\>2\), cluster 2, the datapoint with the largest negative Hungarian correlation score\), as shown in the right panel of Figure[9](https://arxiv.org/html/2609.16860#S6.F9)\. Clearly, the quality of the Procrustes analysis depends on the quality of the Hungarian alignment of the base word clusters and the reduplication clusters\.
The magnitude of the residuals also corresponds with the distribution of the clusters in the left panel of Figure[9](https://arxiv.org/html/2609.16860#S6.F9)\. In Cluster 3, the brown points in the upper\-left region, the circles and triangles are well mixed, suggesting that the base\-words and rotated reduplications occupy highly similar regions of the semantic space\. By contrast, the light green points of Cluster 6 are more dispersed\. The triangles are concentrated more centrally, while many of the circles are distributed to the left or to the right, suggesting that the base\-word and rotated reduplication vectors do not align even after Procrustes rotation\.
Both Hungarian matching correlations and the Procrustes residuals can be seen as measures of semantic transparency, as they gauge the extent to which the meaning of a reduplicative construction remains predictable from its base word\. A high transparency is expected to be reflected not only in the Hungarian correlation and in the Procrustes residual, but also in a greater degree of mixing between the base\-word and rotated reduplication vectors in the semantic space, as visualized in the left panel of Figure[9](https://arxiv.org/html/2609.16860#S6.F9)\.
Cluster 3 \(represented by the dark\-brown points in the upper\-left region\) shows the highest degree of semantic transparency\. In this cluster, the reduplicative construction preserves the core semantic features of the base word, while elaborating one of its inherent semantic dimensions through nominal plurality, event iteration and scalar intensification\. For example, the base noun compound 枝叶zhi1\-ye4collectively refers to ‘branches and leaves’, whereas 枝枝叶叶zhi1\-zhi1\-ye4\-ye4‘many branches and leaves’ expresses greater plurality of the same entities\. These two words show a high degree of semantic transparency, with a cosine similarity of 0\.781, because both refer to branches and leaves and just differ in referential plurality\. As for the event iteration, the base word 明灭ming2\-mie4denotes an action of ‘light on and off’, whereas 明明灭灭ming2\-ming2\-mie4\-mie4‘keep flickering on and off’ foregrounds this action as more recurrent\. Their cosine similarity of 0\.866 indicates that reduplication preserves the core action meaning while amplifying its event iteration\. The base word 模糊mo2\-hu2‘vague’ and its reduplicative construction 模模糊糊mo2\-mo2\-hu2\-hu2‘very vague’ show a high degree of semantic transparency, with a cosine similarity of 0\.782\. Both forms denote a lack of perceptual or cognitive clarity, but the reduplication describes the state at a higher degree on the scale of indistinctness, as in 模糊的印象mo2\-hu2 de yin4\-xiang4‘a vague impression’ and 模模糊糊的印象mo2\-mo2\-hu2\-hu2 de yin4\-xiang4‘remarkably vague impression’\.
By contrast, Cluster 6 has the largest Procrustes residual\. The reduplicative constructions in this cluster often develop discourse\-pragmatic uses that are not fully predictable from the lexical semantics of their bases\. For example, the base verb 睡觉shui4\-jiao4‘sleep’ denotes a sleeping event, whereas its ABAB form 睡觉睡觉shui4\-jiao4\-shui4\-jiao4‘time to sleep’ functions as a directive\. The base 走开zou3\-kai1‘walk away’ denotes a motion event, as in 他转身走开了ta1 zhuan3\-shen1 zou3\-kai1 le‘He turned and walked away’\. By contrast, 走开走开zou3\-kai1\-zou3\-kai1‘go away’ is typically used as an independent utterance to perform a directive speech act\.
Cluster 4 \(purple dots\) contains two subgroups, one dominated by base words and the other by reduplicative constructions\. The cluster has an intermediate degree of cross\-space alignment, with a Procrustes residual of 0\.520\. The reduplicative constructions generally preserve the core lexical meanings of their bases\. Many base words in this cluster function as gradable adjectives and can be modified by degree expressions such as 很hen3‘very’ and 更geng4‘more’\. However, their reduplicative counterparts generally resist such modification\(see, e\.g\.[Zhu, 2003](https://arxiv.org/html/2609.16860#bib.bib17);[Wang, 2023](https://arxiv.org/html/2609.16860#bib.bib5)\)\. For example, 高兴gao1\-xing4‘happy’ can occur in 很高兴hen3 gao1\-xing4‘very happy’ or 更高兴geng4 gao1\-xing4‘happier’, while 高高兴兴gao1\-gao1\-xing4\-xing4normally cannot occur with these degree modifiers\. At the same time, the reduplicative construction shows a stronger preference for adverbial functions, as in 高高兴兴地回家gao1\-gao1\-xing4\-xing4 de hui2\-jia1‘go home cheerfully’\. In other words, the separation between base words and reduplications within cluster 4 does not so much reflect a change in lexical meaning, but rather different preferences for use as adjectives or adverbs\.
In summary, we used Procrustes rotation to align the semantic spaces of the base words and reduplicative constructions\. The overlap between the aligned spaces reflects the general semantic transparency of reduplication, whereas the remaining differences bring its semantic effects\. In highly transparent regions, reduplicative constructions largely preserve the core meanings of their bases while amplifying scalar degree, referential plurality, or event iteration\. By contrast, the more diffuse regions with larger Procrustes residuals comprise reduplications that are less semantically transparent with respect to their base words, or that have different pragmatic functions\.
## 7General discussion
The present study reports a quantitative investigation of Mandarin AABB and ABAB reduplicative constructions using distributional semantics\.
Mandarin reduplicative constructions have a range of subtle shades of meaning that have been documented in the literature\([Chao, 1968](https://arxiv.org/html/2609.16860#bib.bib1);[Hua, 2003](https://arxiv.org/html/2609.16860#bib.bib4);[Zhu, 1998](https://arxiv.org/html/2609.16860#bib.bib41);[Zhu, 2003](https://arxiv.org/html/2609.16860#bib.bib17);[Zhang, 2015](https://arxiv.org/html/2609.16860#bib.bib16)\), but that are far from straightforward to characterize precisely\. We show that considerable headway can be made for semantic and pragmatic profiling of Mandarin reduplication using embeddings from distributional semantics\.
We first examined whether reduplicative constructions form distinct clusters in semantic space and whether the AABB and ABAB constructions show different semantic distributions\. We then profiled these constructions with respect to word category, valence, self\-relatedness, and sound\-related meaning\. For the reduplicative constructions with identifiable base words, we further investigated the semantic shifts of reduplication using shift vectors\. Finally, Hungarian matching and Procrustes analysis were used to compare the semantic spaces of the base words and the reduplicative constructions to assess semantic transparency\.
A first important result is that several generalizations previously discussed in the descriptive literature on Mandarin reduplication are recovered in the embedding space\. AABB and ABAB constructions show systematic semantic differences: they occupy partly different regions of semantic space and have different semantic profiles\. AABB constructions are more strongly associated with affective, plural, and sound\-related meanings, whereas ABAB constructions are more strongly associated with verbal, interactional, and self\-relevant meanings\. The two patterns can also be distinguished with high accuracy from both their embeddings and their shift vectors\.
This result goes beyond the qualitative descriptions by making the differences between AABB and ABAB more explicit and measurable\. Distributional semantics makes it possible to ask not only whether the two patterns differ, but also where in semantic space the differences arise, how strong they are, and which semantic properties contribute to them\.
The semantic profiling analyses further clarify that the meaning of a Mandarin reduplicative construction cannot be reduced to a single semantic category\. For example, 开开心心kai1\-kai1\-xin1\-xin1‘very happy’ is strongly positive, relatively adjective\- and adverb\-like, moderately self\-relevant, and only weakly sound\-related\. This example illustrates that several semantic tendencies may be present in the same construction at the same time\. In fact, the previous literature on Mandarin reduplication offers a bewilderingly wide range of descriptions, including intensification, plurality, iteration, vividness, affect, tentativeness, politeness, and speaker stance\. High\-dimensional embeddings succeed in capturing all these different facets of meaning jointly\. The many descriptions in the literature are not in conflict, but highlight different aspects of the same high\-dimensional semantic space\. Semantic profiling makes it possible to obtain precise but also linguistically meaningful interpretations of how reduplications structure the semantic space\.
Shift vectors specify the semantic change from a base word to its derived reduplicative construction\. Seven types of semantic shifts have been identified in the literature: intensification, iteration, tentativeness, positive affect, plurality, stance reinforcement, and onomatopoeic depiction\. These types of semantic shifts are distributed differently across AABB and ABAB constructions\. Some types of shift are strongly associated with one particular pattern\. For example, tentative meanings were found only among ABAB constructions, and the valence of positive affect is dominant among AABB constructions\.
A Procrustes analysis clarified that the overall organization of the base\-word space is largely preserved in the reduplication space\. However, this preservation is not equally strong across the semantic space\. In some regions, base words and their corresponding reduplicative constructions align closely\. Here, reduplication mainly elaborates a semantic property that is already present in the base\. For example, 山水shan1\-shui3‘mountains and waters’ and 山山水水shan1\-shan1\-shui3\-shui3‘mountains and waters everywhere; landscape’ share the same core meaning\. Here, the reduplicative construction adds a stronger plural or collective interpretation\. In other regions of semantic space, the alignment of base words and reduplications is weaker\. For example, 走开zou3\-kai1‘walk away’ denotes a motion event, whereas 走开走开zou3\-kai1\-zou3\-kai1‘Go away\!’ is typically used as a directive\. In this example, reduplication introduces additional discourse\-pragmatic meanings that are not an intrinsic part of the lexical meaning of the base\.
The Procrustes analysis was carried out using the centroids of 10 clusters of base words and 10 clusters of reduplications, which were established independently using the k\-means algorithm\. We used Hungarian matching to align as best as possible the base word clusters with the reduplication clusters, and then established the Procrustes rotation on the basis of the centroids of these clusters\. Both the Hungarian matching and the Procrustes analysis provide new ways for assessing the transparency quantitatively\. Hungarian matching measures how well independently obtained base\-word and reduplication clusters correspond to each other, whereas the Procrustes residuals measure how much mismatch remains after alignment\. The two measures are correlated: clusters with better Hungarian matching also tend to have smaller Procrustes residuals\. Taken together, the two measures make it possible to identify regions in which reduplicative constructions are more, or less, transparent with respect to their base words\.
In the present study, we made use of static embeddings that assign a fixed meaning to each reduplicated word\. These static embeddings cannot distinguish between the different meanings and discourse functions that one and the same form may have in different contexts\. They are ‘blends’ of these different senses and functions that are dominated by those senses and functions that are most frequent\. As a consequence, the semantic profiles reported in the present study are approximate, and will benefit considerably by moving from fixed embeddings to contextualized embeddings, the embeddings that can be obtained with large language models applied to word tokens in their discourse context\.
In conclusion, distributional semantics makes it possible to see what remains stable and what changes when a Mandarin disyllabic base word is used in the reduplicative construction\. It thereby provides a quantitative perspective on the semantic systematicities and the semantic versatility of Mandarin reduplication\. The present study is offered in the hope that the combination of semantic profiling, the investigation of shift vectors, and semantic\-space alignment may prove useful for studying other word\-formation and construction\-formation processes\.
## Data availability statement
The data, analysis code, and interactive figures are available through an[anonymous OSF repository](https://osf.io/42p7e/overview?view_only=d026af95c60a4d68a91ea62f03947c52)\.
## References
- Boleda \(2020\)G\. BoledaDistributional semantics and linguistic theory\.Annual Review of Linguistics6,pp\. 213–234\.External Links:[Document](https://dx.doi.org/10.1146/annurev-linguistics-011619-030303),1905\.01896Cited by:[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Chao \(1968\)Y\. R\. ChaoA grammar of spoken chinese\.University of California Press,Berkeley and Los Angeles\.Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§1](https://arxiv.org/html/2609.16860#S1.p3.1),[§7](https://arxiv.org/html/2609.16860#S7.p2.1)\.
- Chuanget al\.\(2022\)Y\. Chuang, D\. Brown, R\. H\. Baayen, and R\. EvansParadigm gaps are associated with weird “distributional semantics” properties: russian defective nouns and their case and number paradigms\.The Mental Lexicon17\(3\),pp\. 395–421\.External Links:[Document](https://dx.doi.org/10.1075/ml.22013.chu)Cited by:[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Drozdet al\.\(2016\)A\. Drozd, A\. Gladkova, and S\. MatsuokaWord embeddings, analogies, and machine learning: beyond king \- man \+ woman = queen\.InProceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers,Y\. Matsumoto and R\. Prasad \(Eds\.\),pp\. 3519–3530\.External Links:[Link](https://aclanthology.org/C16-1332)Cited by:[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Ekman \(1992\)P\. EkmanAn argument for basic emotions\.Cognition and Emotion6\(3\-4\),pp\. 169–200\.External Links:[Document](https://dx.doi.org/10.1080/02699939208411068)Cited by:[§4\.4](https://arxiv.org/html/2609.16860#S4.SS4.p1.1)\.
- Gaver \(1993\)W\. W\. GaverWhat in the world do we hear? an ecological approach to auditory event perception\.Ecological Psychology5\(1\),pp\. 1–29\.External Links:[Document](https://dx.doi.org/10.1207/s15326969eco0501%5F1)Cited by:[§4\.6](https://arxiv.org/html/2609.16860#S4.SS6.p1.1)\.
- Hua \(2003\)Y\. Hua汉语重叠研究 \[A study of reduplication in chinese\]\.Hunan People’s Publishing House,Changsha\.Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§1](https://arxiv.org/html/2609.16860#S1.p3.1),[§7](https://arxiv.org/html/2609.16860#S7.p2.1)\.
- Jorge\-Botanaet al\.\(2018\)G\. Jorge\-Botana, R\. Olmos, and J\. M\. LuzónWord maturity indices with latent semantic analysis: why, when, and where is procrustes rotation applied?\.WIREs Cognitive Science9\(1\),pp\. e1457\.External Links:[Document](https://dx.doi.org/10.1002/wcs.1457)Cited by:[§6](https://arxiv.org/html/2609.16860#S6.p2.1)\.
- Kuhn \(1955\)H\. W\. KuhnThe hungarian method for the assignment problem\.Naval Research Logistics Quarterly2\(1–2\),pp\. 83–97\.External Links:[Document](https://dx.doi.org/10.1002/nav.3800020109)Cited by:[§6](https://arxiv.org/html/2609.16860#S6.p3.1)\.
- Li and Thompson \(1981\)C\. N\. Li and S\. A\. ThompsonMandarin Chinese: a functional reference grammar\.University of California Press,Berkeley\.External Links:[Document](https://dx.doi.org/10.1525/9780520352858)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Lu and Müller \(2026\)Y\. Lu and S\. MüllerDeliminative verbal reduplication in Mandarin Chinese\.Journal of Linguistics,pp\. 1–38\.Note:First published onlineExternal Links:[Document](https://dx.doi.org/10.1017/S0022226725101047)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§1](https://arxiv.org/html/2609.16860#S1.p3.1)\.
- MacQueen \(1967\)J\. B\. MacQueenSome methods of classification and analysis of multivariate observations\.InProc\. of 5th berkeley symposium on math\. stat\. and prob\.,pp\. 281–297\.Cited by:[§3](https://arxiv.org/html/2609.16860#S3.p3.1)\.
- Marelli and Baroni \(2015\)M\. Marelli and M\. BaroniAffixation in semantic space: modeling morpheme meanings with compositional distributional semantics\.Psychological Review122\(3\),pp\. 485–515\.External Links:[Document](https://dx.doi.org/10.1037/a0039267)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Melloni and Basciano \(2018\)C\. Melloni and B\. BascianoReduplication across boundaries: the case of Mandarin\.InThe lexeme in descriptive and theoretical morphology,O\. Bonami, G\. Boyé, G\. Dal, H\. Giraudo, and F\. Namer \(Eds\.\),pp\. 325–363\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.1407013)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p3.1)\.
- Mikolovet al\.\(2013a\)T\. Mikolov, I\. Sutskever, K\. Chen, G\. S\. Corrado, and J\. DeanDistributed representations of words and phrases and their compositionality\.Advances in neural information processing systems26\.Cited by:[§2](https://arxiv.org/html/2609.16860#S2.p1.1)\.
- Mikolovet al\.\(2013b\)T\. Mikolov, W\. Yih, and G\. ZweigLinguistic regularities in continuous space word representations\.InProceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,L\. Vanderwende, H\. Daumé III, and K\. Kirchhoff \(Eds\.\),Stroudsburg, PA,pp\. 746–751\.External Links:[Link](https://aclanthology.org/N13-1090)Cited by:[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Mohiuddin and Joty \(2020\)T\. Mohiuddin and S\. JotyUnsupervised word translation with adversarial autoencoder\.Computational Linguistics46\(2\),pp\. 257–288\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00374)Cited by:[§6](https://arxiv.org/html/2609.16860#S6.p2.1)\.
- Nikolaevet al\.\(2022\)A\. Nikolaev, Y\. Chuang, and R\. H\. BaayenA generating model for finnish nominal inflection using distributional semantics\.The Mental Lexicon17\(3\),pp\. 368–394\.External Links:[Document](https://dx.doi.org/10.1075/ml.22008.nik)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Oksanenet al\.\(2022\)J\. Oksanen, G\. L\. Simpson, F\. G\. Blanchet, R\. Kindt, P\. Legendre, P\. R\. Minchin, R\. B\. O’Hara, P\. Solymos, M\. H\. H\. Stevens, E\. Szoecs, H\. Wagner, M\. Barbour, M\. Bedward, B\. Bolker, D\. Borcard, G\. Carvalho, M\. Chirico, M\. De Caceres, S\. Durand, H\. B\. A\. Evangelista, R\. FitzJohn, M\. Friendly, B\. Furneaux, G\. Hannigan, M\. O\. Hill, L\. Lahti, D\. McGlinn, M\. Ouellette, E\. Ribeiro Cunha, T\. Smith, A\. Stier, C\. J\. F\. Ter Braak, and J\. WeedonVegan: community ecology package\.Note:R package version 2\.6\-4Cited by:[§6](https://arxiv.org/html/2609.16860#S6.p5.1)\.
- Perek and Hilpert \(2017\)F\. Perek and M\. HilpertA distributional semantic approach to the periodization of change in the productivity of constructions\.International Journal of Corpus Linguistics22\(4\),pp\. 490–520\.External Links:[Document](https://dx.doi.org/10.1075/ijcl.16128.per)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Shafaei\-Bajestanet al\.\(2024\)E\. Shafaei\-Bajestan, M\. Moradipour\-Tari, P\. Uhrig, and R\. H\. BaayenThe pluralization palette: unveiling semantic clusters in english nominal pluralization through distributional semantics\.Morphology34,pp\. 369–413\.External Links:[Document](https://dx.doi.org/10.1007/s11525-024-09428-9)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Shafaei\-Bajestanet al\.\(2022\)E\. Shafaei\-Bajestan, P\. Uhrig, and R\. H\. BaayenMaking sense of spoken plurals\.The Mental Lexicon17\(3\),pp\. 337–367\.External Links:[Document](https://dx.doi.org/10.1075/ml.22011.sha)Cited by:[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Shen and Baayen \(2022a\)T\. Shen and R\. H\. BaayenAdjective–noun compounds in Mandarin: a study on productivity\.Corpus Linguistics and Linguistic Theory18\(3\),pp\. 543–572\.External Links:[Document](https://dx.doi.org/10.1515/cllt-2020-0059)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Shen and Baayen \(2022b\)T\. Shen and R\. H\. BaayenProductivity and semantic transparency: an exploration of word formation in Mandarin Chinese\.The Mental Lexicon17\(3\),pp\. 458–479\.External Links:[Document](https://dx.doi.org/10.1075/ml.22009.she)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- Songet al\.\(2018\)Y\. Song, S\. Shi, J\. Li, and H\. ZhangDirectional skip\-gram: explicitly distinguishing left and right context for word embeddings\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 \(Short Papers\),New Orleans, Louisiana,pp\. 175–180\.Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p5.1),[§2](https://arxiv.org/html/2609.16860#S2.p1.1)\.
- Stupak and Baayen \(2022\)I\. V\. Stupak and R\. H\. BaayenAn inquiry into the semantic transparency and productivity of german particle verbs and derivational affixation\.The Mental Lexicon17\(3\),pp\. 422–457\.External Links:[Document](https://dx.doi.org/10.1075/ml.22012.stu)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§5](https://arxiv.org/html/2609.16860#S5.p1.1)\.
- van der Maaten and Hinton \(2008\)L\. van der Maaten and G\. HintonVisualizing data using t\-SNE\.Journal of Machine Learning Research9,pp\. 2579–2605\.Cited by:[§3](https://arxiv.org/html/2609.16860#S3.p4.1)\.
- Venables and Ripley \(2002\)W\. N\. Venables and B\. D\. RipleyModern applied statistics with s\.4 edition,Springer,New York\.Cited by:[§3](https://arxiv.org/html/2609.16860#S3.p3.1)\.
- Wang \(2023\)C\. WangA syntactic derivation of the reduplication patterns and their interpretation in Mandarin\.Natural Language & Linguistic Theory41\(4\),pp\. 847–877\.External Links:[Document](https://dx.doi.org/10.1007/s11049-022-09549-y)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§6](https://arxiv.org/html/2609.16860#S6.p14.1)\.
- Westbury and Hollis \(2019\)C\. Westbury and G\. HollisConceptualizing syntactic categories as semantic categories: unifying part\-of\-speech identification and semantics using co\-occurrence vector averaging\.Behavior Research Methods51,pp\. 1371–1398\.External Links:[Document](https://dx.doi.org/10.3758/s13428-018-1118-4)Cited by:[§4\.1\.1](https://arxiv.org/html/2609.16860#S4.SS1.SSS1.p1.1)\.
- Westburyet al\.\(2015\)C\. Westbury, J\. Keith, B\. B\. Briesemeister, M\. J\. Hofmann, and A\. M\. JacobsAvoid violence, rioting, and outrage; approach celebration, delight, and strength: using large text corpora to compute valence, arousal, and the basic emotions\.The Quarterly Journal of Experimental Psychology68\(8\),pp\. 1599–1622\.External Links:[Document](https://dx.doi.org/10.1080/17470218.2014.970204)Cited by:[§4\.1\.1](https://arxiv.org/html/2609.16860#S4.SS1.SSS1.p1.1)\.
- Westbury and Wurm \(2022\)C\. Westbury and L\. H\. WurmIs it you you’re looking for? personal relevance as a principal component of semantics\.The Mental Lexicon17\(1\),pp\. 1–33\.External Links:[Document](https://dx.doi.org/10.1075/ml.20031.wes)Cited by:[§4\.1\.1](https://arxiv.org/html/2609.16860#S4.SS1.SSS1.p1.1)\.
- Xiaoet al\.\(2009\)R\. Xiao, P\. Rayson, and T\. McEneryA frequency dictionary of mandarin chinese: core vocabulary for learners\.Routledge,London\.Cited by:[§4\.1\.2](https://arxiv.org/html/2609.16860#S4.SS1.SSS2.p1.1)\.
- Xuet al\.\(2008\)L\. Xu, H\. Lin, Y\. Pan, H\. Ren, and J\. ChenQinggan cihui benti de gouzao \[constructing the affective lexicon ontology\]\.Qingbao Xuebao \[Journal of the China Society for Scientific and Technical Information\]27\(2\),pp\. 180–185\(Chinese\)\.External Links:[Document](https://dx.doi.org/10.3969/j.issn.1000-0135.2008.02.004)Cited by:[§4\.1\.2](https://arxiv.org/html/2609.16860#S4.SS1.SSS2.p1.1)\.
- Yang and Baayen \(2025\)Y\. Yang and R\. H\. BaayenComparing the semantic structures of the lexicons of Mandarin and English\.Language and Cognition17,pp\. e10\.External Links:[Document](https://dx.doi.org/10.1017/langcog.2024.47)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1),[§1](https://arxiv.org/html/2609.16860#S1.p5.1),[§3](https://arxiv.org/html/2609.16860#S3.p1.1),[§6](https://arxiv.org/html/2609.16860#S6.p2.1)\.
- Yang and Baayen \(2026\)Y\. Yang and R\. H\. BaayenA quantitative study of measure words in Mandarin Chinese\.Corpus Linguistics and Linguistic Theory\.Note:Advance online publicationExternal Links:[Document](https://dx.doi.org/10.1515/cllt-2025-0014)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Zhanet al\.\(2019\)W\. Zhan, R\. Guo, B\. Chang, Y\. Chen, and L\. ChenThe building of the CCL corpus: its design and implementation\.Corpus Linguistics6\(1\),pp\. 71–86\.Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p5.1),[§2](https://arxiv.org/html/2609.16860#S2.p1.1)\.
- Zhang \(2015\)N\. N\. ZhangThe morphological expression of plurality and pluractionality in mandarin\.Lingua165,pp\. 1–27\.External Links:[Document](https://dx.doi.org/10.1016/j.lingua.2015.07.001)Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p3.1),[§7](https://arxiv.org/html/2609.16860#S7.p2.1)\.
- Zhu \(1982\)D\. Zhu语法讲义 \[Lectures on grammar\]\.The Commercial Press,Beijing\.Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p1.1)\.
- Zhu \(1998\)J\. Zhu动词重叠式的语法意义 \[The grammatical meaning of verbal reduplication\]\.中国语文 \[Studies of the Chinese Language\]\(5\),pp\. 378–386\.Cited by:[§7](https://arxiv.org/html/2609.16860#S7.p2.1)\.
- Zhu \(2003\)J\. Zhu形容词重叠式的语法意义 \[The grammatical meaning of adjectival reduplication\]\.语文研究 \[Linguistic Research\]\(3\),pp\. 9–17\.Cited by:[§1](https://arxiv.org/html/2609.16860#S1.p3.1),[§6](https://arxiv.org/html/2609.16860#S6.p14.1),[§7](https://arxiv.org/html/2609.16860#S7.p2.1)\.Similar Articles
Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis
This paper presents a case study using unsupervised articulatory probing to examine how self-supervised speech models encode phonetic features across Mandarin sub-dialects, finding that salient features like labiality remain stable while finer spectral distinctions show dialect-dependent variation.
LLMs for automatic annotation of Mandarin narrative transcripts
This paper evaluates LLMs for automatically annotating narrative macrostructure in spoken Mandarin, finding that the best model achieves near-human reliability while reducing annotation time by 65%, though performance degrades on semantically complex or lexically diverse narratives.
An ERP Study on Recursive Locative Processing in Mandarin-Speaking Children with Autism
This ERP study examines recursive locative processing in Mandarin-speaking children with autism, finding reduced early predictive engagement and increased semantic integration demands in the ASD group.
The Proxy Presumption: From Semantic Embeddings to Valid Social Measures
This paper critiques the 'Proxy Presumption' in NLP, where geometric embedding properties are incorrectly equated with social constructs. It introduces the Construct Validity Protocol and Counterfactual Neutralization methods to ensure rigorous validation of social measures derived from semantic embeddings.
Psychological Constructs in Shared Semantic Space
This paper proposes a framework using Supervised Semantic Differential to represent psychological constructs as directions in a shared word-embedding space, enabling comparison across different measurement instruments and research traditions.