Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models

arXiv cs.CL Papers

Summary

This paper studies how humans and large language models linguistically accommodate each other during multi-turn conversations, finding that LLMs overconverge to user style while humans accommodate LLMs no differently than humans.

arXiv:2605.29278v1 Announce Type: new Abstract: As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We present a large-scale study of linguistic convergence in human-LLM dialogue, examining how humans and LLMs accommodate each other's linguistic style during multi-turn conversations. Using an asymmetric convergence metric on WildChat, a corpus of real-world ChatGPT transcripts, we find that while LLMs significantly overconverge toward their users on both function word and open-class features across eight languages, human convergence rates in this setting are broadly consistent with human-human baselines. These findings suggest that accommodation in human-LLM dialogue is asymmetric: while LLMs dramatically overfit to their users' style, humans linguistically accommodate LLMs no differently than they would another person.
Original Article
View Cached Full Text

Cached at: 05/29/26, 09:17 AM

# Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models
Source: [https://arxiv.org/html/2605.29278](https://arxiv.org/html/2605.29278)
Terra Blevins Khoury College of Computer Sciences Northeastern University, Boston, MA t\.blevins@northeastern\.edu

###### Аннотация

As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question\. We present a large\-scale study of linguistic convergence in human\-LLM dialogue, examining how humans and LLMs accommodate each other’s linguistic style during multi\-turn conversations\. Using an asymmetric convergence metric on WildChat, a corpus of real\-world ChatGPT transcripts, we find that while LLMs significantly overconverge toward their users on both function word and open\-class features across eight languages, human convergence rates in this setting are broadly consistent with human\-human baselines\. These findings suggest that accommodation in human\-LLM dialogue is asymmetric: while LLMs dramatically overfit to their users’ style, humans linguistically accommodate LLMs no differently than they would another person\.

Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models

Terra BlevinsKhoury College of Computer SciencesNortheastern University, Boston, MAt\.blevins@northeastern\.edu

## 1Introduction

Large language models \(LLMs\) can produce text indistinguishable from human language in dialogue settings, passing the Turing test\(Jones and Bergen,[2025](https://arxiv.org/html/2605.29278#bib.bib18)\)\. As these models are increasingly deployed in professional and personal settings, LLMs become new linguistic actors in natural language settings that users can interact with \(almost\) as naturally as they would another person\. Yet despite their growing presence, little is known about how interacting with LLMs shapes the linguistic behavior of their human interlocutors\.

We approach this problem through the lens oflinguistic accommodation, the tendency of speakers to adapt their language based on their interlocutor’s over the course of a conversation\. This linguistic adaptation is largely unconscious and manifests across a wide range of linguistic features, from phonology and syntax to lexical choices\(Giles et al\.,[1991](https://arxiv.org/html/2605.29278#bib.bib15)\), and humans have been observed exhibiting some accommodation behaviors to facilitate communication with dialogue systemsBrennan \([1996](https://arxiv.org/html/2605.29278#bib.bib9)\)\. In this work, we build on prior computational methods for studying accommodation in dialogue systems\(Ireland et al\.,[2011](https://arxiv.org/html/2605.29278#bib.bib17); Danescu\-Niculescu\-Mizil and Lee,[2011](https://arxiv.org/html/2605.29278#bib.bib13)\)to analyze whether the language of people and modelsconvergesduring human\-LLM interactions\.

We use this framework to ask:

R1: Do LLMs converge to the style of their users over the course of a multi\-turn conversation?

R2: Do human users converge to the style of their LLM interlocutor?

R3: Does human accommodation of LLMs differ from their accommodation of human interlocutors?

R1is partially addressed in prior work:Kandra et al\. \([2025](https://arxiv.org/html/2605.29278#bib.bib20)\)test whether two LLMs syntactically converge over time, andBlevins et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib6)\)evaluate whether LLMs stylistically accommodate their context when asked to replace turns in a preexisting conversation\. However, these settings are both synthetic evaluations using model\-model interactions and turn\-replacement prompts, respectively; no studies have examined the per\-turn coordination of both humans and LLMs in a realistic chatbot setting across multiple languages\. Moreover, the human side of these interactions remains underexplored, with open questions about if and how users adapt their language when conversing with LLMs\.

We examine these questions with WildChat\(Zhao et al\.,[2024](https://arxiv.org/html/2605.29278#bib.bib33)\), a corpus of human\-LLM conversations drawn from the transcripts of opt\-in ChatGPT interactions, to compare the accommodation behavior of LLMs and humans across eight languages on multiple forms of lexical convergence, contrasting the human behavior in this setting with that observed in human\-human dialogue\. Our analysis reveals that LLMs accommodate users significantly more than humans accommodate LLMs, consistent with prior synthetic evaluationsBlevins et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib6)\)\. Surprisingly, we also find that while users accommodate less than LLMs in WildChat dialogues, human accommodation rates in human\-LLM interaction are broadly consistent with human\-human dialogue\. This asymmetry holds for all features and languages examined, suggesting that the accommodation of both humans and LLMs in this setting is linguistically and cross\-lingually robust\. These findings suggest that humans, at least linguistically, accommodate modern LLMs no differently than they would another person, despite the asymmetric and overaccommodating behavior exhibited by the LLM\.

## 2Related Work

Accommodation is a well\-documented phenomenon where speakers adapt their language to their interlocutor, manifesting across features from phonology to lexical choice\(Giles et al\.,[1991](https://arxiv.org/html/2605.29278#bib.bib15); Brennan and Clark,[1996](https://arxiv.org/html/2605.29278#bib.bib10)\)\. We focus on linguistic convergence, where a speaker’s style becomes more similar to their interlocutor’s over timeNiederhoffer and Pennebaker \([2002](https://arxiv.org/html/2605.29278#bib.bib27)\)\.Ireland et al\. \([2011](https://arxiv.org/html/2605.29278#bib.bib17)\)andDanescu\-Niculescu\-Mizil and Lee \([2011](https://arxiv.org/html/2605.29278#bib.bib13)\)develop computational methods for measuring convergence in large corpora, enabling the study of accommodation at scale, andWard and Litman \([2007](https://arxiv.org/html/2605.29278#bib.bib31)\)extends this to lexical entrainment, showing that accommodation on content words can be measured automatically\. These methods have been broadly applied to human\-human dialogue settings\(Mukherjee and Liu,[2012](https://arxiv.org/html/2605.29278#bib.bib24); Bawa et al\.,[2018](https://arxiv.org/html/2605.29278#bib.bib3); Berdičevskis and Erbro,[2023](https://arxiv.org/html/2605.29278#bib.bib4), i\.a\.\)\.

Recently, these methods have been used to study LLM accommodation, albeit in artificial settings\.Kandra et al\. \([2025](https://arxiv.org/html/2605.29278#bib.bib20)\)find that LLMs syntactically converge in model\-to\-model interactions, andBlevins et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib6)\)show that LLMs accommodate their context when completing turns in existing human conversations\. We extend these findings by studying accommodation in preexisting human\-LLM interactions, enabling a direct comparison of LLM and human accommodation in this setting\. Concurrently to our work,Chen et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib11)\)apply a bidirectional within\-vs\-between design to English GPT\-4o conversations and find a similar LLM\-User asymmetry\. We corroborate these prior findings and show that LLMs also overconverge in this deployed, realistic setting\.

A parallel line of work on computers as social actors shows that people unconsciously apply social rules to computers\(Nass et al\.,[1994](https://arxiv.org/html/2605.29278#bib.bib26); Nass and Moon,[2000](https://arxiv.org/html/2605.29278#bib.bib25)\)\. However, the case for user accommodation of computers is mixed:Brennan \([1996](https://arxiv.org/html/2605.29278#bib.bib9)\)finds that users entrain to computer\-generated referring expressions, but recent work on human\-model interaction indicates that users adopt different linguistic behaviors with LLMsBhatt and Rios \([2021](https://arxiv.org/html/2605.29278#bib.bib5)\); Zhang and Yu \([2025](https://arxiv.org/html/2605.29278#bib.bib32)\)\. We directly address this question, finding that human accommodation is broadly stable across LLM\- and human\-human settings\.

## 3Methods and Experimental Setup

To study how humans and LLMs accommodate each other, we analyze conversations from real\-world corpora across multiple languages\.

Таблица 1:Dataset statistics for all experimental settings\. WildChat represents human\-LLM dialogue, while DailyDialog and Ubuntu cover human\-human conversations\.Таблица 2:Convergence scores for WildChat \(EN\), DailyDialog, and Ubuntu on averaged NOUN and LIWC scores, as well as each function word class in LIWC\. All values are significant atp<0\.05p<0\.05unless underlined\.Δ\\Delta= WildChat LLM−\-User\. DailyDialog scores are averaged across both speakers\.### 3\.1Data

Our primary analysis is conducted on WildchatZhao et al\. \([2024](https://arxiv.org/html/2605.29278#bib.bib33)\), a corpus of approximately one million real\-world conversations between human users and ChatGPT \(GPT\-3\.5\-Turbo and GPT\-4\) spanning many topics and languages, making it well\-suited for studying organic human\-LLM interactions\. For our study, we filter these conversations to those with 2 to 10 \(human, LLM\) turn pairs\. We then sample up to 5,000 conversations per language across 8 languages \(English, French, Spanish, Portuguese, Italian, Russian, Chinese, and Turkish\), using stratified sampling by turn count to ensure a range of conversation lengths in our sample\. Included languages must occur frequently \(∼2\.5​k\\sim 2\.5kor more conversations after filtering\) and have the language\-specific resources needed to calculate convergence features\.

As a counterpoint to our study of LLM\-human conversations, we also consider speaker behavior in human\-human dialogue settings\. Given the inherent paired structure of chatbot conversations, we focus on dyadic settings, or conversations with two speakers, and consider two English datasets that meet this criteria: DailyDialogLi et al\. \([2017](https://arxiv.org/html/2605.29278#bib.bib21)\), a widely used dialogue dataset of everyday conversations, and the original conversations in the Ubuntu Dialogue CorpusLowe et al\. \([2015](https://arxiv.org/html/2605.29278#bib.bib22)\), a corpus of technical support conversations from the Ubuntu IRC channel that provides a task\-oriented human\-human baseline representing a common LLM use case\.111Due to their release dates, there is little risk of machine\-generated content in either corpus\.We follow the same preprocessing steps as on WildChat for human\-human baselines, requiring all turns between speakers to be paired\. Table[1](https://arxiv.org/html/2605.29278#S3.T1)provides our data statistics after filtering\.

### 3\.2Convergence Metrics

Unlike prior work on accommodation in LLMs\(e\.g\., Blevins et al\.,[2026](https://arxiv.org/html/2605.29278#bib.bib6)\), we study the separate behaviors of humans and LLMs in interactive chatbot settings, thus requiring an asymmetric measure to disentangle the effect of conversational partners on each other\. We adopt the convergence metric fromDanescu\-Niculescu\-Mizil and Lee \([2011](https://arxiv.org/html/2605.29278#bib.bib13)\), which measures the degree to which one speaker accommodates another on a binary featuret\. For exchanges whereAAinitiates andBBresponds, letata^\{t\}andb→atb^\{t\}\_\{\\rightarrow a\}be indicator variables forAA’s utterances andBB’s replies exhibiting featurett\. Then the coordination ofBBtoAAis defined asConv​\(B→A,t\)\\text\{Conv\}\(B\\rightarrow A,t\):

P​\(b→at=1∣at=1\)−P​\(b→at=1\)P\(b^\{t\}\_\{\\rightarrow a\}=1\\mid a^\{t\}=1\)\-P\(b^\{t\}\_\{\\rightarrow a\}=1\)\(1\)Positive scores indicateBBusesttabove their baseline rate whenAAdoes \(convergence\); negative scores indicatedivergence\. Conv is averaged across all turns in the sample, unless otherwise stated\.

We measure convergence on two feature types:

LIWC function wordsare frequent, largely unconscious markers that are sensitive to accommodation effects in human dialogue\(Niederhoffer and Pennebaker,[2002](https://arxiv.org/html/2605.29278#bib.bib27)\)\. FollowingDanescu\-Niculescu\-Mizil and Lee \([2011](https://arxiv.org/html/2605.29278#bib.bib13)\)andIreland et al\. \([2011](https://arxiv.org/html/2605.29278#bib.bib17)\), we measure convergence on function word categories from the Linguistic Inquiry and Word Count \(LIWC\) lexicon\(Chung and Pennebaker,[2012](https://arxiv.org/html/2605.29278#bib.bib12)\)\. We use language\-specific LIWC 2007 dictionaries \(Appendix[A](https://arxiv.org/html/2605.29278#A1)\) for the languages in our study, reporting per\-category and mean scores\.222Function word coverage varies slightly by language due to differences in grammar and dictionary versions — for example, articles are absent in Chinese and Turkish\.

NOUN lemmasreflect lexical entrainmentBrennan and Clark \([1996](https://arxiv.org/html/2605.29278#bib.bib10)\), or the tendency of speakers to adopt each other’s specific vocabulary choices when discussing a concept\. FollowingWard and Litman \([2007](https://arxiv.org/html/2605.29278#bib.bib31)\), we extract the 100 most frequent noun lemmas per corpus after filtering to nouns with at least one WordNet synonym\(Fellbaum,[1998](https://arxiv.org/html/2605.29278#bib.bib14)\)when possible to ensure each lemma represents a lexical choice\.333WordNet filtering is available for EN, FR, ES, PT, IT, and ZH via NLTK OMW 1\.4Bond and Foster \([2013](https://arxiv.org/html/2605.29278#bib.bib7)\); RU and TR use frequency\-only filtering due to missing coverage\.We report mean convergence score across selected nouns in the vocabulary; Appendix Table[5](https://arxiv.org/html/2605.29278#A2.T5)reports which lemmas are selected in each setting\.

## 4Results

Table[2](https://arxiv.org/html/2605.29278#S3.T2)reports our experimental results on English WildChat, as well as the two human\-human baselines\. Generally, humans and LLMs both exhibit significant convergence over random chance\. Furthermore, we observe that:

#### Human speakers accommodate LLMs and humans similarly

We see broadly consistent convergence scores across the human speaker settings ofWildChat User,DailyDialogspeakers, and bothUbuntusettings \(which are split due to the different roles of the user asking for help and the expert assisting the user\): mean LIWC convergence scores for human speakers range from \.022–\.036 across all three settings, and noun convergence scores from \.104–\.141\. This indicates that not only are speakers generally consistent across different dialogue settings and roles in our chosen corpora, but also that their behavior towards LLM interlocutors is similar on the considered metrics\. One exception where speakers accommodate LLMsmoreis on negation:WildChatusers show notably higher negation convergence \(\.081\) than the non\-significant scores ofDailyDialog\(\-\.005\) orUbuntu\(\.012\), which is potentially driven by the \(much\) larger convergence of LLMs on the same word class \(\.213\)\.

#### LLMs overconverge to their user

We find that LLMs significantlyoverconvergetowards their users, both relative to the users they are speaking with and to human\-human accommodation rates in our two baseline corpora\. This is consistent with prior workBlevins et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib6)\)and further shows that LLM mirroring of user style is excessive relative to human behavior even in realistic, deployed settings\. Specifically, mean convergence of LLMs on LIWC is roughly twice the user rate \(\.068\.068vs\.035\.035\), with relative increases of LLM scores over users on individual features ranging from22%22\\%on prepositions to164%164\\%on negating function words \(an outlier that contrasts sharply with the minimal negation convergence seen in both human\-human baselines\)\. LLM convergence on nouns is even more pronounced: ChatGPT achieves a mean noun convergence of\.442\.442, compared to scores of\.141\.141for users in WildChat and\.104\.104to\.132\.132across the human\-human baselines\. As LLM noun convergence occurs at over three times the user rate, this indicates ChatGPT is even more sensitive to user lexical choice beyond function words\.

![Refer to caption](https://arxiv.org/html/2605.29278v1/x1.png)Рис\. 1:Mean \(a\) LIWC and \(b\) NOUN scores for LLMs and users across eight languages in WildChat\. LLMs consistently converge more than users across all languages; the reported values indicateΔ\\Delta= LLM \- User\.
#### LLM convergence is robust across languages

The observed asymmetry between user and LLM convergence in WildChat holds across the eight considered languages \(Figure[1](https://arxiv.org/html/2605.29278#S4.F1)\)\.444Full per\-language results are in Appendix Table[4](https://arxiv.org/html/2605.29278#A2.T4)\.Despite typological differences across the eight languages and \(presumed\) variation in ChatGPT training data coverage, the asymmetry is remarkably consistent\. ChatGPT’s average LIWC convergence scores range from\.048\.048\(PT\) to\.135\.135\(IT\), while user scores remain stable between\.016\.016\(ZH\) and\.052\.052\(IT\); relatively, ChatGPT converges between95%95\\%\-213%213\\%above user rates across languages\. Notably, English, which is often the best\-resourced language for LLM training, falls in the middle of this range rather than at the top or bottom for any individual feature, suggesting the asymmetry is not simply a function of model quality on a given language\. These same patterns hold for NOUN convergence, with LLM scores exceeding user rates by214%−330%214\\%\-330\\%across all languages; this increase in convergence on nouns also reinforces the finding that LLMs are particularly responsive to user choices of open\-class lexical items\.

![Refer to caption](https://arxiv.org/html/2605.29278v1/x2.png)Рис\. 2:Mean LIWC convergence scores per turn position for LLMs and users across eight WildChat languages\. Significant trends are indicated with an∗\*, saturated lines, and reported correlation value\(p<0\.05p<0\.05\)\.
#### Per\-turn convergence is consistent over time

While our prior experiments consider the overall convergence rates in human\-LLM dialogue, it remains an open question whether these rates change over the course of conversations\. We therefore perform a temporal analysis to compare the per\-turn convergence rates of users and LLMs; Figure[2](https://arxiv.org/html/2605.29278#S4.F2)reports the mean LIWC scores at each turn in the conversations, while Appendix Figure[3](https://arxiv.org/html/2605.29278#A2.F3)shows the corresponding NOUN results\.555We do not include cases where there are<100<100utterances \(e\.g\., on TR turns 9 and 10, as well as IT, PT turn 10\)\.We also fit a linear regression of convergence score on turn position for each setting to test whether scores change significantly over time, finding that 10 of 32 cases show a significant linear trend \(p<0\.05p<0\.05\), indicated with their slope values in each figure\.

Overall, we find no consistent patterns, with most fluctuations appearing as noise or barely significant \(only five cases remain so under a stricter p\-value of0\.010\.01\)\. Significant slopes are also distributed across both positive and negative trends: RU users and models increasingly converge on LIWC while PT speakers both diverge, suggesting that even real trends in the data reflect language\-specific variation rather than a systematic tendency for either speaker type to accommodate more or less over time\. Human\-human baselines are similarly flat on LIWC across turns, with no significant trends in either DailyDialog or Ubuntu\. However, DailyDialog NOUN scores show a significant decreasing trend \(p<\.01p<\.01\), which may be due to the domain \(language learners\) or increased data sparsity on later turns\.

These findings somewhat contrast withChen et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib11)\), who study English GPT\-4o conversations \(n=1319n=1319\) and report progressive user convergence on personal pronouns across turns\. The discrepancy may be due to methodological differences: they measure whether speakers become increasingly similar to their own conversational partner relative to random conversations, whereas we measure local turn\-by\-turn coordination\.666Note that stable coordination under our metric does not preclude the increasing similarity thatChen et al\. \([2026](https://arxiv.org/html/2605.29278#bib.bib11)\)report\. If speakers consistently coordinate throughout a conversation, their language will cumulatively converge even without an increase in the per\-turn coordination rate\.It may also reflect a setting\-specific case, given that this effect is limited to personal pronouns and does not replicate across their other features, or the broader set of languages and corpus sizes examined here\.

## 5Conclusion

We study linguistic convergence in human\-LLM dialogue, finding that LLMs significantly overconverge toward their users on both function and open\-class lexical features, while human accommodation rates remain stable across LLM and human\-human settings\. This asymmetry holds consistently across eight languages and both feature types, suggesting it is a robust property of human\-LLM interaction\.

Notably, the conclusions we draw about human behavior differ fromBhatt and Rios \([2021](https://arxiv.org/html/2605.29278#bib.bib5)\): they find that humans accommodate bots differently than other humans, which they attribute to their systems often producing incoherent responses that disrupt natural conversation, unlike the more coherent outputs of modern LLMs\. Based on the consensus that linguistic accommodation is a largely unconscious process\(Giles et al\.,[1991](https://arxiv.org/html/2605.29278#bib.bib15); Pickering and Garrod,[2004](https://arxiv.org/html/2605.29278#bib.bib28)\), we hypothesize that LLM fluency acts as a bottleneck: once the model produces sufficiently coherent language in dialogue, these processes occur as they do in human\-human dialogue, regardless of the interlocutor’s nature\. As model quality improves and LLMs are increasingly integrated into everyday communication, understanding how humans subconsciously treat these systems will be critical for anticipating the long\-term effects of LLMs on natural language\.

## Limitations

Our analysis is limited to a single LLM provider, as WildChat conversations are drawn exclusively from OpenAI’s ChatGPT \(GPT\-3\.5\-Turbo and GPT\-4\); thus, it is unclear whether the observed overconvergence generalizes to other LLM families, though it is consistent with prior work\(Blevins et al\.,[2026](https://arxiv.org/html/2605.29278#bib.bib6); Kandra et al\.,[2025](https://arxiv.org/html/2605.29278#bib.bib20)\)\. Relying exclusively on ChatGPT introduces the confound of post\-training alignment, which we do not attempt to disentangle from the base pretraining objective but leave for future work\. Additionally, although we study eight languages, our human\-human baselines are English\-only, which limits the cross\-lingual scope of our findings\. LIWC dictionary coverage also varies across languages, with some categories absent entirely for certain languages \(e\.g\., RU, TR, and ZH all do not have an article category\), which may affect cross\-lingual comparability of our per\-feature LIWC results\. WordNet’s OMW coverage is similarly limited by its lack of support for RU and TR, which may affect our NOUN results on those languages\. Beyond WordNet, our noun lemma vocabularies are still corpus\-specific and may reflect domain biases of their underlying data \(for example, the prevalence of technical terms in the vocabulary for Ubuntu\), which could skew noun convergence scores based on the conversational setting\. Future work should examine convergence behavior across a broader range of LLM families and deployment settings and develop human\-human dialogue baselines in languages beyond English to fully enable a cross\-lingual comparison of human and LLM accommodation behavior\.

## Список литературы

- Agosti and Rellini \(2007\)Alberto Agosti and Alessandra Rellini\. 2007\.The italian liwc dictionary\.*Austin, TX: LIWC\. Net*\.
- Balage Filho et al\. \(2013\)Pedro Balage Filho, Thiago Alexandre Salgueiro Pardo, and Sandra Aluísio\. 2013\.An evaluation of the brazilian portuguese liwc dictionary for sentiment analysis\.In*Proceedings of the 9th Brazilian Symposium in Information and Human Language Technology*\.
- Bawa et al\. \(2018\)Anshul Bawa, Monojit Choudhury, and Kalika Bali\. 2018\.Accommodation of conversational code\-choice\.In*Proceedings of the Third Workshop on Computational Approaches to Linguistic Code\-Switching*, pages 82–91\.
- Berdičevskis and Erbro \(2023\)Aleksandrs Berdičevskis and Viktor Erbro\. 2023\.You say tomato, i say the same: A large\-scale study of linguistic accommodation in online communities\.In*Proceedings of the 24th Nordic Conference on Computational Linguistics \(NoDaLiDa\)*, pages 415–424\.
- Bhatt and Rios \(2021\)Paras Bhatt and Anthony Rios\. 2021\.Detecting bot\-generated text by characterizing linguistic accommodation in human\-bot interactions\.In*Findings of the Association for Computational Linguistics: ACL\-IJCNLP 2021*, pages 3235–3247\.
- Blevins et al\. \(2026\)Terra Blevins, Susanne Schmalwieser, and Benjamin Roth\. 2026\.[Do language models accommodate their users? a study of linguistic convergence](https://doi.org/10.18653/v1/2026.eacl-long.34)\.In*Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 791–807, Rabat, Morocco\. Association for Computational Linguistics\.
- Bond and Foster \(2013\)Francis Bond and Ryan Foster\. 2013\.Linking and extending an open multilingual wordnet\.In*Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 1352–1362\.
- Boyd et al\. \(2022\)Ryan L Boyd, Ashwini Ashokkumar, Sarah Seraj, and James W Pennebaker\. 2022\.The development and psychometric properties of liwc\-22\.*Austin, TX: University of Texas at Austin*, 10\(1\-47\):6\.
- Brennan \(1996\)Susan E Brennan\. 1996\.Lexical entrainment in spontaneous dialog\.*Proceedings of ISSD*, 96:41–44\.
- Brennan and Clark \(1996\)Susan E Brennan and Herbert H Clark\. 1996\.Conceptual pacts and lexical choice in conversation\.*Journal of experimental psychology: Learning, memory, and cognition*, 22\(6\):1482\.
- Chen et al\. \(2026\)Pengbo Chen, Huining Guan, and Eui Jun Jeong\. 2026\.Who accommodates whom? bidirectional linguistic accommodation and progressive interpersonal convergence in human–ai conversations\.*Behavioral Sciences*\.
- Chung and Pennebaker \(2012\)Cindy K Chung and James W Pennebaker\. 2012\.Linguistic inquiry and word count \(liwc\): pronounced “luke,”… and other useful facts\.In*Applied natural language processing: Identification, investigation and resolution*, pages 206–229\. IGI Global Scientific Publishing\.
- Danescu\-Niculescu\-Mizil and Lee \(2011\)Cristian Danescu\-Niculescu\-Mizil and Lillian Lee\. 2011\.Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs\.In*Proceedings of the 2nd workshop on cognitive modeling and computational linguistics*, pages 76–87\.
- Fellbaum \(1998\)Christiane Fellbaum\. 1998\.*WordNet: An electronic lexical database*\.MIT press\.
- Giles et al\. \(1991\)Howard Giles, Nikolas Coupland, and Justine Coupland\. 1991\.Accommodation theory: Communication, context, and consequence\.*Contexts of accommodation: Developments in applied sociolinguistics*, 1:1–68\.
- Honnibal et al\. \(2020\)Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, and 1 others\. 2020\.spacy: Industrial\-strength natural language processing in python\.
- Ireland et al\. \(2011\)Molly E Ireland, Richard B Slatcher, Paul W Eastwick, Lauren E Scissors, Eli J Finkel, and James W Pennebaker\. 2011\.Language style matching predicts relationship initiation and stability\.*Psychological science*, 22\(1\):39–44\.
- Jones and Bergen \(2025\)Cameron R Jones and Benjamin K Bergen\. 2025\.Large language models pass the turing test\.*arXiv preprint arXiv:2503\.23674*\.
- Kailer and Chung \(2011\)Andreas Kailer and Cindy K Chung\. 2011\.The russian liwc2007 dictionary\.*Austin, TX: LIWC\. net*\.
- Kandra et al\. \(2025\)Florian Kandra, Vera Demberg, and Alexander Koller\. 2025\.[LLMs syntactically adapt their language use to their conversational partner](https://doi.org/10.18653/v1/2025.acl-short.68)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 873–886, Vienna, Austria\. Association for Computational Linguistics\.
- Li et al\. \(2017\)Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu\. 2017\.Dailydialog: A manually labelled multi\-turn dialogue dataset\.In*Proceedings of the Eighth International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pages 986–995\.
- Lowe et al\. \(2015\)Ryan Lowe, Nissan Pow, Iulian Vlad Serban, and Joelle Pineau\. 2015\.The Ubuntu Dialogue Corpus: A large dataset for research in unstructured multi\-turn dialogue systems\.In*Proceedings of the 16th annual meeting of the special interest group on discourse and dialogue*, pages 285–294\.
- Lyu et al\. \(2022\)Yuwen Lyu, Julian Chun\-Chung Chow, Ji\-Jen Hwang, Zhi Li, Cheng Ren, and Jungui Xie\. 2022\.Psychological well\-being of left\-behind children in china: text mining of the social media website zhihu\.*International journal of environmental research and public health*, 19\(4\):2127\.
- Mukherjee and Liu \(2012\)Arjun Mukherjee and Bing Liu\. 2012\.Analysis of linguistic style accommodation in online debates\.In*Proceedings of COLING 2012*, pages 1831–1846\.
- Nass and Moon \(2000\)Clifford Nass and Youngme Moon\. 2000\.Machines and mindlessness: Social responses to computers\.*Journal of social issues*, 56\(1\):81–103\.
- Nass et al\. \(1994\)Clifford Nass, Jonathan Steuer, and Ellen R Tauber\. 1994\.Computers are social actors\.In*Proceedings of the SIGCHI conference on Human factors in computing systems*, pages 72–78\.
- Niederhoffer and Pennebaker \(2002\)Kate G Niederhoffer and James W Pennebaker\. 2002\.Linguistic style matching in social interaction\.*Journal of language and social psychology*, 21\(4\):337–360\.
- Pickering and Garrod \(2004\)Martin J Pickering and Simon Garrod\. 2004\.Toward a mechanistic psychology of dialogue\.*Behavioral and brain sciences*, 27\(2\):169–190\.
- Piolat et al\. \(2011\)Annie Piolat, RJ Booth, Cindy K Chung, M Davids, and JW Pennebaker\. 2011\.The french dictionary for liwc: Modalities of construction and examples of use\.*Psychologie française*, 56\(3\):145–159\.
- Ramírez\-Esparza et al\. \(2007\)Nairán Ramírez\-Esparza, James W Pennebaker, Florencia Andrea García, and Raquel Suriá\. 2007\.La psicología del uso de las palabras: Un programa de computadora que analiza textos en español\.*Revista mexicana de psicología*, 24\(1\):85–99\.
- Ward and Litman \(2007\)Arthur Ward and Diane J Litman\. 2007\.Automatically measuring lexical and acoustic/prosodic convergence in tutorial dialog corpora\.In*SLaTE*, pages 57–60\.
- Zhang and Yu \(2025\)Fulei Zhang and Zhou Yu\. 2025\.Mind the gap: Linguistic divergence and adaptation strategies in human\-llm assistant vs\. human\-human interactions\.*arXiv preprint arXiv:2510\.02645*\.
- Zhao et al\. \(2024\)Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng\. 2024\.Wildchat: 1m chatgpt interaction logs in the wild\.In*The Twelfth International Conference on Learning Representations*\.

## Приложение AMore Experimental Details

We use LIWC dictionaries to obtain inventories of words that fall into each function word class measured as a proxy for linguistic convergence in prior workIreland et al\. \([2011](https://arxiv.org/html/2605.29278#bib.bib17)\); Danescu\-Niculescu\-Mizil and Lee \([2011](https://arxiv.org/html/2605.29278#bib.bib13)\): personal pronouns \(ppron\), impersonal pronouns \(ipron\), articles, conjunctions \(conj\), prepositions \(prep\), auxiliary verbs \(auxvb\), adverbs, negations \(negate\), and quantifiers \(quant\)\. In non\-English experimental settings, we use language\-specific dictionaries that were constructed in the vein of LIWC\-2007: SpanishRamírez\-Esparza et al\. \([2007](https://arxiv.org/html/2605.29278#bib.bib30)\); FrenchPiolat et al\. \([2011](https://arxiv.org/html/2605.29278#bib.bib29)\); \(Brazilian\) PortugueseBalage Filho et al\. \([2013](https://arxiv.org/html/2605.29278#bib.bib2)\); ItalianAgosti and Rellini \([2007](https://arxiv.org/html/2605.29278#bib.bib1)\); RussianKailer and Chung \([2011](https://arxiv.org/html/2605.29278#bib.bib19)\); TurkishBoyd et al\. \([2022](https://arxiv.org/html/2605.29278#bib.bib8)\); and Simplified ChineseLyu et al\. \([2022](https://arxiv.org/html/2605.29278#bib.bib23)\)\. We note that we do not include other languages, such as German, whose dictionaries are based on different versions of LIWC, as these have lower overlap with the word\-class categories considered in this paper\.

We report the lemma inventories used to calculate NOUN scores for all settings in Table[5](https://arxiv.org/html/2605.29278#A2.T5)\. While these inventories are automatically extracted from the data, they are mostly reasonable for their target language; however, the Turkish vocabulary contains some English tokens \(e\.g\.,the,and,for\), likely due to code\-switching in WildChat Turkish conversations and the lack of WordNet filtering for Turkish, which in other languages serves to remove low\-quality noun candidates\. To obtain our tokens and noun lemmas for our analysis, we process all text using a language\-appropriate SpaCy pipelineHonnibal et al\. \([2020](https://arxiv.org/html/2605.29278#bib.bib16)\); we use a community\-hosted pipeline for Turkish\.777[https://huggingface\.co/turkish\-nlp\-suite/tr\_core\_news\_md](https://huggingface.co/turkish-nlp-suite/tr_core_news_md)\.

We use the following data and software resources: WildChatZhao et al\. \([2024](https://arxiv.org/html/2605.29278#bib.bib33)\)\(ODC\-BY license\), DailyDialogLi et al\. \([2017](https://arxiv.org/html/2605.29278#bib.bib21)\)\(CC BY\-NC\-SA 4\.0\), spaCyHonnibal et al\. \([2020](https://arxiv.org/html/2605.29278#bib.bib16)\)\(MIT license\), WordNetFellbaum \([1998](https://arxiv.org/html/2605.29278#bib.bib14)\)\(Princeton WordNet License\), and OMW 1\.4 \(CC BY 4\.0\)\. LIWC 2007 dictionariesChung and Pennebaker \([2012](https://arxiv.org/html/2605.29278#bib.bib12)\)are used under a purchased academic license and are not redistributed\. The Ubuntu Dialogue CorpusLowe et al\. \([2015](https://arxiv.org/html/2605.29278#bib.bib22)\)is derived from publicly available Ubuntu IRC chat logs; while commonly used, no explicit license is associated with the dataset\. Specifically, we use WildChat, which has been de\-identified using Microsoft Presidio and custom rules to identify and remove PII across various data types in English, Chinese, Russian, French, Spanish, German, Portuguese, Italian, Japanese, and Korean\. DailyDialog and the Ubuntu Dialogue Corpus are widely used academic datasets that do not contain identifying information\. We do not collect any new data in this work\.

Таблица 3:Linear trend slopes andpp\-values for per\-turn convergence scores across all settings\.p∗<0\.05\{\}^\{\*\}p<0\.05\. DD and Ubuntu report only user/speaker scores since there is no LLM speaker\.
## Приложение BAdditional Results

Таблица 4:Per\-language convergence scores WildChat on averaged NOUN and LIWC scores, as well as each function word class in LIWC\. All values are significant atp<0\.05p<0\.05unless underlined\.Δ\\Delta= LLM−\-User\. Cells marked – indicate categories absent from that languageś LIWC dictionary\.Table[4](https://arxiv.org/html/2605.29278#A2.T4)presents the full results for WildChat on all languages and features, while Figure[3](https://arxiv.org/html/2605.29278#A2.F3)reports the per\-turn convergence results for WildChat on NOUNs\. We also report the full statistical analysis of the per\-turn trends \(Table[3](https://arxiv.org/html/2605.29278#A1.T3)\)\.

![Refer to caption](https://arxiv.org/html/2605.29278v1/x3.png)Рис\. 3:Mean NOUN convergence scores per turn position for LLMs and users across eight WildChat languages\. Significant trends are indicated with an∗\*, saturated lines, and reported correlation value\(p<0\.05p<0\.05\)\. One notable outlier is ZH in this setting: both human and LLM speakers converge less over time, with LLMs in particular sharply mirroring their user less on later turns\.Таблица 5:Top\-100 noun vocabularies per language and dataset, ranked by corpus frequency after WordNet synonym filtering\.

Similar Articles

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

arXiv cs.CL

This paper introduces a register-aware linguistic evaluation framework to assess how human-like large language models (LLMs) are by comparing the distribution of 67 lexico-grammatical features between human and LLM-generated texts using Maximum Mean Discrepancy. Experiments across seven instruction-tuned open-source models and five registers show that no model perfectly matches human baselines, and closeness to human language varies by register rather than model size.

Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

arXiv cs.CL

This paper evaluates the abilities of large language models (LLMs) and multimodal LLMs for addressee detection, turn-change prediction, and next speaker prediction in multi-party meeting conversations. Results show text-based LLMs outperform supervised models and humans in next speaker prediction, while multimodal LLMs improve over text-only models in other tasks but remain below human performance.