Linguistic Monoculture in LLM-Assisted Language Use
摘要
This paper introduces a mathematical framework to study how reliance on shared LLMs for writing may reduce population-level linguistic diversity, analyzing fixed, recursive, and personalized interaction mechanisms and characterizing equilibria and convergence rates.
arXiv:2607.27134v1 Announce Type: new
Abstract: Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.
查看缓存全文
缓存时间: 2026/07/31 04:01
# Linguistic Monoculture in LLM-Assisted Language Use
Source: [https://arxiv.org/html/2607.27134](https://arxiv.org/html/2607.27134)
###### Abstract
Writing and communication are increasingly mediated by large language models \(LLMs\) that are being used to draft, revise and polish text\. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population\-level variation in linguistic form, a phenomenon we refer to as*linguistic monoculture*\. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and co\-evolve through repeated interaction\. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author\-specific and population\-level feedback\. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author–model equilibria with nonzero linguistic diversity\. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style\. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity\. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long\-run diversity outcomes\.
#### Keywords:
Algorithmic monoculture, Linguistic monoculture, Price of Monoculture
## 1Introduction
Large language models \(LLMs\) are increasingly used to draft, revise, and polish text\. Although such assistance can improve clarity, reduce errors, and help authors meet institutional expectations, widespread reliance on shared models may pull users toward similar model\-mediated linguistic patterns\. Moreover, LLM\-generated or LLM\-assisted text may enter future training and personalization pipelines, creating a coupled population\-level system in which authors adapt to model outputs and models may in turn adapt to language already influenced by earlier models\. Repeated interaction with shared generative systems may therefore reduce variation in language through which people develop, communicate and distinguish ideas\[[36](https://arxiv.org/html/2607.27134#bib.bib12),[32](https://arxiv.org/html/2607.27134#bib.bib11)\]\.
We call this reduction in population\-level variation across lexical choices, syntactic constructions, discourse markers, and other stylistic patterns as*linguistic monoculture*\. Recent empirical work provides evidence that LLM assistance can homogenize human text, ideas, expressions, and stylistic choices across users\[[37](https://arxiv.org/html/2607.27134#bib.bib3),[32](https://arxiv.org/html/2607.27134#bib.bib11),[13](https://arxiv.org/html/2607.27134#bib.bib10),[3](https://arxiv.org/html/2607.27134#bib.bib18),[31](https://arxiv.org/html/2607.27134#bib.bib15),[23](https://arxiv.org/html/2607.27134#bib.bib9)\]\. Large\-scale studies have also documented shifts in word frequencies and LLM\-associated stylistic markers in scientific articles and abstracts\[[26](https://arxiv.org/html/2607.27134#bib.bib25),[17](https://arxiv.org/html/2607.27134#bib.bib17),[18](https://arxiv.org/html/2607.27134#bib.bib16),[28](https://arxiv.org/html/2607.27134#bib.bib23)\]\. This concern is especially consequential in academia, where scientific writing helps make concepts legible, establish accepted methods, and shape what intellectual communities recognize as a contribution\[[4](https://arxiv.org/html/2607.27134#bib.bib5),[38](https://arxiv.org/html/2607.27134#bib.bib6),[22](https://arxiv.org/html/2607.27134#bib.bib7),[27](https://arxiv.org/html/2607.27134#bib.bib8)\]\. We therefore ask:
*Under what forms of repeated author–model interaction does LLM assistance drive a population toward a shared linguistic norm, and when can heterogeneous preferences or personalization preserve diversity?*
Our analysis focuses on diversity in linguistic\-feature distributions rather than convergence in semantic content, reasoning, or intellectual perspective\. Linguistic variation is nevertheless a distinct and measurable dimension of expression through which authors establish voice, structure arguments and make distinctions legible to others\[[20](https://arxiv.org/html/2607.27134#bib.bib36),[30](https://arxiv.org/html/2607.27134#bib.bib24),[15](https://arxiv.org/html/2607.27134#bib.bib35),[14](https://arxiv.org/html/2607.27134#bib.bib34)\]\. At the same time, not all convergence is undesirable: shared conventions can improve clarity, reduce errors, and make writing easier to understand and evaluate\[[22](https://arxiv.org/html/2607.27134#bib.bib7),[7](https://arxiv.org/html/2607.27134#bib.bib29)\]\. The relevant question is therefore when shared AI assistance produces useful standardization and when it causes an excessive loss of population\-level linguistic diversity\.
To answer this question, we develop a reduced\-form framework representing authors and prompt\-conditioned LLM output distributions over linguistic features\. It isolates how sharing, deployment\-level feedback, and personalization affect population\-level diversity, measured by average pairwise Jensen–Shannon \(JS\) divergence between author distributions\. We examine three author–LLM interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model updated recursively from author outputs, and personalized models updated through author\-specific and population\-level feedback\. These mechanisms isolate whether the model is shared, whether it is recursively updated, and whether its feedback is personalized\. We then endogenize conformity to a shared linguistic norm as a strategic choice and compare individually optimal conformity with the socially optimal level and quantify the loss of linguistic diversity through the*price of monoculture*\.
Our analysis of interaction dynamics initially treats conformity as exogenous\. In practice, however, adopting a shared linguistic norm may be a strategic choice: standardized language can improve clarity, legibility and perceived fluency, and may be rewarded by reviewers, readers, or institutions\[[19](https://arxiv.org/html/2607.27134#bib.bib41),[5](https://arxiv.org/html/2607.27134#bib.bib40)\]\. Yet such private benefits may come at the cost of an author’s distinctiveness preferences and reduce population\-level linguistic diversity\. We therefore compare individually optimal conformity with the socially optimal level\. In detail, our contributions are as follows:111All proofs and omitted details are in the Appendix\.
1. 1\.A reduced\-form framework for LLM\-mediated dynamics\.We formalize authors and prompt\-conditioned LLM output distributions over linguistic features under three deployment mechanisms: a fixed shared system, a recursively updated shared system, and recursively updated personalized systems \([Section2](https://arxiv.org/html/2607.27134#S2)\)\.
2. 2\.Linguistic convergence and determinants of diversity\.We establish convergence and convergence\-rate bounds for all three mechanisms and characterize how conformity pressure, author\-specific preferences, recursive feedback, and personalization determine long\-run population\-level diversity\. Shared models can drive diversity to low levels, whereas author\-specific preferences and personalization can preserve diversity \([Section3](https://arxiv.org/html/2607.27134#S3)\)\.
3. 3\.Strategic conformity and the price of monoculture\.We show that individually rational conformity can exceed the social optimum because authors may not internalize the value their distinctiveness provides to others; we bound the price of monoculture in a symmetric regime and identify when it diverges \([Section4](https://arxiv.org/html/2607.27134#S4)\)\.
4. 4\.Quantitative comparison of interaction mechanisms\.Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long\-run diversity outcomes across controlled parameter regimes\. Paired runs on identical initializations enable controlled comparison of the interaction mechanisms \([Section5](https://arxiv.org/html/2607.27134#S5)and Appendix[F](https://arxiv.org/html/2607.27134#A6)\)\.
Related work\.Our work connects research on LLM\-induced homogenization, opinion dynamics, recursive model training, algorithmic monoculture, and communication accommodation\. Building on classical averaging models, we couple author adaptation to shared or personalized models updated from author outputs, and analyze linguistic\-diversity equilibria and strategic welfare\.
*LLM\-induced homogenization\.*Empirical studies find that LLM assistance can homogenize written expression, ideas, authorial voice, and cultural style across users\[[32](https://arxiv.org/html/2607.27134#bib.bib11),[3](https://arxiv.org/html/2607.27134#bib.bib18),[31](https://arxiv.org/html/2607.27134#bib.bib15),[1](https://arxiv.org/html/2607.27134#bib.bib4),[2](https://arxiv.org/html/2607.27134#bib.bib14)\]\. Analyses of academic writing and presentations similarly document increased use of LLM\-associated linguistic markers\[[17](https://arxiv.org/html/2607.27134#bib.bib17),[18](https://arxiv.org/html/2607.27134#bib.bib16),[26](https://arxiv.org/html/2607.27134#bib.bib25),[28](https://arxiv.org/html/2607.27134#bib.bib23)\]\. These studies motivate our problem; we complement them by formalizing how homogenization can emerge through repeated author–LLM interaction\.
*Opinion dynamics\.*Our work builds on classical averaging models\[[12](https://arxiv.org/html/2607.27134#bib.bib43),[16](https://arxiv.org/html/2607.27134#bib.bib44),[21](https://arxiv.org/html/2607.27134#bib.bib45)\], using them as a reduced\-form representation of an algorithmic intermediary whose output distribution may be updated from the population it influences\. This introduces two deployment choices absent from human\-only models: whether the intermediary is shared or personalized, and whether author outputs affect its future behavior\. These choices determine whether the system converges to one shared norm or to a family of author\-specific equilibria\.
*Model collapse, knowledge collapse, and recursive feedback\.*Recursive training on model\-generated data can degrade generative models and reduce coverage of the tails of the original data distribution\[[35](https://arxiv.org/html/2607.27134#bib.bib20)\]\. Related work studies how AI\-mediated information environments may narrow the range of available knowledge and loss of tails\[[34](https://arxiv.org/html/2607.27134#bib.bib42),[39](https://arxiv.org/html/2607.27134#bib.bib39)\]\. This literature centers on the quality or epistemic diversity of model outputs, whereas we study the linguistic diversity of human authors who adapt to and may shape a shared model\.
*Algorithmic monoculture\.*Kleinberg and Raghavan \[[24](https://arxiv.org/html/2607.27134#bib.bib21)\]show that reliance on common algorithmic advice can be individually beneficial while reducing social welfare\. Recent work byKleinberget al\.\[[25](https://arxiv.org/html/2607.27134#bib.bib13)\]quantifies this loss in matching markets, proving a tight price\-of\-anarchy bound of factor22\. Our work studies a linguistic analogue: shared LLM assistance may improve individual legibility while reducing population\-level linguistic diversity\.
*Communication accommodation and personalization\.*Communication accommodation theory studies how speakers adjust their language toward or away from interlocutors\[[5](https://arxiv.org/html/2607.27134#bib.bib40),[19](https://arxiv.org/html/2607.27134#bib.bib41)\], including in human–LLM interaction\[[9](https://arxiv.org/html/2607.27134#bib.bib22)\]\. Related language\-game models examine how shared conventions emerge through peer\-to\-peer interaction\[[10](https://arxiv.org/html/2607.27134#bib.bib19),[11](https://arxiv.org/html/2607.27134#bib.bib1)\]\. In contrast, we study an LLM as a shared or personalized algorithmic interlocutor and characterize when personalization can preserve diversity across authors\.
## 2A Mathematical Framework for LLM\-Assisted Language Use
We use\[j\]\[j\]to denote\{1,…,j\}\\\{1,\\ldots,j\\\}forj∈ℕj\\in\\mathbb\{N\}, andΔm−1:=\{x∈ℝ≥0m:∑k=1mxk=1\}\\Delta^\{m\-1\}:=\\\{x\\in\\mathbb\{R\}\_\{\\geq 0\}^\{m\}:\\sum\_\{k=1\}^\{m\}x\_\{k\}=1\\\}to denote an\(m−1\)\(m\-1\)\-dimensional probability simplex\. Letℱ=\[m\]\\mathcal\{F\}=\[m\]be a common set of linguistic features encoding lexical choices, syntactic constructions, discourse markers or other stylistic patterns\. We represent linguistic style as a probability distribution overℱ\\mathcal\{F\}\. For each authori∈\[n\]i\\in\[n\]and timet∈\{0,…,T\}t\\in\\\{0,\\ldots,T\\\}, letpit∈Δm−1p\_\{i\}^\{t\}\\in\\Delta^\{m\-1\}denote authorii’s linguistic\-style distribution, with initial distributionpi0p\_\{i\}^\{0\}\. Letqt∈Δm−1q^\{t\}\\in\\Delta^\{m\-1\}denote the prompt\-conditioned linguistic\-style distribution induced by an LLM system at timettunder a fixed class \(or distribution\) of writing prompts; if personalized to authorii, we writeqitq\_\{i\}^\{t\}\. This is an output\-level abstraction rather than a model\-parameter representation\. For notational uniformity, in shared\-model settings we writeqit:=qtq\_\{i\}^\{t\}:=q^\{t\}for every authorii\.
At each time stept∈\{0,…,T−1\}t\\in\\\{0,\\ldots,T\-1\\\}, authoriiprompts the LLM, observes LLM\-generated linguistic suggestions and updates their own style as
pit\+1=\(1−αi\)pit⏟Styleretention\+αi𝒜i\(qit,zit\)⏟Model\- andcontext\-driven update,p\_\{i\}^\{t\+1\}=\\underbrace\{\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\}\_\{\\begin\{subarray\}\{c\}\\text\{Style\}\\\\ \\text\{retention\}\\end\{subarray\}\}\+\\underbrace\{\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q\_\{i\}^\{t\},z\_\{i\}^\{t\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{Model\- and\}\\\\ \\text\{context\-driven update\}\\end\{subarray\}\},\(1\)whereαi∈\[0,1\]\\alpha\_\{i\}\\in\[0,1\]measures the rate at which authoriiadapts their linguistic style\. The operator𝒜i:Δm−1×𝒵→Δm−1\\mathcal\{A\}\_\{i\}:\\Delta^\{m\-1\}\\times\\mathcal\{Z\}\\to\\Delta^\{m\-1\}describes how the author incorporates the model’s output given an author\-specific adaptation inputzit∈𝒵z\_\{i\}^\{t\}\\in\\mathcal\{Z\}, which may encode external incentives, preferred style, or other contextual information\.
The induced output distribution may itself evolve when author\-produced text is incorporated into future training data, a recursive feedback mechanism studied in the model\-collapse literature\[[35](https://arxiv.org/html/2607.27134#bib.bib20)\]\. We model this in reduced form as
qit\+1=βqit⏟Modelretention\+\(1−β\)ℬ\(p1t\+1,…,pnt\+1\)⏟Feedback fromauthor outputs,q\_\{i\}^\{t\+1\}=\\underbrace\{\\beta q\_\{i\}^\{t\}\}\_\{\\begin\{subarray\}\{c\}\\text\{Model\}\\\\ \\text\{retention\}\\end\{subarray\}\}\+\\underbrace\{\(1\-\\beta\)\\,\\mathcal\{B\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{Feedback from\}\\\\ \\text\{author outputs\}\\end\{subarray\}\},\(2\)whereβ∈\[0,1\]\\beta\\in\[0,1\]is the retention parameter: largerβ\\betaplaces more weight on the past model distribution\. The operatorℬ\\mathcal\{B\}maps the updated author distributions to the distribution of training data used to update the model; we instantiate this to be the weighted population averageℬ\(p1t\+1,…,pnt\+1\)=∑i=1nwipit\+1\\mathcal\{B\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\)=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{t\+1\}withwi≥0w\_\{i\}\\geq 0,∑iwi=1\\sum\_\{i\}w\_\{i\}=1\.
Our main object of interest is population\-level linguistic diversity, measured by the average pairwise Jensen–Shannon divergence\[[29](https://arxiv.org/html/2607.27134#bib.bib2)\]among author distributions:
Dt\\displaystyle D^\{t\}=1n\(n−1\)∑i≠jJS\(pit,pjt\),\\displaystyle=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\),whereJS\(p,q\)=12KL\(p∥x\)\+12KL\(q∥x\)JS\(p,q\)=\\frac\{1\}\{2\}KL\(p\\\|x\)\+\\frac\{1\}\{2\}KL\(q\\\|x\)withx=p\+q2x=\\frac\{p\+q\}\{2\}; smaller values indicate convergence toward a common linguistic norm\. When the model distribution evolves, we also track the average author–model divergence
Mt=1n∑i=1nJS\(pit,qit\),M^\{t\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}JS\(p\_\{i\}^\{t\},q\_\{i\}^\{t\}\),which measures the alignment between each author and the model they interact with; in shared\-model settings,Mt→0M^\{t\}\\to 0corresponds to a shared human\-LLM linguistic equilibrium\. Finally, when each author interacts with a personalized model, we track diversity among the personalized models,
Qt=1n\(n−1\)∑i≠jJS\(qit,qjt\),Q^\{t\}=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(q\_\{i\}^\{t\},q\_\{j\}^\{t\}\),whereQt→0Q^\{t\}\\to 0means all personalized model distributions become indistinguishable\. In the personalized setting,DtD^\{t\}measures diversity among authors,QtQ^\{t\}diversity among personalized models andMtM^\{t\}alignment between each author and their model\. Since Jensen\-Shannon divergence is symmetric and finite for probability distributions, it is a natural choice for all three measures\.[AppendixA](https://arxiv.org/html/2607.27134#A1)in Appendix[A](https://arxiv.org/html/2607.27134#A1)justifies our choice to represent linguistic style as a distribution rather than a point in a high\-dimensional feature space\.
### 2\.1Author\-LLM interaction mechanisms
We formalize three interaction mechanisms \(IMs\) that introduce increasing degrees of author–model coupling\.
###### Interaction Mechanism 1\(Shared model with fixed distribution\)\.
All authors interact with the same model, whose linguistic\-style distribution remains fixed over time,*i\.e\.*,qt=q0q^\{t\}=q^\{0\}for allt∈\{0,…,T\}t\\in\\\{0,\\ldots,T\\\}\. The linguistic\-style distribution of authoriievolves as
pit\+1=\(1−αi\)pit\+αi𝒜i\(q0,zit\)\.p\_\{i\}^\{t\+1\}=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\}\)\.
###### Interaction Mechanism 2\(Shared model with recursive updates\)\.
All authors interact with the same model, but the model distribution evolves over time\. The coupled author–model dynamics evolve as
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αi𝒜i\(qt,zit\),\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q^\{t\},z\_\{i\}^\{t\}\),qt\+1\\displaystyle q^\{t\+1\}=βqt\+\(1−β\)ℬ\(p1t\+1,…,pnt\+1\),\\displaystyle=\\beta\\,q^\{t\}\+\(1\-\\beta\)\\,\\mathcal\{B\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\),withℬ\(p1t\+1,…,pnt\+1\)=∑j=1nwjpjt\+1\\mathcal\{B\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\)=\\sum\_\{j=1\}^\{n\}w\_\{j\}p\_\{j\}^\{t\+1\}as the weighted\-average instantiation, wherewj≥0w\_\{j\}\\geq 0and∑j=1nwj=1\\sum\_\{j=1\}^\{n\}w\_\{j\}=1\.
###### Interaction Mechanism 3\(Personalized models with recursive updates\)\.
Each authoriiinteracts with an author\-specific personalized model with linguistic\-style distributionqitq\_\{i\}^\{t\}, initialized with a common base distributionqi0=q0q\_\{i\}^\{0\}=q^\{0\}for alli∈\[n\]i\\in\[n\]\. The coupled author–model dynamics evolve as
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αi𝒜i\(qit,zit\),\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q\_\{i\}^\{t\},z\_\{i\}^\{t\}\),qit\+1\\displaystyle q\_\{i\}^\{t\+1\}=βqit\+\(1−β\)ℬi\(p1t\+1,…,pnt\+1\),\\displaystyle=\\beta\\,q\_\{i\}^\{t\}\+\(1\-\\beta\)\\,\\mathcal\{B\}\_\{i\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\),where the feedback operatorℬi\\mathcal\{B\}\_\{i\}mixes authorii’s own updated distribution with the population\-level update:
ℬi\(p1t\+1,…,pnt\+1\)=ρpit\+1\+\(1−ρ\)∑j=1nwjpjt\+1,\\mathcal\{B\}\_\{i\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\)=\\rho\\,p\_\{i\}^\{t\+1\}\+\(1\-\\rho\)\\sum\_\{j=1\}^\{n\}w\_\{j\}p\_\{j\}^\{t\+1\},withρ∈\[0,1\]\\rho\\in\[0,1\],wj≥0w\_\{j\}\\geq 0,∑j=1nwj=1\\sum\_\{j=1\}^\{n\}w\_\{j\}=1\. Equivalently,
qit\+1=βqit\+γpit\+1\+δ∑j=1nwjpjt\+1,q\_\{i\}^\{t\+1\}=\\beta\\,q\_\{i\}^\{t\}\+\\gamma\\,p\_\{i\}^\{t\+1\}\+\\delta\\sum\_\{j=1\}^\{n\}w\_\{j\}p\_\{j\}^\{t\+1\},whereγ=\(1−β\)ρ\\gamma=\(1\-\\beta\)\\rho,δ=\(1−β\)\(1−ρ\)\\delta=\(1\-\\beta\)\(1\-\\rho\), andβ\+γ\+δ=1\\beta\+\\gamma\+\\delta=1\. The parameterρ\\rhodetermines the degree of personalization: whenρ=1\\rho=1, the model update depends only on authorii’s own distribution; whenρ=0\\rho=0, all personalized models receive the same population\-level feedback\.
In Appendix[E](https://arxiv.org/html/2607.27134#A5), we investigate Interaction Mechanism 4, which partitions a single population into three subpopulations, each following the update rule of one of IMs 1\-3\.
## 3Convergence of Linguistic Distributions
We analyze the three interaction mechanisms in increasing order of coupling, establishing convergence as well as convergence\-rate bounds\.
### 3\.1A shared model with fixed distribution
In[Mechanism1](https://arxiv.org/html/2607.27134#Thmmechanism1)the model distribution is fixed, soqt=q0q^\{t\}=q^\{0\}for allttand authorii’s linguistic\-style distribution evolves aspit\+1=\(1−αi\)pit\+αi𝒜i\(q0,zit\)p\_\{i\}^\{t\+1\}=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\}\)\. First, we consider the model\-induced adaptation to be homogeneous across authors and time,𝒜i\(q0,zit\+1\)=𝒜i\(q0,zit\)=𝒜\(q0,z\)\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\+1\}\)=\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\}\)=\\mathcal\{A\}\(q^\{0\},z\)for alliiandtt, so all authors are influenced by the same fixed model distribution and have the same adaptation input\.
###### Proposition 3\.1\.
In[Mechanism1](https://arxiv.org/html/2607.27134#Thmmechanism1), suppose the model\-induced adaptation is homogeneous across authors and time,*i\.e\.*,𝒜i\(q0,zit\)=𝒜\(q0,z\)\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\}\)=\\mathcal\{A\}\(q^\{0\},z\)for every authoriiand timett\. Letαmin=mini∈\[n\]αi\>0\\alpha\_\{\\min\}=\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\>0be the minimum adaptation rate\. Then, for everyε\>0\\varepsilon\>0, it is sufficient to taket≥log\(2/ε\)αmint\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\\alpha\_\{\\min\}\}to guaranteeDt≤εD^\{t\}\\leq\\varepsilon\.
The homogeneity assumption is strong: all authors are pulled toward the same distribution\. Thus[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1)gives a baseline showing that common adaptation and rewards cause diversity to decay exponentially, at a rate set by the slowest\-adapting author\. Our next result relaxes the homogeneity assumption with author\-specific adaptation\.
###### Proposition 3\.2\.
Consider[Mechanism1](https://arxiv.org/html/2607.27134#Thmmechanism1)with fixed model distributionqt=q0q^\{t\}=q^\{0\}\. Suppose that the adaptation operator is time\-invariant but author\-specific,*i\.e\.*, for each authorii,𝒜i\(q0,zit\)=ai\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\}\)=a\_\{i\}for alltt, whereai∈Δm−1a\_\{i\}\\in\\Delta^\{m\-1\}, and letαmin=mini∈\[n\]αi\>0\\alpha\_\{\\min\}=\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\>0\. Then each author’s linguistic\-style distribution converges to its author\-specific adapted distribution,pit→aip\_\{i\}^\{t\}\\to a\_\{i\}, exponentially fast at a rate controlled byαmin\\alpha\_\{\\min\}\. Consequently,
Dt→D∞:=1n\(n−1\)∑i≠jJS\(ai,aj\),D^\{t\}\\to D^\{\\infty\}:=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(a\_\{i\},a\_\{j\}\),andDt→0D^\{t\}\\to 0if and only ifai=aja\_\{i\}=a\_\{j\}for alli,ji,j\.
When adaptation is author\-specific but fixed over time, authors do not collapse to a single linguistic\-style distribution: each converges to their own adaptation target, and population\-level diversity converges to the diversity among these targets—author\-specific incentives preserve diversity even under repeated exposure to the same fixed model\. IfD∞\>0D^\{\\infty\}\>0, the process stabilizes at a positive level of diversity\.
Interplay between diversity and conformity\.To interpolate between these regimes, let each author balance a conformity incentive pulling toward the shared model\-induced norm against a diversity incentive preserving author\-specific style\. For the remainder of this section, letri∈Δm−1r\_\{i\}\\in\\Delta^\{m\-1\}denote authorii’s fixed preferred linguistic\-style distribution, and setzit=riz\_\{i\}^\{t\}=r\_\{i\}for alltt\. We model the adaptation as
𝒜i\(q0,ri\)=\(1−λi\)ri\+λiq0,λi∈\[0,1\],\\mathcal\{A\}\_\{i\}\(q^\{0\},r\_\{i\}\)=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{0\},\\qquad\\lambda\_\{i\}\\in\[0,1\],whereλi\\lambda\_\{i\}measures the strength of the conformity incentive:λi=0\\lambda\_\{i\}=0preserves the preferred stylerir\_\{i\}, whileλi=1\\lambda\_\{i\}=1fully adopts the shared model distributionq0q^\{0\}\.
###### Proposition 3\.3\.
Consider[Mechanism1](https://arxiv.org/html/2607.27134#Thmmechanism1)with fixed model distributionqt=q0q^\{t\}=q^\{0\}\. Suppose each author has an author\-specific preferred distributionri∈Δm−1r\_\{i\}\\in\\Delta^\{m\-1\}, and the adaptation operator is
𝒜i\(q0,ri\)=\(1−λi\)ri\+λiq0,λi∈\[0,1\]\.\\mathcal\{A\}\_\{i\}\(q^\{0\},r\_\{i\}\)=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{0\},\\qquad\\lambda\_\{i\}\\in\[0,1\]\.Then each author distribution converges toai\(λi\):=\(1−λi\)ri\+λiq0a\_\{i\}\(\\lambda\_\{i\}\):=\(1\-\\lambda\_\{i\}\)\\,r\_\{i\}\+\\lambda\_\{i\}q^\{0\}\. Precisely, ifαmin:=mini∈\[n\]αi\>0\\alpha\_\{\\min\}:=\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\>0then for everyε\>0\\varepsilon\>0, it is sufficient to taket≥log\(2/ε\)αmint\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\\alpha\_\{\\min\}\}to guarantee thatmaxi∈\[n\]‖pit−ai\(λi\)‖1≤ε\.\\max\_\{i\\in\[n\]\}\\\|p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\\|\_\{1\}\\leq\\varepsilon\.Consequently,
Dt→D∞\(λ\):=1n\(n−1\)∑i≠jJS\(ai\(λi\),aj\(λj\)\)\.D^\{t\}\\to D^\{\\infty\}\(\\lambda\):=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\\left\(a\_\{i\}\(\\lambda\_\{i\}\),a\_\{j\}\(\\lambda\_\{j\}\)\\right\)\.In particular,Dt→0D^\{t\}\\to 0if and only ifai\(λi\)=aj\(λj\)a\_\{i\}\(\\lambda\_\{i\}\)=a\_\{j\}\(\\lambda\_\{j\}\)for alli,ji,j\. A sufficient condition isλi=1\\lambda\_\{i\}=1for allii, in which caseai\(λi\)=q0a\_\{i\}\(\\lambda\_\{i\}\)=q^\{0\}for every authorii\.
The proof establishes that for a common conformity levelλi=λ\\lambda\_\{i\}=\\lambda, the limiting diversity satisfiesD∞\(λ\)=O\(1−λ\)D^\{\\infty\}\(\\lambda\)=O\(1\-\\lambda\)\. The parameterλ\\lambdarepresents institutional or reward\-based pressure to conform to a shared model\-induced linguistic norm: whenλ\\lambdais small, author\-specific preferences dominate and diversity persists; whenλ\\lambdais large, the shared norm dominates and diversity collapses\. Even with author\-specific adaptation, monoculture can emerge if the reward structure places sufficiently high weight on conformity\.
### 3\.2A shared model with recursive updates
We next allow the shared model to be retrained on author outputs\. WritingPt\+1:=ℬ\(p1t\+1,…,pnt\+1\)=∑i=1nwipit\+1P^\{t\+1\}:=\\mathcal\{B\}\(p\_\{1\}^\{t\+1\},\\ldots,p\_\{n\}^\{t\+1\}\)=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{t\+1\}, the coupled dynamics of[Mechanism2](https://arxiv.org/html/2607.27134#Thmmechanism2)are
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αi𝒜i\(qt,zit\),\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q^\{t\},z\_\{i\}^\{t\}\),qt\+1\\displaystyle q^\{t\+1\}=βqt\+\(1−β\)Pt\+1\.\\displaystyle=\\beta q^\{t\}\+\(1\-\\beta\)\\,P^\{t\+1\}\.If the adaptation target does not depend on the evolving model distribution \(𝒜i\(qt,zit\)=a\\mathcal\{A\}\_\{i\}\(q^\{t\},z\_\{i\}^\{t\}\)=aoraia\_\{i\}\), the author dynamics are unchanged from the fixed\-model setting and[Sections3\.1](https://arxiv.org/html/2607.27134#S3.SS1)and[3\.1](https://arxiv.org/html/2607.27134#S3.SS1)apply verbatim: recursive feedback alone does not force monoculture when authors are pulled toward fixed targets\. Recursive updates matter only when the adaptation operator depends non\-trivially onqtq^\{t\}\. We therefore consider the conformity\-mixture adaptation
𝒜i\(qt,ri\)=\(1−λi\)ri\+λiqt,λi∈\[0,1\],\\mathcal\{A\}\_\{i\}\(q^\{t\},r\_\{i\}\)=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{t\},\\qquad\\lambda\_\{i\}\\in\[0,1\],where the author’s preferred distributionrir\_\{i\}is fixed over time andλi\\lambda\_\{i\}now measures conformity to the*current*model distribution\.
###### Proposition 3\.4\.
Consider[Mechanism2](https://arxiv.org/html/2607.27134#Thmmechanism2)with𝒜i\(qt,ri\)=\(1−λi\)ri\+λiqt\\mathcal\{A\}\_\{i\}\(q^\{t\},r\_\{i\}\)=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{t\}andλi∈\[0,1\)\\lambda\_\{i\}\\in\[0,1\)\. Suppose thatαi∈\(0,1\]\\alpha\_\{i\}\\in\(0,1\]for every authorii,β∈\[0,1\)\\beta\\in\[0,1\), andPt\+1=∑i=1nwipit\+1P^\{t\+1\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{t\+1\},wi≥0w\_\{i\}\\geq 0,∑i=1nwi=1\\sum\_\{i=1\}^\{n\}w\_\{i\}=1\. Then the coupled author–model dynamics converge to an equilibrium\(p1∗,…,pn∗,q∗\)\(p\_\{1\}^\{\\ast\},\\ldots,p\_\{n\}^\{\\ast\},q^\{\\ast\}\)satisfying
pi∗=\(1−λi\)ri\+λiq∗p\_\{i\}^\{\\ast\}=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{\\ast\}for every authori, andq∗=∑i=1nwi\(1−λi\)ri1−∑i=1nwiλi\.\\text\{for every author $i$, and\}\\quad q^\{\\ast\}=\\frac\{\\sum\_\{i=1\}^\{n\}w\_\{i\}\(1\-\\lambda\_\{i\}\)r\_\{i\}\}\{1\-\\sum\_\{i=1\}^\{n\}w\_\{i\}\\lambda\_\{i\}\}\.Consequently,Dt→D∗:=1n\(n−1\)∑i≠jJS\(pi∗,pj∗\)\.\\text\{Consequently,\}\\quad D^\{t\}\\to D^\{\\ast\}:=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{\\ast\},p\_\{j\}^\{\\ast\}\)\.In particular, the limiting diversity is the diversity among the equilibrium author distributionspi∗p\_\{i\}^\{\\ast\}, rather than the diversity among the fixed\-model targets\(1−λi\)ri\+λiq0\(1\-\\lambda\_\{i\}\)\\,r\_\{i\}\+\\lambda\_\{i\}\\,q^\{0\}\.
The proof also shows that convergence is exponential, with a worst\-case time bound oft≥log\(2/ε\)\(1−β\)mini∈\[n\]αi\(1−λi\)t\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\(1\-\\beta\)\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\(1\-\\lambda\_\{i\}\)\}, which is weaker than the fixed\-model bound, controlled byαmin\\alpha\_\{\\min\}alone: the moving targetqtq^\{t\}can slow the guaranteed rate, though the bound need not be tight in typical instances \(see[Section5](https://arxiv.org/html/2607.27134#S5)\)\. The conditionλi<1\\lambda\_\{i\}<1ensures that each author retains some pull toward their preferred distributionrir\_\{i\}; ifλi=1\\lambda\_\{i\}=1for all authors, the fixed\-point formula forq∗q^\{\\ast\}becomes singular and that case must be treated separately\.
The weaker rate bound should not be read as recursive feedback dampening homogenization: recursion changes the equilibrium, anchoring authors at the endogenousq∗q^\{\\ast\}rather than the exogenousq0q^\{0\}\.[Section3\.3](https://arxiv.org/html/2607.27134#S3.SS3)makes the resulting comparison precise\.
### 3\.3Personalized models with recursive updates
Personalization changes what the system converges*to*: because each model update includes an author\-specific feedback component, the limiting model distributionsqi∗q\_\{i\}^\{\\ast\}can remain distinct\. The system may therefore converge to a family of author\-specific equilibria rather than a single shared norm\. The parameterρ\\rhocontrols the extent of personalization\.
###### Proposition 3\.5\.
Consider[Mechanism3](https://arxiv.org/html/2607.27134#Thmmechanism3)with author and model updates
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αi𝒜i\(qit,ri\),\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q\_\{i\}^\{t\},r\_\{i\}\),qit\+1\\displaystyle q\_\{i\}^\{t\+1\}=βqit\+γpit\+1\+δ∑j=1nwjpjt\+1,\\displaystyle=\\beta q\_\{i\}^\{t\}\+\\gamma p\_\{i\}^\{t\+1\}\+\\delta\\sum\_\{j=1\}^\{n\}w\_\{j\}p\_\{j\}^\{t\+1\},whereαi∈\(0,1\]\\alpha\_\{i\}\\in\(0,1\],𝒜i\(qit,ri\)=\(1−λi\)ri\+λiqit\\mathcal\{A\}\_\{i\}\(q\_\{i\}^\{t\},r\_\{i\}\)=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q\_\{i\}^\{t\}withλi∈\[0,1\)\\lambda\_\{i\}\\in\[0,1\),wj≥0w\_\{j\}\\geq 0,∑j=1nwj=1\\sum\_\{j=1\}^\{n\}w\_\{j\}=1, andβ\+γ\+δ=1\\beta\+\\gamma\+\\delta=1\. Suppose thatβ∈\[0,1\)\\beta\\in\[0,1\), and letρ=γ1−β∈\[0,1\]\\rho=\\frac\{\\gamma\}\{1\-\\beta\}\\in\[0,1\]\. Then the coupled personalized author–model dynamics converge to an equilibrium\(p1∗,…,pn∗,q1∗,…,qn∗\)\(p\_\{1\}^\{\\ast\},\\ldots,p\_\{n\}^\{\\ast\},q\_\{1\}^\{\\ast\},\\ldots,q\_\{n\}^\{\\ast\}\)with
pi∗=ηiri\+\(1−ηi\)P∗,qi∗=ρηiri\+\(1−ρηi\)P∗,p\_\{i\}^\{\\ast\}=\\eta\_\{i\}r\_\{i\}\+\(1\-\\eta\_\{i\}\)P^\{\\ast\},\\quad q\_\{i\}^\{\\ast\}=\\rho\\eta\_\{i\}r\_\{i\}\+\(1\-\\rho\\eta\_\{i\}\)P^\{\\ast\},whereηi=1−λi1−ρλi\\eta\_\{i\}=\\frac\{1\-\\lambda\_\{i\}\}\{1\-\\rho\\lambda\_\{i\}\}andP∗=∑i=1nwiηiri∑i=1nwiηiP^\{\\ast\}=\\frac\{\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}r\_\{i\}\}\{\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}\}\. Consequently,
Dt→D∗:=1n\(n−1\)∑i≠jJS\(pi∗,pj∗\),D^\{t\}\\to D^\{\\ast\}:=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{\\ast\},p\_\{j\}^\{\\ast\}\),Qt→Q∗:=1n\(n−1\)∑i≠jJS\(qi∗,qj∗\)\.Q^\{t\}\\to Q^\{\\ast\}:=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(q\_\{i\}^\{\\ast\},q\_\{j\}^\{\\ast\}\)\.
Whenρ\>0\\rho\>0, each equilibrium model distributionqi∗q\_\{i\}^\{\\ast\}retains an author\-specific component proportional toρηi\\rho\\eta\_\{i\}, so both author diversityD∗D^\{\\ast\}and personalized\-model diversityQ∗Q^\{\\ast\}can remain bounded away from zero; whenρ=0\\rho=0, all personalized models collapse to the same population\-level distribution\. Personalization thus does not prevent convergence—it changes its object, from a single shared norm to a family of distinct author–model equilibria\.
Because JS depends on both pairwise separation and the cloud’s location in the simplex \([AppendixA](https://arxiv.org/html/2607.27134#A1)\), we isolate pairwise geometry using the translation\-invariant quadratic diversity
D^\(p1,…,pn\):=12n\(n−1\)∑i≠j‖pi−pj‖22,\\widehat\{D\}\(p\_\{1\},\\ldots,p\_\{n\}\):=\\frac\{1\}\{2n\(n\-1\)\}\\sum\_\{i\\neq j\}\\\|p\_\{i\}\-p\_\{j\}\\\|\_\{2\}^\{2\},and writeD^∞\\widehat\{D\}\_\{\\infty\}for its value at equilibrium\.
###### Proposition 3\.6\.
Supposeλi=λ∈\[0,1\)\\lambda\_\{i\}=\\lambda\\in\[0,1\)for every authorii, and letη=1−λ1−ρλ\\eta=\\frac\{1\-\\lambda\}\{1\-\\rho\\lambda\}\. Then, for every pairi≠ji\\neq j,
ai\(λ\)−aj\(λ\)=\(1−λ\)\(ri−rj\)under IM 1,a\_\{i\}\(\\lambda\)\-a\_\{j\}\(\\lambda\)=\(1\-\\lambda\)\(r\_\{i\}\-r\_\{j\}\)\\quad\\text\{under IM~1\},pi∗−pj∗=\(1−λ\)\(ri−rj\)under IM 2,p\_\{i\}^\{\\ast\}\-p\_\{j\}^\{\\ast\}=\(1\-\\lambda\)\(r\_\{i\}\-r\_\{j\}\)\\quad\\text\{under IM~2\},pi∗−pj∗=η\(ri−rj\)under IM 3\.p\_\{i\}^\{\\ast\}\-p\_\{j\}^\{\\ast\}=\\eta\(r\_\{i\}\-r\_\{j\}\)\\quad\\text\{under IM~3\}\.Consequently, IMs 1 and 2 have identical limiting quadratic diversity, while
D^∞\(IM3\)=D^∞\(IM1\)\(1−ρλ\)2≥D^∞\(IM1\),\\widehat\{D\}\_\{\\infty\}\(\\mathrm\{IM~3\}\)=\\frac\{\\widehat\{D\}\_\{\\infty\}\(\\mathrm\{IM~1\}\)\}\{\(1\-\\rho\\lambda\)^\{2\}\}\\geq\\widehat\{D\}\_\{\\infty\}\(\\mathrm\{IM~1\}\),with equality if and only ifρλ=0\\rho\\lambda=0\.
Under common conformity, recursion relocates the limiting author cloud from the exogenousq0q^\{0\}to the endogenousq∗q^\{\\ast\}without changing its pairwise geometry\. Hence any IM 1–IM 2 gap in limiting JS diversity arises only from JS’s position dependence\. Personalization instead expands pairwise geometry relative to IMs 1–2 by1/\(1−ρλ\)1/\(1\-\\rho\\lambda\); heterogeneousλi\\lambda\_\{i\}can additionally create a genuine IM 1–IM 2 geometric gap \(Appendix[F](https://arxiv.org/html/2607.27134#A6)\)\.
Summary\.These results treat the conformity parametersλi\\lambda\_\{i\}as exogenous: they characterize what happens when authors have a fixed tendency to adopt the model\-induced norm, but not why authors would choose to have that tendency\.
## 4Strategic Conformity and the Price of Monoculture
In[Section3](https://arxiv.org/html/2607.27134#S3), the conformity parameterλi\\lambda\_\{i\}quantifies*how much*authoriiis influenced by the model\-induced linguistic norm, but not*why*\. We now endogenize this parameter in the setting of[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1): each author choosesλi\\lambda\_\{i\}to trade off individual rewards of conforming—such as legibility, perceived polish and alignment with reviewer expectations—against the loss of distinctiveness that conformity entails\. This turns the dynamics into a strategic game, and we ask the linguistic analogue of the question raised by algorithmic monoculture\[[24](https://arxiv.org/html/2607.27134#bib.bib21)\]:*can individually rational linguistic adaptation produce collectively inefficient homogenization?*Under the fixed\-model quadratic game with orthogonal signatures analyzed below, the answer is yes: the unique Nash equilibrium conformity level weakly exceeds the socially optimal level for every author, and strictly exceeds it for any author who conforms when others value distinctiveness\. This gap is the value of an author’s distinctiveness to others: a payoff externality no individual internalizes and one that harms even authors who optimally choose not to conform \([Section4\.1](https://arxiv.org/html/2607.27134#S4.SS1)\)\. In a symmetric regime, the resulting price of monoculture can be arbitrarily large across instances\.
### 4\.1Conformity as a choice
We work in the setting of[Mechanism1](https://arxiv.org/html/2607.27134#Thmmechanism1),[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1): the model distribution is fixed atq0q^\{0\}, each authoriihas a preferred distributionri∈Δm−1r\_\{i\}\\in\\Delta^\{m\-1\}and the adaptation operator is𝒜i\(q0,ri\)=\(1−λi\)ri\+λiq0\\mathcal\{A\}\_\{i\}\(q^\{0\},r\_\{i\}\)=\(1\-\\lambda\_\{i\}\)\\,r\_\{i\}\+\\lambda\_\{i\}\\,q^\{0\}\. For any conformity profileλ=\(λ1,…,λn\)∈\[0,1\]n\\lambda=\(\\lambda\_\{1\},\\dots,\\lambda\_\{n\}\)\\in\[0,1\]^\{n\},[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1)shows that the author dynamics converge exponentially toai\(λi\)=\(1−λi\)ri\+λiq0a\_\{i\}\(\\lambda\_\{i\}\)=\(1\-\\lambda\_\{i\}\)\\,r\_\{i\}\+\\lambda\_\{i\}\\,q^\{0\}\. We use this convergence result to define a strategic game over long\-run writing policies\. Each author first chooses a fixed conformity levelλi\\lambda\_\{i\}, which determines how strongly they rely on the model\-induced norm\. The interaction dynamics in[eq\.1](https://arxiv.org/html/2607.27134#S2.E1)then unfold and payoffs are evaluated at the limiting profile\(a1\(λ1\),…,an\(λn\)\)\(a\_\{1\}\(\\lambda\_\{1\}\),\\dots,a\_\{n\}\(\\lambda\_\{n\}\)\)\. Thus the game is played over long\-run writing styles and not over individual time\-step updates\. This timescale separation is justified by exponential convergence: when the evaluation horizon is long relative to the convergence time, authors’ distributions spend most of the horizon close to their limiting profile, so payoffs evaluated at the limiting profile approximate long\-run average payoffs\.
For brevity, we writeui:=ri−q0u\_\{i\}:=r\_\{i\}\-q^\{0\}for authorii’s*signature vector*—the direction and magnitude of their idiosyncrasy relative to the model’s linguistic norm—andσi:=1−λi∈\[0,1\]\\sigma\_\{i\}:=1\-\\lambda\_\{i\}\\in\[0,1\]for their*retained distinctiveness*, so that
ai\(λi\)−q0=σiui,ai\(λi\)−ri=−λiui,a\_\{i\}\(\\lambda\_\{i\}\)\-q^\{0\}=\\sigma\_\{i\}\\,u\_\{i\},\\quad a\_\{i\}\(\\lambda\_\{i\}\)\-r\_\{i\}=\-\\lambda\_\{i\}\\,u\_\{i\},ai\(λi\)−aj\(λj\)=σiui−σjuj\.a\_\{i\}\(\\lambda\_\{i\}\)\-a\_\{j\}\(\\lambda\_\{j\}\)=\\sigma\_\{i\}u\_\{i\}\-\\sigma\_\{j\}u\_\{j\}\.
Payoffs\.Given a conformity profileλ\\lambda, authorii’s utility is
Ui\(λi,λ−i\)=\\displaystyle U\_\{i\}\(\\lambda\_\{i\},\\lambda\_\{\-i\}\)\\;=−bi2‖ai\(λi\)−q0‖22⏟legibility reward−ci2‖ai\(λi\)−ri‖22⏟authenticity cost\\displaystyle\\;\\underbrace\{\-\\,\\frac\{b\_\{i\}\}\{2\}\\,\\big\\\|a\_\{i\}\(\\lambda\_\{i\}\)\-q^\{0\}\\big\\\|\_\{2\}^\{2\}\}\_\{\\text\{legibility reward\}\}\\;\\underbrace\{\-\\,\\frac\{c\_\{i\}\}\{2\}\\,\\big\\\|a\_\{i\}\(\\lambda\_\{i\}\)\-r\_\{i\}\\big\\\|\_\{2\}^\{2\}\}\_\{\\text\{authenticity cost\}\}\+θi2\(n−1\)∑j≠i‖ai\(λi\)−aj\(λj\)‖22⏟distinctiveness payoff,\\displaystyle\+\\;\\underbrace\{\\frac\{\\theta\_\{i\}\}\{2\(n\-1\)\}\\sum\_\{j\\neq i\}\\big\\\|a\_\{i\}\(\\lambda\_\{i\}\)\-a\_\{j\}\(\\lambda\_\{j\}\)\\big\\\|\_\{2\}^\{2\}\}\_\{\\text\{distinctiveness payoff\}\},\(3\)wherebi,θi≥0b\_\{i\},\\theta\_\{i\}\\geq 0andci\>0c\_\{i\}\>0\. The three terms represent, respectively, the benefit of proximity to the model\-induced normq0q^\{0\}, the private cost of deviating from the author’s preferred stylerir\_\{i\}and the value of being distinguishable from other authors\. The authenticity cost makes full conformity costly, while the distinctiveness payoff depends on the realized styles of others and couples with the author’s choices\.
Payoffs in[section4\.1](https://arxiv.org/html/2607.27134#S4.Ex33)use squared Euclidean distance rather than the Jensen–Shannon divergence used forDtD^\{t\}; JS is locally quadratic \([AppendixA](https://arxiv.org/html/2607.27134#A1)\)\. We therefore evaluate long\-run diversity using the quadratic measure introduced in[Section3](https://arxiv.org/html/2607.27134#S3),D^∞\(λ\):=D^\(a1\(λ1\),…,an\(λn\)\)\\widehat\{D\}\_\{\\infty\}\(\\lambda\):=\\widehat\{D\}\(a\_\{1\}\(\\lambda\_\{1\}\),\\ldots,a\_\{n\}\(\\lambda\_\{n\}\)\), which matches the payoff geometry and yields closed\-form equilibria\. We further make the following structural assumption\.
###### Definition 1\.
A population has*orthogonal signatures*if⟨ui,uj⟩=0\\langle u\_\{i\},u\_\{j\}\\rangle=0for alli≠ji\\neq j, withdi2:=‖ui‖22\>0d\_\{i\}^\{2\}:=\\\|u\_\{i\}\\\|\_\{2\}^\{2\}\>0for allii\.
Orthogonality is a tractable benchmark yielding closed\-form welfare results\. Signatures lie in them−1m\-1\-dimensional tangent space of the simplex, so[Definition1](https://arxiv.org/html/2607.27134#Thmdefinition1)requiresn≤m−1n\\leq m\-1: the benchmark describes populations no larger than the feature dimension\. This is not restrictive whenFFis a rich inventory of lexical and syntactic markers, but it does couple population size to feature\-set size\. Appendix[D](https://arxiv.org/html/2607.27134#A4)relaxes orthogonality\.
Under orthogonality,‖ai−aj‖22=σi2di2\+σj2dj2\\\|a\_\{i\}\-a\_\{j\}\\\|\_\{2\}^\{2\}=\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\+\\sigma\_\{j\}^\{2\}d\_\{j\}^\{2\}, and substituting in[section4\.1](https://arxiv.org/html/2607.27134#S4.Ex33)gives the separable form
Ui\(λ\)=di22\[\(θi−bi\)σi2−ci\(1−σi\)2\]\+θi2\(n−1\)∑j≠iσj2dj2,U\_\{i\}\(\\lambda\)=\\frac\{d\_\{i\}^\{2\}\}\{2\}\\Big\[\(\\theta\_\{i\}\\\!\-\\\!b\_\{i\}\)\\,\\sigma\_\{i\}^\{2\}\\\!\-\\\!c\_\{i\}\\,\(1\-\\sigma\_\{i\}\)^\{2\}\\Big\]\+\\frac\{\\theta\_\{i\}\}\{2\(n\-1\)\}\\sum\_\{j\\neq i\}\\sigma\_\{j\}^\{2\}\\,d\_\{j\}^\{2\},\(4\)and the long\-run diversity simplifies to
D^∞\(λ\)=1n∑i=1nσi2di2\.\\widehat\{D\}\_\{\\infty\}\(\\lambda\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\sigma\_\{i\}^\{2\}\\,d\_\{i\}^\{2\}\.\(5\)
The coupling in[eq\.4](https://arxiv.org/html/2607.27134#S4.E4)is purely through the second term: authorii’s choice does not change their own best response, but it does change everyone else’s*payoff*\.
###### Lemma 4\.1\.
Under orthogonal signatures, for everyi≠ji\\neq j,
∂Uj∂λi=−θjn−1σidi2≤0,\\frac\{\\partial U\_\{j\}\}\{\\partial\\lambda\_\{i\}\}\\;=\\;\-\\,\\frac\{\\theta\_\{j\}\}\{n\-1\}\\,\\sigma\_\{i\}\\,d\_\{i\}^\{2\}\\;\\leq\\;0,with strict inequality wheneverθj\>0\\theta\_\{j\}\>0andλi<1\\lambda\_\{i\}<1\. In particular, an increase in any author’s conformity strictly reduces the payoff of every other author who places positive value on distinctiveness, including authors who choose not to conform\.
Pairwise contrast is a shared resource: the distance‖ai−aj‖2\\\|a\_\{i\}\-a\_\{j\}\\\|^\{2\}enters bothUiU\_\{i\}andUjU\_\{j\}, so when authoriimoves toward the norm they consume contrast that authorjjwas also drawing on\.[Section4\.1](https://arxiv.org/html/2607.27134#S4.SS1)also resolves the strategic question for non\-adapters: an author for whom resisting is optimal still bears the cost of others’ adaptation, because the pool of styles against which their distinctiveness is measured collapses towardq0q^\{0\}\. There is no insulation in holding out\.
Equilibrium exceeds optimal conformity\.Next we compare the conformity levels chosen by self\-interested authors at Nash equilibrium with those chosen by a social planner that maximizes total welfare\. The next theorem shows that authors over\-conform relative to the social optimum, which leads to weakly lower long\-run diversity and weakly lower welfare at equilibrium\.
###### Theorem 4\.2\.
Consider the game with payoffs in[Section4\.1](https://arxiv.org/html/2607.27134#S4.Ex33)under orthogonal signatures, withci\>0c\_\{i\}\>0for alliiand letθ¯−i:=1n−1∑j≠iθj\\bar\{\\theta\}\_\{\-i\}:=\\frac\{1\}\{n\-1\}\\sum\_\{j\\neq i\}\\theta\_\{j\}denote the average distinctiveness value of the other authors\. Then:
- \(i\)Every author has a strictly dominant strategy, so the game has a unique Nash equilibriumλNE\\lambda^\{\\mathrm\{NE\}\}, given by λiNE=\(bi−θi\)\+\(bi−θi\)\+\+ci,where\(x\)\+:=max\{0,x\}\.\\lambda\_\{i\}^\{\\mathrm\{NE\}\}\\;=\\;\\frac\{\(b\_\{i\}\-\\theta\_\{i\}\)\_\{\+\}\}\{\\,\(b\_\{i\}\-\\theta\_\{i\}\)\_\{\+\}\+c\_\{i\}\\,\},\\quad\\text\{where \}\(x\)\_\{\+\}:=\\max\\\{0,x\\\}\.
- \(ii\)Utilitarian welfareW\(λ\):=∑i=1nUi\(λ\)W\(\\lambda\):=\\sum\_\{i=1\}^\{n\}U\_\{i\}\(\\lambda\)is maximized at the unique profileλSO\\lambda^\{\\mathrm\{SO\}\}given by λiSO=\(bi−θi−θ¯−i\)\+\(bi−θi−θ¯−i\)\+\+ci\.\\lambda\_\{i\}^\{\\mathrm\{SO\}\}\\;=\\;\\frac\{\\big\(b\_\{i\}\-\\theta\_\{i\}\-\\bar\{\\theta\}\_\{\-i\}\\big\)\_\{\+\}\}\{\\,\\big\(b\_\{i\}\-\\theta\_\{i\}\-\\bar\{\\theta\}\_\{\-i\}\\big\)\_\{\+\}\+c\_\{i\}\\,\}\.
- \(iii\)For every author,λiNE≥λiSO\\lambda\_\{i\}^\{\\mathrm\{NE\}\}\\geq\\lambda\_\{i\}^\{\\mathrm\{SO\}\}, with strict inequality if and only ifλiNE\>0\\lambda\_\{i\}^\{\\mathrm\{NE\}\}\>0andθ¯−i\>0\\bar\{\\theta\}\_\{\-i\}\>0\. That is, every author who conforms at all over\-conforms relative to the social optimum, and consequentlyD^∞\(λNE\)≤D^∞\(λSO\)\\widehat\{D\}\_\{\\infty\}\(\\lambda^\{\\mathrm\{NE\}\}\)\\leq\\widehat\{D\}\_\{\\infty\}\(\\lambda^\{\\mathrm\{SO\}\}\)andW\(λNE\)≤W\(λSO\)W\(\\lambda^\{\\mathrm\{NE\}\}\)\\leq W\(\\lambda^\{\\mathrm\{SO\}\}\)\.
Authorii’s privately optimal conformity treats their distinctiveness as worthθi\\theta\_\{i\}, while the planner values it atθi\+θ¯−i\\theta\_\{i\}\+\\bar\{\\theta\}\_\{\-i\}, because every other author also derives contrast from authorii’s signature\. Each author’s over\-conformity increases with how much the rest of the community values contrast with them\. Because the wedge depends on the averageθ¯−i\\bar\{\\theta\}\_\{\-i\}, it need not vanish in large populations\. Notably, since the equilibrium is in dominant strategies, the inefficiency is driven purely by the payoff externality of[Section4\.1](https://arxiv.org/html/2607.27134#S4.SS1)—exactly as in algorithmic monoculture\[[24](https://arxiv.org/html/2607.27134#bib.bib21)\], where individually optimal reliance on a shared algorithm reduces aggregate welfare without any agent erring\.
### 4\.2The price of monoculture
We quantify the diversity loss caused by strategic conformity via the price of monoculture \(PoM\\mathrm\{PoM\}\), which compares the long\-run diversity of the socially optimal conformity profile with that of Nash equilibrium\. We evaluatePoM\\mathrm\{PoM\}in a symmetric setting, where all authors have the same conformity reward, authenticity cost, value for distinctiveness and distance from the model norm\. In this case, thePoM\\mathrm\{PoM\}has a closed\-form expression and reveals three regimes: no inefficiency, pure deadweight conformity and interior over\-conformity with an arbitrarily large diversity loss\.
###### Definition 2\.
The*price of monoculture*of an instance is the ratio of socially optimal to equilibrium long\-run diversity,
PoM:=D^∞\(λSO\)D^∞\(λNE\)=∑i\(σiSO\)2di2∑i\(σiNE\)2di2≥1\.\\mathrm\{PoM\}\\;:=\\;\\frac\{\\widehat\{D\}\_\{\\infty\}\(\\lambda^\{\\mathrm\{SO\}\}\)\}\{\\widehat\{D\}\_\{\\infty\}\(\\lambda^\{\\mathrm\{NE\}\}\)\}\\;=\\;\\frac\{\\sum\_\{i\}\(\\sigma\_\{i\}^\{\\mathrm\{SO\}\}\)^\{2\}d\_\{i\}^\{2\}\}\{\\sum\_\{i\}\(\\sigma\_\{i\}^\{\\mathrm\{NE\}\}\)^\{2\}d\_\{i\}^\{2\}\}\\;\\geq\\;1\.
###### Corollary 4\.3\.
Supposebi=bb\_\{i\}=b,ci=cc\_\{i\}=c,θi=θ\\theta\_\{i\}=\\theta, anddi2=d2d\_\{i\}^\{2\}=d^\{2\}for allii\. Then:
- \(i\)Ifb≤θb\\leq\\theta:λNE=λSO=0\\lambda^\{\\mathrm\{NE\}\}=\\lambda^\{\\mathrm\{SO\}\}=0andPoM=1\\mathrm\{PoM\}=1\. Distinctiveness is at least as valuable as conformity privately, so no inefficiency arises\.
- \(ii\)Ifθ<b≤2θ\\theta<b\\leq 2\\theta:λNE=b−θb\+c−θ\>0\\lambda^\{\\mathrm\{NE\}\}=\\frac\{b\-\\theta\}\{b\+c\-\\theta\}\>0whileλSO=0\\lambda^\{\\mathrm\{SO\}\}=0, and PoM=\(b\+c−θc\)2\.\\mathrm\{PoM\}=\\left\(\\frac\{b\+c\-\\theta\}\{c\}\\right\)^\{\\\!2\}\.Every author rationally conforms, yet the planner would have*no one*conform: equilibrium conformity is pure deadweight\.
- \(iii\)Ifb\>2θb\>2\\theta: both levels are interior,λNE=b−θb\+c−θ\>λSO=b−2θb\+c−2θ\\lambda^\{\\mathrm\{NE\}\}=\\frac\{b\-\\theta\}\{b\+c\-\\theta\}\>\\lambda^\{\\mathrm\{SO\}\}=\\frac\{b\-2\\theta\}\{b\+c\-2\\theta\}, and PoM=\(b\+c−θb\+c−2θ\)2=\(1\+θb\+c−2θ\)2,\\mathrm\{PoM\}=\\left\(\\frac\{b\+c\-\\theta\}\{\\,b\+c\-2\\theta\\,\}\\right\)^\{\\\!2\}=\\left\(1\+\\frac\{\\theta\}\{\\,b\+c\-2\\theta\\,\}\\right\)^\{\\\!2\},which is increasing inθ\\theta, equals11atθ=0\\theta=0, and can diverge along sequences withb\+c−2θ↓0b\+c\-2\\theta\\downarrow 0\. The per\-author welfare loss at equilibrium is 1n\(W\(λSO\)−W\(λNE\)\)=d22⋅c2θ2\(b\+c−2θ\)\(b\+c−θ\)2,\\frac\{1\}\{n\}\\Big\(W\(\\lambda^\{\\mathrm\{SO\}\}\)\-W\(\\lambda^\{\\mathrm\{NE\}\}\)\\Big\)=\\frac\{d^\{2\}\}\{2\}\\cdot\\frac\{c^\{2\}\\theta^\{2\}\}\{\(b\+c\-2\\theta\)\(b\+c\-\\theta\)^\{2\}\},which is strictly positive wheneverθ\>0\\theta\>0\.
Three observations are worth highlighting\. First, the welfare loss vanishes atθ=0\\theta=0and is quadratic to leading order near zero; when distinctiveness has no value, conformity is simply beneficial standardization\. Thus,θ\\thetadistinguishes benign standardization from harmful monoculture\. Second, in regime \(iiii\), conformity is individually rational but socially wasteful: every author conforms although the planner prefers none\. Third, sinceb\>2θb\>2\\thetaimpliesb\+c−2θ\>cb\+c\-2\\theta\>c, every regime satisfiesPoM≤\(1\+θ/c\)2\\mathrm\{PoM\}\\leq\(1\+\\theta/c\)^\{2\}\. The inefficiency is finite for each fixed instance but not uniformly bounded: it can diverge wheneverθ/c→∞\\theta/c\\to\\infty, including asc↓0c\\downarrow 0withθ\\thetafixed\. Appendix[A](https://arxiv.org/html/2607.27134#A1)discusses externalities on readers \([AppendixA](https://arxiv.org/html/2607.27134#A1)\) and strategic conformity under recursive model updates \([AppendixA](https://arxiv.org/html/2607.27134#A1)\)\.
## 5Quantitative Comparison and Illustrations
We simulaten=100n=100authors overm=10m=10abstract linguistic features for100100independent runs ofT=200T=200steps \(see Appendix[F](https://arxiv.org/html/2607.27134#A6)for details\)\.[Figure1](https://arxiv.org/html/2607.27134#S5.F1)reports the mean and standard deviation of population\-level linguistic diversity across runs\. Panel \(a\) compares the three interaction mechanisms usingβ=0\.25\\beta=0\.25and, for IM 3,ρ=2/3\\rho=2/3\(γ=0\.5\\gamma=0\.5andδ=0\.25\\delta=0\.25\)\. Across all mechanisms, LLM assistance initially reduces diversity, after whichDtD^\{t\}stabilizes at a mechanism\-specific level\. Under the baseline parameterization, recursive updates of the shared model \(IM 2\) produce a faster decline inDtD^\{t\}and a lower limiting diversity than IM 1\. By[Section3\.3](https://arxiv.org/html/2607.27134#S3.SS3), under common conformity the mechanisms have identical limiting pairwise geometry, so any JS gap is purely positional\. In our baseline, heterogeneousλi\\lambda\_\{i\}additionally produces the quadratic gap isolated in Appendix[F](https://arxiv.org/html/2607.27134#A6)\. Panel \(b\) examines personalization in IM 3\. Holdingβ=0\.25\\beta=0\.25fixed, we varyρ\\rhoand setγ=\(1−β\)ρ\\gamma=\(1\-\\beta\)\\rhoandδ=\(1−β\)\(1−ρ\)\\delta=\(1\-\\beta\)\(1\-\\rho\), thereby reallocating update weight from population\-level to author\-specific feedback\. Diversity atT=200T=200increases monotonically withρ\\rho, from approximately0\.0570\.057atρ=0\\rho=0to0\.1710\.171atρ=1\\rho=1, showing that stronger personalization can preserve population\-level linguistic diversity\. Additional details and experiments are reported in Appendix[F](https://arxiv.org/html/2607.27134#A6)\.
Figure 1:Population\-level linguistic diversity under LLM assistance\. Panel \(a\) shows the evolution ofDtD^\{t\}under IM 1–3, with time displayed on alog\(1\+t\)\\log\(1\+t\)scale and tick labels reporting the original time steps\. Panel \(b\) shows IM 3 diversity atT=200T=200as personalizationρ\\rhoincreases\. Lines report means over100100runs and shading denotes±1\\pm 1standard deviation; larger values indicate greater diversity\.
## 6Conclusions, Limitations and Future Work
We formalized linguistic monoculture as a population\-level dynamical process over author and model distributions\. Under our mechanisms, shared models can pull authors toward a common norm, while author\-specific preferences and personalization can preserve diversity\. Within our utility model, equilibrium conformity can exceed the social optimum, yielding a potentially unbounded price of monoculture in some regimes\. When distinctiveness has little value, convergence can instead be efficient\.
Limitations and future work\.Our work introduces a mathematical framework for formalizing linguistic monoculture in LLM\-assisted language use\. The framework is stylized by design: it focuses on a limited set of interaction mechanisms that make it possible to analyze how linguistic diversity evolves under different forms of author–model interaction\. As with many theoretical models, this tractability relies on simplifying assumptions, which allow us to obtain explicit convergence, equilibrium and welfare guarantees\. These assumptions provide a foundation for identifying and formally analyzing mechanisms of linguistic convergence, while leaving richer behavioral and empirical refinements for future work\. We therefore discuss limitations of the framework and outline directions for connecting it more closely to empirical settings and broadening its applicability\.
Heterogeneous non\-LLM language exposure\.Our framework isolates LLM\-mediated adaptation from other influences, such as collaborators, reading research articles, reviewers, disciplinary norms, broader media environments and other human factors that shape language use\. Thus, our results characterize convergence under controlled settings rather than a complete model of linguistic change under real\-world conditions\. Future work should incorporate heterogeneous non\-LLM exposures, social networks, recursive norm\-shaping incentives, and empirical calibration of conformity and personalization parameters from longitudinal corpora of LLM\-assisted writing\.
Fixed conformity, quadratic payoff and orthogonal signaturesIn[Section4](https://arxiv.org/html/2607.27134#S4), we make three simplifying assumptions to obtain closed\-form expressions for equilibrium, welfare, and the price of monoculture\. First, each author chooses a fixed conformity levelλi\\lambda\_\{i\}, which determines their writing policy and remains unchanged over time\. In practice, authors may update their conformity levels in response to experience, feedback, or changes in model behavior\. Second, the analysis uses quadratic payoffs\. Finally, the main\-text closed\-form results assume orthogonal author signatures, permitting at mostm−1m\-1nonzero signatures inΔm−1\\Delta^\{m\-1\}\. These assumptions isolate the conformity externality, but real signatures may be correlated and distinctiveness need not be valued quadratically\. Extending the framework to dynamic conformity choices, richer similarity structures, and alternative payoff models is an important direction for future work\.
Empirical validationOur theoretical analysis is supported by simulations intended to illustrate the qualitative behavior predicted by our framework, rather than to provide empirically calibrated forecasts\. The parameter choices allow us to compare interaction mechanisms under controlled conditions and should not be interpreted as estimates of linguistic convergence\. A natural next step is to verify these findings in controlled longitudinal studies of LLM\-assisted writing, measuring how authors’ linguistic\-style distributions evolve under fixed, recursively updated and personalized assistance\. In this sense, our framework provides a foundation for future empirical work by identifying the interaction mechanisms, parameters, and diversity measures that such studies could estimate or experimentally vary\.
Acknowledgments and Funding\.Suhas Thejaswi acknowledges support from the Technology Industries of Finland Centennial Foundation through a grant awarded to Aalto University\. The authors declare that the funding source does not create any conflict of interest in relation to this work\.
GenAI usage statement\.LLMs were used to edit and polish author\-written text, and to assist with implementation and code refinement\.
## References
- \[1\]M\. Abdulhai, I\. White, Y\. Wan, I\. Qureshi, J\. Leibo, M\. Kleiman\-Weiner, and N\. Jaques\(2026\)How LLMs distort our written language\.arXiv preprint arXiv:2603\.18161\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[2\]D\. Agarwal, M\. Naaman, and A\. Vashistha\(2025\)AI suggestions homogenize writing toward western styles and diminish cultural nuances\.InProceedings of the CHI conference on human factors in computing systems,pp\. 1–21\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[3\]B\. R\. Anderson, J\. H\. Shah, and M\. Kreminski\(2024\)Homogenization effects of large language models on human creative ideation\.InProceedings of the 16th conference on creativity & cognition,pp\. 413–425\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[4\]C\. Bazerman\(1988\)Shaping written knowledge: the genre and activity of the experimental article in science\.University of Wisconsin Press,Madison, WI\.External Links:ISBN 978\-0299116941Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1)\.
- \[5\]A\. Bell\(1984\)Language style as audience design\.Language in Society13\(2\),pp\. 145–204\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p13.1),[§1](https://arxiv.org/html/2607.27134#S1.p6.1)\.
- \[6\]D\. Biber and S\. Conrad\(2019\)Register, genre, and style\.Cambridge University Press\.Cited by:[Remark A\.1](https://arxiv.org/html/2607.27134#A1.1.p1.1.1)\.
- \[7\]L\. M\. Bietti and A\. Bangerter\(2026\)Will the widespread use of large language models in scientific writing undermine scientists’ critical thinking?\.PLoS biology24\(6\),pp\. e3003801\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p4.1)\.
- \[8\]M\. Brucks and O\. Toubia\(2025\)Prompt architecture induces methodological artifacts in large language models\.PLOS one20\(4\),pp\. e0319159\.Cited by:[Remark A\.1](https://arxiv.org/html/2607.27134#A1.1.p1.1.1)\.
- \[9\]R\. Chang and H\. Wang\(2025\)Communication Accommodation Between Large Language Models and Users Across Cultures \(Student Abstract\)\.InAAAI Conference on Artificial Intelligence,pp\. 29331–29333\.External Links:[Document](https://dx.doi.org/10.1609/AAAI.V39I28.35241),[Link](https://mlanthology.org/aaai/2025/chang2025aaai-communication/)Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p13.1)\.
- \[10\]G\. Chen and Y\. Lou\(2019\)Naming game\.Switzerland: Springer International Publishing\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p13.1)\.
- \[11\]B\. De Vylder and K\. Tuyls\(2006\)How to reach linguistic consensus: a proof of convergence for the naming game\.Journal of theoretical biology242\(4\),pp\. 818–831\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p13.1)\.
- \[12\]M\. H\. DeGroot\(1974\)Reaching a consensus\.Journal of the American Statistical association69\(345\),pp\. 118–121\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p10.1)\.
- \[13\]A\. R\. Doshi and O\. P\. Hauser\(2024\)Generative AI enhances individual creativity but reduces the collective diversity of novel content\.Science Advances10\(28\),pp\. eadn5290\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1)\.
- \[14\]N\. J\. Enfield\(2015\)Linguistic relativity from reference to agency\.Annual Review of Anthropology44,pp\. 207–224\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p4.1)\.
- \[15\]N\. Evans and S\. C\. Levinson\(2009\)The myth of language universals: language diversity and its importance for cognitive science\.Behavioral and Brain Sciences32\(5\),pp\. 429–448\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p4.1)\.
- \[16\]N\. E\. Friedkin and E\. C\. Johnsen\(1990\)Social influence and opinions\.Journal of Mathematical Sociology15\(3\-4\),pp\. 193–206\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p10.1)\.
- \[17\]M\. Geng, C\. Chen, Y\. Wu, Y\. Wan, P\. Zhou, and D\. Chen\(2025\)The impact of large language models in academia: from writing to speaking\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 19303–19319\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[18\]M\. Geng and R\. Trotta\(2025\)Human\-LLM coevolution: evidence from academic writing\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 12689–12696\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[19\]H\. Giles\(2016\)Communication accommodation theory: negotiating personal relationships and social identities across contexts\.Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p13.1),[§1](https://arxiv.org/html/2607.27134#S1.p6.1)\.
- \[20\]J\. J\. Gumperz and S\. C\. Levinson \(Eds\.\)\(1996\)Rethinking linguistic relativity\.Studies in the Social and Cultural Foundations of Language, Vol\.17,Cambridge University Press,Cambridge\.External Links:ISBN 9780521567067Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p4.1)\.
- \[21\]R\. Hegselmann and U\. Krause\(2002\)Opinion dynamics and bounded confidence: models, analysis and simulation\.J\. Artif\. Soc\. Soc\. Simul\.5\(3\)\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p10.1)\.
- \[22\]K\. Hyland\(2009\)Academic discourse: english in a global context\.Continuum,London\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p4.1)\.
- \[23\]M\. Jakesch, A\. Bhat, D\. Buschek, L\. Zalmanson, and M\. Naaman\(2023\)Co\-writing with opinionated language models affects users’ views\.InProceedings of the CHI conference on human factors in computing systems,pp\. 1–15\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1)\.
- \[24\]J\. Kleinberg and M\. Raghavan\(2021\)Algorithmic monoculture and social welfare\.Proceedings of the National Academy of Sciences118\(22\),pp\. e2018340118\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p12.1),[§4\.1](https://arxiv.org/html/2607.27134#S4.SS1.p10.5),[§4](https://arxiv.org/html/2607.27134#S4.p1.3)\.
- \[25\]R\. Kleinberg, E\. Sinanaj, and É\. Tardos\(2026\)Price of anarchy of algorithmic monoculture\.arXiv preprint arXiv:2604\.00444\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p12.1)\.
- \[26\]D\. Kobak, R\. González\-Márquez, E\. Horvát, and J\. Lause\(2025\)Delving into LLM\-assisted writing in biomedical publications through excess vocabulary\.Science Advances11\(27\),pp\. eadt3813\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[27\]B\. Latour and S\. Woolgar\(1986\)Laboratory life: the construction of scientific facts\.2nd edition,Princeton University Press,Princeton, NJ\.External Links:ISBN 9780691028323Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1)\.
- \[28\]W\. Liang, Y\. Zhang, Z\. Wu, H\. Lepp, W\. Ji, X\. Zhao, H\. Cao, S\. Liu, S\. He, Y\. Cui,et al\.\(2025\)Quantifying large language model usage in scientific papers\.Nature Human Behaviour9\(12\),pp\. 2599–2609\.External Links:[Document](https://dx.doi.org/10.1038/s41562-025-02273-8)Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[29\]J\. Lin\(1991\)Divergence measures based on the shannon entropy\.IEEE Transactions on Information theory37\(1\),pp\. 145–151\.Cited by:[§2](https://arxiv.org/html/2607.27134#S2.p4.8)\.
- \[30\]J\. A\. Lucy\(1997\)Linguistic relativity\.Annual review of anthropology26\(1\),pp\. 291–312\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p4.1)\.
- \[31\]K\. Moon, A\. E\. Green, and K\. Kushlev\(2025\)Homogenizing effect of large language models \(llms\) on creative diversity: an empirical comparison of human and chatgpt writing\.Computers in Human Behavior: Artificial Humans6,pp\. 100207\.External Links:[Document](https://dx.doi.org/10.1016/j.chbah.2025.100207)Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[32\]V\. Padmakumar and H\. He\(2024\)Does writing with language models reduce content diversity?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 642–669\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p1.1),[§1](https://arxiv.org/html/2607.27134#S1.p2.1),[§1](https://arxiv.org/html/2607.27134#S1.p9.1)\.
- \[33\]V\. N\. Pescuma, D\. Serova, J\. Lukassek, A\. Sauermann, R\. Schäfer, A\. Adli, F\. Bildhauer, M\. Egg, K\. Hülk, A\. Ito,et al\.\(2023\)Situating language register across the ages, languages, modalities, and cultural aspects: evidence from complementary methods\.Frontiers in Psychology13,pp\. 964658\.Cited by:[Remark A\.1](https://arxiv.org/html/2607.27134#A1.1.p1.1.1)\.
- \[34\]A\. J\. Peterson\(2025\)AI and the problem of knowledge collapse\.AI & Society40\(5\),pp\. 3249–3269\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p11.1)\.
- \[35\]I\. Shumailov, Z\. Shumaylov, Y\. Zhao, N\. Papernot, R\. Anderson, and Y\. Gal\(2024\)AI models collapse when trained on recursively generated data\.Nature631\(8022\),pp\. 755–759\.Cited by:[Remark A\.5](https://arxiv.org/html/2607.27134#A1.5.p1.6.5),[§1](https://arxiv.org/html/2607.27134#S1.p11.1),[§2](https://arxiv.org/html/2607.27134#S2.p3.7)\.
- \[36\]Z\. Sourati, F\. Karimi\-Malekabadi, M\. Ozcan, C\. McDaniel, A\. Ziabari, J\. Trager, A\. Tak, M\. Chen, F\. Morstatter, and M\. Dehghani\(2025\)The shrinking landscape of linguistic diversity in the age of large language models\.arXiv preprint arXiv:2502\.11266\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p1.1)\.
- \[37\]Z\. Sourati, A\. S\. Ziabari, and M\. Dehghani\(2026\)The homogenizing effect of large language models on human expression and thought\.Trends in Cognitive Sciences\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1)\.
- \[38\]J\. M\. Swales\(1990\)Genre analysis: english in academic and research settings\.Cambridge University Press,Cambridge, UK\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p2.1)\.
- \[39\]D\. Wright, S\. Masud, J\. Moore, S\. Yadav, M\. Antoniak, P\. E\. Christensen, C\. Y\. Park, and I\. Augenstein\(2025\)Epistemic diversity and knowledge collapse in large language models\.arXiv preprint arXiv:2510\.04226\.Cited by:[§1](https://arxiv.org/html/2607.27134#S1.p11.1)\.
## Appendix ARemarks and Further Clarifications
This section provides additional clarifications and remarks that are omitted from the main text due to space constraints\.
## Appendix BOmitted Proofs from Section[3](https://arxiv.org/html/2607.27134#S3)
See[3\.1](https://arxiv.org/html/2607.27134#S3.SS1)
###### Proof\.
Leta:=𝒜\(q0,z\)a:=\\mathcal\{A\}\(q^\{0\},z\)\. By assumption, the update rule becomes
pit\+1=\(1−αi\)pit\+αia\.p\_\{i\}^\{t\+1\}=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,a\.By subtractingaaon both sides followed by induction ontt, we get,
pit\+1−a\\displaystyle p\_\{i\}^\{t\+1\}\-a=\(1−αi\)\(pit−a\)\\displaystyle=\(1\-\\alpha\_\{i\}\)\(\\,p\_\{i\}^\{t\}\-a\)pit−a\\displaystyle p\_\{i\}^\{t\}\-a=\(1−αi\)t\(pi0−a\)\\displaystyle=\(1\-\\alpha\_\{i\}\)^\{t\}\\,\(p\_\{i\}^\{0\}\-a\)Takingℓ1\\ell\_\{1\}\-norms and relying on the fact that bothpi0p\_\{i\}^\{0\}andaaare probability distributions, so theirℓ1\\ell\_\{1\}distance is at most22, we have
‖pit−a‖1\\displaystyle\\\|p\_\{i\}^\{t\}\-a\\\|\_\{1\}=\(1−αi\)t‖\(pi0−a\)‖1≤2\(1−αi\)t\.\\displaystyle=\(1\-\\alpha\_\{i\}\)^\{t\}\\,\\\|\\left\(p\_\{i\}^\{0\}\-a\\right\)\\\|\_\{1\}\\leq 2\\,\(1\-\\alpha\_\{i\}\)^\{t\}\.Using the standard inequality1−x≤e−x1\-x\\leq e^\{\-x\}forx∈\[0,1\]x\\in\[0,1\],\(1−αi\)t≤e−αit\(1\-\\alpha\_\{i\}\)^\{t\}\\leq e^\{\-\\alpha\_\{i\}t\}, we get
‖pit−a‖1≤2e−αit\.\\\|p\_\{i\}^\{t\}\-a\\\|\_\{1\}\\leq 2e^\{\-\\alpha\_\{i\}t\}\.Now consider any pair of authorsi,ji,j\. By triangle inequality and using the bound above, it follows that,
‖pit−pjt‖1\\displaystyle\\\|p\_\{i\}^\{t\}\-p\_\{j\}^\{t\}\\\|\_\{1\}≤‖pit−a‖1\+‖pjt−a‖1\\displaystyle\\leq\\\|p\_\{i\}^\{t\}\-a\\\|\_\{1\}\+\\\|p\_\{j\}^\{t\}\-a\\\|\_\{1\}≤2e−αit\+2e−αjt≤4e−αmint\.\\displaystyle\\leq 2e^\{\-\\alpha\_\{i\}t\}\+2e^\{\-\\alpha\_\{j\}t\}\\leq 4e^\{\-\\alpha\_\{\\min\}\\,t\}\.Now, using the standard boundJS\(pit,pjt\)≤12‖pit−pjt‖1JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\leq\\frac\{1\}\{2\}\\\|p\_\{i\}^\{t\}\-p\_\{j\}^\{t\}\\\|\_\{1\}, we obtainJS\(pit,pjt\)≤2e−αmintJS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\leq 2e^\{\-\\alpha\_\{\\min\}\\,t\}\. And, averaging over all ordered pairsi≠ji\\neq j, we get
Dt=1n\(n−1\)∑i≠jJS\(pit,pjt\)≤2e−αmint\.D^\{t\}=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\leq 2e^\{\-\\alpha\_\{\\min\}t\}\.Therefore, to ensureDt≤εD^\{t\}\\leq\\varepsilon, it is sufficient that
2e−αmint≤ε⟹t≥log\(2/ε\)αmin=O\(log\(1/ε\)αmin\),\\displaystyle 2e^\{\-\\alpha\_\{\\min\}\\,t\}\\leq\\varepsilon\\implies t\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\\alpha\_\{\\min\}\}=O\\left\(\\frac\{\\log\(1/\\varepsilon\)\}\{\\alpha\_\{\\min\}\}\\right\),which completes the proof\. ∎
See[3\.1](https://arxiv.org/html/2607.27134#S3.SS1)
###### Proof of[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1)\.
For each authorii, defineai:=𝒜i\(q0,zit\)a\_\{i\}:=\\mathcal\{A\}\_\{i\}\(q^\{0\},z\_\{i\}^\{t\}\), which is fixed over time by assumption\. The update rule becomes
pit\+1=\(1−αi\)pit\+αiai\.p\_\{i\}^\{t\+1\}=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,a\_\{i\}\.Subtractingaia\_\{i\}from both sides, followed by induction gives
pit\+1−ai\\displaystyle p\_\{i\}^\{t\+1\}\-a\_\{i\}=\(1−αi\)\(pit−ai\)\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,\(p\_\{i\}^\{t\}\-a\_\{i\}\)pit−ai\\displaystyle p\_\{i\}^\{t\}\-a\_\{i\}=\(1−αi\)t\(pi0−ai\)\.\\displaystyle=\(1\-\\alpha\_\{i\}\)^\{t\}\\,\(p\_\{i\}^\{0\}\-a\_\{i\}\)\.Takingℓ1\\ell\_\{1\}\-norms and using the fact that bothpi0p\_\{i\}^\{0\}andaia\_\{i\}are probability distributions whoseℓ1\\ell\_\{1\}distance is at most22, gives
‖pit−ai‖1≤2\(1−αi\)t≤2e−αit≤2e−αmint\.\\\|p\_\{i\}^\{t\}\-a\_\{i\}\\\|\_\{1\}\\leq 2\\,\(1\-\\alpha\_\{i\}\)^\{t\}\\leq 2\\,e^\{\-\\alpha\_\{i\}t\}\\leq 2\\,e^\{\-\\alpha\_\{\\min\}t\}\.Thuspit→aip\_\{i\}^\{t\}\\to a\_\{i\}for every authorii, ast→∞t\\to\\infty\. Since Jensen–Shannon divergence is continuous on the probability simplex, for every pairi,ji,j,
JS\(pit,pjt\)→JS\(ai,aj\)\.JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\to JS\(a\_\{i\},a\_\{j\}\)\.Averaging over all ordered pairsi≠ji\\neq j, we obtain
Dt=1n\(n−1\)∑i≠jJS\(pit,pjt\)→1n\(n−1\)∑i≠jJS\(ai,aj\)=D∞\.D^\{t\}=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\to\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(a\_\{i\},a\_\{j\}\)=D^\{\\infty\}\.ThereforeDt→0D^\{t\}\\to 0ifai=aja\_\{i\}=a\_\{j\}for alli,ji,j\. Conversely, sinceJS\(ai,aj\)=0JS\(a\_\{i\},a\_\{j\}\)=0if and only ifai=aja\_\{i\}=a\_\{j\}, we haveD∞=0D^\{\\infty\}=0only when all limiting adapted distributions are identical\.
Convergence of linguistic diversity\.The time steps necessary for convergence of authors to their own limits withinε\>0\\varepsilon\>0,*i\.e\.*‖pit−ai‖1≤ε\\\|p\_\{i\}^\{t\}\-a\_\{i\}\\\|\_\{1\}\\leq\\varepsilon, follows from the analysis in[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1), and it is enough to requiret≥log\(2/ε\)αmint\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\\alpha\_\{\\min\}\}\. However, ifD∞\>0D^\{\\infty\}\>0, then for anyε<D∞\\varepsilon<D^\{\\infty\}, there is no convergence time toDt≤εD^\{t\}\\leq\\varepsilon\. The process stabilizes at a positive level of linguistic diversity\. ∎
See[3\.1](https://arxiv.org/html/2607.27134#S3.SS1)
###### Proof of[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1)\.
Usingai\(λi\):=\(1−λi\)ri\+λiq0a\_\{i\}\(\\lambda\_\{i\}\):=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{0\}, the update rule of the author’s distribution can be written as
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αiai\(λi\)\\displaystyle=\(1\-\\alpha\_\{i\}\)p\_\{i\}^\{t\}\+\\alpha\_\{i\}a\_\{i\}\(\\lambda\_\{i\}\)pit\+1−ai\(λi\)\\displaystyle p\_\{i\}^\{t\+1\}\-a\_\{i\}\(\\lambda\_\{i\}\)=\(1−αi\)\(pit−ai\(λi\)\)\\displaystyle=\(1\-\\alpha\_\{i\}\)\\left\(p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\right\)pit−ai\(λi\)\\displaystyle p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)=\(1−αi\)t\(pi0−ai\(λi\)\)\\displaystyle=\(1\-\\alpha\_\{i\}\)^\{t\}\\left\(p\_\{i\}^\{0\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\right\)‖pit−ai\(λi\)‖1\\displaystyle\\\|p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\\|\_\{1\}=\(1−αi\)t‖pi0−ai\(λi\)‖1,\\displaystyle=\(1\-\\alpha\_\{i\}\)^\{t\}\\\|p\_\{i\}^\{0\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\\|\_\{1\},where the second line subtractsai\(λi\)a\_\{i\}\(\\lambda\_\{i\}\)on both sides, the third follows by induction, and the fourth takesℓ1\\ell\_\{1\}\-norms\. Since bothpi0p\_\{i\}^\{0\}andai\(λi\)a\_\{i\}\(\\lambda\_\{i\}\)are probability distributions, theirℓ1\\ell\_\{1\}\-distance is at most22\. Therefore, using the standard inequality1−x≤e−x1\-x\\leq e^\{\-x\}forx∈\[0,1\]x\\in\[0,1\]andαmin:=mini∈\[n\]αi\\alpha\_\{\\min\}:=\\min\_\{i\\in\[n\]\}\\alpha\_\{i\},
‖pit−ai\(λi\)‖1\\displaystyle\\\|p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\\|\_\{1\}≤2\(1−αi\)t≤2e−αit≤2e−αmint,\\displaystyle\\leq 2\(1\-\\alpha\_\{i\}\)^\{t\}\\leq 2e^\{\-\\alpha\_\{i\}t\}\\leq 2e^\{\-\\alpha\_\{\\min\}t\},maxi∈\[n\]‖pit−ai\(λi\)‖1\\displaystyle\\max\_\{i\\in\[n\]\}\\\|p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\\|\_\{1\}≤2e−αmint\.\\displaystyle\\leq 2e^\{\-\\alpha\_\{\\min\}t\}\.To guaranteemaxi∈\[n\]‖pit−ai\(λi\)‖1≤ε\\max\_\{i\\in\[n\]\}\\\|p\_\{i\}^\{t\}\-a\_\{i\}\(\\lambda\_\{i\}\)\\\|\_\{1\}\\leq\\varepsilon, it is sufficient that2e−αmint≤ε2e^\{\-\\alpha\_\{\\min\}t\}\\leq\\varepsilon, equivalentlyt≥log\(2/ε\)αmint\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\\alpha\_\{\\min\}\}\. This proves the claimed convergence\-time bound\.
Sincepit→ai\(λi\)p\_\{i\}^\{t\}\\to a\_\{i\}\(\\lambda\_\{i\}\)for every authorii, and Jensen–Shannon divergence is continuous on the probability simplex,
JS\(pit,pjt\)→JS\(ai\(λi\),aj\(λj\)\)\.JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\to JS\(a\_\{i\}\(\\lambda\_\{i\}\),a\_\{j\}\(\\lambda\_\{j\}\)\)\.Averaging over all ordered pairsi≠ji\\neq jgives
Dt→D∞\(λ\)=1n\(n−1\)∑i≠jJS\(ai\(λi\),aj\(λj\)\)\.D^\{t\}\\to D^\{\\infty\}\(\\lambda\)=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\\left\(a\_\{i\}\(\\lambda\_\{i\}\),a\_\{j\}\(\\lambda\_\{j\}\)\\right\)\.Because every Jensen–Shannon term is nonnegative andJS\(p,q\)=0JS\(p,q\)=0if and only ifp=qp=q, we haveD∞\(λ\)=0D^\{\\infty\}\(\\lambda\)=0if and only ifai\(λi\)=aj\(λj\)a\_\{i\}\(\\lambda\_\{i\}\)=a\_\{j\}\(\\lambda\_\{j\}\)for alli,ji,j\. In particular, ifλi=1\\lambda\_\{i\}=1for allii, thenai\(λi\)=q0a\_\{i\}\(\\lambda\_\{i\}\)=q^\{0\}for everyii, and thereforeDt→0D^\{t\}\\to 0\.
Finally, suppose thatλi=λ\\lambda\_\{i\}=\\lambdafor all authors\. For any0≤λ1<λ2≤10\\leq\\lambda\_\{1\}<\\lambda\_\{2\}\\leq 1, define
c:=1−λ21−λ1∈\[0,1\]\.c:=\\frac\{1\-\\lambda\_\{2\}\}\{1\-\\lambda\_\{1\}\}\\in\[0,1\]\.Then, for every authorii,
ai\(λ2\)=cai\(λ1\)\+\(1−c\)q0\.a\_\{i\}\(\\lambda\_\{2\}\)=c\\,a\_\{i\}\(\\lambda\_\{1\}\)\+\(1\-c\)q^\{0\}\.By joint convexity of Jensen–Shannon divergence,
JS\(ai\(λ2\),aj\(λ2\)\)≤cJS\(ai\(λ1\),aj\(λ1\)\)\.JS\\bigl\(a\_\{i\}\(\\lambda\_\{2\}\),a\_\{j\}\(\\lambda\_\{2\}\)\\bigr\)\\leq c\\,JS\\bigl\(a\_\{i\}\(\\lambda\_\{1\}\),a\_\{j\}\(\\lambda\_\{1\}\)\\bigr\)\.Averaging over all ordered pairs shows thatD∞\(λ\)D^\{\\infty\}\(\\lambda\)is nonincreasing in the common conformity levelλ\\lambda\.
Moreover,
ai\(λ\)−aj\(λ\)=\(1−λ\)\(ri−rj\),‖ai\(λ\)−aj\(λ\)‖1=\(1−λ\)‖ri−rj‖1\.a\_\{i\}\(\\lambda\)\-a\_\{j\}\(\\lambda\)=\(1\-\\lambda\)\(r\_\{i\}\-r\_\{j\}\),\\qquad\\\|a\_\{i\}\(\\lambda\)\-a\_\{j\}\(\\lambda\)\\\|\_\{1\}=\(1\-\\lambda\)\\\|r\_\{i\}\-r\_\{j\}\\\|\_\{1\}\.UsingJS\(p,q\)≤12‖p−q‖1JS\(p,q\)\\leq\\frac\{1\}\{2\}\\\|p\-q\\\|\_\{1\}, we obtain
D∞\(λ\)≤\(1−λ\)12n\(n−1\)∑i≠j‖ri−rj‖1\.D^\{\\infty\}\(\\lambda\)\\leq\(1\-\\lambda\)\\frac\{1\}\{2n\(n\-1\)\}\\sum\_\{i\\neq j\}\\\|r\_\{i\}\-r\_\{j\}\\\|\_\{1\}\.ThusD∞\(λ\)→0D^\{\\infty\}\(\\lambda\)\\to 0asλ→1\\lambda\\to 1, andD∞\(λ\)=O\(1−λ\)D^\{\\infty\}\(\\lambda\)=O\(1\-\\lambda\)\. ∎
See[3\.2](https://arxiv.org/html/2607.27134#S3.SS2)
###### Proof of[Section3\.2](https://arxiv.org/html/2607.27134#S3.SS2)\.
At equilibriumpit\+1=pit=pi∗p\_\{i\}^\{t\+1\}=p\_\{i\}^\{t\}=p\_\{i\}^\{\\ast\}andqt=q∗q^\{t\}=q^\{\\ast\}\. Sinceαi\>0\\alpha\_\{i\}\>0, the author distributions satisfy
pi∗\\displaystyle p\_\{i\}^\{\\ast\}=\(1−αi\)pi∗\+αi\(\(1−λi\)ri\+λiq∗\)\\displaystyle=\(1\-\\alpha\_\{i\}\)p\_\{i\}^\{\\ast\}\+\\alpha\_\{i\}\\left\(\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{\\ast\}\\right\)αipi∗\\displaystyle\\alpha\_\{i\}\\,p\_\{i\}^\{\\ast\}=αi\(\(1−λi\)ri\+λiq∗\)\\displaystyle=\\alpha\_\{i\}\\left\(\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{\\ast\}\\right\)pi∗\\displaystyle p\_\{i\}^\{\\ast\}=\(1−λi\)ri\+λiq∗\\displaystyle=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{\\ast\}\(6\)Moreover, at equilibrium the model update satisfies
q∗\\displaystyle q^\{\\ast\}=βq∗\+\(1−β\)P∗,\\displaystyle=\\beta q^\{\\ast\}\+\(1\-\\beta\)P^\{\\ast\},q∗\\displaystyle q^\{\\ast\}=P∗=∑i=1nwipi∗\\displaystyle=P^\{\\ast\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{\\ast\}q∗\\displaystyle q^\{\\ast\}=∑i=1nwi\(\(1−λi\)ri\+λiq∗\)\\displaystyle=\\sum\_\{i=1\}^\{n\}w\_\{i\}\\bigl\(\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q^\{\\ast\}\\bigr\)q∗\\displaystyle q^\{\\ast\}=∑i=1nwi\(1−λi\)ri\+\(∑i=1nwiλi\)q∗,\\displaystyle=\\sum\_\{i=1\}^\{n\}w\_\{i\}\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\left\(\\sum\_\{i=1\}^\{n\}w\_\{i\}\\lambda\_\{i\}\\right\)q^\{\\ast\},where the third line substitutes[AppendixB](https://arxiv.org/html/2607.27134#A2.Ex72)\. Sinceλi<1\\lambda\_\{i\}<1for everyii, and sincewi≥0w\_\{i\}\\geq 0with∑iwi=1\\sum\_\{i\}w\_\{i\}=1, we have∑i=1nwiλi<1\\sum\_\{i=1\}^\{n\}w\_\{i\}\\lambda\_\{i\}<1\. Therefore,
q∗\\displaystyle q^\{\\ast\}=∑i=1nwi\(1−λi\)ri1−∑i=1nwiλi\.\\displaystyle=\\frac\{\\sum\_\{i=1\}^\{n\}w\_\{i\}\(1\-\\lambda\_\{i\}\)\\,r\_\{i\}\}\{1\-\\sum\_\{i=1\}^\{n\}w\_\{i\}\\lambda\_\{i\}\}\.\(7\)Equations \([B](https://arxiv.org/html/2607.27134#A2.Ex72)\) and \([7](https://arxiv.org/html/2607.27134#A2.E7)\) give the claimed expressions forpi∗p\_\{i\}^\{\\ast\}andq∗q^\{\\ast\}at equilibrium\. It remains to show that the distributions converge to this equilibrium\. Define the deviations from equilibrium byxit:=pit−pi∗x\_\{i\}^\{t\}:=p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}andyt:=qt−q∗y^\{t\}:=q^\{t\}\-q^\{\\ast\}\. Using the update rule forpit\+1p\_\{i\}^\{t\+1\}and the equilibrium identity forpi∗p\_\{i\}^\{\\ast\},
xit\+1\\displaystyle x\_\{i\}^\{t\+1\}=pit\+1−pi∗=\(1−αi\)\(pit−pi∗\)\+αiλi\(qt−q∗\)\\displaystyle=p\_\{i\}^\{t\+1\}\-p\_\{i\}^\{\\ast\}=\(1\-\\alpha\_\{i\}\)\(p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}\)\+\\alpha\_\{i\}\\lambda\_\{i\}\(q^\{t\}\-q^\{\\ast\}\)=\(1−αi\)xit\+αiλiyt,\\displaystyle=\(1\-\\alpha\_\{i\}\)x\_\{i\}^\{t\}\+\\alpha\_\{i\}\\lambda\_\{i\}y^\{t\},‖xit\+1‖1\\displaystyle\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}≤\(1−αi\)‖xit‖1\+αiλi‖yt‖1\.\\displaystyle\\leq\(1\-\\alpha\_\{i\}\)\\,\\\|x\_\{i\}^\{t\}\\\|\_\{1\}\+\\alpha\_\{i\}\\,\\lambda\_\{i\}\\\|y^\{t\}\\\|\_\{1\}\.LetRt:=max\{maxi∈\[n\]‖xit‖1,‖yt‖1\}R^\{t\}:=\\max\\left\\\{\\max\_\{i\\in\[n\]\}\\\|x\_\{i\}^\{t\}\\\|\_\{1\},\\,\\\|y^\{t\}\\\|\_\{1\}\\right\\\}andμ:=mini∈\[n\]αi\(1−λi\)\\mu:=\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\\,\(1\-\\lambda\_\{i\}\)\. Sinceαi\>0\\alpha\_\{i\}\>0andλi<1\\lambda\_\{i\}<1for allii, we haveμ\>0\\mu\>0\. Then, for everyii,
‖xit\+1‖1≤\(1−αi\)Rt\+αiλiRt=\(1−αi\(1−λi\)\)Rt≤\(1−μ\)Rt\.\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}\\leq\(1\-\\alpha\_\{i\}\)R^\{t\}\+\\alpha\_\{i\}\\lambda\_\{i\}R^\{t\}=\\left\(1\-\\alpha\_\{i\}\(1\-\\lambda\_\{i\}\)\\right\)R^\{t\}\\leq\(1\-\\mu\)R^\{t\}\.Next, consider the model deviationyt\+1=qt\+1−q∗y^\{t\+1\}=q^\{t\+1\}\-q^\{\\ast\}\. Using the update ruleqt\+1=βqt\+\(1−β\)Pt\+1q^\{t\+1\}=\\beta q^\{t\}\+\(1\-\\beta\)P^\{t\+1\}and the equilibrium identityq∗=βq∗\+\(1−β\)P∗q^\{\\ast\}=\\beta q^\{\\ast\}\+\(1\-\\beta\)P^\{\\ast\}, we get
yt\+1=β\(qt−q∗\)\+\(1−β\)\(Pt\+1−P∗\)\.y^\{t\+1\}=\\beta\(q^\{t\}\-q^\{\\ast\}\)\+\(1\-\\beta\)\(P^\{t\+1\}\-P^\{\\ast\}\)\.UsingPt\+1=∑i=1nwipit\+1P^\{t\+1\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{t\+1\}andP∗=∑i=1nwipi∗P^\{\\ast\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{\\ast\}, we havePt\+1−P∗=∑i=1nwixit\+1\.P^\{t\+1\}\-P^\{\\ast\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}x\_\{i\}^\{t\+1\}\.Substituting back and using‖xit\+1‖1≤\(1−μ\)Rt\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}\\leq\(1\-\\mu\)R^\{t\}together with∑iwi=1\\sum\_\{i\}w\_\{i\}=1,
yt\+1\\displaystyle y^\{t\+1\}=βyt\+\(1−β\)∑i=1nwixit\+1\\displaystyle=\\beta y^\{t\}\+\(1\-\\beta\)\\sum\_\{i=1\}^\{n\}w\_\{i\}x\_\{i\}^\{t\+1\}‖yt\+1‖1\\displaystyle\\\|y^\{t\+1\}\\\|\_\{1\}≤β‖yt‖1\+\(1−β\)∑i=1nwi‖xit\+1‖1\\displaystyle\\leq\\beta\\\|y^\{t\}\\\|\_\{1\}\+\(1\-\\beta\)\\sum\_\{i=1\}^\{n\}w\_\{i\}\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}≤\(β\+\(1−β\)\(1−μ\)\)Rt=\(1−\(1−β\)μ\)Rt\.\\displaystyle\\leq\\bigl\(\\beta\+\(1\-\\beta\)\(1\-\\mu\)\\bigr\)R^\{t\}=\\bigl\(1\-\(1\-\\beta\)\\mu\\bigr\)R^\{t\}\.Letκ:=1−\(1−β\)μ\\kappa:=1\-\(1\-\\beta\)\\mu\. Sinceβ<1\\beta<1andμ\>0\\mu\>0, we haveκ<1\\kappa<1\. Alsoκ≥1−μ\\kappa\\geq 1\-\\mu, becauseκ=1−\(1−β\)μ≥1−μ\\kappa=1\-\(1\-\\beta\)\\mu\\geq 1\-\\mu\. Therefore, combining the bounds forxit\+1x\_\{i\}^\{t\+1\}andyt\+1y^\{t\+1\}, we getRt\+1≤κRtR^\{t\+1\}\\leq\\kappa R^\{t\}, and applying this recursively yieldsRt≤κtR0R^\{t\}\\leq\\kappa^\{t\}R^\{0\}\. Sinceκ<1\\kappa<1, it follows thatRt→0R^\{t\}\\to 0ast→∞t\\to\\infty\. Hence,pit→pi∗p\_\{i\}^\{t\}\\to p\_\{i\}^\{\\ast\}for everyii, andqt→q∗q^\{t\}\\to q^\{\\ast\}\. This establishes the convergence\.
Since Jensen–Shannon divergence is continuous on the probability simplex,JS\(pit,pjt\)→JS\(pi∗,pj∗\)JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\to JS\(p\_\{i\}^\{\\ast\},p\_\{j\}^\{\\ast\}\)for every pairi,ji,j, and averaging over all ordered pairsi≠ji\\neq jyields
Dt=1n\(n−1\)∑i≠jJS\(pit,pjt\)→1n\(n−1\)∑i≠jJS\(pi∗,pj∗\)=D∗\.D^\{t\}=\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\to\\frac\{1\}\{n\(n\-1\)\}\\sum\_\{i\\neq j\}JS\(p\_\{i\}^\{\\ast\},p\_\{j\}^\{\\ast\}\)=D^\{\\ast\}\.
Convergence time\.The convergence is exponential\. Since all initial and equilibrium quantities are probability distributions,R0≤2R^\{0\}\\leq 2, so
max\{maxi‖pit−pi∗‖1,‖qt−q∗‖1\}≤2κt\.\\max\\left\\\{\\max\_\{i\}\\\|p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}\\\|\_\{1\},\\\|q^\{t\}\-q^\{\\ast\}\\\|\_\{1\}\\right\\\}\\leq 2\\kappa^\{t\}\.Consequently, to guaranteemaxi‖pit−pi∗‖1≤ε\\max\_\{i\}\\\|p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}\\\|\_\{1\}\\leq\\varepsilonand‖qt−q∗‖1≤ε\\\|q^\{t\}\-q^\{\\ast\}\\\|\_\{1\}\\leq\\varepsilon, it is sufficient to take
t≥log\(2/ε\)\(1−β\)mini∈\[n\]αi\(1−λi\)\.t\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\(1\-\\beta\)\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\(1\-\\lambda\_\{i\}\)\}\.Thus the worst\-case convergence guarantee for the recursive\-update setting is weaker than for the fixed\-model setting, where the corresponding bound is controlled only byαmin\\alpha\_\{\\min\}; the bound need not be tight, and empirically the recursive dynamics can homogenize faster \([Section5](https://arxiv.org/html/2607.27134#S5)\)\. This completes the proof\. ∎
See[3\.3](https://arxiv.org/html/2607.27134#S3.SS3)
###### Proof of[Section3\.3](https://arxiv.org/html/2607.27134#S3.SS3)\.
First, we characterize the equilibrium, and then show that the dynamics converge to it\.
Characterizing the equilibrium\.At equilibrium, for every authorii, we havepit\+1=pit=pi∗p\_\{i\}^\{t\+1\}=p\_\{i\}^\{t\}=p\_\{i\}^\{\\ast\}andqit\+1=qit=qi∗q\_\{i\}^\{t\+1\}=q\_\{i\}^\{t\}=q\_\{i\}^\{\\ast\}\. The author update gives
pi∗\\displaystyle p\_\{i\}^\{\\ast\}=\(1−αi\)pi∗\+αi\(\(1−λi\)ri\+λiqi∗\)\\displaystyle=\(1\-\\alpha\_\{i\}\)p\_\{i\}^\{\\ast\}\+\\alpha\_\{i\}\\bigl\(\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q\_\{i\}^\{\\ast\}\\bigr\)pi∗\\displaystyle p\_\{i\}^\{\\ast\}=\(1−λi\)ri\+λiqi∗\\displaystyle=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}q\_\{i\}^\{\\ast\}sinceαi\>0\\alpha\_\{i\}\>0\.Next, the personalized model update gives, withP∗:=∑j=1nwjpj∗P^\{\\ast\}:=\\sum\_\{j=1\}^\{n\}w\_\{j\}p\_\{j\}^\{\\ast\},
qi∗\\displaystyle q\_\{i\}^\{\\ast\}=βqi∗\+γpi∗\+δP∗\\displaystyle=\\beta q\_\{i\}^\{\\ast\}\+\\gamma p\_\{i\}^\{\\ast\}\+\\delta P^\{\\ast\}\(1−β\)qi∗\\displaystyle\(1\-\\beta\)q\_\{i\}^\{\\ast\}=γpi∗\+δP∗\\displaystyle=\\gamma p\_\{i\}^\{\\ast\}\+\\delta P^\{\\ast\}qi∗\\displaystyle q\_\{i\}^\{\\ast\}=γ1−βpi∗\+δ1−βP∗\.\\displaystyle=\\frac\{\\gamma\}\{1\-\\beta\}p\_\{i\}^\{\\ast\}\+\\frac\{\\delta\}\{1\-\\beta\}P^\{\\ast\}\.Usingγ\+δ=1−β\\gamma\+\\delta=1\-\\betaandρ:=γ1−β\\rho:=\\frac\{\\gamma\}\{1\-\\beta\}, this becomes
qi∗=ρpi∗\+\(1−ρ\)P∗\.q\_\{i\}^\{\\ast\}=\\rho p\_\{i\}^\{\\ast\}\+\(1\-\\rho\)P^\{\\ast\}\.We now solve forpi∗p\_\{i\}^\{\\ast\}\. Substituting the expression forqi∗q\_\{i\}^\{\\ast\}into the equilibrium equation forpi∗p\_\{i\}^\{\\ast\},
pi∗\\displaystyle p\_\{i\}^\{\\ast\}=\(1−λi\)ri\+λi\(ρpi∗\+\(1−ρ\)P∗\)\\displaystyle=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}\\bigl\(\\rho p\_\{i\}^\{\\ast\}\+\(1\-\\rho\)P^\{\\ast\}\\bigr\)\(1−ρλi\)pi∗\\displaystyle\(1\-\\rho\\lambda\_\{i\}\)p\_\{i\}^\{\\ast\}=\(1−λi\)ri\+λi\(1−ρ\)P∗\\displaystyle=\(1\-\\lambda\_\{i\}\)r\_\{i\}\+\\lambda\_\{i\}\(1\-\\rho\)P^\{\\ast\}pi∗\\displaystyle p\_\{i\}^\{\\ast\}=1−λi1−ρλiri\+λi\(1−ρ\)1−ρλiP∗,\\displaystyle=\\frac\{1\-\\lambda\_\{i\}\}\{1\-\\rho\\lambda\_\{i\}\}r\_\{i\}\+\\frac\{\\lambda\_\{i\}\(1\-\\rho\)\}\{1\-\\rho\\lambda\_\{i\}\}P^\{\\ast\},where the division is valid sinceλi<1\\lambda\_\{i\}<1andρ∈\[0,1\]\\rho\\in\[0,1\]imply1−ρλi\>01\-\\rho\\lambda\_\{i\}\>0\. Definingηi:=1−λi1−ρλi\\eta\_\{i\}:=\\frac\{1\-\\lambda\_\{i\}\}\{1\-\\rho\\lambda\_\{i\}\}, we have1−ηi=λi\(1−ρ\)1−ρλi1\-\\eta\_\{i\}=\\frac\{\\lambda\_\{i\}\(1\-\\rho\)\}\{1\-\\rho\\lambda\_\{i\}\}, and therefore
pi∗=ηiri\+\(1−ηi\)P∗\.p\_\{i\}^\{\\ast\}=\\eta\_\{i\}r\_\{i\}\+\(1\-\\eta\_\{i\}\)P^\{\\ast\}\.It remains to identifyP∗P^\{\\ast\}\. By definition,P∗=∑i=1nwipi∗P^\{\\ast\}=\\sum\_\{i=1\}^\{n\}w\_\{i\}p\_\{i\}^\{\\ast\}\. Substituting the expression forpi∗p\_\{i\}^\{\\ast\}and using∑iwi=1\\sum\_\{i\}w\_\{i\}=1,
P∗\\displaystyle P^\{\\ast\}=∑i=1nwiηiri\+\(1−∑i=1nwiηi\)P∗\\displaystyle=\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}r\_\{i\}\+\\left\(1\-\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}\\right\)P^\{\\ast\}\(∑i=1nwiηi\)P∗\\displaystyle\\left\(\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}\\right\)P^\{\\ast\}=∑i=1nwiηiri\\displaystyle=\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}r\_\{i\}P∗\\displaystyle P^\{\\ast\}=∑i=1nwiηiri∑i=1nwiηi,\\displaystyle=\\frac\{\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}r\_\{i\}\}\{\\sum\_\{i=1\}^\{n\}w\_\{i\}\\eta\_\{i\}\},where the division is valid sinceλi<1\\lambda\_\{i\}<1impliesηi\>0\\eta\_\{i\}\>0, hence∑iwiηi\>0\\sum\_\{i\}w\_\{i\}\\eta\_\{i\}\>0\. Finally, substitutingpi∗=ηiri\+\(1−ηi\)P∗p\_\{i\}^\{\\ast\}=\\eta\_\{i\}r\_\{i\}\+\(1\-\\eta\_\{i\}\)P^\{\\ast\}intoqi∗=ρpi∗\+\(1−ρ\)P∗q\_\{i\}^\{\\ast\}=\\rho p\_\{i\}^\{\\ast\}\+\(1\-\\rho\)P^\{\\ast\},
qi∗\\displaystyle q\_\{i\}^\{\\ast\}=ρηiri\+ρ\(1−ηi\)P∗\+\(1−ρ\)P∗=ρηiri\+\(1−ρηi\)P∗\.\\displaystyle=\\rho\\eta\_\{i\}r\_\{i\}\+\\rho\(1\-\\eta\_\{i\}\)P^\{\\ast\}\+\(1\-\\rho\)P^\{\\ast\}=\\rho\\eta\_\{i\}r\_\{i\}\+\(1\-\\rho\\eta\_\{i\}\)P^\{\\ast\}\.This proves the claimed form of the equilibrium\.
Convergence to equilibrium\.Define the deviationsxit:=pit−pi∗x\_\{i\}^\{t\}:=p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}andyit:=qit−qi∗y\_\{i\}^\{t\}:=q\_\{i\}^\{t\}\-q\_\{i\}^\{\\ast\}\. Using the author update and the equilibrium identity,
xit\+1\\displaystyle x\_\{i\}^\{t\+1\}=pit\+1−pi∗=\(1−αi\)\(pit−pi∗\)\+αiλi\(qit−qi∗\)\\displaystyle=p\_\{i\}^\{t\+1\}\-p\_\{i\}^\{\\ast\}=\(1\-\\alpha\_\{i\}\)\(p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}\)\+\\alpha\_\{i\}\\lambda\_\{i\}\(q\_\{i\}^\{t\}\-q\_\{i\}^\{\\ast\}\)=\(1−αi\)xit\+αiλiyit,\\displaystyle=\(1\-\\alpha\_\{i\}\)x\_\{i\}^\{t\}\+\\alpha\_\{i\}\\lambda\_\{i\}y\_\{i\}^\{t\},‖xit\+1‖1\\displaystyle\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}≤\(1−αi\)‖xit‖1\+αiλi‖yit‖1\.\\displaystyle\\leq\(1\-\\alpha\_\{i\}\)\\\|x\_\{i\}^\{t\}\\\|\_\{1\}\+\\alpha\_\{i\}\\lambda\_\{i\}\\\|y\_\{i\}^\{t\}\\\|\_\{1\}\.DefineRt:=max\{maxi‖xit‖1,maxi‖yit‖1\}R^\{t\}:=\\max\\left\\\{\\max\_\{i\}\\\|x\_\{i\}^\{t\}\\\|\_\{1\},\\,\\max\_\{i\}\\\|y\_\{i\}^\{t\}\\\|\_\{1\}\\right\\\}andμ:=mini∈\[n\]αi\(1−λi\)\\mu:=\\min\_\{i\\in\[n\]\}\\alpha\_\{i\}\(1\-\\lambda\_\{i\}\)\. Sinceαi\>0\\alpha\_\{i\}\>0andλi<1\\lambda\_\{i\}<1for everyii, we haveμ\>0\\mu\>0\. Then
‖xit\+1‖1≤\(1−αi\(1−λi\)\)Rt≤\(1−μ\)Rt\.\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}\\leq\\bigl\(1\-\\alpha\_\{i\}\(1\-\\lambda\_\{i\}\)\\bigr\)R^\{t\}\\leq\(1\-\\mu\)R^\{t\}\.Next, consider the model deviations\. From the model update and the equilibrium identity,
yit\+1\\displaystyle y\_\{i\}^\{t\+1\}=βyit\+γxit\+1\+δ∑j=1nwjxjt\+1\\displaystyle=\\beta y\_\{i\}^\{t\}\+\\gamma x\_\{i\}^\{t\+1\}\+\\delta\\sum\_\{j=1\}^\{n\}w\_\{j\}x\_\{j\}^\{t\+1\}‖yit\+1‖1\\displaystyle\\\|y\_\{i\}^\{t\+1\}\\\|\_\{1\}≤β‖yit‖1\+γ‖xit\+1‖1\+δ∑j=1nwj‖xjt\+1‖1\\displaystyle\\leq\\beta\\\|y\_\{i\}^\{t\}\\\|\_\{1\}\+\\gamma\\\|x\_\{i\}^\{t\+1\}\\\|\_\{1\}\+\\delta\\sum\_\{j=1\}^\{n\}w\_\{j\}\\\|x\_\{j\}^\{t\+1\}\\\|\_\{1\}≤\(β\+\(γ\+δ\)\(1−μ\)\)Rt\\displaystyle\\leq\\left\(\\beta\+\(\\gamma\+\\delta\)\(1\-\\mu\)\\right\)R^\{t\}=\(β\+\(1−β\)\(1−μ\)\)Rt=\(1−\(1−β\)μ\)Rt,\\displaystyle=\\left\(\\beta\+\(1\-\\beta\)\(1\-\\mu\)\\right\)R^\{t\}=\\left\(1\-\(1\-\\beta\)\\mu\\right\)R^\{t\},using‖xjt\+1‖1≤\(1−μ\)Rt\\\|x\_\{j\}^\{t\+1\}\\\|\_\{1\}\\leq\(1\-\\mu\)R^\{t\},∑jwj=1\\sum\_\{j\}w\_\{j\}=1, andγ\+δ=1−β\\gamma\+\\delta=1\-\\beta\. Letκ:=1−\(1−β\)μ\\kappa:=1\-\(1\-\\beta\)\\mu\. Sinceβ<1\\beta<1andμ\>0\\mu\>0, we haveκ<1\\kappa<1; moreoverκ≥1−μ\\kappa\\geq 1\-\\mu\. Therefore, the bounds for bothxit\+1x\_\{i\}^\{t\+1\}andyit\+1y\_\{i\}^\{t\+1\}implyRt\+1≤κRtR^\{t\+1\}\\leq\\kappa R^\{t\}, and recursivelyRt≤κtR0R^\{t\}\\leq\\kappa^\{t\}R^\{0\}\. Sinceκ<1\\kappa<1, we haveRt→0R^\{t\}\\to 0, sopit→pi∗p\_\{i\}^\{t\}\\to p\_\{i\}^\{\\ast\}andqit→qi∗q\_\{i\}^\{t\}\\to q\_\{i\}^\{\\ast\}for every authorii, establishing convergence\.
Convergence time\.Since allpi0,pi∗,qi0,qi∗p\_\{i\}^\{0\},p\_\{i\}^\{\\ast\},q\_\{i\}^\{0\},q\_\{i\}^\{\\ast\}are probability distributions, theirℓ1\\ell\_\{1\}\-distances are at most22\. HenceR0≤2R^\{0\}\\leq 2andRt≤2κtR^\{t\}\\leq 2\\kappa^\{t\}\. To guaranteemaxi‖pit−pi∗‖1≤ε\\max\_\{i\}\\\|p\_\{i\}^\{t\}\-p\_\{i\}^\{\\ast\}\\\|\_\{1\}\\leq\\varepsilonandmaxi‖qit−qi∗‖1≤ε\\max\_\{i\}\\\|q\_\{i\}^\{t\}\-q\_\{i\}^\{\\ast\}\\\|\_\{1\}\\leq\\varepsilon, it is sufficient that2κt≤ε2\\kappa^\{t\}\\leq\\varepsilon, equivalentlyt≥log\(2/ε\)−logκt\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\-\\log\\kappa\}\. Sinceκ=1−\(1−β\)μ\\kappa=1\-\(1\-\\beta\)\\muand−log\(1−z\)≥z\-\\log\(1\-z\)\\geq zforz∈\(0,1\)z\\in\(0,1\), it is sufficient to take
t≥log\(2/ε\)\(1−β\)miniαi\(1−λi\)\.t\\geq\\frac\{\\log\(2/\\varepsilon\)\}\{\(1\-\\beta\)\\min\_\{i\}\\alpha\_\{i\}\(1\-\\lambda\_\{i\}\)\}\.Finally, since Jensen–Shannon divergence is continuous on the probability simplex, for every pairi,ji,j,
JS\(pit,pjt\)→JS\(pi∗,pj∗\),JS\(qit,qjt\)→JS\(qi∗,qj∗\)\.JS\(p\_\{i\}^\{t\},p\_\{j\}^\{t\}\)\\to JS\(p\_\{i\}^\{\\ast\},p\_\{j\}^\{\\ast\}\),\\qquad JS\(q\_\{i\}^\{t\},q\_\{j\}^\{t\}\)\\to JS\(q\_\{i\}^\{\\ast\},q\_\{j\}^\{\\ast\}\)\.Averaging over all ordered pairsi≠ji\\neq j, we obtainDt→D∗D^\{t\}\\to D^\{\\ast\}andQt→Q∗Q^\{t\}\\to Q^\{\\ast\}\. This completes the proof\. ∎
See[3\.3](https://arxiv.org/html/2607.27134#S3.SS3)
###### Proof\.
For IM 1,[Section3\.1](https://arxiv.org/html/2607.27134#S3.SS1)givesai\(λ\)=\(1−λ\)ri\+λq0a\_\{i\}\(\\lambda\)=\(1\-\\lambda\)r\_\{i\}\+\\lambda q^\{0\}, henceai\(λ\)−aj\(λ\)=\(1−λ\)\(ri−rj\)a\_\{i\}\(\\lambda\)\-a\_\{j\}\(\\lambda\)=\(1\-\\lambda\)\(r\_\{i\}\-r\_\{j\}\)\. Under[Section3\.2](https://arxiv.org/html/2607.27134#S3.SS2),pi∗=\(1−λ\)ri\+λq∗p\_\{i\}^\{\\ast\}=\(1\-\\lambda\)r\_\{i\}\+\\lambda q^\{\\ast\}, giving the same difference for IM 2\. Under[Section3\.3](https://arxiv.org/html/2607.27134#S3.SS3),ηi=η\\eta\_\{i\}=\\etafor everyii, henceP∗=∑iwiriP^\{\\ast\}=\\sum\_\{i\}w\_\{i\}r\_\{i\}andpi∗=ηri\+\(1−η\)P∗p\_\{i\}^\{\\ast\}=\\eta r\_\{i\}\+\(1\-\\eta\)P^\{\\ast\}, sopi∗−pj∗=η\(ri−rj\)p\_\{i\}^\{\\ast\}\-p\_\{j\}^\{\\ast\}=\\eta\(r\_\{i\}\-r\_\{j\}\)\. SinceD^\\widehat\{D\}is homogeneous of degree two in pairwise differences andη/\(1−λ\)=1/\(1−ρλ\)\\eta/\(1\-\\lambda\)=1/\(1\-\\rho\\lambda\), the diversity relation follows\. ∎
## Appendix COmitted Proofs from Section[4](https://arxiv.org/html/2607.27134#S4)
See[4\.1](https://arxiv.org/html/2607.27134#S4.SS1)
###### Proof\.
From \([4](https://arxiv.org/html/2607.27134#S4.E4)\),λi\\lambda\_\{i\}entersUjU\_\{j\}\(j≠ij\\neq i\) only through the termθj2\(n−1\)σi2di2\\frac\{\\theta\_\{j\}\}\{2\(n\-1\)\}\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\. Differentiating withσi=1−λi\\sigma\_\{i\}=1\-\\lambda\_\{i\}gives∂Uj/∂λi=−θjn−1σidi2\\partial U\_\{j\}/\\partial\\lambda\_\{i\}=\-\\frac\{\\theta\_\{j\}\}\{n\-1\}\\sigma\_\{i\}d\_\{i\}^\{2\}\. ∎
See[4\.2](https://arxiv.org/html/2607.27134#S4.Thmtheorem2)
###### Proof\.
*\(i\)*By \([4](https://arxiv.org/html/2607.27134#S4.E4)\),UiU\_\{i\}is additively separable: the only term involvingλi\\lambda\_\{i\}is
hi\(σi\):=di22\[\(θi−bi\)σi2−ci\(1−σi\)2\],h\_\{i\}\(\\sigma\_\{i\}\):=\\frac\{d\_\{i\}^\{2\}\}\{2\}\\Big\[\(\\theta\_\{i\}\-b\_\{i\}\)\\sigma\_\{i\}^\{2\}\-c\_\{i\}\(1\-\\sigma\_\{i\}\)^\{2\}\\Big\],which is independent ofλ−i\\lambda\_\{\-i\}\. Hence the maximizer ofhih\_\{i\}overσi∈\[0,1\]\\sigma\_\{i\}\\in\[0,1\]is a dominant strategy\. We have
hi′\(σi\)=di2\[\(θi−bi\)σi\+ci\(1−σi\)\]=di2\[ci−\(bi\+ci−θi\)σi\]\.h\_\{i\}^\{\\prime\}\(\\sigma\_\{i\}\)=d\_\{i\}^\{2\}\\Big\[\(\\theta\_\{i\}\-b\_\{i\}\)\\sigma\_\{i\}\+c\_\{i\}\(1\-\\sigma\_\{i\}\)\\Big\]=d\_\{i\}^\{2\}\\Big\[c\_\{i\}\-\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\)\\sigma\_\{i\}\\Big\]\.*Casebi\>θib\_\{i\}\>\\theta\_\{i\}\.*Thenbi\+ci−θi\>ci\>0b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\>c\_\{i\}\>0,hih\_\{i\}is strictly concave, and the unconstrained maximizerσi∗=ci/\(bi\+ci−θi\)∈\(0,1\)\\sigma\_\{i\}^\{\\ast\}=c\_\{i\}/\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\)\\in\(0,1\)is interior, givingλiNE=1−σi∗=\(bi−θi\)/\(bi\+ci−θi\)\\lambda\_\{i\}^\{\\mathrm\{NE\}\}=1\-\\sigma\_\{i\}^\{\\ast\}=\(b\_\{i\}\-\\theta\_\{i\}\)/\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\)\.*Casebi≤θib\_\{i\}\\leq\\theta\_\{i\}\.*Thenhi′\(σi\)=di2\[ci\(1−σi\)\+\(θi−bi\)σi\]≥0h\_\{i\}^\{\\prime\}\(\\sigma\_\{i\}\)=d\_\{i\}^\{2\}\\big\[c\_\{i\}\(1\-\\sigma\_\{i\}\)\+\(\\theta\_\{i\}\-b\_\{i\}\)\\sigma\_\{i\}\\big\]\\geq 0on\[0,1\]\[0,1\], strictly positive on\[0,1\)\[0,1\), sohih\_\{i\}is maximized atσi=1\\sigma\_\{i\}=1, i\.e\.λiNE=0\\lambda\_\{i\}^\{\\mathrm\{NE\}\}=0\. Both cases match the stated formula, and uniqueness of the maximizer in each case gives uniqueness of the equilibrium\.
*\(ii\)*Summing \([4](https://arxiv.org/html/2607.27134#S4.E4)\) overiiand collecting, for eachii, the coefficient ofσi2di2/2\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}/2—namely\(θi−bi\)\(\\theta\_\{i\}\-b\_\{i\}\)fromUiU\_\{i\}andθjn−1\\frac\{\\theta\_\{j\}\}\{n\-1\}from eachUjU\_\{j\}withj≠ij\\neq i, the latter summing toθ¯−i\\bar\{\\theta\}\_\{\-i\}—we obtain
W\(λ\)=∑i=1ndi22\[\(θi\+θ¯−i−bi\)σi2−ci\(1−σi\)2\]\.W\(\\lambda\)\\;=\\;\\sum\_\{i=1\}^\{n\}\\frac\{d\_\{i\}^\{2\}\}\{2\}\\Big\[\\big\(\\theta\_\{i\}\+\\bar\{\\theta\}\_\{\-i\}\-b\_\{i\}\\big\)\\,\\sigma\_\{i\}^\{2\}\\;\-\\;c\_\{i\}\\,\(1\-\\sigma\_\{i\}\)^\{2\}\\Big\]\.This is again additively separable across authors, and each summand has exactly the formhih\_\{i\}withθi\\theta\_\{i\}replaced byθi\+θ¯−i\\theta\_\{i\}\+\\bar\{\\theta\}\_\{\-i\}\. The argument of part \(i\) applied verbatim with this replacement yields the statedλiSO\\lambda\_\{i\}^\{\\mathrm\{SO\}\}and its uniqueness\.
*\(iii\)*Define, forx≥0x\\geq 0,
φi\(x\):=\(bi−θi−x\)\+\(bi−θi−x\)\+\+ci,\\varphi\_\{i\}\(x\):=\\frac\{\(b\_\{i\}\-\\theta\_\{i\}\-x\)\_\{\+\}\}\{\\,\(b\_\{i\}\-\\theta\_\{i\}\-x\)\_\{\+\}\+c\_\{i\}\\,\},so thatλiNE=φi\(0\)\\lambda\_\{i\}^\{\\mathrm\{NE\}\}=\\varphi\_\{i\}\(0\)andλiSO=φi\(θ¯−i\)\\lambda\_\{i\}^\{\\mathrm\{SO\}\}=\\varphi\_\{i\}\(\\bar\{\\theta\}\_\{\-i\}\)\. The mapy↦y/\(y\+ci\)y\\mapsto y/\(y\+c\_\{i\}\)is strictly increasing on\[0,∞\)\[0,\\infty\)andx↦\(bi−θi−x\)\+x\\mapsto\(b\_\{i\}\-\\theta\_\{i\}\-x\)\_\{\+\}is nonincreasing, strictly decreasing while positive; henceφi\\varphi\_\{i\}is nonincreasing, strictly decreasing while positive\. This givesλiSO≤λiNE\\lambda\_\{i\}^\{\\mathrm\{SO\}\}\\leq\\lambda\_\{i\}^\{\\mathrm\{NE\}\}with strictness exactly whenφi\(0\)\>0\\varphi\_\{i\}\(0\)\>0andθ¯−i\>0\\bar\{\\theta\}\_\{\-i\}\>0\. The diversity comparison follows from \([5](https://arxiv.org/html/2607.27134#S4.E5)\), sinceσiNE≤σiSO\\sigma\_\{i\}^\{\\mathrm\{NE\}\}\\leq\\sigma\_\{i\}^\{\\mathrm\{SO\}\}for everyii; the welfare comparison holds becauseλSO\\lambda^\{\\mathrm\{SO\}\}maximizesWW\. ∎
See[4\.2](https://arxiv.org/html/2607.27134#S4.SS2)
###### Proof\.
Parts \(i\) and \(ii\) and the equilibrium expressions in \(iii\) follow from[Theorem4\.2](https://arxiv.org/html/2607.27134#S4.Thmtheorem2)withθ¯−i=θ\\bar\{\\theta\}\_\{\-i\}=\\theta; thePoM\\mathrm\{PoM\}expressions follow from \([5](https://arxiv.org/html/2607.27134#S4.E5)\), since in the symmetric caseD^∞=σ2d2\\widehat\{D\}\_\{\\infty\}=\\sigma^\{2\}d^\{2\}andσNE=cb\+c−θ\\sigma^\{\\mathrm\{NE\}\}=\\frac\{c\}\{b\+c\-\\theta\}, whileσSO=1\\sigma^\{\\mathrm\{SO\}\}=1in regime \(ii\) andσSO=cb\+c−2θ\\sigma^\{\\mathrm\{SO\}\}=\\frac\{c\}\{b\+c\-2\\theta\}in regime \(iii\)\.
For the welfare loss in \(iii\), writeA:=b\+c−2θ\>0A:=b\+c\-2\\theta\>0\. From the proof of[Theorem4\.2](https://arxiv.org/html/2607.27134#S4.Thmtheorem2)\(ii\), per\-author welfare at a symmetric profileσ\\sigmaisd22w\(σ\)\\frac\{d^\{2\}\}\{2\}\\,w\(\\sigma\)with
w\(σ\)=\(2θ−b\)σ2−c\(1−σ\)2=−Aσ2\+2cσ−c\.w\(\\sigma\)=\(2\\theta\-b\)\\sigma^\{2\}\-c\(1\-\\sigma\)^\{2\}=\-A\\sigma^\{2\}\+2c\\sigma\-c\.Thenw\(σSO\)=w\(c/A\)=c2/A−cw\(\\sigma^\{\\mathrm\{SO\}\}\)=w\(c/A\)=c^\{2\}/A\-c, while withσNE=c/\(A\+θ\)\\sigma^\{\\mathrm\{NE\}\}=c/\(A\+\\theta\),
w\(σNE\)=−Ac2\(A\+θ\)2\+2c2A\+θ−c=c2\(A\+2θ\)\(A\+θ\)2−c\.w\(\\sigma^\{\\mathrm\{NE\}\}\)=\-\\frac\{Ac^\{2\}\}\{\(A\+\\theta\)^\{2\}\}\+\\frac\{2c^\{2\}\}\{A\+\\theta\}\-c=\\frac\{c^\{2\}\(A\+2\\theta\)\}\{\(A\+\\theta\)^\{2\}\}\-c\.Subtracting,
w\(σSO\)−w\(σNE\)=c2⋅\(A\+θ\)2−A\(A\+2θ\)A\(A\+θ\)2=c2θ2A\(A\+θ\)2,w\(\\sigma^\{\\mathrm\{SO\}\}\)\-w\(\\sigma^\{\\mathrm\{NE\}\}\)=c^\{2\}\\cdot\\frac\{\(A\+\\theta\)^\{2\}\-A\(A\+2\\theta\)\}\{A\(A\+\\theta\)^\{2\}\}=\\frac\{c^\{2\}\\,\\theta^\{2\}\}\{A\(A\+\\theta\)^\{2\}\},since\(A\+θ\)2−A\(A\+2θ\)=θ2\(A\+\\theta\)^\{2\}\-A\(A\+2\\theta\)=\\theta^\{2\}\. SubstitutingA=b\+c−2θA=b\+c\-2\\thetaandA\+θ=b\+c−θA\+\\theta=b\+c\-\\thetacompletes the proof\. ∎
## Appendix DCorrelated Author Signatures
We now relax the orthogonality assumption in Definition 1 and characterize the strategic\-conformity game for arbitrary author signatures\. Recall that
ui:=ri−q0,di2:=∥ui∥22,σi:=1−λi,u\_\{i\}:=r\_\{i\}\-q^\{0\},\\qquad d\_\{i\}^\{2\}:=\\lVert u\_\{i\}\\rVert\_\{2\}^\{2\},\\qquad\\sigma\_\{i\}:=1\-\\lambda\_\{i\},so that the limiting style of authoriisatisfies
ai\(λi\)−q0=σiui\.a\_\{i\}\(\\lambda\_\{i\}\)\-q^\{0\}=\\sigma\_\{i\}u\_\{i\}\.Let
gij:=⟨ui,uj⟩g\_\{ij\}:=\\langle u\_\{i\},u\_\{j\}\\rangleand letG=\(gij\)i,j∈\[n\]G=\(g\_\{ij\}\)\_\{i,j\\in\[n\]\}denote the Gram matrix of the author signatures\. Thus,gii=di2g\_\{ii\}=d\_\{i\}^\{2\}, whilegijg\_\{ij\}measures the alignment between the directions in which authorsiiandjjdiffer from the shared model\-induced norm\. Orthogonal signatures correspond togij=0g\_\{ij\}=0for everyi≠ji\\neq j\.
###### Proposition D\.1\(Strategic conformity with correlated signatures\)\.
Consider the game with payoffs in Equation \(3\), without imposing orthogonality\.
1. 1\.For every authorii, Ui\(σ\)=\\displaystyle U\_\{i\}\(\\sigma\)=\{\}di22\[\(θi−bi\)σi2−ci\(1−σi\)2\]\\displaystyle\\frac\{d\_\{i\}^\{2\}\}\{2\}\\left\[\(\\theta\_\{i\}\-b\_\{i\}\)\\sigma\_\{i\}^\{2\}\-c\_\{i\}\(1\-\\sigma\_\{i\}\)^\{2\}\\right\]\+θi2\(n−1\)∑j≠idj2σj2−θin−1σi∑j≠igijσj\.\\displaystyle\+\\frac\{\\theta\_\{i\}\}\{2\(n\-1\)\}\\sum\_\{j\\neq i\}d\_\{j\}^\{2\}\\sigma\_\{j\}^\{2\}\-\\frac\{\\theta\_\{i\}\}\{n\-1\}\\sigma\_\{i\}\\sum\_\{j\\neq i\}g\_\{ij\}\\sigma\_\{j\}\.\(8\)The corresponding quadratic long\-run diversity is D∞\(λ\)=1n∑i=1ndi2σi2−2n\(n−1\)∑i<jgijσiσj\.D^\{\\infty\}\(\\lambda\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}d\_\{i\}^\{2\}\\sigma\_\{i\}^\{2\}\-\\frac\{2\}\{n\(n\-1\)\}\\sum\_\{i<j\}g\_\{ij\}\\sigma\_\{i\}\\sigma\_\{j\}\.\(9\)
2. 2\.For every pairi≠ji\\neq j, ∂Uj∂λi=θjn−1\(σjgij−σidi2\)\.\\frac\{\\partial U\_\{j\}\}\{\\partial\\lambda\_\{i\}\}=\\frac\{\\theta\_\{j\}\}\{n\-1\}\\left\(\\sigma\_\{j\}g\_\{ij\}\-\\sigma\_\{i\}d\_\{i\}^\{2\}\\right\)\.\(10\)Consequently, an increase in authorii’s conformity imposes a nonpositive externality on authorjjif and only if σidi2≥σjgij\.\\sigma\_\{i\}d\_\{i\}^\{2\}\\geq\\sigma\_\{j\}g\_\{ij\}\.\(11\)In particular, the externality is nonpositive at every conformity profile whenevergij≤0g\_\{ij\}\\leq 0\.
3. 3\.Define Ai:=bi\+ci−θi\.A\_\{i\}:=b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\.IfAi\>0A\_\{i\}\>0, authorii’s payoff is strictly concave inσi\\sigma\_\{i\}, conditional onσ−i\\sigma\_\{\-i\}, and their unique best response is BRi\(σ−i\)=Π\[0,1\]\[ci−θi\(n−1\)di2∑j≠igijσjAi\],\\operatorname\{BR\}\_\{i\}\(\\sigma\_\{\-i\}\)=\\Pi\_\{\[0,1\]\}\\left\[\\frac\{c\_\{i\}\-\\dfrac\{\\theta\_\{i\}\}\{\(n\-1\)d\_\{i\}^\{2\}\}\\sum\_\{j\\neq i\}g\_\{ij\}\\sigma\_\{j\}\}\{A\_\{i\}\}\\right\],\(12\)whereΠ\[0,1\]\\Pi\_\{\[0,1\]\}denotes projection onto\[0,1\]\[0,1\]\. If, in addition, maxi∈\[n\]θi\(n−1\)di2Ai∑j≠i\|gij\|<1,\\max\_\{i\\in\[n\]\}\\frac\{\\theta\_\{i\}\}\{\(n\-1\)d\_\{i\}^\{2\}A\_\{i\}\}\\sum\_\{j\\neq i\}\|g\_\{ij\}\|<1,\(13\)then the joint best\-response map is a contraction and the game has a unique Nash equilibrium\. If the Nash equilibrium is interior, its retained\-distinctiveness vector satisfies NσNE=h,N\\sigma^\{\\mathrm\{NE\}\}=h,\(14\)where hi:=cidi2h\_\{i\}:=c\_\{i\}d\_\{i\}^\{2\}and Nii=di2\(bi\+ci−θi\),Nij=θin−1gij,i≠j\.N\_\{ii\}=d\_\{i\}^\{2\}\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\),\\qquad N\_\{ij\}=\\frac\{\\theta\_\{i\}\}\{n\-1\}g\_\{ij\},\\quad i\\neq j\.\(15\)
4. 4\.Let θ¯−i:=1n−1∑j≠iθj\.\\bar\{\\theta\}\_\{\-i\}:=\\frac\{1\}\{n\-1\}\\sum\_\{j\\neq i\}\\theta\_\{j\}\.Utilitarian welfare can be written as W\(σ\)=C\+h⊤σ−12σ⊤Sσ,W\(\\sigma\)=C\+h^\{\\top\}\\sigma\-\\frac\{1\}\{2\}\\sigma^\{\\top\}S\\sigma,\(16\)whereCCis independent ofσ\\sigma,hi=cidi2h\_\{i\}=c\_\{i\}d\_\{i\}^\{2\}, and Sii=di2\(bi\+ci−θi−θ¯−i\),S\_\{ii\}=d\_\{i\}^\{2\}\\left\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\-\\bar\{\\theta\}\_\{\-i\}\\right\),\(17\)Sij=θi\+θjn−1gij,i≠j\.S\_\{ij\}=\\frac\{\\theta\_\{i\}\+\\theta\_\{j\}\}\{n\-1\}g\_\{ij\},\\qquad i\\neq j\.\(18\)IfSSis positive definite, welfare is strictly concave and has a unique maximizer on\[0,1\]n\[0,1\]^\{n\}\. If the maximizer is interior, it is characterized by SσSO=h\.S\\sigma^\{\\mathrm\{SO\}\}=h\.\(19\)
###### Proof\.
For anyi≠ji\\neq j,
∥ai−aj∥22\\displaystyle\\lVert a\_\{i\}\-a\_\{j\}\\rVert\_\{2\}^\{2\}=∥σiui−σjuj∥22\\displaystyle=\\lVert\\sigma\_\{i\}u\_\{i\}\-\\sigma\_\{j\}u\_\{j\}\\rVert\_\{2\}^\{2\}=σi2di2\+σj2dj2−2σiσjgij\.\\displaystyle=\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\+\\sigma\_\{j\}^\{2\}d\_\{j\}^\{2\}\-2\\sigma\_\{i\}\\sigma\_\{j\}g\_\{ij\}\.\(20\)Substituting Equation \([20](https://arxiv.org/html/2607.27134#A4.E20)\) into Equation \(3\), the terms involvingσi2di2\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}appear once for everyj≠ij\\neq i, giving
θi2\(n−1\)∑j≠iσi2di2=θi2σi2di2\.\\frac\{\\theta\_\{i\}\}\{2\(n\-1\)\}\\sum\_\{j\\neq i\}\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}=\\frac\{\\theta\_\{i\}\}\{2\}\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\.Collecting this term with the legibility and authenticity terms gives the first line of Equation \([8](https://arxiv.org/html/2607.27134#A4.E8)\)\. The remaining squared\-signature and cross terms give the second line\.
For quadratic diversity, summing Equation \([20](https://arxiv.org/html/2607.27134#A4.E20)\) over unordered pairs gives
∑i<j\(σi2di2\+σj2dj2\)=\(n−1\)∑iσi2di2\.\\sum\_\{i<j\}\\left\(\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\+\\sigma\_\{j\}^\{2\}d\_\{j\}^\{2\}\\right\)=\(n\-1\)\\sum\_\{i\}\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\.Substitution into the definition of quadratic long\-run diversity gives Equation \([9](https://arxiv.org/html/2607.27134#A4.E9)\)\.
To obtain the externality formula, consideri≠ji\\neq j\. The only part ofUjU\_\{j\}that depends onλi\\lambda\_\{i\}is the pairwise term involving authorsiiandjj\. Since
∂σi∂λi=−1,\\frac\{\\partial\\sigma\_\{i\}\}\{\\partial\\lambda\_\{i\}\}=\-1,we have
∂Uj∂λi\\displaystyle\\frac\{\\partial U\_\{j\}\}\{\\partial\\lambda\_\{i\}\}=θj2\(n−1\)∂∂λi∥σjuj−σiui∥22\\displaystyle=\\frac\{\\theta\_\{j\}\}\{2\(n\-1\)\}\\frac\{\\partial\}\{\\partial\\lambda\_\{i\}\}\\lVert\\sigma\_\{j\}u\_\{j\}\-\\sigma\_\{i\}u\_\{i\}\\rVert\_\{2\}^\{2\}=θjn−1⟨σjuj−σiui,ui⟩\\displaystyle=\\frac\{\\theta\_\{j\}\}\{n\-1\}\\left\\langle\\sigma\_\{j\}u\_\{j\}\-\\sigma\_\{i\}u\_\{i\},u\_\{i\}\\right\\rangle=θjn−1\(σjgij−σidi2\)\.\\displaystyle=\\frac\{\\theta\_\{j\}\}\{n\-1\}\\left\(\\sigma\_\{j\}g\_\{ij\}\-\\sigma\_\{i\}d\_\{i\}^\{2\}\\right\)\.\(21\)This proves Equations \([10](https://arxiv.org/html/2607.27134#A4.E10)\) and \([11](https://arxiv.org/html/2607.27134#A4.E11)\)\. Ifgij≤0g\_\{ij\}\\leq 0, bothσjgij\\sigma\_\{j\}g\_\{ij\}and−σidi2\-\\sigma\_\{i\}d\_\{i\}^\{2\}are nonpositive, so the externality is nonpositive\.
Differentiating Equation \([8](https://arxiv.org/html/2607.27134#A4.E8)\) with respect toσi\\sigma\_\{i\}gives
∂Ui∂σi=di2\[ci−\(bi\+ci−θi\)σi\]−θin−1∑j≠igijσj\.\\frac\{\\partial U\_\{i\}\}\{\\partial\\sigma\_\{i\}\}=d\_\{i\}^\{2\}\\left\[c\_\{i\}\-\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\)\\sigma\_\{i\}\\right\]\-\\frac\{\\theta\_\{i\}\}\{n\-1\}\\sum\_\{j\\neq i\}g\_\{ij\}\\sigma\_\{j\}\.\(22\)Moreover,
∂2Ui∂σi2=−di2Ai\.\\frac\{\\partial^\{2\}U\_\{i\}\}\{\\partial\\sigma\_\{i\}^\{2\}\}=\-d\_\{i\}^\{2\}A\_\{i\}\.Thus,Ai\>0A\_\{i\}\>0implies strict concavity inσi\\sigma\_\{i\}\. Solving Equation \([22](https://arxiv.org/html/2607.27134#A4.E22)\) and projecting the resulting unconstrained maximizer onto the feasible interval gives Equation \([12](https://arxiv.org/html/2607.27134#A4.E12)\)\.
Projection onto a closed interval is nonexpansive\. Hence, for any two profilesσ\\sigmaandσ′\\sigma^\{\\prime\},
\|BRi\(σ−i\)−BRi\(σ−i′\)\|\\displaystyle\\left\|\\operatorname\{BR\}\_\{i\}\(\\sigma\_\{\-i\}\)\-\\operatorname\{BR\}\_\{i\}\(\\sigma^\{\\prime\}\_\{\-i\}\)\\right\|≤θi\(n−1\)di2Ai∑j≠i\|gij\|\|σj−σj′\|\\displaystyle\\leq\\frac\{\\theta\_\{i\}\}\{\(n\-1\)d\_\{i\}^\{2\}A\_\{i\}\}\\sum\_\{j\\neq i\}\|g\_\{ij\}\|\\left\|\\sigma\_\{j\}\-\\sigma^\{\\prime\}\_\{j\}\\right\|≤θi\(n−1\)di2Ai∑j≠i\|gij\|∥σ−σ′∥∞\.\\displaystyle\\leq\\frac\{\\theta\_\{i\}\}\{\(n\-1\)d\_\{i\}^\{2\}A\_\{i\}\}\\sum\_\{j\\neq i\}\|g\_\{ij\}\|\\lVert\\sigma\-\\sigma^\{\\prime\}\\rVert\_\{\\infty\}\.\(23\)Condition \([13](https://arxiv.org/html/2607.27134#A4.E13)\) therefore makes the joint best\-response map a contraction in the sup norm\. Banach’s fixed\-point theorem gives existence and uniqueness of the Nash equilibrium\. At an interior equilibrium, Equation \([22](https://arxiv.org/html/2607.27134#A4.E22)\) is zero for everyii, which is exactly the linear systemNσNE=hN\\sigma^\{\\mathrm\{NE\}\}=h\.
For welfare, each unordered pair\{i,j\}\\\{i,j\\\}appears in both authors’ distinctiveness payoffs, with total coefficient
θi\+θj2\(n−1\)\.\\frac\{\\theta\_\{i\}\+\\theta\_\{j\}\}\{2\(n\-1\)\}\.Therefore,
W\(σ\)=\\displaystyle W\(\\sigma\)=\{\}−12∑i\[bidi2σi2\+cidi2\(1−σi\)2\]\\displaystyle\-\\frac\{1\}\{2\}\\sum\_\{i\}\\left\[b\_\{i\}d\_\{i\}^\{2\}\\sigma\_\{i\}^\{2\}\+c\_\{i\}d\_\{i\}^\{2\}\(1\-\\sigma\_\{i\}\)^\{2\}\\right\]\+∑i<jθi\+θj2\(n−1\)∥σiui−σjuj∥22\.\\displaystyle\+\\sum\_\{i<j\}\\frac\{\\theta\_\{i\}\+\\theta\_\{j\}\}\{2\(n\-1\)\}\\lVert\\sigma\_\{i\}u\_\{i\}\-\\sigma\_\{j\}u\_\{j\}\\rVert\_\{2\}^\{2\}\.\(24\)Expanding the squared distances and collecting coefficients gives Equations \([16](https://arxiv.org/html/2607.27134#A4.E16)\)–\([18](https://arxiv.org/html/2607.27134#A4.E18)\)\. IfSSis positive definite, the Hessian of welfare is−S\-S, so welfare is strictly concave\. Its interior first\-order condition is
∇W\(σ\)=h−Sσ=0,\\nabla W\(\\sigma\)=h\-S\\sigma=0,which gives Equation \([19](https://arxiv.org/html/2607.27134#A4.E19)\)\. ∎
###### Corollary D\.2\(Over\-conformity with nonpositively aligned signatures\)\.
Suppose
gij≤0for everyi≠j\.g\_\{ij\}\\leq 0\\qquad\\text\{for every \}i\\neq j\.\(25\)Assume that the Nash equilibrium and social optimum are interior and that
di2\(bi\+ci−θi\)\>θin−1∑j≠i\|gij\|d\_\{i\}^\{2\}\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\)\>\\frac\{\\theta\_\{i\}\}\{n\-1\}\\sum\_\{j\\neq i\}\|g\_\{ij\}\|\(26\)and
di2\(bi\+ci−θi−θ¯−i\)\>1n−1∑j≠i\(θi\+θj\)\|gij\|d\_\{i\}^\{2\}\\left\(b\_\{i\}\+c\_\{i\}\-\\theta\_\{i\}\-\\bar\{\\theta\}\_\{\-i\}\\right\)\>\\frac\{1\}\{n\-1\}\\sum\_\{j\\neq i\}\(\\theta\_\{i\}\+\\theta\_\{j\}\)\|g\_\{ij\}\|\(27\)for every authorii\. Then the Nash equilibrium and social optimum are unique and
σiSO≥σiNEfor everyi\.\\sigma\_\{i\}^\{\\mathrm\{SO\}\}\\geq\\sigma\_\{i\}^\{\\mathrm\{NE\}\}\\qquad\\text\{for every \}i\.\(28\)Equivalently,
λiNE≥λiSOfor everyi\.\\lambda\_\{i\}^\{\\mathrm\{NE\}\}\\geq\\lambda\_\{i\}^\{\\mathrm\{SO\}\}\\qquad\\text\{for every \}i\.\(29\)The inequality for authoriiis strict whenever
σiNE\>0andθ¯−i\>0\.\\sigma\_\{i\}^\{\\mathrm\{NE\}\}\>0\\qquad\\text\{and\}\\qquad\\bar\{\\theta\}\_\{\-i\}\>0\.\(30\)Consequently,
D∞\(λNE\)≤D∞\(λSO\)\.D^\{\\infty\}\(\\lambda^\{\\mathrm\{NE\}\}\)\\leq D^\{\\infty\}\(\\lambda^\{\\mathrm\{SO\}\}\)\.\(31\)
###### Proof\.
Under Equation \([25](https://arxiv.org/html/2607.27134#A4.E25)\), bothNNandSShave nonpositive off\-diagonal entries\. Conditions \([26](https://arxiv.org/html/2607.27134#A4.E26)\) and \([27](https://arxiv.org/html/2607.27134#A4.E27)\) give positive, strictly dominant diagonal entries\. Consequently,NNandSSare nonsingularMM\-matrices and have entrywise nonnegative inverses\.
Furthermore,S≤NS\\leq Nentrywise\. On the diagonal,
Sii=Nii−di2θ¯−i≤Nii\.S\_\{ii\}=N\_\{ii\}\-d\_\{i\}^\{2\}\\bar\{\\theta\}\_\{\-i\}\\leq N\_\{ii\}\.Fori≠ji\\neq j,
Sij−Nij=θjn−1gij≤0\.S\_\{ij\}\-N\_\{ij\}=\\frac\{\\theta\_\{j\}\}\{n\-1\}g\_\{ij\}\\leq 0\.BecauseσNE≥0\\sigma^\{\\mathrm\{NE\}\}\\geq 0,
SσNE≤NσNE=h\.S\\sigma^\{\\mathrm\{NE\}\}\\leq N\\sigma^\{\\mathrm\{NE\}\}=h\.Multiplication by the nonnegative matrixS−1S^\{\-1\}gives
σNE≤S−1h=σSO,\\sigma^\{\\mathrm\{NE\}\}\\leq S^\{\-1\}h=\\sigma^\{\\mathrm\{SO\}\},which proves Equations \([28](https://arxiv.org/html/2607.27134#A4.E28)\) and \([29](https://arxiv.org/html/2607.27134#A4.E29)\)\.
For strictness, observe that
h−SσNE=\(N−S\)σNE\.h\-S\\sigma^\{\\mathrm\{NE\}\}=\(N\-S\)\\sigma^\{\\mathrm\{NE\}\}\.Under Equation \([25](https://arxiv.org/html/2607.27134#A4.E25)\), the matrixN−SN\-Sis entrywise nonnegative\. Itsiith diagonal contribution is
di2θ¯−iσiNE,d\_\{i\}^\{2\}\\bar\{\\theta\}\_\{\-i\}\\sigma\_\{i\}^\{\\mathrm\{NE\}\},which is strictly positive whenever Equation \([30](https://arxiv.org/html/2607.27134#A4.E30)\) holds\. BecauseS−1S^\{\-1\}is entrywise nonnegative and has strictly positive diagonal entries, it follows that
σiSO\>σiNE,\\sigma\_\{i\}^\{\\mathrm\{SO\}\}\>\\sigma\_\{i\}^\{\\mathrm\{NE\}\},or equivalently,
λiNE\>λiSO\.\\lambda\_\{i\}^\{\\mathrm\{NE\}\}\>\\lambda\_\{i\}^\{\\mathrm\{SO\}\}\.
Finally, whengij≤0g\_\{ij\}\\leq 0, each pairwise squared distance
∥σiui−σjuj∥22=σi2di2\+σj2dj2−2σiσjgij\\lVert\\sigma\_\{i\}u\_\{i\}\-\\sigma\_\{j\}u\_\{j\}\\rVert\_\{2\}^\{2\}=\\sigma\_\{i\}^\{2\}d\_\{i\}^\{2\}\+\\sigma\_\{j\}^\{2\}d\_\{j\}^\{2\}\-2\\sigma\_\{i\}\\sigma\_\{j\}g\_\{ij\}is nondecreasing in each ofσi\\sigma\_\{i\}andσj\\sigma\_\{j\}on\[0,1\]2\[0,1\]^\{2\}\. The componentwise inequalityσSO≥σNE\\sigma^\{\\mathrm\{SO\}\}\\geq\\sigma^\{\\mathrm\{NE\}\}therefore implies Equation \([31](https://arxiv.org/html/2607.27134#A4.E31)\)\. ∎
## Appendix EA Heterogeneous Population with Multiple Interaction Mechanisms
This section introduces a fourth interaction mechanism that combines Mechanisms 1, 2 and 3 within a single population by partitioning authors into subpopulations, each following the update rule of its assigned mechanism\.
###### Interaction Mechanism 4\(A heterogeneous population with multiple interaction mechanisms\)\.
The population is partitioned into three disjoint subpopulations\[n\]=ℐ1∪ℐ2∪ℐ3,\[n\]=\\mathcal\{I\}\_\{1\}\\cup\\mathcal\{I\}\_\{2\}\\cup\\mathcal\{I\}\_\{3\},where authors inℐ1\\mathcal\{I\}\_\{1\},ℐ2\\mathcal\{I\}\_\{2\}andℐ3\\mathcal\{I\}\_\{3\}follow Interaction Mechanisms 1, 2 and 3, respectively\. Letq0q^\{0\}denote the common base model distribution\. Authors inℐ1\\mathcal\{I\}\_\{1\}interact with a shared model whose distribution remains fixed atq0q\_\{0\}\. Their linguistic\-style distributions evolve as
pit\+1=\(1−αi\)pit\+αi𝒜i\(q0,ri\),i∈ℐ1\.p\_\{i\}^\{t\+1\}=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(q^\{0\},r\_\{i\}\),\\qquad i\\in\\mathcal\{I\}\_\{1\}\.Authors inℐ2\\mathcal\{I\}\_\{2\}interact with a shared model whose linguistic style distributionqt\{q\}^\{t\}is recursively updated using feedback from authors inℐ2\\mathcal\{I\}\_\{2\}\. Their coupled author–model dynamics evolve as
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αi𝒜i\(qt,ri\),i∈ℐ2,\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\,\\mathcal\{A\}\_\{i\}\(\{q\}^\{t\},r\_\{i\}\),\\qquad i\\in\\mathcal\{I\}\_\{2\},qt\+1\\displaystyle q^\{t\+1\}=βqt\+\(1−β\)ℬ\(\(pjt\+1\)j∈ℐ2\),\\displaystyle=\\beta\\,q^\{t\}\+\(1\-\\beta\)\\,\\mathcal\{B\}\\big\(\(p\_\{j\}^\{t\+1\}\)\_\{j\\in\\mathcal\{I\}\_\{2\}\}\\big\),whereℬ\\mathcal\{B\}aggregates the updated author distributions inℐ2\\mathcal\{I\}\_\{2\}\. In particular, we use the weighted\-average update
ℬ\(\(pjt\+1\)j∈ℐ2\)=∑j∈ℐ2wjpjt\+1,wj≥0,∑j∈ℐ2wj=1\.\\mathcal\{B\}\\big\(\(p\_\{j\}^\{t\+1\}\)\_\{j\\in\\mathcal\{I\}\_\{2\}\}\\big\)=\\sum\_\{j\\in\\mathcal\{I\}\_\{2\}\}w\_\{j\}p\_\{j\}^\{t\+1\},\\qquad w\_\{j\}\\geq 0,\\quad\\sum\_\{j\\in\\mathcal\{I\}\_\{2\}\}w\_\{j\}=1\.Each authori∈ℐ3i\\in\\mathcal\{I\}\_\{3\}interacts with an author\-specific personalized modelqitq\_\{i\}^\{t\}, initialized from a common base distributionqi0=q0q\_\{i\}^\{0\}=q^\{0\}\. Their coupled author–model dynamics are
pit\+1\\displaystyle p\_\{i\}^\{t\+1\}=\(1−αi\)pit\+αi𝒜i\(qit,ri\),i∈ℐ3,\\displaystyle=\(1\-\\alpha\_\{i\}\)\\,p\_\{i\}^\{t\}\+\\alpha\_\{i\}\\mathcal\{A\}\_\{i\}\(q\_\{i\}^\{t\},r\_\{i\}\),\\qquad i\\in\\mathcal\{I\}\_\{3\},qit\+1\\displaystyle q\_\{i\}^\{t\+1\}=βqit\+γpit\+1\+δ∑j∈ℐ3wjpjt\+1,i∈ℐ3,\\displaystyle=\\beta\\,q\_\{i\}^\{t\}\+\\gamma\\,p\_\{i\}^\{t\+1\}\+\\delta\\,\\sum\_\{j\\in\\mathcal\{I\}\_\{3\}\}w\_\{j\}p\_\{j\}^\{t\+1\},\\qquad i\\in\\mathcal\{I\}\_\{3\},whereβ,γ,δ≥0\\beta,\\gamma,\\delta\\geq 0,β\+γ\+δ=1\\beta\+\\gamma\+\\delta=1,wj≥0w\_\{j\}\\geq 0, and∑j∈ℐ3wj=1\\sum\_\{j\\in\\mathcal\{I\}\_\{3\}\}w\_\{j\}=1\. The three subpopulations evolve according to their respective update rules, with no additional coupling acrossℐ1\\mathcal\{I\}\_\{1\},ℐ2\\mathcal\{I\}\_\{2\}andℐ3\\mathcal\{I\}\_\{3\}\. Thus,[Mechanism4](https://arxiv.org/html/2607.27134#Thmmechanism4)captures a heterogeneous population in which different groups use LLM assistance through different interaction mechanisms\.
## Appendix FAdditional Experimental Results
This appendix reports robustness and diagnostic experiments complementing Figure[1](https://arxiv.org/html/2607.27134#S5.F1)\. We study sensitivity to population size, Dirichlet concentration, conformity heterogeneity, and the diversity metric, together with auxiliary convergence measures, personalization dynamics, and a heterogeneous population whose subgroups follow IM 1–3\. Our implementation222LLMs \(*e\.g\.*, ChatGPT Codex, CoPilot\) were used to assist with implementation and code refinement\. All generated code was reviewed and verified by the authors, who remain responsible for its correctness\.is written inPythonprogramming language using standard numerical libraries\. All experiments were executed on a commodity laptop\. The source code is available as open\-source under a modest license agreement\.333https://github\.com/suhastheju/llm\-monoculture\-code
Unless stated otherwise, we simulaten=100n=100authors overm=10m=10abstract linguistic features forT=200T=200steps\. We initializepi0p\_\{i\}^\{0\},q0q^\{0\}, andrir\_\{i\}independently fromDirichlet\(1,…,1\)\\operatorname\{Dirichlet\}\(1,\\ldots,1\)and use the same initialization across mechanisms within each run\. We independently sampleαi∼Uniform\[0\.05,0\.50\)\\alpha\_\{i\}\\sim\\operatorname\{Uniform\}\[0\.05,0\.50\)andλi∼Uniform\[0,1\)\\lambda\_\{i\}\\sim\\operatorname\{Uniform\}\[0,1\), and use uniform author weightswi=1/nw\_\{i\}=1/n\. The baseline parameters areβ=0\.25\\beta=0\.25andρ=2/3\\rho=2/3\(equivalently,γ=0\.5\\gamma=0\.5andδ=0\.25\\delta=0\.25\)\. All panels report means over 100 independent paired runs, with shading denoting±1\\pm 1sample standard deviation\. Time is displayed on alog\(1\+t\)\\log\(1\+t\)scale to resolve the rapid initial transient; tick labels report the original time steps\. The master random seed is123456789123456789\. The absolute plateau levels and the ordering between IM 1 and IM 2 depend on the initialization and parameter prior\. The experiments therefore illustrate mechanism\-specific dynamics in controlled parameter regimes rather than establish a universal ordering\. Alongside the Jensen–Shannon diversityDtD^\{t\}, we report the translation\-invariant quadratic diagnostic
D~t:=m8n\(n−1\)∑i≠j∥pit−pjt∥22\.\\widetilde\{D\}^\{t\}:=\\frac\{m\}\{8n\(n\-1\)\}\\sum\_\{i\\neq j\}\\\|p\_\{i\}^\{t\}\-p\_\{j\}^\{t\}\\\|\_\{2\}^\{2\}\.This quantity depends only on the pairwise difference vectorspit−pjtp\_\{i\}^\{t\}\-p\_\{j\}^\{t\}, so a common translation of an author cloud leaves it unchanged\. The factorm/8m/8matches the second\-order JS scaling when pairwise midpoints lie at the barycenter\. We useD~t\\widetilde\{D\}^\{t\}only as a diagnostic; the paper’s primary diversity measure remainsDtD^\{t\}\.
#### Dirichlet concentration and diversity metric\.
We first vary the initialization concentrationa∈\{0\.1,1,5\}a\\in\\\{0\.1,1,5\\\}inDirichlet\(a,…,a\)\\operatorname\{Dirichlet\}\(a,\\ldots,a\)\. The ordering IM 2<<IM 1<<IM 3 holds atT=200T=200under bothDtD^\{t\}andD~t\\widetilde\{D\}^\{t\}for every value ofaa\. For the baselinea=1a=1, the quadratic endpoints are approximately0\.0670\.067for IM 2,0\.0830\.083for IM 1, and0\.1070\.107for IM 3\. Thus, the observed endpoint ordering is also present under the quadratic diagnostic and is not solely an artifact of the positional sensitivity of Jensen–Shannon divergence\.
Figure 2:Linguistic\-diversity trajectoriesDtD^\{t\}under IM 1–3 forDirichlet\(a,…,a\)\\operatorname\{Dirichlet\}\(a,\\ldots,a\)initialization\. The endpoint ordering IM 2<<IM 1<<IM 3 holds at every concentration\. Lines show means over 100 runs, and shading denotes±1\\pm 1sample standard deviation\.
#### Conformity heterogeneity\.
To separate positional and geometric effects, we sampleλi∼Uniform\[0\.5−w,0\.5\+w\]\\lambda\_\{i\}\\sim\\operatorname\{Uniform\}\[0\.5\-w,0\.5\+w\]forw∈\{0,0\.25,0\.5\}w\\in\\\{0,0\.25,0\.5\\\}\. Atw=0w=0, IM 1 and IM 2 have identical limiting pairwise difference vectors, and their quadratic gap is numerically zero \(8\.9×10−98\.9\\times 10^\{\-9\}\), while their JS gap is approximately0\.00710\.0071\. The latter therefore reflects JS’s dependence on location within the simplex\. As conformity becomes heterogeneous, the quadratic gap rises to approximately0\.00410\.0041atw=0\.25w=0\.25and0\.01630\.0163atw=0\.5w=0\.5, showing that heterogeneous responses to different anchors create a genuine difference in pairwise geometry\.
Figure 3:IM 1 minus IM 2 endpoint diversity as conformity heterogeneity increases\. Under common conformity, the quadratic gap vanishes although the JS gap remains positive; with heterogeneous conformity, a genuine geometric gap emerges under both metrics\.
#### Population\-size robustness\.
Figure[4](https://arxiv.org/html/2607.27134#A6.F4)compares the three interaction mechanisms forn=100n=100andn=500n=500\. The qualitative ordering is stable across the two population sizes: recursive shared feedback produces the lowest long\-run diversity, while personalized feedback preserves the most diversity\. Variability across runs is smaller atn=500n=500, as expected sinceDtD^\{t\}averages overO\(n2\)O\(n^\{2\}\)author pairs\.
Figure 4:Population\-size robustness of linguistic diversityDtD^\{t\}under IM 1–3 forn=100n=100\(top\) andn=500n=500\(bottom\)\. Lines show means over 100 runs and shading denotes±1\\pm 1standard deviation\.
#### Auxiliary convergence measures\.
Figure[5](https://arxiv.org/html/2607.27134#A6.F5)reports author–model divergenceMtM^\{t\}and personalized\-model diversityQtQ^\{t\}\. Because IM 1 and IM 2 use a single shared model, their model diversity is identically zero and is omitted from panel \(b\)\. Under IM 3,MtM^\{t\}rapidly falls to a small positive level whileQtQ^\{t\}remains positive, indicating that personalized models become distinct while remaining closely aligned with their respective authors\. The limiting model diversity is smaller than the limiting author diversity \(Q⋆≈0\.041Q^\{\\star\}\\approx 0\.041versusD⋆≈0\.095D^\{\\star\}\\approx 0\.095\), consistent with[Section3\.3](https://arxiv.org/html/2607.27134#S3.SS3): eachqi⋆q\_\{i\}^\{\\star\}retains the author\-specific componentrir\_\{i\}with weightρηi<ηi\\rho\\eta\_\{i\}<\\eta\_\{i\}\.
Figure 5:Auxiliary convergence measures under the baseline parameters: \(a\) author–model divergenceMtM^\{t\}for IM 1–3 and \(b\) personalized\-model diversityQtQ^\{t\}for IM 3\. Lines indicate mean over 100 runs and shading denotes±1\\pm 1standard deviation\.
#### Personalization and transient dynamics\.
To complement the endpoint comparison in Figure[1](https://arxiv.org/html/2607.27134#S5.F1), Figure[6](https://arxiv.org/html/2607.27134#A6.F6)shows the full IM 3 trajectory forρ∈\{0,2/3,1\}\\rho\\in\\\{0,2/3,1\\\}\. We holdβ=0\.25\\beta=0\.25fixed and setγ=\(1−β\)ρ\\gamma=\(1\-\\beta\)\\rhoandδ=\(1−β\)\(1−ρ\)\\delta=\(1\-\\beta\)\(1\-\\rho\)\. Increasingρ\\rhotherefore reallocates feedback weight from the shared population to the associated author\.
The endpoints provide a check on[Section3\.3](https://arxiv.org/html/2607.27134#S3.SS3)\. Atρ=0\\rho=0the personalized update reduces exactly to IM 2, and the observedD200≈0\.057D^\{200\}\\approx 0\.057matches the IM 2 plateau\. Atρ=1\\rho=1we haveηi=1\\eta\_\{i\}=1, sopi⋆=qi⋆=rip\_\{i\}^\{\\star\}=q\_\{i\}^\{\\star\}=r\_\{i\}; sincerir\_\{i\}andpi0p\_\{i\}^\{0\}are drawn from the same Dirichlet, diversity should return to approximately its initial level, and the observed0\.1710\.171is consistent withD0≈0\.175D^\{0\}\\approx 0\.175\.
All mechanisms exhibit a non\-monotone transient:DtD^\{t\}undershoots its limit neart≈5t\\approx 5before partially recovering\. Because personalized models are initialized at the common base distributionq0q^\{0\}, every author is initially pulled toward the same point; theqitq\_\{i\}^\{t\}differentiate only later, allowing authors to drift back toward their preferredrir\_\{i\}\. The recovery scales withρ\\rho, consistent with this account\.
Figure 6:IM 3 linguistic\-diversity trajectories for three personalization levels\. Largerρ\\rhopreserves more author\-specific feedback and produces higher long\-runDtD^\{t\}\. Lines show mean over100100runs and shading denotes±1\\pm 1standard deviation\.
#### Heterogeneous interaction mechanisms\.
For IM 4, we assign proportions0\.330\.33,0\.330\.33, and0\.340\.34of the population to IM 1, IM 2, and IM 3, respectively, using assignment seed123456789123456789\. Each subgroup follows its own update rule\. In particular, the population\-level component received by personalized models is computed only from authors in the IM 3 subgroup, as specified in[AppendixE](https://arxiv.org/html/2607.27134#A5); there is no additional cross\-subpopulation feedback\. All three quantities stabilize; the positive limiting value ofQtQ^\{t\}reflects differences among the author\-facing models used across and within the three subgroups\.
Figure 7:Evolution of \(a\) author diversityDtD^\{t\}, \(b\) diversity among author\-facing model distributionsQtQ^\{t\}, and \(c\) author–model divergenceMtM^\{t\}under heterogeneous IM 4\. Lines show mean over 100 runs and shading denotes±1\\pm 1standard deviation\.相似文章
多语言中数学推理的LLM参数:共享还是独立?
本文提出了一种跨语言的LLM数学推理机制分析,发现数学相关参数在不同语言之间存在部分重叠,主要集中于中间层。英语拥有最大规模的数学相关参数集,而低资源语言则拥有较小的参数集。
对齐更优,多样性下降?分析两代大语言模型的语法与词汇特征
这篇学术论文分析了两代大语言模型与人类撰写新闻文本相比的句法和词汇多样性,发现较新的对齐模型表现出多样性降低的现象。
实际环境中的多语言多模态大语言模型:面向低资源语言的构建
本教程论文概述了如何为低资源语言构建多语言多模态大语言模型,涵盖数据创建、模型对齐、微调和评估,重点提供实用方案和动手资源。
适应是双向的:研究人类与语言模型之间的语言趋同
本文研究了在多轮对话中人类与大型语言模型之间的语言适应性,发现LLM过度趋同于用户风格,而人类适应LLM的方式与适应其他人类并无不同。
为了内容而内容
作者探讨了LLM如何影响编码和日常语言中的用词,发现LLM偏好的词汇在编程会话和Google Trends中出现的频率均有所增加,这引发了人们对人类开始采用LLM写作风格的担忧。