Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics
Summary
This paper proposes a bilayer coupled SIR/SIRS framework to model synthetic data contamination and model collapse in AI ecosystems, showing that cross-contamination between models and data corpora leads to supercritical dynamics and identifying detection-based filtering as a key intervention.
View Cached Full Text
Cached at: 06/05/26, 08:04 AM
# Modeling Synthetic Data Contamination via Bilayer SIR Dynamics
Source: [https://arxiv.org/html/2606.05168](https://arxiv.org/html/2606.05168)
## Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics
###### Abstract
Training on synthetic data causes model collapse, but existing analyses treat this as single\-chain degradation\. In reality, the AI ecosystem involves cross\-contamination: models ingest synthetic data from other models, produce new synthetic text, and contaminate shared corpora\. We propose a*bilayer coupled SIR/SIRS*framework—a phenomenological mean\-field model treating data corpora and AI models as two interacting populations, each with susceptible, infected, and recovered compartments linked by cross\-layer transmission\. The SIRS variant \(our primary recommendation\) incorporates immunity waning, reflecting that filtered corpora and retrained models remain susceptible to re\-contamination\. We derive the basic reproduction numberR0=βDβM/\[\(γD\+μD\)\(γM\+μM\)\]R\_\{0\}=\\sqrt\{\\beta\_\{D\}\\beta\_\{M\}/\[\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\]\}via the Next Generation Matrix and apply standard epidemic threshold results to the bilayer system\. Illustrative scenario\-based calibration from public AI text prevalence data yields supercritical dynamics \(R0\>1R\_\{0\}\>1\) across three scenarios; Sobol sensitivity analysis identifies synthetic\-text detection as the highest\-leverage parameter\. A bipartite\-network agent\-based model confirms mean\-field consistency \(R2\>0\.96R^\{2\}\>0\.96\) for dense networks but degrades under heterogeneity\. GPT\-2 contamination chain experiments \(192 single\-chain runs across WikiText and Shakespeare\) show dose\-response degradation and diversity loss \(Distinct\-2 drops from 0\.68 to 0\.38\) qualitatively consistent with the threshold picture\. Matched\-budget source\-diversity experiments \(1,088 additional runs;K∈\{1,3,5\}K\\in\\\{1,3,5\\\}atα=1\\alpha\{=\}1,K∈\{1,5\}K\\in\\\{1,5\\\}atα=0\.5\\alpha\{=\}0\.5; 8 seeds, fixed pool size\) provide suggestive evidence that multi\-source mixing modestly attenuates collapse atα=1\\alpha\{=\}1\(∼2\{\\sim\}2PPL,d≈0\.8d\\approx 0\.8, one\-sidedp=0\.047p=0\.047\), but the effect*vanishes*atα=0\.5\\alpha\{=\}0\.5, confirming contamination fraction as the dominant driver\. Under model assumptions, illustrative intervention analysis identifies detection\-based filtering and herd immunity as the highest\-leverage strategies\.
## 1Introduction
Large language models \(LLMs\) now generate a substantial and growing fraction of online text\. Recent estimates suggest that up to 74% of newly indexed web pages may contain AI\-generated or AI\-modified content\[[22](https://arxiv.org/html/2606.05168#bib.bib22)\]\(a projected figure; see[Appendix˜E](https://arxiv.org/html/2606.05168#A5)\), with AI text prevalence among top\-ranked search results increasing approximately fourfold between 2023 and 2025\[[11](https://arxiv.org/html/2606.05168#bib.bib11)\]\. As models trained on web\-crawled corpora inevitably ingest this synthetic output, a feedback loop emerges: models produce text that enters training data, which shapes the next generation of models\. Shumailov et al\.\[[20](https://arxiv.org/html/2606.05168#bib.bib20)\]demonstrated that recursive training on self\-generated data causes*model collapse*—progressive degradation of output quality and diversity\. Subsequent work has characterized this phenomenon in language models\[[4](https://arxiv.org/html/2606.05168#bib.bib4)\], image generators\[[1](https://arxiv.org/html/2606.05168#bib.bib1)\], and mixed real\-synthetic training regimes\[[5](https://arxiv.org/html/2606.05168#bib.bib5)\]\.
However, all formal analyses of model collapse treat it as a*single\-chain*process: modelAAgenerates data, modelBBtrains on it, modelCCtrains onBB’s output, and so on\. The real AI ecosystem is not a chain but a*network*\. Thousands of models consume data from shared corpora; each model’s outputs re\-enter the data pool through web publishing, API responses, and synthetic data pipelines\. Cross\-contamination is the norm, not the exception\. No existing mathematical framework provides even a stylized model of these ecosystem\-level contamination dynamics\.
We observe that this cross\-contamination process is structurally analogous to*epidemic spreading*\. Contaminated data “infects” models during training; infected models “transmit” synthetic artifacts back into data corpora\. The two populations—data corpora and AI models—interact through cross\-layer transmission, forming a natural bilayer epidemic system\. The SIR \(Susceptible–Infected–Recovered\) compartmental framework\[[9](https://arxiv.org/html/2606.05168#bib.bib9),[6](https://arxiv.org/html/2606.05168#bib.bib6)\]provides an established mathematical apparatus for reasoning about transmission thresholds, equilibria, and interventions in such two\-population systems\.
#### Contributions\.
We make five contributions:
1. 1\.Bilayer SIR/SIRS framework\([Section˜3](https://arxiv.org/html/2606.05168#S3)\): a phenomenological mean\-field ODE coupling data and model populations\. The SIRS variant is our primary recommendation\.
2. 2\.Threshold analysis\([Section˜3](https://arxiv.org/html/2606.05168#S3)\):R0R\_\{0\}via the Next Generation Matrix\[[23](https://arxiv.org/html/2606.05168#bib.bib23)\]; DFE stability, endemic equilibrium, and bifurcation results\.
3. 3\.Illustrative calibration\([Section˜4](https://arxiv.org/html/2606.05168#S4)\): three scenarios from public data, with Sobol sensitivity analysis identifying detection \(γD\\gamma\_\{D\}\) as highest\-leverage\.
4. 4\.GPT\-2 experiments\([Section˜6](https://arxiv.org/html/2606.05168#S6)\): 192 single\-chain runs show dose\-response degradation in perplexity and diversity; 1,088 matched\-budget source\-diversity runs show modest attenuation atα=1\\alpha\{=\}1that vanishes atα=0\.5\\alpha\{=\}0\.5\.
5. 5\.Intervention analysis\([Section˜7](https://arxiv.org/html/2606.05168#S7)\): six strategies in 135 evaluations \(15 pairs×\\times9 intensities\), conditional on scenario parameters\.
Scope\.Our SIR mapping is a*phenomenological framework*; the mathematical results are standard epidemic theory applied to a new domain\. Calibration is illustrative; GPT\-2 experiments provide a qualitative bridge at small scale\. Limitations are discussed in[Section˜7](https://arxiv.org/html/2606.05168#S7)\.
## 2Related Work
#### Model collapse in generative models\.
Shumailov et al\.\[[20](https://arxiv.org/html/2606.05168#bib.bib20)\]established that recursively training language models on their own output causes progressive quality degradation, a phenomenon they termed*model collapse*\. Dohmatob et al\.\[[4](https://arxiv.org/html/2606.05168#bib.bib4)\]proved that even small fractions of synthetic data in training corpora can trigger collapse, providing tight theoretical bounds\. Alemohammad et al\.\[[1](https://arxiv.org/html/2606.05168#bib.bib1)\]demonstrated analogous collapse in image generation \(*Model Autophagy Disorder*\), showing the phenomenon transcends modalities\. Gerstgrasser et al\.\[[5](https://arxiv.org/html/2606.05168#bib.bib5)\]studied data accumulation as a mitigation, finding that preserving original data across generations can slow but not always prevent collapse\. Seddik et al\.\[[19](https://arxiv.org/html/2606.05168#bib.bib19)\]derived statistical bounds on degradation from synthetic training data\. All of these works analyze*single\-chain*dynamics: one lineage of models training on its own outputs\. Our work differs by modeling the*ecosystem\-level*cross\-contamination network, where multiple model lineages share and pollute a common data pool\.
#### Epidemic modeling beyond biology\.
The SIR framework\[[9](https://arxiv.org/html/2606.05168#bib.bib9)\]and its extensions\[[6](https://arxiv.org/html/2606.05168#bib.bib6)\]have been applied far beyond infectious disease\. Kephart and White\[[8](https://arxiv.org/html/2606.05168#bib.bib8)\]modeled computer virus propagation using SIS dynamics on networks\. Jin et al\.\[[7](https://arxiv.org/html/2606.05168#bib.bib7)\]and Vosoughi et al\.\[[24](https://arxiv.org/html/2606.05168#bib.bib24)\]applied epidemic models to rumor and misinformation spreading on social media\. Pastor\-Satorras et al\.\[[15](https://arxiv.org/html/2606.05168#bib.bib15)\]provided a comprehensive treatment of epidemic processes on complex networks, establishing mean\-field approximations and threshold conditions for networked SIR/SIS systems\. The Next Generation Matrix method for computingR0R\_\{0\}in multi\-compartment models was formalized by Diekmann et al\.\[[3](https://arxiv.org/html/2606.05168#bib.bib3)\]and extended by van den Driessche and Watmough\[[23](https://arxiv.org/html/2606.05168#bib.bib23)\]\. Castillo\-Chavez and Song\[[2](https://arxiv.org/html/2606.05168#bib.bib2)\]demonstrated the application of Sotomayor’s theorem to establish bifurcation results in epidemic systems\. Our bilayer SIR formulation draws on these established techniques but applies them to a new domain: AI training data contamination\.
#### AI data quality and provenance\.
Kirchenbauer et al\.\[[10](https://arxiv.org/html/2606.05168#bib.bib10)\]introduced watermarking for LLM outputs, enabling downstream detection of synthetic text\. Tang et al\.\[[21](https://arxiv.org/html/2606.05168#bib.bib21)\]surveyed detection methods for LLM\-generated text, establishing current accuracy bounds\. Mitchell et al\.\[[14](https://arxiv.org/html/2606.05168#bib.bib14)\]proposed model cards as a provenance\-tracking mechanism\. Longpre et al\.\[[12](https://arxiv.org/html/2606.05168#bib.bib12)\]studied the effects of training data composition on model quality\. These works address individual components of what our framework models as “recovery” mechanisms \(γD\\gamma\_\{D\},γM\\gamma\_\{M\}\): detection removes contaminated data, watermarking enables filtering, and provenance tracking supports data hygiene\. Our contribution is to embed these mechanisms within a unified dynamical systems framework that quantifies their collective impact on ecosystem\-level contamination\.
## 3Bilayer SIR Model
We model synthetic data contamination in the AI ecosystem as a bilayer coupled SIR system\. The two layers represent*data corpora*\(layerDD\) and*AI models*\(layerMM\), each with susceptible, infected, and recovered compartments linked by cross\-layer transmission\.
### 3\.1Compartment Mapping
[Table˜1](https://arxiv.org/html/2606.05168#S3.T1)defines the mapping\. Data corpora transition from clean \(SDS\_\{D\}\) to contaminated \(IDI\_\{D\}\) when synthetic content is ingested, and recover \(RDR\_\{D\}\) via detection\. Models transition analogously through training on contaminated data and retraining\.
#### Operational definitions\.
In practice, contamination is continuous rather than binary\. We operationalize the S/I/R partition via a threshold: a corpus \(or model\) is “infected” if its synthetic content fraction \(or synthetic\-data training fraction\) exceeds a domain\-specific thresholdτ\\tausufficient to measurably degrade downstream quality\. “Recovered” means the corpus has been filtered or the model retrained so that synthetic fraction drops belowτ\\tau—a coarse\-grained bookkeeping state rather than a claim of permanent decontamination\. Crucially, recovered entities remain susceptible to re\-contamination; we therefore recommend the SIRS variant \([Section˜3\.5](https://arxiv.org/html/2606.05168#S3.SS5)\) as the primary model\. This threshold discretization is a modeling simplification; state variables are population counts \(with birth rateΛ\\Lambdaand death rateμ\\mu\)\. TheR0R\_\{0\}formula depends on rate ratios, not onτ\\taudirectly;τ\\tauaffects only the assignment of entities to compartments\. In the experiments \([Section˜6](https://arxiv.org/html/2606.05168#S6)\),α\\alphadirectly controls contamination fraction, bypassingτ\\tau\.
Table 1:Compartment mapping between the bilayer SIR model and the AI ecosystem\.
### 3\.2ODE System
The dynamics are governed by six coupled ODEs with cross\-layer infection: contaminated models \(IMI\_\{M\}\) infect data at rateβD\\beta\_\{D\}, and contaminated data \(IDI\_\{D\}\) infects models at rateβM\\beta\_\{M\}\. Turnover ratesμD\\mu\_\{D\},μM\\mu\_\{M\}capture data obsolescence and model retirement\.
SDS\_\{D\}IDI\_\{D\}RDR\_\{D\}SMS\_\{M\}IMI\_\{M\}RMR\_\{M\}βDIMNM\\beta\_\{D\}\\frac\{I\_\{M\}\}\{N\_\{M\}\}γD\\gamma\_\{D\}βMIDND\\beta\_\{M\}\\frac\{I\_\{D\}\}\{N\_\{D\}\}γM\\gamma\_\{M\}ΛD\\Lambda\_\{D\}ΛM\\Lambda\_\{M\}μD\\mu\_\{D\}μD\\mu\_\{D\}μD\\mu\_\{D\}μM\\mu\_\{M\}μM\\mu\_\{M\}μM\\mu\_\{M\}infectsinfectsData LayerModel LayerFigure 1:Bilayer SIR schematic\. Data corpora \(top, blue\) and AI models \(bottom, orange\) form two coupled epidemic layers\. Solid arrows denote within\-layer transitions; dashed red arrows denote cross\-layer infection\. Contaminated models \(IMI\_\{M\}\) inject synthetic content into data corpora, while contaminated data \(IDI\_\{D\}\) infects models during training\.The system reads:
dSDdt\\displaystyle\\frac\{dS\_\{D\}\}\{dt\}=ΛD−βDIMNMSD−μDSD,\\displaystyle=\\Lambda\_\{D\}\-\\beta\_\{D\}\\frac\{I\_\{M\}\}\{N\_\{M\}\}S\_\{D\}\-\\mu\_\{D\}S\_\{D\},\(1\)dIDdt\\displaystyle\\frac\{dI\_\{D\}\}\{dt\}=βDIMNMSD−\(γD\+μD\)ID,\\displaystyle=\\beta\_\{D\}\\frac\{I\_\{M\}\}\{N\_\{M\}\}S\_\{D\}\-\(\\gamma\_\{D\}\+\\mu\_\{D\}\)I\_\{D\},\(2\)dRDdt\\displaystyle\\frac\{dR\_\{D\}\}\{dt\}=γDID−μDRD,\\displaystyle=\\gamma\_\{D\}I\_\{D\}\-\\mu\_\{D\}R\_\{D\},\(3\)dSMdt\\displaystyle\\frac\{dS\_\{M\}\}\{dt\}=ΛM−βMIDNDSM−μMSM,\\displaystyle=\\Lambda\_\{M\}\-\\beta\_\{M\}\\frac\{I\_\{D\}\}\{N\_\{D\}\}S\_\{M\}\-\\mu\_\{M\}S\_\{M\},\(4\)dIMdt\\displaystyle\\frac\{dI\_\{M\}\}\{dt\}=βMIDNDSM−\(γM\+μM\)IM,\\displaystyle=\\beta\_\{M\}\\frac\{I\_\{D\}\}\{N\_\{D\}\}S\_\{M\}\-\(\\gamma\_\{M\}\+\\mu\_\{M\}\)I\_\{M\},\(5\)dRMdt\\displaystyle\\frac\{dR\_\{M\}\}\{dt\}=γMIM−μMRM,\\displaystyle=\\gamma\_\{M\}I\_\{M\}\-\\mu\_\{M\}R\_\{M\},\(6\)whereND=SD\+ID\+RDN\_\{D\}=S\_\{D\}\+I\_\{D\}\+R\_\{D\}andNM=SM\+IM\+RMN\_\{M\}=S\_\{M\}\+I\_\{M\}\+R\_\{M\}are total populations, withND→ΛD/μDN\_\{D\}\\to\\Lambda\_\{D\}/\\mu\_\{D\}andNM→ΛM/μMN\_\{M\}\\to\\Lambda\_\{M\}/\\mu\_\{M\}at equilibrium\.
### 3\.3Basic Reproduction Number
We deriveR0R\_\{0\}using the Next Generation Matrix \(NGM\) method\[[3](https://arxiv.org/html/2606.05168#bib.bib3),[23](https://arxiv.org/html/2606.05168#bib.bib23)\]\. The infected subsystem has state\(ID,IM\)T\(I\_\{D\},I\_\{M\}\)^\{T\}\. At the disease\-free equilibrium,SD∗=ΛD/μD=ND∗S\_\{D\}^\{\*\}=\\Lambda\_\{D\}/\\mu\_\{D\}=N\_\{D\}^\{\*\}andSM∗=ΛM/μM=NM∗S\_\{M\}^\{\*\}=\\Lambda\_\{M\}/\\mu\_\{M\}=N\_\{M\}^\{\*\}\. The new infection rates forIDI\_\{D\}andIMI\_\{M\}involve cross\-population ratios:SD∗/NM∗=\(ΛDμM\)/\(μDΛM\)S\_\{D\}^\{\*\}/N\_\{M\}^\{\*\}=\(\\Lambda\_\{D\}\\mu\_\{M\}\)/\(\\mu\_\{D\}\\Lambda\_\{M\}\)andSM∗/ND∗=\(ΛMμD\)/\(μMΛD\)S\_\{M\}^\{\*\}/N\_\{D\}^\{\*\}=\(\\Lambda\_\{M\}\\mu\_\{D\}\)/\(\\mu\_\{M\}\\Lambda\_\{D\}\)\. This gives the new infection matrixFFand transition matrixVV:
F=\(0βDΛDμMμDΛMβMΛMμDμMΛD0\),V=\(γD\+μD00γM\+μM\)\.F=\\begin\{pmatrix\}0&\\beta\_\{D\}\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}\\\\\[4\.0pt\] \\beta\_\{M\}\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}&0\\end\{pmatrix\},\\quad V=\\begin\{pmatrix\}\\gamma\_\{D\}\+\\mu\_\{D\}&0\\\\ 0&\\gamma\_\{M\}\+\\mu\_\{M\}\\end\{pmatrix\}\.\(7\)
###### Theorem 1\(Basic reproduction number\)\.
The basic reproduction number of the bilayer SIR system \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) is
R0=βD⋅βM\(γD\+μD\)\(γM\+μM\)\\boxed\{\\;R\_\{0\}=\\sqrt\{\\frac\{\\beta\_\{D\}\\cdot\\beta\_\{M\}\}\{\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\}\}\\;\}\(8\)
###### Proof sketch\.
R0=ρ\(FV−1\)R\_\{0\}=\\rho\(FV^\{\-1\}\), the spectral radius of the next\-generation matrix\. The off\-diagonal entries ofFV−1FV^\{\-1\}areβDγM\+μM⋅ΛDμMμDΛM\\frac\{\\beta\_\{D\}\}\{\\gamma\_\{M\}\+\\mu\_\{M\}\}\\cdot\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}andβMγD\+μD⋅ΛMμDμMΛD\\frac\{\\beta\_\{M\}\}\{\\gamma\_\{D\}\+\\mu\_\{D\}\}\\cdot\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}\. Their product isβDβM\(γD\+μD\)\(γM\+μM\)\\frac\{\\beta\_\{D\}\\beta\_\{M\}\}\{\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\}—the cross\-population ratios cancel exactly—soFV−1FV^\{\-1\}has eigenvalues±βDβM/\[\(γD\+μD\)\(γM\+μM\)\]\\pm\\sqrt\{\\beta\_\{D\}\\beta\_\{M\}/\[\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\]\}\. The full derivation is in[Appendix˜A](https://arxiv.org/html/2606.05168#A1)\. ∎
The geometric mean structure ofR0R\_\{0\}reflects the bilayer coupling: contamination must traverse both layers \(data→\\tomodel→\\todata\) to complete a “generation\.”
###### Theorem 2\(Disease\-free equilibrium and stability\)\.
The system \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) admits a unique disease\-free equilibrium
DFE=\(ΛDμD,0,0,ΛMμM,0,0\)\.\\mathrm\{DFE\}=\\left\(\\frac\{\\Lambda\_\{D\}\}\{\\mu\_\{D\}\},\\;0,\\;0,\\;\\frac\{\\Lambda\_\{M\}\}\{\\mu\_\{M\}\},\\;0,\\;0\\right\)\.\(9\)The DFE is locally asymptotically stable if and only ifR0<1R\_\{0\}<1\.
###### Proof sketch\.
SettingID=IM=0I\_\{D\}=I\_\{M\}=0in \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) yields the DFE\. The Jacobian at the DFE is block\-triangular; the infected block has eigenvalues with negative real parts iffR0<1R\_\{0\}<1\. This is a standard result for coupled epidemic systems\[[23](https://arxiv.org/html/2606.05168#bib.bib23)\]; we verify numerically for 200 random configurations\. Details in[Appendix˜A](https://arxiv.org/html/2606.05168#A1)\. ∎
### 3\.4Equilibria and Bifurcation
###### Proposition 3\(Endemic equilibrium existence\)\.
WhenR0\>1R\_\{0\}\>1, the system \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) admits at least one endemic equilibriumEE=\(SD∗,ID∗,RD∗,SM∗,IM∗,RM∗\)\\mathrm\{EE\}=\(S\_\{D\}^\{\*\},I\_\{D\}^\{\*\},R\_\{D\}^\{\*\},S\_\{M\}^\{\*\},I\_\{M\}^\{\*\},R\_\{M\}^\{\*\}\)withID∗\>0I\_\{D\}^\{\*\}\>0andIM∗\>0I\_\{M\}^\{\*\}\>0\.
###### Proof sketch\.
Substituting equilibrium conditions reduces the system to a scalar equationh\(ID∗\)=0h\(I\_\{D\}^\{\*\}\)=0\. We showh\(0\)<0h\(0\)<0whenR0\>1R\_\{0\}\>1andh→\+∞h\\to\+\\inftyasID∗→ΛD/μDI\_\{D\}^\{\*\}\\to\\Lambda\_\{D\}/\\mu\_\{D\}, so by the intermediate value theorem at least one positive root exists\. Numerical confirmation: of 200 random parameter configurations, 169 satisfy the preconditionR0\>1R\_\{0\}\>1; all 169 \(100%\) yield a unique positive root, and all 31 configurations withR0≤1R\_\{0\}\\leq 1correctly show no positive root\. Details in[Appendix˜A](https://arxiv.org/html/2606.05168#A1)\. ∎
###### Proposition 4\(Transcritical bifurcation\)\.
The system \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) undergoes a transcritical bifurcation atR0=1R\_\{0\}=1\. AsR0R\_\{0\}increases through 1, the DFE loses stability and exchanges it with the endemic equilibrium\.
###### Proof sketch\.
We apply Sotomayor’s theorem\[[2](https://arxiv.org/html/2606.05168#bib.bib2),[16](https://arxiv.org/html/2606.05168#bib.bib16)\], verifying the three transversality conditions atR0=1R\_\{0\}=1\. Numerical verification: of 200 random configurations, 196 \(98%\) show the expected transcritical behavior—endemic equilibrium present whenR0\>1R\_\{0\}\>1and absent whenR0<1R\_\{0\}<1\. The 4 borderline cases \(R0∈\[0\.99,1\.01\]R\_\{0\}\\in\[0\.99,1\.01\]\) fall within the numerical tolerance of the root finder near the bifurcation point\. Details in[Appendix˜A](https://arxiv.org/html/2606.05168#A1)\. ∎
[Figure˜2](https://arxiv.org/html/2606.05168#S3.F2)shows the ODE trajectory under baseline parameters \(R0=2\.62R\_\{0\}=2\.62\), illustrating convergence to the endemic equilibrium with bothID∗I\_\{D\}^\{\*\}andIM∗I\_\{M\}^\{\*\}positive\.
Figure 2:ODE trajectory under baseline parameters \(R0=2\.62R\_\{0\}=2\.62\)\. Both data and model infection fractions converge to a positive endemic equilibrium, consistent with[Proposition˜3](https://arxiv.org/html/2606.05168#Thmtheorem3)\. The initial transient captures the early growth phase; damped oscillations are visible before settling\.
### 3\.5SIRS Extension \(Primary Recommendation\)
The base SIR assumes permanent immunity, but cleaned corpora are re\-contaminated and retrained models re\-exposed\. We add waning immunity at rateδ\>0\\delta\>0\(RD→SDR\_\{D\}\\to S\_\{D\},RM→SMR\_\{M\}\\to S\_\{M\}\)\. Sinceδ\\deltaaffects onlyR→SR\\to Stransitions \(not the infected subsystem linearized at the DFE\),R0R\_\{0\}is*identical*for SIR and SIRS: all threshold results \(Theorems 1–2, Propositions 3–4\) apply directly\. The key difference is the endemic equilibrium’s approach dynamics:
###### Proposition 5\(SIRS oscillatory dynamics\)\.
Whenδ\>0\\delta\>0andR0\>1R\_\{0\}\>1, the SIRS variant can exhibit damped oscillatory convergence toward the endemic equilibrium\.
Verified numerically on 50 random SIRS configurations \([Section˜A\.5](https://arxiv.org/html/2606.05168#A1.SS5)\)\. This predicts periodic contamination resurgence even after cleanup campaigns\.
#### Why bilayer?
A single\-population model treats contamination as spreading within one pool\. The bilayer’s*cross\-layer coupling*makes explicit that interventions on one layer \(e\.g\., data detection\) reduce infection in the other \(models\): the geometric\-meanR0R\_\{0\}can be driven subcritical by acting on*either*layer \([Appendix˜B](https://arxiv.org/html/2606.05168#A2)\)\. This cross\-layer leverage is the basis for the intervention analysis \([Section˜7](https://arxiv.org/html/2606.05168#S7)\)\.
#### Source\-diversity hypothesis\.
WhenK\>1K\>1models contribute to the contaminated pool, we*hypothesize*that heterogeneous sources attenuate effective contamination\. We model this asβMeff\(K\)=βM/f\(K\)\\beta\_\{M\}^\{\\text\{eff\}\}\(K\)=\\beta\_\{M\}/f\(K\)withf\(1\)=1f\(1\)=1andffincreasing, yieldingR0\(K\)=R0\(1\)/f\(K\)R\_\{0\}\(K\)=R\_\{0\}\(1\)/\\sqrt\{f\(K\)\}\. This is an auxiliary modeling assumption motivated by heterogeneous mixing in epidemic theory\[[6](https://arxiv.org/html/2606.05168#bib.bib6)\], not a derivation from the base ODE; the matched\-budget experiment in[Section˜6\.5](https://arxiv.org/html/2606.05168#S6.SS5)provides a direct empirical test \([Appendix˜C](https://arxiv.org/html/2606.05168#A3)\)\.
## 4Calibration and Sensitivity Analysis
We calibrate the model using public data on AI text prevalence and development practices\. The calibration is*scenario\-based and illustrative*: it shows supercritical dynamics are consistent with current trends, not that the ecosystem’sR0R\_\{0\}has been measured\.
### 4\.1Parameter Estimation
Each parameter is derived from public data under simplifying assumptions:
- •βD≈0\.217\\beta\_\{D\}\\approx 0\.217: Log\-linear regression on 6 AI text prevalence data points \(2023–2025\)\[[22](https://arxiv.org/html/2606.05168#bib.bib22),[11](https://arxiv.org/html/2606.05168#bib.bib11)\], using the approximate relationβD≈r/12\+γD\+μD\\beta\_\{D\}\\approx r/12\+\\gamma\_\{D\}\+\\mu\_\{D\}from SIR linearization \(adequate for scenario\-level estimation; the bilayer growth rate depends jointly on all parameters\)\. CI:\[0\.201,0\.231\]\[0\.201,0\.231\]\. The baseline scenario usesβD=0\.216\\beta\_\{D\}=0\.216\(rounded\)\.
- •γD=0\.099\\gamma\_\{D\}=0\.099: Order\-of\-magnitude product of detection recall \(∼\\sim85%\[[21](https://arxiv.org/html/2606.05168#bib.bib21)\]\) and deployment coverage \(∼\\sim12% of platforms\)\. Actual removal efficacy depends on recall, latency, and platform pipelines—none quantified at ecosystem scale\.
- •βM=0\.340\\beta\_\{M\}=0\.340: Product of training frequency \(fraction of models retrained per month\) and exposure probability \(fraction of training data from web sources\)\.
- •γM=0\.060\\gamma\_\{M\}=0\.060: Estimated clean\-retraining rate \(fraction of contaminated models retrained on verified\-clean data per month\)\.
- •μD=0\.02\\mu\_\{D\}=0\.02,μM=0\.03\\mu\_\{M\}=0\.03: Data obsolescence and model retirement rates, respectively\.
- •ΛD=5\.0\\Lambda\_\{D\}=5\.0,ΛM=3\.0\\Lambda\_\{M\}=3\.0: Creation rates \(corpora and models per month\)\.
### 4\.2Scenario Analysis
We define three scenarios spanning the plausible parameter range \([Table˜2](https://arxiv.org/html/2606.05168#S4.T2)\)\. All three yieldR0\>1R\_\{0\}\>1, though the optimistic scenario is only marginally supercritical\.
Table 2:Calibration scenarios\. All scenarios yieldR0\>1R\_\{0\}\>1\. These are illustrative estimates, not precise ecosystem measurements\. Data sources: AI text prevalence\[[22](https://arxiv.org/html/2606.05168#bib.bib22),[11](https://arxiv.org/html/2606.05168#bib.bib11)\]; detection accuracy\[[21](https://arxiv.org/html/2606.05168#bib.bib21)\]; training practices from public model documentation\.[Figure˜3](https://arxiv.org/html/2606.05168#S4.F3)showsR0R\_\{0\}as a function ofβD\\beta\_\{D\}andβM\\beta\_\{M\}with the three scenario points overlaid\. TheR0=1R\_\{0\}=1contour separates subcritical from supercritical regimes; all three scenarios lie above it\.
Figure 3:R0R\_\{0\}as a function of data infection rate \(βD\\beta\_\{D\}\) and model infection rate \(βM\\beta\_\{M\}\), with other parameters at baseline values\. The thick contour marksR0=1R\_\{0\}=1\. All three calibration scenarios \(markers\) lie in the supercritical region, though the optimistic scenario is near the boundary\.#### Uncertainty propagation\.
The Sobol sensitivity analysis \([Section˜4\.3](https://arxiv.org/html/2606.05168#S4.SS3)\) uses Saltelli’s quasi\-random sampling scheme withN=512N=512base samples over uniform ranges for all 6 parameters, producingN\(2k\+2\)=7,168N\(2k\+2\)=7\{,\}168total evaluations\. Over these samples, we obtain meanR0=2\.94R\_\{0\}=2\.94, median2\.652\.65, andℙ\(R0\>1\)=98\.2%\\mathbb\{P\}\(R\_\{0\}\>1\)=98\.2\\%\. This probability is conditional on assumed parameter ranges—it quantifies the fraction of plausible scenarios that are supercritical, not a frequentist statement about the real ecosystem\.
### 4\.3Sensitivity Analysis
We perform variance\-based global sensitivity analysis using Sobol indices\[[18](https://arxiv.org/html/2606.05168#bib.bib18)\]over all 6 ODE parameters \(βD\\beta\_\{D\},γD\\gamma\_\{D\},βM\\beta\_\{M\},γM\\gamma\_\{M\},μD\\mu\_\{D\},μM\\mu\_\{M\}\) withN=512N=512base samples to identify which parameters most strongly influenceR0R\_\{0\}\. Total\-order sensitivity indices \(STS\_\{T\}\) are:
The detection and removal of contaminated data \(γD\\gamma\_\{D\}\) is the single most influential parameter, followed by the data contamination rate \(βD\\beta\_\{D\}\) and model recovery rate \(γM\\gamma\_\{M\}\)\. Turnover rates \(μD\\mu\_\{D\},μM\\mu\_\{M\}\) contribute negligibly \(ST<0\.04S\_\{T\}<0\.04\)\. This is a*model\-conditional*finding: it holds if the bilayer structure and assumed parameter ranges are approximately correct\. The sensitivity ranking is a property ofR0R\_\{0\}’s algebraic form and the assumed ranges, not an empirically measured elasticity\. Full Sobol analysis details appear in[Appendix˜E](https://arxiv.org/html/2606.05168#A5)\.
## 5Stochastic Consistency Check
The ODE system \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) assumes homogeneous mixing\. We test this via a stochastic agent\-based model \(ABM\) on an explicit bipartite network—a*consistency check*, not independent validation\.
#### Design\.
The ABM operates on a bipartite random graph \(\|𝒟\|=100\|\\mathcal\{D\}\|\{=\}100data nodes,\|ℳ\|=50\|\\mathcal\{M\}\|\{=\}50model nodes, edge probabilitypedge=0\.8p\_\{\\text\{edge\}\}\{=\}0\.8\)\. Each susceptible data node is infected with probabilityβD\\beta\_\{D\}times the traffic\-weighted fraction of infected model neighbors; each susceptible model node with probabilityβM\\beta\_\{M\}times the fraction of infected data neighbors\. Recovery, turnover, and immunity waning follow per\-node Bernoulli trials\. We run 20 realizations per configuration\.
#### ODE–ABM agreement\.
Under baseline parameters \([Figure˜4](https://arxiv.org/html/2606.05168#S5.F4)\), agreement is strong:R2=0\.976R^\{2\}=0\.976\(data layer\),R2=0\.968R^\{2\}=0\.968\(model layer\), NRMSE of 0\.044 and 0\.054\.
Figure 4:ODE \(solid\) vs\. ABM ensemble mean \(dashed,±1\\pm 1std shaded\) for infection fractions atpedge=0\.8p\_\{\\text\{edge\}\}=0\.8\.R2=0\.976R^\{2\}=0\.976\(data\),0\.9680\.968\(model\)\.
#### Where mean\-field breaks down\.
Aspedgep\_\{\\text\{edge\}\}decreases from 0\.9 to 0\.3,R2R^\{2\}degrades from 0\.99 to 0\.84, with non\-monotonic behavior atpedge=0\.7p\_\{\\text\{edge\}\}=0\.7\(R2=0\.61R^\{2\}=0\.61\)\. Superspreader nodes \(20% of models with10×10\\timestraffic\) degrade agreement toR2=0\.29R^\{2\}=0\.29; active detectors break it further\.
#### Threshold verification\.
A 20\-point sweep acrossβD∈\[0\.01,0\.60\]\\beta\_\{D\}\\in\[0\.01,0\.60\]\(spanning subcritical to supercritical regimes\), with 15 realizations each, correctly identifies the qualitative regime in 18/20 cases; the two errors occur nearR0≈1R\_\{0\}\\approx 1\.
The ODE is a useful approximation for dense, homogeneous ecosystems but should not be applied without modification to settings with strong structural heterogeneity\. The dense regime is a reasonable first approximation given that dominant web\-crawl corpora \(Common Crawl, C4\) are shared across most training pipelines, though emerging proprietary data pipelines may push toward sparser regimes\.
## 6Empirical Bridge: GPT\-2 Contamination Chains
We now ask whether actual LLM contamination chains exhibit dynamics qualitatively consistent with the SIR threshold picture\. This section provides an*empirical bridge*, not a model fit: we do not measureS/I/RS/I/Rcompartment fractions or fit ODE parameters to trajectories\. Instead, single\-chain experiments \([Section˜6\.1](https://arxiv.org/html/2606.05168#S6.SS1)–[Section˜6\.4](https://arxiv.org/html/2606.05168#S6.SS4)\) test qualitative predictions—dose\-response, near\-plateau at partial contamination, supercritical growth at full contamination—while matched\-budget experiments \([Section˜6\.5](https://arxiv.org/html/2606.05168#S6.SS5)\) test the bilayer\-specific prediction that multi\-source mixing attenuates collapse\. All experiments use GPT\-2 \(124M\); whether these patterns hold at larger scale is an open question \([Section˜7](https://arxiv.org/html/2606.05168#S7)\)\.
### 6\.1Experimental Design
We construct contamination chains using GPT\-2 \(124M parameters\)\[[17](https://arxiv.org/html/2606.05168#bib.bib17)\]on two domains: WikiText\-103\[[13](https://arxiv.org/html/2606.05168#bib.bib13)\]\(encyclopedic text\) and Tiny Shakespeare \(literary/archaic English\)\.
#### Chain construction\.
GenerationG0G\_\{0\}is fine\-tuned onn=2,000n=2\{,\}000real samples \(WikiText\) orn=1,500n=1\{,\}500\(Shakespeare\)\. For each subsequent generationGtG\_\{t\}\(t=1,…,7t=1,\\ldots,7\), we: \(i\) samplenntexts fromGt−1G\_\{t\-1\}; \(ii\) mix them with fresh real samples at contamination fractionα∈\{0\.00,0\.25,0\.50,0\.75,1\.00\}\\alpha\\in\\\{0\.00,0\.25,0\.50,0\.75,1\.00\\\}; \(iii\) fine\-tune a fresh GPT\-2 checkpoint on the mixture\. Theα=0\\alpha=0chain serves as a*size\-matched control*: it undergoes the same training budget and procedure with purely real data, isolating contamination effects from training\-budget drift\.
#### Protocol\.
Each\(α,generation\)\(\\alpha,\\text\{generation\}\)pair is run with 3 independent seeds\. Training uses 3 epochs, batch size 8, learning rate5×10−55\\times 10^\{\-5\}, and maximum sequence length 128 tokens\. Evaluation computes perplexity on 500 held\-out real samples\. Single\-chain total: 120 runs \(WikiText:5×3×85\\times 3\\times 8\)\+\+72 runs \(Shakespeare:3×3×83\\times 3\\times 8\)=192=192fine\-tuning runs\. The source\-diversity experiments \([Section˜6\.5](https://arxiv.org/html/2606.05168#S6.SS5)\) add 1,088 runs\.
### 6\.2WikiText Results
[Figure˜9](https://arxiv.org/html/2606.05168#A7.F9)\(appendix\) shows perplexity trajectories across 8 generations;[Table˜3](https://arxiv.org/html/2606.05168#S6.T3)summarizes the key numbers\.
Table 3:Empirical results: excess perplexity over theα=0\\alpha=0control \(WikiText and Shakespeare\)\. Growth raterris the slope of excess log\-perplexity ratiolog\(PPLα/PPLcontrol\)\\log\(\\mathrm\{PPL\}\_\{\\alpha\}/\\mathrm\{PPL\}\_\{\\text\{control\}\}\)vs\. generation\.ΔAIC\\Delta\\mathrm\{AIC\}compares linear growth vs\. plateau models \(negative favors growth\)\. For WikiText, all contaminated chains favor growth; the control favors plateau\. Shakespeare rows report excess PPL only \(growth rate and AIC not computed\)\.Domainα\\alphaG0 PPLG7 PPLExcess G1Excess G7rr95% CIΔAIC\\Delta\\mathrm\{AIC\}WikiText0\.0033\.5233\.47—0\.000\.000\[0\.000,0\.000\]\[0\.000,0\.000\]\+2\.0\+2\.00\.2533\.5234\.40\+0\.80\+0\.80\+0\.93\+0\.930\.003\[0\.0022,0\.0029\]\[0\.0022,0\.0029\]−12\.0\-12\.00\.5033\.5235\.91\+1\.87\+1\.87\+2\.44\+2\.440\.007\[0\.0065,0\.0074\]\[0\.0065,0\.0074\]−16\.8\-16\.80\.7533\.5238\.55\+3\.68\+3\.68\+5\.08\+5\.080\.014\[0\.014,0\.015\]\[0\.014,0\.015\]−19\.2\-19\.21\.0033\.52126\.92\+15\.49\+15\.49\+93\.45\+93\.450\.183\[0\.180,0\.186\]\[0\.180,0\.186\]−60\.2\-60\.2Shakespeare0\.0033\.6633\.63—0\.00———0\.5033\.6637\.60—\+4\.0\+4\.0———1\.0033\.66220\.77—\+187\.1\+187\.1———
#### Key findings\.
\(i\) Theα=0\\alpha=0control shows negligible drift \(PPL 33\.52→\\to33\.47\), indicating degradation is contamination\-specific, not a training\-budget artifact\. \(ii\) Excess perplexity increases monotonically withα\\alpha\(dose\-response\)\. \(iii\) Atα=1\.0\\alpha=1\.0, perplexity grows from 33\.52 to 126\.92 \(\+93\.45\+93\.45excess,r=0\.183r=0\.183\), consistent with supercritical dynamics\. \(iv\) Atα<1\\alpha<1, growth is slow \(r<0\.015r<0\.015\) but AIC favors linear growth over plateau, consistent with near\-critical dynamics\. These are qualitative monotonicity results; we do not claim to have identified a sharp critical threshold empirically\.
### 6\.3Cross\-Domain Replication
We replicate on Tiny Shakespeare \(α∈\{0\.0,0\.5,1\.0\}\\alpha\\in\\\{0\.0,0\.5,1\.0\\\},n=1,500n=1\{,\}500, 3 seeds;[Figure˜10](https://arxiv.org/html/2606.05168#A7.F10)in appendix\)\. Shakespeare shows stronger degradation \(6\.6×6\.6\\timesvs\.3\.8×3\.8\\timesatα=1\\alpha\{=\}1\), likely due to lower text redundancy\. The qualitative dose\-response pattern replicates across both domains\.
### 6\.4SIR Mapping and Statistical Analysis
The chain dynamics exhibit qualitative correspondence with SIR dynamics \(*phenomenological analogy*, not fitted parameters\):α\\alphaplays the role of transmission intensity;\(1−α\)\(1\{\-\}\\alpha\)the role of recovery \(real\-data dilution\); theα=0\\alpha\{=\}0control corresponds to the disease\-free state; and continued growth atα=1\\alpha\{=\}1to supercritical dynamics\. We stress that this is a*qualitative*mapping:α\\alphais a mixture fraction whileβ\\betaandγ\\gammaare dynamical rates, so the correspondence is structural rather than quantitative\.
Growth rates \([Table˜3](https://arxiv.org/html/2606.05168#S6.T3)\) use 95% bootstrap CIs \(1,000 resamples\)\.111Growth raterris computed on the control\-corrected ratiolog\(PPLα/PPLcontrol\)\\log\(\\text\{PPL\}\_\{\\alpha\}/\\text\{PPL\}\_\{\\text\{control\}\}\); uncorrected values are lower \(e\.g\., 0\.154 vs\. 0\.183 forα=1\.0\\alpha\{=\}1\.0\)\.AIC comparison favors plateau only forα=0\\alpha=0\(ΔAIC=\+2\.0\\Delta\\mathrm\{AIC\}=\+2\.0\); all contaminated chains favor growth, withα=1\.0\\alpha=1\.0overwhelmingly so \(ΔAIC=−60\.2\\Delta\\mathrm\{AIC\}=\-60\.2\)\. Beyond perplexity, Distinct\-2 bigram diversity drops from 0\.68 to 0\.38 underα=1\.0\\alpha=1\.0\([Figure˜11](https://arxiv.org/html/2606.05168#A7.F11)\), confirming support shrinkage as an independent degradation signal consistent with the “model collapse” characterization\.
### 6\.5Matched\-Budget Source\-Diversity Experiment
The single\-chain experiments show dose\-response degradation that any recursive\-training model would predict\. We design a matched\-budget ablation that isolates a*bilayer\-specific*variable—source diversityKK—testing whether multi\-model contamination dynamics differ from pure self\-training, as the source\-diversity hypothesis \([Section˜3\.5](https://arxiv.org/html/2606.05168#S3.SS5)\) predicts\.
#### Design\.
The synthetic pool has*fixed*sizen=2,000n\{=\}2\{,\}000; each ofKKmodels generates⌊n/K⌋\\lfloor n/K\\rfloorsamples \(remainder assigned to the last model, so the pool totals exactlynn\)\. We sweepK∈\{1,3,5\}K\\in\\\{1,3,5\\\}atα=1\.0\\alpha\{=\}1\.0with 8 seeds and 8 generations, plusα=0\\alpha\{=\}0control—640 runs total \([Section˜G\.8](https://arxiv.org/html/2606.05168#A7.SS8)\)\. Pool size, training budget,α\\alpha, and procedure are identical; onlyKKvaries\.
#### Results\.
[Figure˜5](https://arxiv.org/html/2606.05168#S6.F5)shows thatK=1K\{=\}1\(pure self\-training\) produces the most degradation, whileK=3K\{=\}3andK=5K\{=\}5show a modest reduction\. At generation 7,K=1K\{=\}1reaches excess PPL of\+87\.6\+87\.6,K=3K\{=\}3reaches\+85\.6\+85\.6, andK=5K\{=\}5reaches\+85\.6\+85\.6—a∼2\{\\sim\}2PPL attenuation \(Cohen’sd≈0\.8d\\approx 0\.8\)\. The effect is*not*strictly monotonic \(K=3≈K=5K\{=\}3\\approx K\{=\}5\), and statistical significance is borderline\. One\-sided tests \(pre\-specified direction:Ka\>KbK\_\{a\}\>K\_\{b\}, motivated by the source\-diversity hypothesis\):K=1K\{=\}1vs\.K=5K\{=\}5exact permutationp=0\.047p=0\.047, Wilcoxonp=0\.055p=0\.055; two\-sided: pairedt\(7\)=2\.16t\(7\)=2\.16,p=0\.068p=0\.068, bootstrap CI\[0\.3,3\.6\]\[0\.3,3\.6\]\.K=1K\{=\}1vs\.K=3K\{=\}3: one\-sidedp=0\.14p=0\.14;K=3K\{=\}3vs\.K=5K\{=\}5:p=0\.54p=0\.54\. We regard this as*suggestive*evidence, not strong confirmation\.
Figure 5:Matched\-budget source\-diversity experiment \(α=1\.0\\alpha\{=\}1\.0, pool size fixed atn=2,000n\{=\}2\{,\}000\)\.Left: perplexity trajectories forK∈\{1,3,5\}K\\in\\\{1,3,5\\\}and theα=0\\alpha\{=\}0control \(mean±\\pmstd, 8 seeds\)\.Right: G7 excess PPL vs\.KK\. Self\-training \(K=1K\{=\}1\) degrades most; multi\-source \(K\>1K\{\>\}1\) shows modest attenuation \(∼2\{\\sim\}2PPL,d≈0\.8d\\approx 0\.8\) that saturates byK=3K\{=\}3\.
#### Partial contamination \(α=0\.5\\alpha\{=\}0\.5\)\.
To test whether the diversity buffer persists at realistic contamination levels, we repeat the matched\-budget experiment withK∈\{1,5\}K\\in\\\{1,5\\\}atα=0\.5\\alpha\{=\}0\.5\(448 additional runs, 8 seeds\)\. Atα=0\.5\\alpha\{=\}0\.5, theK=1K\{=\}1vs\.K=5K\{=\}5difference is negligible:\+2\.20\+2\.20vs\.\+2\.18\+2\.18excess PPL \(mean diff=0\.02=0\.02,p=0\.61p=0\.61,d=0\.17d=0\.17, 3/8 seeds consistent\)\. The 50% real\-data component already breaks the self\-reinforcing loop, eliminating any detectable benefit from source diversity\.
#### Interpretation\.
The combined evidence is clear on one point and suggestive on another\.*Clear*: contamination fractionα\\alphais the dominant driver of collapse; atα=0\.5\\alpha\{=\}0\.5, source diversity has no detectable effect\.*Suggestive*: atα=1\\alpha\{=\}1, multi\-source mixing provides a modest buffer \(∼2\{\\sim\}2PPL,d≈0\.8d\\approx 0\.8,p=0\.047p=0\.047one\-sided\), consistent with heterogeneous mixing in epidemic models\[[6](https://arxiv.org/html/2606.05168#bib.bib6)\], but this requires replication with larger models and more seeds to be considered confirmed\. The practical message is robust: real\-world interventions should prioritize detection and filtering \(reducingα\\alpha\) over ecosystem diversity\.
## 7Intervention Analysis and Conclusion
We evaluate six strategies as an*illustrative scenario exercise*\([Table˜6](https://arxiv.org/html/2606.05168#A8.T6)in[Appendix˜H](https://arxiv.org/html/2606.05168#A8)\): each modeled as a parameter change, swept at 9 intensities in 15 pairwise combinations \(135 evaluations\)\. Rankings reflect the model’s algebraic structure, not empirically measured effect sizes\. Under baseline parameters, only watermark\-based filtering and herd immunity achieveR0<1R\_\{0\}<1alone; roughly one\-third of pairwise evaluations reach subcritical dynamics, with detection combinations dominating—consistent withγD\\gamma\_\{D\}as highest\-leverage \([Section˜4\.3](https://arxiv.org/html/2606.05168#S4.SS3)\)\.
#### Limitations\.
\(i\) Stylized mean\-field ODE; degrades under heterogeneity \([Section˜5](https://arxiv.org/html/2606.05168#S5)\)\. A simpler recursion model might explain some empirical trends; the bilayer framework’s added value is cross\-layer coupling and intervention structure\. \(ii\) All numbers \(R0=2\.62R\_\{0\}\{=\}2\.62, 98\.2%, 63%\) are conditional scenario outputs, not ecosystem measurements\. \(iii\) GPT\-2 124M only; scaling to 7B\+ and realistic continual\-pretraining is needed\. \(iv\) SIR mapping is phenomenological; experiments do not measure compartment fractions or fit ODE parameters\. \(v\) Three seeds per single\-chain condition; source\-diversity effect rests on borderline one\-sidedp=0\.047p\{=\}0\.047\. Future: larger models, data\-driven calibration, direct measurement of contamination state variables\.
#### Takeaway\.
The epidemic framing provides a principled vocabulary—R0R\_\{0\}, herd immunity, intervention thresholds—for reasoning about AI ecosystem contamination\. The matched\-budget experiments yield a clear practical message:*reducing contamination fraction \(detection and filtering\) is far more effective than diversifying contamination sources*\. Source diversity shows only a modest buffer atα=1\\alpha\{=\}1and vanishes entirely atα=0\.5\\alpha\{=\}0\.5, confirming thatβD\\beta\_\{D\}andγD\\gamma\_\{D\}—not ecosystem structure—are the intervention levers that matter\.
## References
- \[1\]Sina Alemohammad, Josue Casco\-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, and Richard G Baraniuk\.Self\-Consuming Generative Models Go MAD\.In*Proceedings of the International Conference on Learning Representations \(ICLR\)*, 2024\.
- \[2\]Carlos Castillo\-Chavez and Baojun Song\.Dynamical Models of Tuberculosis and Their Applications\.*Mathematical Biosciences and Engineering*, volume 1, pp\. 361–404, 2004\.
- \[3\]Odo Diekmann, J A P Heesterbeek, and Johan A J Metz\.On the Definition and the Computation of the Basic Reproduction RatioR0R\_\{0\}in Models for Infectious Diseases in Heterogeneous Populations\.*Journal of Mathematical Biology*, volume 28, pp\. 365–382, 1990\.
- \[4\]Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, and Horia Mania\.Strong Model Collapse\.In*Proceedings of the International Conference on Learning Representations \(ICLR\)*, 2025\.
- \[5\]Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, et al\.Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data\.*arXiv preprint arXiv:2404\.01413*, 2024\.
- \[6\]Herbert W Hethcote\.The Mathematics of Infectious Diseases\.*SIAM Review*, volume 42, pp\. 599–653, 2000\.
- \[7\]Fang Jin, Edward Dougherty, Parang Saraf, Yang Cao, and Naren Ramakrishnan\.Epidemiological Modeling of News and Rumors on Twitter\.*Proceedings of the Workshop on Social Network Mining and Analysis*, pp\. 1–9, 2013\.
- \[8\]Jeffrey O Kephart and Steve R White\.Measuring and Modeling Computer Virus Prevalence\.*Proceedings of the IEEE Symposium on Security and Privacy*, pp\. 2–15, 1993\.
- \[9\]William Ogilvy Kermack and Anderson G McKendrick\.A Contribution to the Mathematical Theory of Epidemics\.*Proceedings of the Royal Society of London\. Series A*, volume 115, pp\. 700–721, 1927\.
- \[10\]John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein\.A Watermark for Large Language Models\.In*Proceedings of the International Conference on Machine Learning \(ICML\)*, 2023\.
- \[11\]Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou\.Monitoring AI\-Modified Content at Scale\.2024\.
- \[12\]Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeid, Damien Chan, Andrea Madotto, Colin Raffel, and Harm de Vries\.A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, and Toxicity\.*arXiv preprint arXiv:2305\.13169*, 2024\.
- \[13\]Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher\.Pointer Sentinel Mixture Models\.*arXiv preprint arXiv:1609\.07843*, 2016\.
- \[14\]Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru\.Model Cards for Model Reporting\.In*Proceedings of the Conference on Fairness, Accountability, and Transparency \(FAT\*\)*, pp\. 220–229, 2019\.
- \[15\]Romualdo Pastor\-Satorras, Claudio Castellano, Piet Van Mieghem, and Alessandro Vespignani\.Epidemic Processes in Complex Networks\.*Reviews of Modern Physics*, volume 87, pp\. 925–979, 2015\.
- \[16\]Lawrence Perko\.*Differential Equations and Dynamical Systems*, Springer, 2001\.
- \[17\]Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever\.Language Models are Unsupervised Multitask Learners\.*OpenAI Blog*, 2019\.
- \[18\]Andrea Saltelli, Paola Annoni, Ivano Azzini, Francesca Campolongo, Marco Ratto, and Stefano Tarantola\.Variance Based Sensitivity Analysis of Model Output\. Design and Estimator for the Total Sensitivity Index\.*Computer Physics Communications*, volume 181, pp\. 259–270, 2010\.
- \[19\]Mohamed El Amine Seddik, Suei\-Hai Chen, Soufiane Hayou, Pierre Youssef, and Merouane Debbah\.How Bad is Training on Synthetic Data? A Statistical Analysis\.In*Proceedings of the International Conference on Machine Learning \(ICML\)*, 2024\.
- \[20\]Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal\.AI models collapse when trained on recursively generated data\.*Nature*, volume 631, pp\. 755–759, 2024\.
- \[21\]Ruixiang Tang, Yu\-Neng Chuang, and Xia Hu\.The Science of Detecting LLM\-Generated Text\.*Communications of the ACM*, volume 67, pp\. 81–90, 2024\.
- \[22\]Neil C Thompson and Shuning Ge\.AI\-Generated Content Prevalence in Web Corpora\.2024\.
- \[23\]Pauline van den Driessche and James Watmough\.Reproduction Numbers and Sub\-threshold Endemic Equilibria for Compartmental Models of Disease Transmission\.*Mathematical Biosciences*, volume 180, pp\. 29–48, 2002\.
- \[24\]Soroush Vosoughi, Deb Roy, and Sinan Aral\.The Spread of True and False News Online\.*Science*, volume 359, pp\. 1146–1151, 2018\.
## Appendix AProof Sketches and Numerical Evidence
### A\.1Proof of[Theorem˜1](https://arxiv.org/html/2606.05168#Thmtheorem1): Basic Reproduction Number
We apply the Next Generation Matrix \(NGM\) method of van den Driessche and Watmough\[[23](https://arxiv.org/html/2606.05168#bib.bib23)\]\. The infected compartments arex=\(ID,IM\)Tx=\(I\_\{D\},I\_\{M\}\)^\{T\}\. We decomposex˙=\(ℱ−𝒱\)x\\dot\{x\}=\(\\mathcal\{F\}\-\\mathcal\{V\}\)xnear the DFE, whereℱ\\mathcal\{F\}represents new infections and𝒱\\mathcal\{V\}represents transitions out of infected compartments\.
At the DFE,SD∗=ΛD/μDS\_\{D\}^\{\*\}=\\Lambda\_\{D\}/\\mu\_\{D\},SM∗=ΛM/μMS\_\{M\}^\{\*\}=\\Lambda\_\{M\}/\\mu\_\{M\},ND∗=ΛD/μDN\_\{D\}^\{\*\}=\\Lambda\_\{D\}/\\mu\_\{D\}, andNM∗=ΛM/μMN\_\{M\}^\{\*\}=\\Lambda\_\{M\}/\\mu\_\{M\}\. The linearized new\-infection and transition matrices are:
F\\displaystyle F=\(0βDSD∗NM∗βMSM∗ND∗0\)=\(0βDΛDμMμDΛMβMΛMμDμMΛD0\),\\displaystyle=\\begin\{pmatrix\}0&\\beta\_\{D\}\\frac\{S\_\{D\}^\{\*\}\}\{N\_\{M\}^\{\*\}\}\\\\\[4\.0pt\] \\beta\_\{M\}\\frac\{S\_\{M\}^\{\*\}\}\{N\_\{D\}^\{\*\}\}&0\\end\{pmatrix\}=\\begin\{pmatrix\}0&\\beta\_\{D\}\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}\\\\\[4\.0pt\] \\beta\_\{M\}\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}&0\\end\{pmatrix\},\(10\)V\\displaystyle V=\(γD\+μD00γM\+μM\)\.\\displaystyle=\\begin\{pmatrix\}\\gamma\_\{D\}\+\\mu\_\{D\}&0\\\\ 0&\\gamma\_\{M\}\+\\mu\_\{M\}\\end\{pmatrix\}\.\(11\)
The next\-generation matrix is:
K=FV−1=\(0βDγM\+μM⋅ΛDμMμDΛMβMγD\+μD⋅ΛMμDμMΛD0\)\.K=FV^\{\-1\}=\\begin\{pmatrix\}0&\\frac\{\\beta\_\{D\}\}\{\\gamma\_\{M\}\+\\mu\_\{M\}\}\\cdot\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}\\\\\[4\.0pt\] \\frac\{\\beta\_\{M\}\}\{\\gamma\_\{D\}\+\\mu\_\{D\}\}\\cdot\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}&0\\end\{pmatrix\}\.\(12\)
The eigenvalues ofKKsatisfyλ2=βDβM\(γD\+μD\)\(γM\+μM\)⋅ΛDμMμDΛM⋅ΛMμDμMΛD=βDβM\(γD\+μD\)\(γM\+μM\)\\lambda^\{2\}=\\frac\{\\beta\_\{D\}\\beta\_\{M\}\}\{\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\}\\cdot\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}\\cdot\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}=\\frac\{\\beta\_\{D\}\\beta\_\{M\}\}\{\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\}, where the cross\-population ratios cancel exactly\. By definition,R0=ρ\(K\)R\_\{0\}=\\rho\(K\), the spectral radius:
R0=βD⋅βM\(γD\+μD\)\(γM\+μM\)\.R\_\{0\}=\\sqrt\{\\frac\{\\beta\_\{D\}\\cdot\\beta\_\{M\}\}\{\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\}\}\.\(13\)
Numerical verification: the formula matches the spectral radius computed vianumpy\.linalg\.eigvalsto machine precision \(<10−10<10^\{\-10\}relative error\) across all 200 test configurations\.
#### Implementation note\.
The code computes the NGM usingFFwith off\-diagonal entriesβD\\beta\_\{D\}andβM\\beta\_\{M\}directly, omitting the cross\-population ratiosΛDμM/\(μDΛM\)\\Lambda\_\{D\}\\mu\_\{M\}/\(\\mu\_\{D\}\\Lambda\_\{M\}\)andΛMμD/\(μMΛD\)\\Lambda\_\{M\}\\mu\_\{D\}/\(\\mu\_\{M\}\\Lambda\_\{D\}\)\. This is spectrally equivalent because these ratios cancel in the product of off\-diagonal entries ofFV−1FV^\{\-1\}, so both forms yield the sameR0R\_\{0\}\. The paper presents the fullFFwith ratios for pedagogical clarity; the code uses the simplified form for computational efficiency\. ∎
### A\.2Proof of[Theorem˜2](https://arxiv.org/html/2606.05168#Thmtheorem2): DFE Stability
Existence\.SettingID=IM=0I\_\{D\}=I\_\{M\}=0in \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\), the remaining equations yieldSD∗=ΛD/μDS\_\{D\}^\{\*\}=\\Lambda\_\{D\}/\\mu\_\{D\},RD∗=0R\_\{D\}^\{\*\}=0,SM∗=ΛM/μMS\_\{M\}^\{\*\}=\\Lambda\_\{M\}/\\mu\_\{M\},RM∗=0R\_\{M\}^\{\*\}=0\.
Local stability\.The Jacobian of the full system evaluated at the DFE is:
JDFE=\(−μD000−βDSD∗NM∗00−\(γD\+μD\)00βDSD∗NM∗00γD−μD0000−βMSM∗ND∗0−μM000βMSM∗ND∗00−\(γM\+μM\)00000γM−μM\)\.J\_\{\\mathrm\{DFE\}\}=\\begin\{pmatrix\}\-\\mu\_\{D\}&0&0&0&\-\\beta\_\{D\}\\frac\{S\_\{D\}^\{\*\}\}\{N\_\{M\}^\{\*\}\}&0\\\\ 0&\-\(\\gamma\_\{D\}\+\\mu\_\{D\}\)&0&0&\\beta\_\{D\}\\frac\{S\_\{D\}^\{\*\}\}\{N\_\{M\}^\{\*\}\}&0\\\\ 0&\\gamma\_\{D\}&\-\\mu\_\{D\}&0&0&0\\\\ 0&\-\\beta\_\{M\}\\frac\{S\_\{M\}^\{\*\}\}\{N\_\{D\}^\{\*\}\}&0&\-\\mu\_\{M\}&0&0\\\\ 0&\\beta\_\{M\}\\frac\{S\_\{M\}^\{\*\}\}\{N\_\{D\}^\{\*\}\}&0&0&\-\(\\gamma\_\{M\}\+\\mu\_\{M\}\)&0\\\\ 0&0&0&0&\\gamma\_\{M\}&\-\\mu\_\{M\}\\end\{pmatrix\}\.\(14\)
The eigenvalues include−μD\-\\mu\_\{D\}\(multiplicity 2\),−μM\-\\mu\_\{M\}\(multiplicity 2\), and the two eigenvalues of the2×22\\times 2infected\-subsystem block\. Reading off the\(ID,IM\)\(I\_\{D\},I\_\{M\}\)rows ofJDFEJ\_\{\\mathrm\{DFE\}\}, this block is:
B=\(−\(γD\+μD\)βDSD∗NM∗βMSM∗ND∗−\(γM\+μM\)\)=\(−\(γD\+μD\)βDΛDμMμDΛMβMΛMμDμMΛD−\(γM\+μM\)\)\.B=\\begin\{pmatrix\}\-\(\\gamma\_\{D\}\+\\mu\_\{D\}\)&\\beta\_\{D\}\\frac\{S\_\{D\}^\{\*\}\}\{N\_\{M\}^\{\*\}\}\\\\\[4\.0pt\] \\beta\_\{M\}\\frac\{S\_\{M\}^\{\*\}\}\{N\_\{D\}^\{\*\}\}&\-\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\\end\{pmatrix\}=\\begin\{pmatrix\}\-\(\\gamma\_\{D\}\+\\mu\_\{D\}\)&\\beta\_\{D\}\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}\\\\\[4\.0pt\] \\beta\_\{M\}\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}&\-\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\\end\{pmatrix\}\.\(15\)Note that the off\-diagonal entries carry the cross\-population ratiosSD∗/NM∗S\_\{D\}^\{\*\}/N\_\{M\}^\{\*\}andSM∗/ND∗S\_\{M\}^\{\*\}/N\_\{D\}^\{\*\}, consistent with theFFmatrix above\.
The eigenvalues ofBBsatisfyλ2\+\(γD\+μD\+γM\+μM\)λ\+det\(B\)=0\\lambda^\{2\}\+\(\\gamma\_\{D\}\{\+\}\\mu\_\{D\}\{\+\}\\gamma\_\{M\}\{\+\}\\mu\_\{M\}\)\\lambda\+\\det\(B\)=0, with:
det\(B\)=\(γD\+μD\)\(γM\+μM\)−βDβM⋅ΛDμMμDΛM⋅ΛMμDμMΛD⏟=1\.\\det\(B\)=\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\-\\beta\_\{D\}\\beta\_\{M\}\\cdot\\underbrace\{\\frac\{\\Lambda\_\{D\}\\mu\_\{M\}\}\{\\mu\_\{D\}\\Lambda\_\{M\}\}\\cdot\\frac\{\\Lambda\_\{M\}\\mu\_\{D\}\}\{\\mu\_\{M\}\\Lambda\_\{D\}\}\}\_\{=1\}\.\(16\)The cross\-population ratios cancel in the product\. Both eigenvalues have negative real parts if and only ifdet\(B\)\>0\\det\(B\)\>0, which gives:
\(γD\+μD\)\(γM\+μM\)\>βDβM⇔R0<1\.\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\>\\beta\_\{D\}\\beta\_\{M\}\\iff R\_\{0\}<1\.\(17\)
Numerical verification: for all 200 random configurations \(100%\), the maximum real part of all Jacobian eigenvalues is negative whenR0<1R\_\{0\}<1and positive whenR0\>1R\_\{0\}\>1, with no exceptions\. ∎
### A\.3Proof of[Proposition˜3](https://arxiv.org/html/2606.05168#Thmtheorem3): Endemic Equilibrium Existence
At equilibrium, from \([1](https://arxiv.org/html/2606.05168#S3.E1)\)–\([6](https://arxiv.org/html/2606.05168#S3.E6)\) we can express all compartments in terms ofID∗I\_\{D\}^\{\*\}:
SD∗\\displaystyle S\_\{D\}^\{\*\}=ΛDβDIM∗/NM∗\+μD,\\displaystyle=\\frac\{\\Lambda\_\{D\}\}\{\\beta\_\{D\}I\_\{M\}^\{\*\}/N\_\{M\}^\{\*\}\+\\mu\_\{D\}\},\(18\)IM∗\\displaystyle I\_\{M\}^\{\*\}=βMID∗SM∗\(γM\+μM\)ND∗,\\displaystyle=\\frac\{\\beta\_\{M\}I\_\{D\}^\{\*\}S\_\{M\}^\{\*\}\}\{\(\\gamma\_\{M\}\+\\mu\_\{M\}\)N\_\{D\}^\{\*\}\},\(19\)SM∗\\displaystyle S\_\{M\}^\{\*\}=ΛMβMID∗/ND∗\+μM\.\\displaystyle=\\frac\{\\Lambda\_\{M\}\}\{\\beta\_\{M\}I\_\{D\}^\{\*\}/N\_\{D\}^\{\*\}\+\\mu\_\{M\}\}\.\(20\)
Substituting these into theID˙=0\\dot\{I\_\{D\}\}=0equation and simplifying yields a scalar equationh\(ID∗\)=0h\(I\_\{D\}^\{\*\}\)=0\. We verify:
- •h\(0\)=\(γD\+μD\)\(1−R02\)h\(0\)=\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(1\-R\_\{0\}^\{2\}\)\. WhenR0\>1R\_\{0\}\>1,h\(0\)<0h\(0\)<0\.
- •AsID∗→ΛD/μDI\_\{D\}^\{\*\}\\to\\Lambda\_\{D\}/\\mu\_\{D\}\(the maximum possible\),h\(ID∗\)→\+∞h\(I\_\{D\}^\{\*\}\)\\to\+\\inftybecause the susceptible pool is depleted\.
By the intermediate value theorem, there exists at least oneID∗∈\(0,ΛD/μD\)I\_\{D\}^\{\*\}\\in\(0,\\Lambda\_\{D\}/\\mu\_\{D\}\)withh\(ID∗\)=0h\(I\_\{D\}^\{\*\}\)=0\. Numerical verification: of 200 random parameter configurations, 169 satisfy the preconditionR0\>1R\_\{0\}\>1; all 169 \(100%\) yield exactly one positive root\. The remaining 31 configurations haveR0≤1R\_\{0\}\\leq 1and correctly show no positive root, consistent with the theorem\. This supports \(but does not prove\) uniqueness of the endemic equilibrium\. ∎
### A\.4Proof of[Proposition˜4](https://arxiv.org/html/2606.05168#Thmtheorem4): Transcritical Bifurcation
We apply Sotomayor’s theorem\[[2](https://arxiv.org/html/2606.05168#bib.bib2),[16](https://arxiv.org/html/2606.05168#bib.bib16)\]withβD\\beta\_\{D\}as the bifurcation parameter \(varyingβD\\beta\_\{D\}sweepsR0R\_\{0\}through 1\)\.
AtR0=1R\_\{0\}=1, the JacobianJDFEJ\_\{\\mathrm\{DFE\}\}has a simple zero eigenvalue\. Letwwbe the corresponding right eigenvector andvvthe left eigenvector\. The three Sotomayor conditions are:
1. 1\.vTfβD\(x0,βD∗\)≠0v^\{T\}f\_\{\\beta\_\{D\}\}\(x\_\{0\},\\beta\_\{D\}^\{\*\}\)\\neq 0, wherefβDf\_\{\\beta\_\{D\}\}is the derivative of the vector field with respect toβD\\beta\_\{D\}\.
2. 2\.vT\[DfβD\(x0,βD∗\)\]w≠0v^\{T\}\[Df\_\{\\beta\_\{D\}\}\(x\_\{0\},\\beta\_\{D\}^\{\*\}\)\]w\\neq 0\.
3. 3\.vT\[D2f\(x0,βD∗\)\(w,w\)\]≠0v^\{T\}\[D^\{2\}f\(x\_\{0\},\\beta\_\{D\}^\{\*\}\)\(w,w\)\]\\neq 0\.
We computewwandvvexplicitly from the zero\-eigenvalue structure ofBBatR0=1R\_\{0\}=1and verify all three conditions hold \(the products are nonzero rational functions of the parameters with no cancellations at generic parameter values\)\. This confirms a transcritical bifurcation: asR0R\_\{0\}increases through 1, the DFE loses stability and the endemic equilibrium emerges with positive infected fractions\.
Numerical verification: across 196 random configurations, we verify that the endemic equilibrium exists precisely whenR0\>1R\_\{0\}\>1and vanishes whenR0<1R\_\{0\}<1, consistent with transcritical exchange of stability between DFE and EE atR0=1R\_\{0\}=1\. ∎
### A\.5SIRS Oscillatory Dynamics \([Proposition˜5](https://arxiv.org/html/2606.05168#Thmtheorem5)\)
For the SIRS extension, we add terms\+δRD\+\\delta R\_\{D\}toSD˙\\dot\{S\_\{D\}\},−δRD\-\\delta R\_\{D\}toRD˙\\dot\{R\_\{D\}\}, and similarly\+δRM\+\\delta R\_\{M\}toSM˙\\dot\{S\_\{M\}\},−δRM\-\\delta R\_\{M\}toRM˙\\dot\{R\_\{M\}\}\.
#### Threshold invariance\.
The waning termsδRi\\delta R\_\{i\}do not appear in the infected\-compartment equations \(I˙D\\dot\{I\}\_\{D\},I˙M\\dot\{I\}\_\{M\}\), so the linearization at the DFE—and hence the next\-generation matrix,R0R\_\{0\}, and local DFE stability—is identical to the SIR case \(Theorems 1–2\)\. For the endemic equilibrium \(Proposition 3\): at equilibrium, the SIRS waning terms redistribute population betweenSSandRRbut the infected\-subsystem fixed\-point equationh\(ID∗\)=0h\(I\_\{D\}^\{\*\}\)=0retains the same structure \(withSD∗S\_\{D\}^\{\*\}now depending onδ\\delta\); the IVT argument still yields a positive root whenR0\>1R\_\{0\}\>1\. Numerical verification: all 50 SIRS configurations withR0\>1R\_\{0\}\>1converge to a unique positive endemic equilibrium\. Transcritical bifurcation \(Proposition 4\) atR0=1R\_\{0\}=1is verified numerically: all 50 SIRS configurations show the expected exchange of stability\.
#### Oscillatory dynamics\.
We generated 50 random SIRS configurations withδ∈\[0\.01,0\.2\]\\delta\\in\[0\.01,0\.2\]andR0\>1R\_\{0\}\>1\. For each, we integrated the ODE forT=500T=500time units and analyzed the infected trajectories\. All 50 configurations exhibit damped oscillatory convergence to the endemic equilibrium, with oscillation frequency increasing withδ\\delta\.
The oscillatory behavior arises because waning immunity replenishes the susceptible pool, creating overshoot\-undershoot cycles\. Representative trajectories are shown in[Figure˜6](https://arxiv.org/html/2606.05168#A1.F6)\.
Figure 6:SIRS oscillatory dynamics for a representative configuration \(δ=0\.1\\delta=0\.1,R0=2\.62R\_\{0\}=2\.62\)\. Both data and model infection fractions exhibit damped oscillations before converging to the endemic equilibrium\.
## Appendix BBilayer vs\. Single\-Layer Comparison
A single\-layer SIR model treats all entities \(data and models together\) as one population with a single transmission rateβ\\betaand recovery rateγ\\gamma, givingR01L=β/\(γ\+μ\)R\_\{0\}^\{\\text\{1L\}\}=\\beta/\(\\gamma\+\\mu\)\. The bilayer model separates data and model populations, yieldingR02L=βDβM/\[\(γD\+μD\)\(γM\+μM\)\]R\_\{0\}^\{\\text\{2L\}\}=\\sqrt\{\\beta\_\{D\}\\beta\_\{M\}/\[\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\]\}\.
The key structural difference is that the bilayerR0R\_\{0\}is a*geometric mean*over two layers with*independent*rate parameters\. This has a concrete consequence for intervention analysis: in the single\-layer model, reducingR0R\_\{0\}below 1 requiresβ<γ\+μ\\beta<\\gamma\+\\mu—there is only one “lever\.” In the bilayer model, the system can be driven subcritical by reducing*either*βD\\beta\_\{D\}orβM\\beta\_\{M\}\(or increasingγD\\gamma\_\{D\}orγM\\gamma\_\{M\}\), and the required reduction in any single parameter is smaller because it enters under a square root\. Concretely, at baseline parameters, halvingγD\\gamma\_\{D\}alone \(data detection\) reducesR02LR\_\{0\}^\{\\text\{2L\}\}by a factor of2\\sqrt\{2\}, whereas in a single\-layer model with equivalent aggregate parameters, the same intervention would need to be applied to the entire system\. This cross\-layer leverage—the ability to target the more tractable layer—is the bilayer model’s primary structural advantage and the basis for the intervention analysis in[Section˜7](https://arxiv.org/html/2606.05168#S7)\.
We do not claim the bilayer model fits the experimental data better than a single\-layer model \(neither is fitted\); the advantage is in the*intervention structure*, not in descriptive accuracy for the GPT\-2 experiments\.
## Appendix CSource\-Diversity Extension
WhenK\>1K\>1models contribute to the contaminated data pool, their outputs are drawn fromKKdifferent learned distributions, increasing the effective diversity of the contaminated corpus\. We model this by replacingβM\\beta\_\{M\}withβMeff\(K\)=βM/f\(K\)\\beta\_\{M\}^\{\\text\{eff\}\}\(K\)=\\beta\_\{M\}/f\(K\), wheref\(K\)≥1f\(K\)\\geq 1is a diversity attenuation factor satisfyingf\(1\)=1f\(1\)=1andf\(K\)→f∞f\(K\)\\to f\_\{\\infty\}asK→∞K\\to\\infty\. TheKK\-dependent reproduction number is:
R0\(K\)=βD⋅βM/f\(K\)\(γD\+μD\)\(γM\+μM\)=R0\(1\)f\(K\)\.R\_\{0\}\(K\)=\\sqrt\{\\frac\{\\beta\_\{D\}\\cdot\\beta\_\{M\}/f\(K\)\}\{\(\\gamma\_\{D\}\+\\mu\_\{D\}\)\(\\gamma\_\{M\}\+\\mu\_\{M\}\)\}\}=\\frac\{R\_\{0\}\(1\)\}\{\\sqrt\{f\(K\)\}\}\.\(21\)
This predicts monotonic attenuation: asKKincreases, source diversity grows, effective transmission decreases, andR0R\_\{0\}drops\. The functional form off\(K\)f\(K\)depends on how model diversity scales; a simple parameterization isf\(K\)=1\+clogKf\(K\)=1\+c\\log Kfor somec\>0c\>0, reflecting diminishing marginal diversity\. Our matched\-budget experiment \([Section˜6\.5](https://arxiv.org/html/2606.05168#S6.SS5)\) finds suggestive support atα=1\\alpha\{=\}1\(K=1K\{=\}1vs\.K\>1K\{\>\}1\) but the effect is not strictly monotonic and vanishes atα=0\.5\\alpha\{=\}0\.5; see[Section˜G\.8](https://arxiv.org/html/2606.05168#A7.SS8)for details\.
## Appendix DODE System Details
### D\.1Full Parameter Table
Table 4:Complete parameter table with units and baseline values\.
### D\.2DFE and EE Closed\-Form Expressions
The disease\-free equilibrium is:
DFE=\(ΛDμD,0,0,ΛMμM,0,0\)=\(250,0,0,100,0,0\)\.\\mathrm\{DFE\}=\\left\(\\frac\{\\Lambda\_\{D\}\}\{\\mu\_\{D\}\},\\;0,\\;0,\\;\\frac\{\\Lambda\_\{M\}\}\{\\mu\_\{M\}\},\\;0,\\;0\\right\)=\(250,\\;0,\\;0,\\;100,\\;0,\\;0\)\.\(22\)
At baseline parameters, the endemic equilibrium \(computed numerically\) is approximately:
EE≈\(82\.7,28\.1,139\.2,44\.0,18\.7,37\.4\)\.\\mathrm\{EE\}\\approx\(82\.7,\\;28\.1,\\;139\.2,\\;44\.0,\\;18\.7,\\;37\.4\)\.\(23\)
## Appendix ECalibration Details
### E\.1Data Sources
TheβD\\beta\_\{D\}estimate is based on 6 data points from AI text prevalence studies:
- •2023\-01:∼\\sim5% of web text AI\-generated\[[22](https://arxiv.org/html/2606.05168#bib.bib22)\]\.
- •2023\-06:∼\\sim10%\.
- •2024\-01:∼\\sim20%\.
- •2024\-05:∼\\sim37%\.
- •2025\-01:∼\\sim55%\.
- •2025\-04:∼\\sim74% \(projected estimate based on extrapolation from earlier data points; not an independent measurement\)\.
Log\-linear regression \(log\(f\)\\log\(f\)vs\. time in years, i\.e\.logf=logf0\+rt\\log f=\\log f\_\{0\}\+r\\,t\) yields an exponential growth raterrper year\. Converting to monthly SIR parameters:βD=r/12\+γD\+μD\\beta\_\{D\}=r/12\+\\gamma\_\{D\}\+\\mu\_\{D\}, where ther/12r/12term converts the yearly rate to monthly and the remaining terms account for ongoing recovery and turnover\. The resulting point estimate isβD≈0\.217\\beta\_\{D\}\\approx 0\.217\(the small difference from the rounded 0\.216 is within regression uncertainty\)\.
The 95% CI from regression standard error is\[0\.201,0\.231\]\[0\.201,0\.231\], reflecting uncertainty in the prevalence estimates\.
### E\.2Uncertainty Propagation via Sobol Sampling
Theℙ\(R0\>1\)=98\.2%\\mathbb\{P\}\(R\_\{0\}\>1\)=98\.2\\%reported in[Section˜4](https://arxiv.org/html/2606.05168#S4)is computed from the same Saltelli quasi\-random samples used for the Sobol sensitivity analysis below—not from a separate independent Monte Carlo\. Specifically, Saltelli’s sampling scheme withN=512N=512base samples over 6 parameters producesN\(2k\+2\)=7,168N\(2k\+2\)=7\{,\}168parameter vectors, each drawn from uniform distributions over the rangesβD∈\[0\.10,0\.70\]\\beta\_\{D\}\\in\[0\.10,0\.70\],γD∈\[0\.02,0\.25\]\\gamma\_\{D\}\\in\[0\.02,0\.25\],βM∈\[0\.15,0\.70\]\\beta\_\{M\}\\in\[0\.15,0\.70\],γM∈\[0\.02,0\.20\]\\gamma\_\{M\}\\in\[0\.02,0\.20\],μD∈\[0\.01,0\.05\]\\mu\_\{D\}\\in\[0\.01,0\.05\],μM∈\[0\.01,0\.06\]\\mu\_\{M\}\\in\[0\.01,0\.06\]\.R0R\_\{0\}is evaluated at each sample point; the reported statistics \(mean, median,ℙ\(R0\>1\)\\mathbb\{P\}\(R\_\{0\}\>1\)\) are computed over these 7,168 evaluations\.
### E\.3Sobol Sensitivity Analysis
Full Sobol indices \(first\-orderS1S\_\{1\}and total\-orderSTS\_\{T\}\) computed via Saltelli’s sampling scheme\[[18](https://arxiv.org/html/2606.05168#bib.bib18)\]withN=512N=512base samples over all 6 parameters:
Interactions are modest \(ST−S1<0\.06S\_\{T\}\-S\_\{1\}<0\.06\), indicating thatR0R\_\{0\}is dominated by main effects\. Turnover rates \(μD\\mu\_\{D\},μM\\mu\_\{M\}\) have negligible influence \(ST<0\.04S\_\{T\}<0\.04\)\. The full Sobol bar chart is shown in[Figure˜7](https://arxiv.org/html/2606.05168#A5.F7)\.
Figure 7:Sobol total\-order sensitivity indices forR0R\_\{0\}\. Data recovery rate \(γD\\gamma\_\{D\}\) has the highest leverage\.
## Appendix FABM Implementation Details
### F\.1Network Construction
The bipartite graph is constructed using NetworkX with\|𝒟\|=100\|\\mathcal\{D\}\|=100data nodes and\|ℳ\|=50\|\\mathcal\{M\}\|=50model nodes\. Each data\-model pair is connected independently with probabilitypedgep\_\{\\text\{edge\}\}\. Atpedge=0\.8p\_\{\\text\{edge\}\}=0\.8, each data node has on average 40 model neighbors, and each model node has on average 80 data neighbors\.
### F\.2Agent Rules
At each discrete time step:
1. 1\.Infection \(data nodes\): For each susceptible data node, infection pressure is the traffic\-weighted fraction of infected model neighbors:p=βD⋅\(∑infectedtraffic\)/\(∑alltraffic\)p=\\beta\_\{D\}\\cdot\(\\sum\_\{\\text\{infected\}\}\\text\{traffic\}\)/\(\\sum\_\{\\text\{all\}\}\\text\{traffic\}\)\. Infection occurs with probabilitypp\(Bernoulli trial\)\.
2. 2\.Infection \(model nodes\): For each susceptible model node, infection pressure is the fraction of infected data neighbors:p=βM⋅kI/kp=\\beta\_\{M\}\\cdot k^\{I\}/k, wherekIk^\{I\}is the number of infected data neighbors andkkis the total\.
3. 3\.Recovery: Each infected node recovers with probabilityγi\\gamma\_\{i\}per step\.
4. 4\.Turnover: Each node transitions to susceptible with probabilityμi\\mu\_\{i\}per step\.
5. 5\.Superspreaders: If enabled, a configurable fraction of model nodes have10×10\\timestraffic weight, amplifying their infection pressure on data nodes\.
6. 6\.Detectors: If enabled, detector nodes scan assigned neighbors and recover each infected neighbor with probabilityprecision×coverage\\text\{precision\}\\times\\text\{coverage\}per step\.
20 realizations are run per configuration for 50 time steps \(default\)\. Ensemble means and standard deviations are computed at each time step\.
### F\.3Robustness Sweep Results
Figure 8:ABM threshold verification: sub\-critical vs\. super\-critical classification across 20 parameter configurations\. 18/20 are correctly identified; the two errors occur nearR0≈1R\_\{0\}\\approx 1\.
## Appendix GEmpirical Experiment Details
### G\.1Single\-Chain Perplexity Trajectories
Figure 9:Perplexity over 8 generations for WikiText contamination chains at five contamination fractionsα\\alpha\. Lines show means across 3 seeds; shaded regions are±1\\pm 1std\. Theα=0\\alpha=0control \(gray\) remains flat\. Atα=1\.0\\alpha=1\.0, perplexity grows from 33\.52 to 126\.92 \(\+93\.45\+93\.45excess\)\. Atα<1\\alpha<1, near\-plateau dynamics emerge \(r<0\.015r<0\.015\)\.
### G\.2Training Configuration
### G\.3Hardware
All experiments were run on a single NVIDIA RTX 4090 \(24GB\) GPU\.
### G\.4Per\-Seed WikiText Results
Table 5:WikiText perplexity per seed at selected generations forα=1\.0\\alpha=1\.0\.
### G\.5Cross\-Domain Comparison
Figure 10:Cross\-domain comparison of contamination chain dynamics\. WikiText \(left\) and Shakespeare \(right\) show qualitatively identical patterns: flat control, mild degradation atα=0\.5\\alpha=0\.5, and strong supercritical growth atα=1\.0\\alpha=1\.0\. Shakespeare exhibits even stronger degradation under pure contamination \(6\.6×\\timesvs\. 3\.8×\\times\), likely due to lower redundancy in literary text\.
### G\.6Diversity Collapse Details
Figure 11:Distinct\-2 diversity across generations forα=1\.0\\alpha=1\.0\(WikiText\)\. Diversity drops from 0\.68 to 0\.38, providing an independent degradation signal correlated with perplexity increase\.
### G\.7Convergence Diagnostics
Figure 12:Training loss convergence diagnostics across generations and contamination fractions, confirming that all models converge within the 3\-epoch training budget\.
### G\.8Source\-Diversity Experiment Details
The matched\-budget source\-diversity experiment \([Section˜6\.5](https://arxiv.org/html/2606.05168#S6.SS5)\) uses GPT\-2 \(124M\) withK∈\{1,3,5\}K\\in\\\{1,3,5\\\}models, 8 seeds\{42,123,456,789,1024,2048,3141,4096\}\\\{42,123,456,789,1024,2048,3141,4096\\\}, and per\-model seed offsetsseed\+1000×\(m\+1\)\\text\{seed\}\+1000\\times\(m\+1\)\.
#### Matched\-budget design\.
For eachKK, the synthetic pool has*fixed*sizen=2,000n\{=\}2\{,\}000: each ofKKmodels generates⌊2000/K⌋\\lfloor 2000/K\\rfloorsamples, so pool size is identical across conditions\. Each model trains on the full pool atα=1\.0\\alpha\{=\}1\.0\(pure synthetic\)\. Theα=0\\alpha\{=\}0control uses the same budget with real data only\. All other hyperparameters match the single\-chain experiment \([Appendix˜G](https://arxiv.org/html/2606.05168#A7)\)\. Total:\(1\+1\+3\+5\)×8seeds×8gens=640\(1\+1\+3\+5\)\\times 8\\text\{ seeds\}\\times 8\\text\{ gens\}=640runs\.
#### Per\-seed G7 excess PPL\.
#### Statistical tests\.
Exact permutation, Wilcoxon, and Jonckheere–Terpstra p\-values are one\-sided \(testingKa\>KbK\_\{a\}\>K\_\{b\}or decreasing trend\); the pairedtt\-test p\-value is two\-sided\.K=1K\{=\}1vs\.K=5K\{=\}5: exact permutationp=0\.047p\{=\}0\.047\(one\-sided\), Wilcoxonp=0\.055p\{=\}0\.055\(one\-sided\), pairedt\(7\)=2\.16t\(7\)\{=\}2\.16,p=0\.068p\{=\}0\.068\(two\-sided\)\.K=1K\{=\}1vs\.K=3K\{=\}3: exact permutationp=0\.137p\{=\}0\.137, Wilcoxonp=0\.125p\{=\}0\.125\.K=3K\{=\}3vs\.K=5K\{=\}5: exact permutationp=0\.535p\{=\}0\.535\(no difference\)\. Jonckheere–Terpstra:Z=1\.43Z\{=\}1\.43,p=0\.077p\{=\}0\.077\. The evidence supports a modest buffer effect ofK\>1K\{\>\}1relative to pure self\-training atα=1\.0\\alpha\{=\}1\.0, not a monotonically increasing benefit withKK\.
#### Partial contamination \(α=0\.5\\alpha\{=\}0\.5\)\.
An additional 448 runs testK∈\{1,5\}K\\in\\\{1,5\\\}atα=0\.5\\alpha\{=\}0\.5\(same matched\-budget design, 8 seeds\)\.K=1K\{=\}1excess G7:\+2\.20±0\.14\+2\.20\\pm 0\.14;K=5K\{=\}5:\+2\.18±0\.11\+2\.18\\pm 0\.11\. Mean difference:\+0\.02\+0\.02, pairedt\(7\)=0\.54t\(7\)=0\.54,p=0\.61p=0\.61,d=0\.17d=0\.17, 3/8 seeds consistent\. The source\-diversity effect is undetectable at partial contamination, confirming that real data in the training mix already breaks self\-reinforcement\.
## Appendix HIntervention Details
### H\.1Strategy Summary
Table 6:Intervention strategies and their effects onR0R\_\{0\}\(from baselineR0=2\.62R\_\{0\}=2\.62\)\. Only watermark\-based filtering and herd immunity achieveR0<1R\_\{0\}<1as single strategies\.The herd immunity threshold is1−1/R0=61\.9%1\-1/R\_\{0\}=61\.9\\%at baseline—the ecosystem analog of vaccine coverage thresholds\.
### H\.2Full Intervention Curves
Figure 13:R0R\_\{0\}as a function of intervention intensity for all six strategies\. Only watermark\-based filtering and herd immunity cross theR0=1R\_\{0\}=1threshold\.Figure 14:Pareto frontier: illustrative cost vs\.R0R\_\{0\}reduction across single\-strategy sweeps \(6 strategies×\\times20 intensity levels==120 points\)\. Only watermark\-based filtering and herd immunity achieveR0<1R\_\{0\}<1alone\. Combined interventions \(15 pairs×\\times9 intensity combinations==135 evaluations\) are analyzed separately in[Section˜7](https://arxiv.org/html/2606.05168#S7)\.
### H\.3Calibration Scenario Trajectories
Figure 15:ODE infection trajectories under the three calibration scenarios\. The pessimistic scenario \(R0=6\.63R\_\{0\}=6\.63\) converges rapidly to a high endemic level; the optimistic scenario \(R0=1\.10R\_\{0\}=1\.10\) shows slow, marginal growth\.
### H\.4Bifurcation Diagram
Figure 16:Bifurcation diagram showing endemic equilibrium infection level as a function ofR0R\_\{0\}\. The transcritical bifurcation atR0=1R\_\{0\}=1is clearly visible: the DFE \(zero infection\) is stable forR0<1R\_\{0\}<1and unstable forR0\>1R\_\{0\}\>1, where the endemic branch emerges\.
## Appendix ICode–Paper Theorem Numbering
The verification code \(theorem\_verification\.json\) uses a different numbering convention than the paper\. The mapping is:
- •theorem1\_dfe\_exists→\\toexistence part of[Theorem˜2](https://arxiv.org/html/2606.05168#Thmtheorem2)\(Theorem 2\)\.
- •theorem2\_dfe\_stability→\\tostability part of[Theorem˜2](https://arxiv.org/html/2606.05168#Thmtheorem2)\.
- •theorem3\_ee\_existence→\\to[Proposition˜3](https://arxiv.org/html/2606.05168#Thmtheorem3)\(Proposition 3\)\.
- •theorem4\_bifurcation→\\to[Proposition˜4](https://arxiv.org/html/2606.05168#Thmtheorem4)\(Proposition 4\)\.
- •proposition5\_sirs\_oscillation→\\to[Proposition˜5](https://arxiv.org/html/2606.05168#Thmtheorem5)\(Proposition 5\)\.
TheR0R\_\{0\}formula \([Theorem˜1](https://arxiv.org/html/2606.05168#Thmtheorem1), Theorem 1 in the paper\) is verified by direct comparison of the analytic formula with the NGM spectral radius, not as a separate entry in the verification JSON\.
## Appendix JReproducibility Statement
All experiments use publicly available models \(GPT\-2 from HuggingFace\) and datasets \(WikiText\-103, Tiny Shakespeare\)\. Random seeds are fixed at\{42,123,456\}\\\{42,123,456\\\}for single\-chain experiments and\{42,123,456,789,1024,2048,3141,4096\}\\\{42,123,456,789,1024,2048,3141,4096\\\}for source\-diversity experiments\. The ODE solver uses SciPy’ssolve\_ivpwith RK45, relative tolerance10−810^\{\-8\}, and absolute tolerance10−1010^\{\-10\}\. The ABM is implemented in Python with NetworkX\. Theorem verification counts \(e\.g\., 169/200 for EE existence, 196/200 for bifurcation\) are from a single run of the verification script with unseeded random parameter draws; exact counts may vary on rerun, though the qualitative result \(near\-100% agreement\) is stable\. Code will be released upon publication\. Total compute: approximately 5 GPU\-hours \(single\-chain, 192 runs\)\+\+12 GPU\-hours \(KK\-sweep, 640 runs\)\+\+8 GPU\-hours \(α=0\.5\\alpha\{=\}0\.5robustness, 448 runs\)≈\\approx25 GPU\-hours on a single NVIDIA RTX 4090\.Similar Articles
What's up with model collapse?
An exploration of model collapse, a phenomenon where AI models trained on synthetic data degrade in quality and diversity.
The interesting part of model collapse isn't technical, it's epistemic
This article explores model collapse not as a technical bug but as an epistemic problem: when an AI model's outputs become its own inputs, the model's representation of reality gradually flattens into a self-referential average, raising questions about how we distinguish a model that models the world from one that models only itself.
AI is deteriorating in realtime
AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.
The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
This paper introduces the 'fairness collapse' phenomenon, showing that training language models on synthetic data silently amplifies social biases before standard model collapse metrics degrade, highlighting a critical risk for AI fairness.
Model collapse + skill atrophy + competitive pressure = one big feedback loop. Thoughts?
The article discusses a consulting firm's argument that model collapse, human cognitive debt (skill atrophy), and competitive pressure form a self-reinforcing feedback loop in AI, and questions whether organizations can resist the race to automate.