The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
Summary
This paper introduces the 'fairness collapse' phenomenon, showing that training language models on synthetic data silently amplifies social biases before standard model collapse metrics degrade, highlighting a critical risk for AI fairness.
View Cached Full Text
Cached at: 08/06/26, 07:45 AM
# Bias Amplification in Language Models Trained on Synthetic Data
Source: [https://arxiv.org/html/2608.04268](https://arxiv.org/html/2608.04268)
## The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
Irina Proskurina1,2Antoine Gourru1Julien Velcin3 1Laboratoire Hubert Curien, UMR CNRS 5516, Saint\-Étienne, France 2Université Claude Bernard Lyon 1, Université Lumière Lyon 2, ERIC 3École Centrale de Lyon, LIRIS, CNRS UMR 5205 irina\.proskurina@univ\-lyon2\.fr
###### Abstract
Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation\. As synthetic content increasingly contaminates the training corpora of language models, this raises critical concerns about the use of open data in continued pretraining\. Although previous work has demonstrated model collapse in language models, it remains unclear whether exposure to synthetic data amplifies or attenuates the social biases already present in pretrained models\. Because language models are known to reproduce and amplify demographic stereotypes, recursive training on self\-generated data may create a self\-reinforcing feedback loop in which biased associations become progressively stronger across generations\. We call this hypothesized phenomenon “*fairness collapse*”\. In this work, we construct controlled training regimes in which models are repeatedly trained on synthetic data using the Bias in Bios dataset\. Across experiments, we observe a consistent and concerning pattern: fairness degradation emerges before substantial degradation is reflected by standard language\-modeling metrics\. This result highlights a critical risk associated with synthetic data contamination in language model training: bias can increase silently before strong indicators of model collapse become apparent\.
The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data
Irina Proskurina1,2Antoine Gourru1Julien Velcin31Laboratoire Hubert Curien, UMR CNRS 5516, Saint\-Étienne, France2Université Claude Bernard Lyon 1, Université Lumière Lyon 2, ERIC3École Centrale de Lyon, LIRIS, CNRS UMR 5205irina\.proskurina@univ\-lyon2\.fr
## 1Introduction
Human dataweb corpora, feedbackLLM traininggenerationttBias amplificationSynthetic datamodel\-generated textbiased outputs re\-enterthe training distributionFairness Collapse LoopFigure 1:Fairness collapse as a recursive bias self\-amplification process: model\-generated text is reintroduced into future training corpora, potentially causing distribution shift, reducing diversity, and amplifying social biases across generations\.As machine\-generated text proliferates on the webDolezalet al\.\([2026](https://arxiv.org/html/2608.04268#bib.bib4)\), language models are exposed to increasing amounts of synthetic content in both training corpora and annotation pipelines, potentially affecting model reliability and evaluation\. Outputs generated by models trained on such data mixtures may subsequently re\-enter future training corpora, increasing the proportion of artificially generated data\. This recursive process, in which model\-generated data are repeatedly reused for training, has been shown to degrade language\-generation quality through a phenomenon known asmodel collapse\(Shumailovet al\.,[2023](https://arxiv.org/html/2608.04268#bib.bib42); Dohmatobet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib16)\)\. Model collapse refers to a distributional shift driven by the progressive loss of low\-probability events, which can be measured using standard language\-modeling metrics and downstream task performance\.
While model collapse characterizes distributional degradation in generated text, it does not address whether synthetic training data also alters model responses to socially sensitive inputs, particularly those involving demographic attributes such as gender\. Prior work has shown that pretrained language models encode stereotypical associations involving gender and professional roles\(Nadeemet al\.,[2021](https://arxiv.org/html/2608.04268#bib.bib30)\), and that these demographic biases can surface in model\-generated text\(Shenget al\.,[2019](https://arxiv.org/html/2608.04268#bib.bib10)\)\. One canonical example is the disproportionate association of “nurse” with women and “engineer” with men\(Bolukbasiet al\.,[2016](https://arxiv.org/html/2608.04268#bib.bib51); Zhaoet al\.,[2018b](https://arxiv.org/html/2608.04268#bib.bib7)\), which has measurable downstream consequences in applications such as hiring and candidate screening\(Wanget al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib6); Wilson and Caliskan,[2024](https://arxiv.org/html/2608.04268#bib.bib5); Armstronget al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib11)\)\. In cases where synthetic data amplify these gendered associations, their repeated inclusion in subsequent training corpora may compound existing biases, independent of observable degradation in language modeling performance\.
In this paper, we investigate the interaction between model collapse and fairness drift under repeated continued pretraining on synthetic data\. Using the Bias in Bios corpus of professional biographies annotated with gender and occupation\(De\-Arteagaet al\.,[2019](https://arxiv.org/html/2608.04268#bib.bib12)\), we first characterize the distributional differences between human\-written and synthetic biographies across generation regimes, seed lengths, and decoding temperatures\. We then conduct controlled, continued\-pretraining experiments under both iterative and recursive synthetic\-data\-generation regimes and examine how fairness and general model performance metrics evolve across training iterations\.
#### Our main finding is that fairness degradation emerges before severe language\-model degradation\.
Across four controlled continued\-pretraining settings, spanning iterative and recursive regimes with seeded and few\-shot generation strategies, exposure to synthetic data consistently amplifies gender\-occupation bias before conventional indicators of model collapse become apparent\.111Anonymous code is available at[https://anonymous\.4open\.science/r/fairness\-collapse\-llamafactory\-C51B](https://anonymous.4open.science/r/fairness-collapse-llamafactory-C51B)\.Across all settings, fairness metrics deteriorate despite continued improvements in perplexity and comparatively stable downstream language\-modeling performance\. These results provide empirical evidence for a distinct early\-stage failure mode, which we refer to asfairness collapse\.
## 2Related Work
#### Model collapse
Recent work has examinedmodel collapsein autophagous training regimes, in which generative models are repeatedly trained on data produced by earlier model generations, across both vision and language modeling settings\(Shumailovet al\.,[2023](https://arxiv.org/html/2608.04268#bib.bib42); Seddiket al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib43)\)\.Shumailovet al\.\([2023](https://arxiv.org/html/2608.04268#bib.bib42)\)define model collapse as a degenerative phenomenon in which repeatedly training models on data generated by earlier models causes progressive deviation from the original data distribution\. Subsequent work has studied recursive and self\-consuming training regimes\(Alemohammadet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib45); Kazdanet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib35)\), including language\-model settings where synthetic\-text training reduces lexical, syntactic, and semantic diversity across generations\(Guoet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib44)\)\. Theoretical analyses further relate collapse to changes in scaling behavior and the progressive loss of distributional tails\(Dohmatobet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib16)\), while work on mixed\-data regimes shows that maintaining sufficient human data can prevent or attenuate collapse\(Gerstgrasseret al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib32); Kazdanet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib35)\)\. Later works suggest that collapse depends on the data\-generation approaches and proportion of human data retained during training\(Ferbachet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib46); Zhuet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib33)\)\.
#### Biases in language models
Language models trained on web\-scale corpora encode statistical associations between demographic attributes and professional roles\(Bolukbasiet al\.,[2016](https://arxiv.org/html/2608.04268#bib.bib51); Zhaoet al\.,[2018a](https://arxiv.org/html/2608.04268#bib.bib50)\)\. Such associations can lead to systematic disparities in downstream applications and to allocational harms\(Crawford,[2017](https://arxiv.org/html/2608.04268#bib.bib49)\), including in candidate screening, resume ranking, job recommendation, and professional profiling\(De\-Arteagaet al\.,[2019](https://arxiv.org/html/2608.04268#bib.bib12); Armstronget al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib11); Wanget al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib6); Wilson and Caliskan,[2024](https://arxiv.org/html/2608.04268#bib.bib5); Anet al\.,[2025](https://arxiv.org/html/2608.04268#bib.bib14)\)\. Biases in pre\-trained models also contribute to stereotyping language generation\(Nangiaet al\.,[2020](https://arxiv.org/html/2608.04268#bib.bib31); Nadeemet al\.,[2021](https://arxiv.org/html/2608.04268#bib.bib30); Marchiori Manerbaet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib29)\)\. Bias mitigation approaches include post\-hoc interventions, such as decoding\-time self\-debiasing\(Schicket al\.,[2021](https://arxiv.org/html/2608.04268#bib.bib28)\), inference\-time activation steering\(Liet al\.,[2025](https://arxiv.org/html/2608.04268#bib.bib27)\), and representation editing methods that remove linearly recoverable protected\-attribute information from hidden representations\(Ravfogelet al\.,[2020](https://arxiv.org/html/2608.04268#bib.bib26)\)\. A complementary line of work proposes debiasing language\-model training objectives, including counterfactual sentence\-level debiasing\(Lianget al\.,[2020](https://arxiv.org/html/2608.04268#bib.bib25)\), name\-based regularization for occupation classification\(Romanovet al\.,[2019](https://arxiv.org/html/2608.04268#bib.bib24)\), and fairness\-oriented representation regularization during fine\-tuning\(Xuet al\.,[2025](https://arxiv.org/html/2608.04268#bib.bib23)\)\.
Human𝒟\\mathcal\{D\}ℳθ⋆\\mathcal\{M\}\_\{\\theta^\{\\star\}\}Human\-onlytrainℳbase\\mathcal\{M\}\_\{base\}Synthetic𝒟t∗\\mathcal\{D\}^\{\*\}\_\{t\}Refreshedℳbase\\mathcal\{M\}\_\{base\}ℳθt\+1\\mathcal\{M\}\_\{\\theta\_\{t\+1\}\}Iterative traininggeneratetraingenerateℳθt\\mathcal\{M\}\_\{\\theta\_\{t\}\}Synthetic𝒟t∗\\mathcal\{D\}^\{\*\}\_\{t\}ℳθt\+1\\mathcal\{M\}\_\{\\theta\_\{t\+1\}\}Recursive traininggeneratetrainrepeatFigure 2:Overview of the three controlled continued\-pretraining regimes used in our experiments\. We begin with the human\-written corpus𝒟=\{\(xi,yi,ai\)\}i=1N\\mathcal\{D\}=\\\{\(x\_\{i\},y\_\{i\},a\_\{i\}\)\\\}\_\{i=1\}^\{N\}and the human\-trained checkpointℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}\. Inhuman\-only training, the model is repeatedly trained on𝒟\\mathcal\{D\}, without generating synthetic biographies\. Initerative regeneration, the model from the previous iteration,ℳθt\\mathcal\{M\}\_\{\\theta\_\{t\}\}, is used to generate a new synthetic dataset𝒟tsyn\\mathcal\{D\}^\{\\mathrm\{syn\}\}\_\{t\}, but the next model is obtained by training a newly initialized base modelℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}\. Inrecursive contamination,ℳθt\\mathcal\{M\}\_\{\\theta\_\{t\}\}generates𝒟tsyn\\mathcal\{D\}^\{\\mathrm\{syn\}\}\_\{t\}and serves as the initialization for the next training iteration, yieldingℳθt\+1\\mathcal\{M\}\_\{\\theta\_\{t\+1\}\}\.Unlike prior work that characterizes model collapse primarily using distributional and language\-modeling metrics, we examine whether repeated exposure to synthetic data amplifies or mitigates existing social biases in language models\.
## 3Data Generation and Training Regimes
In this section, we present the synthetic data generation protocol\. We consider three training strategies: \(1\) training only on human\-written data, \(2\) iterative training, where a model is first trained on synthetic data and then used to generate a new synthetic dataset for training a newly initialized base model, and \(3\) contaminated/recursive training, where the model is continuously trained on its own generated data\. For strategies \(2\) and \(3\), we use either seeded or few\-shot generation, which we describe in detail later\. An overview of these training regimes is provided in[Figure 2](https://arxiv.org/html/2608.04268#S2.F2)\. We also provide examples of generated biographies in[Appendix B](https://arxiv.org/html/2608.04268#A2)\.
### 3\.1Human\-Written Data
We use theBias in Bioscorpus\(De\-Arteagaet al\.,[2019](https://arxiv.org/html/2608.04268#bib.bib12)\)of human\-written biographies spanning 28 professions\. Let us denote the human\-written training dataset as𝒟=\{\(xi,yi,ai\)\}i=1N\\mathcal\{D\}=\\\{\(x\_\{i\},y\_\{i\},a\_\{i\}\)\\\}\_\{i=1\}^\{N\}, wherexix\_\{i\}is a biography,yiy\_\{i\}is its associated profession label andaia\_\{i\}denotes the gender \(0for male,11for female\)\. The full corpus contains around 300k biographies and has been widely used to study occupational gender bias in classification tasks\. The profession distribution in the original dataset is highly imbalanced, often introducing spurious training artifacts\. We therefore construct a balanced training subset to isolate the effect of recursive training, resulting in a corpus containing27,752biographies with a uniform profession distribution, totaling approximately2M tokens\. Additional details are provided in[Appendix A](https://arxiv.org/html/2608.04268#A1)\.
Figure 3:Lexical diversity \(left\), relative continuation Wasserstein distance \(ΔW\\Delta W\) from model\-generated continuations to the input seed and the Wasserstein distance from human\-written continuations to the same seed \(center\), and Wasserstein distance between model\-generated and human\-written continuations \(right\) for the Qwen model\.
### 3\.2Synthetic Data Generation
We generate synthetic biographies using two approaches: \(i\) few\-shot generation and \(ii\) seeded generation\.
#### Few\-shot generation
In the few\-shot setting, synthetic biographies are generated using a prompt \(see[Appendix A](https://arxiv.org/html/2608.04268#A1)\) that instructs the model to write a single\-paragraph professional biography for a given profession\. To approximate the distribution of real data, we prepend the prompt with a small number \(K=3K=3\) of human\-written biographies randomly sampled from𝒟\\mathcal\{D\}for each profession\. This few\-shot context is resampled at each generation step\. For each professionpip\_\{i\}, we generate the same number of synthetic biographies as in the corresponding subset of the balanced human\-written dataset\.
#### Seeded generation
In seeded generation, synthetic biographies are conditioned directly on prefixes extracted from human\-written biographies\. Given a biographyxi∈𝒟x\_\{i\}\\in\\mathcal\{D\}, we define a seedsis\_\{i\}as the firstkktokens ofxix\_\{i\}\. We writexi=\(si,xi∖s\)x\_\{i\}=\(s\_\{i\},x\_\{i\}^\{\\setminus s\}\), wherexi∖sx\_\{i\}^\{\\setminus s\}denotes the remaining suffix\.
The seedsis\_\{i\}is used as the prompt for generation, and the model produces a continuation conditioned on this prefix\. The final synthetic biography is formed by concatenating the human\-written prefix with the model\-generated continuation\. By varying the seed lengthkkfrom 5 to 50 tokens in increments of 5, we directly control the proportion of model\-generated content within each synthetic biography\. This range provides broad coverage of the dataset: a seed length of 50 tokens exceeds the length of more than 75% of the biographies in the balanced training corpus\.
### 3\.3Synthetic\-Data Feedback Regimes
We perform continued pretraining under two synthetic\-data feedback regimes\. In both settings, we start from the modelℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}obtained after training the base model on the balanced human\-written corpus𝒟\\mathcal\{D\}\. The synthetic portion of the training data is regenerated at each iteration, but the regimes differ in how the model is initialized for training: either from the human\-trained checkpointℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}or from the checkpoint obtained at the previous iteration, as shown in[Figure 2](https://arxiv.org/html/2608.04268#S2.F2)\.
#### Iterative regeneration regime
Synthetic data are first generated using the human\-trained modelℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}\. At iterationtt, the synthetic dataset is regenerated using the model obtained from the previous iteration,ℳθt\\mathcal\{M\}\_\{\\theta\_\{t\}\}\. However,ℳθt\\mathcal\{M\}\_\{\\theta\_\{t\}\}is used only as the generator: the next training run is initialized from the base model rather than from the previous checkpointℳθt\\mathcal\{M\}\_\{\\theta\_\{t\}\}\. This setup allows us to isolate the effect of regenerated synthetic data from cumulative parameter updates across generations\.
#### Recursive contamination regime
Training is initialized from the human\-trained modelℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}at the first iteration and then continues from the checkpoint obtained at the previous iteration\. After each iteration, synthetic data from earlier steps are discarded, and the updated modelℳθt\+1\\mathcal\{M\}\_\{\\theta\_\{t\+1\}\}is used to generate a new synthetic dataset for the next training step\. The next training iteration is initialized fromℳθt\+1\\mathcal\{M\}\_\{\\theta\_\{t\+1\}\}rather than from the human\-trained checkpointℳθ0\\mathcal\{M\}\_\{\\theta\_\{0\}\}\. Although previous synthetic datasets are not reused directly, their influence persists through the model parameters\. This creates a feedback loop in which the model is repeatedly trained on data derived from its own outputs\.
## 4Generated Data Evaluation
In this section, we present the evaluation protocol used to analyze synthetic\-data quality and generation parameters\. In all our experiments, we use theQwen2\.5\-0\.5Bmodel\.222[https://hf\.co/Qwen/Qwen2\.5\-0\.5B](https://hf.co/Qwen/Qwen2.5-0.5B)
### 4\.1Evaluation Protocol
We evaluate seeded and few\-shot synthetic biographies across decoding temperatures using two complementary measures: Wasserstein distance between embedding distributions of human\-written and generated biographies, and lexical diversity measured with MTLD\. Sentence embeddings are computed using Qwen3\-Embedding\-0\.6B\(Enevoldsenet al\.,[2025](https://arxiv.org/html/2608.04268#bib.bib41)\), which ranks among the top\-performing models on the MMTEB benchmark\(Enevoldsenet al\.,[2025](https://arxiv.org/html/2608.04268#bib.bib41)\)\. The analysis is conducted at four decoding temperatures \(T∈0\.3,0\.6,0\.9,1\.2T\\in\{0\.3,0\.6,0\.9,1\.2\}\) for seeded generation\.
### 4\.2Metrics Description
#### Wasserstein distance \(Sinkhorn\-2\)\.
To compare the distributions of human\-written and generated biographies in embedding space, we compute the entropically regularized22\-Wasserstein distance between embedding setsX=\{xi\}i=1nX=\\\{x\_\{i\}\\\}\_\{i=1\}^\{n\}andY=\{yj\}j=1m\.Y=\\\{y\_\{j\}\\\}\_\{j=1\}^\{m\}\.Intuitively, the Wasserstein distance measures the minimal transport cost required to align the two embedding distributions\. We use the Sinkhorn\-regularized formulationW2,ε\(X,Y\)W\_\{2,\\varepsilon\}\(X,Y\)with squared Euclidean transport costCij=‖xi−yj‖22\.C\_\{ij\}=\\\|x\_\{i\}\-y\_\{j\}\\\|\_\{2\}^\{2\}\.Additional details on the optimal\-transport formulation are provided in[Appendix A](https://arxiv.org/html/2608.04268#A1)\. Wasserstein distances are computed using thePOTlibrary\(Flamaryet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib40)\)\.
#### Relative continuation distance\.
Beyond comparing generated and human\-written biographies globally, we also analyze how continuations evolve relative to their conditioning seeds\. LetSS,XhumanX\_\{\\text\{human\}\}, andXmodelX\_\{\\text\{model\}\}denote the embeddings of seed prefixes, human\-written continuations, and model\-generated continuations, respectively\. We define the relative continuation distance as
ΔW=W2,ε\(Xmodel,S\)−W2,ε\(Xhuman,S\),\\Delta W=W\_\{2,\\varepsilon\}\(X\_\{\\text\{model\}\},S\)\-W\_\{2,\\varepsilon\}\(X\_\{\\text\{human\}\},S\),where values close to zero indicate human\-like continuation behavior\. Negative values indicate continuations that remain overly anchored to the seed, whereas positive values indicate semantic drift beyond the variability observed in human\-written continuations\.
#### Lexical diversity\.
To measure surface\-level diversity, we report the Measure of Textual Lexical Diversity \(MTLD;McCarthy and Jarvis \([2010](https://arxiv.org/html/2608.04268#bib.bib2)\)\)\. MTLD estimates lexical diversity from the stability of the type\-token ratio throughout a text and is substantially less sensitive to text length than the raw type\-token ratio \(TTR\)\.
### 4\.3Results of Generated Data Evaluation
#### Seeded generation exhibits conservative continuation behavior\.
[Figure 3](https://arxiv.org/html/2608.04268#S3.F3)shows the semantic and lexical properties of seeded generations across decoding temperatures and seed lengths\. The center panel reports the relative continuation distanceΔW\\Delta W\. Across most temperatures and seed lengths,ΔW\\Delta Wremains negative, indicating that generated continuations stay semantically closer to the conditioning seed than human\-written continuations\. This effect is strongest at low temperatures \(0\.30\.3and0\.60\.6\), suggesting that more deterministic decoding produces continuations that remain anchored to the seed contexts\. Increasing the decoding temperature progressively reduces this tendency to remain close to the seed context\. AtT=1\.2T=1\.2,ΔW\\Delta Wbecomes positive for longer seeds, indicating semantic drift beyond the variability observed in human\-written continuations\. Overall, decoding temperature controls the model’s semantic extrapolation regime, ranging from conservative continuation at low temperatures to over\-extrapolation at higher temperatures\.
#### Moderate temperatures provide the closest match to human\-written biographies\.
[Figure 3](https://arxiv.org/html/2608.04268#S3.F3)\(right\) reports the Wasserstein distance between model\-generated and human\-written continuations\. Distances remain relatively stable across seed lengths, indicating that seeded generations preserve the alignment with human\-written distribution even as the proportion of model\-generated content increases\. However, decoding temperature consistently affects semantic fidelity: moderate temperatures \(0\.60\.6and0\.90\.9\) produce the lowest Wasserstein distances overall, whereas high temperature \(T=1\.2T=1\.2\) leads to larger deviations from the human\-written distribution\. Increasing the seed length slightly reduces the Wasserstein distance, though the effect remains modest compared with the impact of decoding temperature\.
#### Few\-shot prompting increases diversity at the cost of semantic control\.
From the MTLD evaluation results, we find that low\-temperature seeded generation yields substantially lower lexical diversity than the human\-written corpus, while MTLD increases with decoding temperature, approaching the human\-written baseline atT=0\.9T=0\.9and exceeding the baseline atT=1\.2T=1\.2\. At the same time, few\-shot prompting achieves similarly high lexical diversity atT=0\.9T=0\.9, but also yields consistently larger Wasserstein distances\. These results suggest that seeded continuation provides stronger semantic control, while few\-shot prompting increases lexical variability at the cost of greater semantic drift\.
## 5Bias Amplification Evaluation
Table 1:Perplexity, MMLU, Bias\-in\-Bios, CrowS\-Pairs, and SOFA evaluation results for models trained underrecursive synthetic\-data feedback\. PPL is computed on the training data used at the corresponding iteration\. Bias\-in\-Bios reports signed gender\-conditioned NLL gap, occupation\-prediction accuracy, and aggregate equal\-opportunity gap across professions\. CrowS\-Pairs reports likelihood gap and stereotype preference on the gender subset; SOFA reports aggregate social\-fairness bias\.### 5\.1Experimental Setup
In this section, we describe the experimental protocol used to evaluate ourfairness collapsehypothesis\. In all experiments, we use Qwen2\.5\-0\.5B for both synthetic\-data generation and continued pretraining\. Across all controlled training experiments, we keep the model architecture and training hyperparameters the same\. Training hyperparameters are provided in[Appendix A](https://arxiv.org/html/2608.04268#A1)\. We use the Bias in Bios dataset presented in §[3\.1](https://arxiv.org/html/2608.04268#S3.SS1)for training, seeded generation, and few\-shot prompting\. We follow prior work and split𝒟\\mathcal\{D\}into training, validation, and test subsets\. Synthetic biographies are generated from prefixes sampled from the training split using either seeded continuation or few\-shot prompting\. All evaluation is performed on held\-out human\-written biographies from the test split\. We generate the synthetic biographies using a seed of length3030andT=0\.9T=0\.9, as this yields the best trade\-off between similarity to human\-written text and lexical diversity\.
### 5\.2Evaluation Metrics
#### Occupation prediction accuracy\.
We evaluate occupational bias using a controlled pairwise classification task over the 28 professions in Bias in Bios\. Each profession is paired with a fixed semantically related alternative provided in[Appendix A](https://arxiv.org/html/2608.04268#A1)\(e\.g\.,nurse↔\\leftrightarrowphysician\)\. Given a biographyxx, the model receives a multiple\-choice prompt containing the biography together with the gold profession and its paired alternative\. The order of the two options is randomized to avoid positional bias\. Predictions are obtained by comparing the next\-token logits of the candidate answers\.
#### Classification bias\.
Our primary fairness metric is Equality of Opportunity GAP \(EO GAP\), which measures disparities in true positive rates across demographic groups conditioned on the correct profession label\. For each professionc∈Cc\\in C, we defineEOc=TPRc,0−TPRc,1EO\_\{c\}=\\mathrm\{TPR\}\_\{c,0\}\-\\mathrm\{TPR\}\_\{c,1\}, whereTPRc,a=p\(y^=c∣y=c,a\)\\mathrm\{TPR\}\_\{c,a\}=p\(\\hat\{y\}=c\\mid y=c,a\), anda=0a=0anda=1a=1denote male and female biographies, respectively\. To summarize profession\-level disparities, we report the aggregate GAP score, which is the average of the squared values ofEOcEO\_\{c\}\. Lower GAP values correspond to smaller demographic disparities\.
#### Likelihood asymmetry\.
Beyond classification metrics, we analyze demographic asymmetries directly in the language\-model distribution using signed conditional likelihood gaps\. For each profession, we compute the average negative log\-likelihood separately for male\- and female\-associated biographies and report the difference between the two groups\. Positive values indicate that female\-associated biographies are assigned a lower likelihood, whereas negative values indicate that male\-associated biographies are assigned a lower likelihood\. This metric captures token\-level likelihood asymmetries between demographic groups within the same profession\.
#### General language\-model evaluation\.
To monitor broader language\-model degradation during recursive training, we evaluate all models on MMLU\(Hendryckset al\.,[2021](https://arxiv.org/html/2608.04268#bib.bib3)\)\. MMLU measures multitask accuracy across a diverse set of domains, including science, law, medicine, and humanities\. We report average accuracy over the benchmark as a proxy for general language\-modeling capabilities under repeated synthetic\-data exposure\.
#### Additional fairness benchmarks\.
Beyond Bias in Bios, we evaluate models on external fairness benchmarks based on likelihood comparisons\.CrowS\-Pairs\(Nangiaet al\.,[2020](https://arxiv.org/html/2608.04268#bib.bib31)\)measures stereotypical bias using minimally different sentence pairs, and reports the percentage of cases where the model assigns a higher likelihood to the stereotypical sentence\. We also evaluate the gender subset ofSoFA\(Marchiori Manerbaet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib29)\), which measures likelihood variation across gendered stereotype probes after normalization by the likelihood of the identity term itself\. Higher SoFA scores indicate larger demographic likelihood asymmetries\.
Table 2:Perplexity, MMLU, Bias\-in\-Bios, CrowS\-Pairs, and SOFA evaluation results for models trained underiterative synthetic\-data regeneration\. PPL is computed on the training data used at the corresponding iteration\. Bias\-in\-Bios reports signed gender\-conditioned NLL gap, occupation\-prediction accuracy, and aggregate equal\-opportunity gap across professions\. CrowS\-Pairs reports likelihood gap and stereotype preference on the gender subset; SOFA reports aggregate social\-fairness bias\.Figure 4:Trajectory of profession\-level fairness and lexical\-diversity shift across contaminated \(recursive\) seeded training iterations\. The x\-axis shows the difference in MTLD from human\-written biographies, and the y\-axis shows the profession\-specific equal opportunity gap\.
### 5\.3Results
Results for all metrics and both synthetic\-data generation strategies are presented in[Table 1](https://arxiv.org/html/2608.04268#S5.T1)and[Table 2](https://arxiv.org/html/2608.04268#S5.T2)\.[Figure 4](https://arxiv.org/html/2608.04268#S5.F4)illustrates the profession\-level drift of EO GAP across recursive training iterations\. For completeness, we report the full results for the human\-only training regime in[Table 15](https://arxiv.org/html/2608.04268#A2.T15)\.
#### Fairness degradation precedes classical model collapse\.
[Table 1](https://arxiv.org/html/2608.04268#S5.T1)and[Table 2](https://arxiv.org/html/2608.04268#S5.T2)demonstrate that exposure to synthetic data amplifies demographic disparities before strong degradation in standard language\-modeling metrics becomes apparent\. Under contaminated seeded training, the EO GAP increases from13\.1813\.18at Iteration 0 to19\.3819\.38at Iteration 5, while perplexity simultaneously improves from16\.0716\.07to10\.5510\.55\. At the same time, MMLU accuracy decreases more gradually, from42\.1442\.14to32\.0832\.08\. Similar trends are observed for few\-shot generation, where fairness metrics deteriorate despite comparatively stable occupation prediction accuracy during early iterations, with SOFA score increasing from0\.50\.5to0\.780\.78\.
This pattern suggests that recursive synthetic\-data exposure progressively amplifies demographic asymmetries independently of model collapse, even as next\-token prediction under the training distribution continues to improve\.
#### Recursive training amplifies existing demographic associations\.
[Figure 4](https://arxiv.org/html/2608.04268#S5.F4)illustrates the evolution of profession\-level EO gaps together with lexical\-diversity shifts across recursive seeded iterations\. Each trajectory corresponds to the evolution of a profession across training iterations relative to the human\-written corpus\.
Several professions exhibit strong directional drift in EO space\. Female\-associated professions such asnurseprogressively move toward increasingly negative EO values, while male\-associated professions such asprofessordrift toward increasingly positive EO values\. Similar amplification patterns are observed forpersonal trainer,software engineer, andinterior designer\. These trajectories indicate that recursive training reinforces demographic associations already present in the original corpus rather than introducing arbitrary noise\.
Importantly, these fairness shifts occur alongside comparatively modest changes in lexical diversity scores\. Most trajectories remain concentrated within a relatively narrow MTLD range while exhibiting large movements in EO space\. This suggests that fairness degradation can accumulate even in the absence of dramatic surface\-level distributional collapse\. A finer\-grained gender\-conditioned analysis in[Figure 7](https://arxiv.org/html/2608.04268#A2.F7)shows that aggregate MTLD does not fully characterize these changes: across the six reported professions, male\-female MTLD differences are most pronounced fornurse,rapper, andsoftware engineer\. This further distinguishes aggregate lexical\-diversity shifts from the group\-conditioned disparities captured by EO GAP\.
#### Likelihood asymmetries increase across recursive iterations\.
The likelihood\-based metrics further support this amplification effect\. Under contaminated recursive training \([Table 1](https://arxiv.org/html/2608.04268#S5.T1)\), both NLL\-GAP and CrowS\-Pairs stereotype preference generally increase across iterations, indicating progressively stronger demographic asymmetries in token\-level likelihoods\. The SoFA benchmark exhibits a similar trend, with bias scores increasing from0\.5090\.509at Iteration 0 to1\.0201\.020under seeded recursive training\. These effects emerge even as perplexity decreases, suggesting that overall language\-modeling improvements can mask growing demographic distortions within the model distribution\.
#### Iterative regeneration partially mitigates fairness collapse\.
The iterative regeneration regime exhibits slower and less consistent fairness degradation than contaminated recursive training\. As shown in[Table 2](https://arxiv.org/html/2608.04268#S5.T2), EO GAP still increases under iterative training, but the amplification remains weaker and less monotonic\. For example, seeded iterative training reaches a maximum EO GAP of18\.3418\.34, compared to20\.3720\.37under recursive contamination\.
This comparison suggests that fairness collapse is driven primarily by exposure to biased synthetic data, since bias amplification appears in both iterative and recursive settings\. However, the effect of adding artificially generated data is stronger and more monotonic under recursive training, where synthetic\-data exposure is compounded by repeatedly updating the same model across iterations\.
#### Fairness collapse as an early\-warning signal\.
Taken together, these results show that fairness degradation emerges earlier than conventional indicators of model collapse such as MMLU or occupation\-classification accuracy\. Synthetic\-data exposure progressively amplifies demographic associations even when perplexity improves and overall language\-modeling performance remains stable\.
## 6Conclusion
In this work, we investigated how repeated continued pretraining on synthetic data affects fairness in language models\. Using controlled regeneration regimes on the Bias in Bios dataset, we compared iterative and recursive training strategies under seeded and few\-shot synthetic generation\.
Our experiments reveal a consistent pattern: fairness degradation emerges substantially earlier than conventional indicators of model collapse\. Recursive exposure to synthetic data amplifies demographic disparities even when perplexity improves and general language\-modeling performance remains comparatively stable\. This effect is strongest under recursive training, where models repeatedly train on their own generated outputs\.
We describe this phenomenon asfairness collapse: an early\-stage failure mode in which demographic biases silently accumulate under recursive synthetic\-data exposure\. Altogether, these results highlight a broader risk for model training as synthetic text becomes increasingly incorporated into large\-scale corpora\.
## Limitations
This study has several limitations that should be considered when interpreting the scope and generality of the results\. First, all experiments are conducted on the Bias in Bios dataset, which focuses on gender and occupation associations in English biographies\. Although this benchmark is widely used for fairness evaluation, it captures only a limited subset of demographic biases and social contexts\. Future work should evaluate whether fairness collapse generalizes to other sensitive attributes, languages, and domains\.
Second, our experiments are performed on relatively small decoder\-only language models\. While this controlled setup enables repeated regeneration experiments at manageable computational cost, larger frontier\-scale models may exhibit different robustness or collapse dynamics under recursive synthetic\-data exposure\.
Third, our synthetic\-data pipeline relies on seeded continuation and few\-shot prompting under fixed decoding configurations\. Alternative generation strategies, filtering procedures, or synthetic\-data curation methods may substantially affect both model\-collapse dynamics and fairness behavior\.
Finally, our analysis primarily focuses on likelihood\-based and classification\-based fairness metrics\. Although these metrics reveal consistent demographic drift under recursive training, they do not fully capture broader social harms, downstream deployment risks, or intersectional effects that may arise in real\-world applications\.
## References
- S\. Alemohammad, J\. Casco\-Rodriguez, L\. Luzi, A\. I\. Humayun, H\. Babaei, D\. LeJeune, A\. Siahkoohi, and R\. Baraniuk \(2024\)Self\-consuming generative models go mad\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 53581–53608\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- Measuring gender and racial biases in large language models: intersectional evidence from automated resume evaluation\.PNAS nexus4\(3\),pp\. pgaf089\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- L\. Armstrong, A\. Liu, S\. MacNeil, and D\. Metaxa \(2024\)The silicon ceiling: auditing gpt’s race and gender biases in hiring\.InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization,pp\. 1–18\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- T\. Bolukbasi, K\. Chang, J\. Y\. Zou, V\. Saligrama, and A\. T\. Kalai \(2016\)Man is to computer programmer as woman is to homemaker? debiasing word embeddings\.Advances in neural information processing systems29\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Crawford \(2017\)The trouble with bias\. keynote at neurips\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- M\. De\-Arteaga, A\. Romanov, H\. Wallach, J\. Chayes, C\. Borgs, A\. Chouldechova, S\. Geyik, K\. Kenthapadi, and A\. T\. Kalai \(2019\)Bias in bios: a case study of semantic representation bias in a high\-stakes setting\.Inproceedings of the Conference on Fairness, Accountability, and Transparency,pp\. 120–128\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p3.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.04268#S3.SS1.p1.6)\.
- E\. Dohmatob, Y\. Feng, P\. Yang, F\. Charton, and J\. Kempe \(2024\)A tale of tails: model collapse as a change of scaling laws\.arXiv preprint arXiv:2402\.07043\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p1.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Dolezal, S\. Alam, M\. Graham, and M\. Bohacek \(2026\)The impact of ai\-generated text on the internet\.arXiv preprint arXiv:2604\.26965\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p1.1)\.
- K\. Enevoldsen, I\. Chung, I\. Kerboua, M\. Kardos, A\. Mathur, D\. Stap, J\. Gala, W\. Siblini, D\. Krzemiński, G\. I\. Winata,et al\.\(2025\)Mmteb: massive multilingual text embedding benchmark\.arXiv preprint arXiv:2502\.13595\.Cited by:[§4\.1](https://arxiv.org/html/2608.04268#S4.SS1.p1.1)\.
- D\. Ferbach, Q\. Bertrand, J\. Bose, and G\. Gidel \(2024\)Self\-consuming generative models with curated data provably optimize human preferences\.InThe Thirty\-eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=cyv0LkIaoH)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- R\. Flamary, C\. Vincent\-Cuaz, N\. Courty, A\. Gramfort, O\. Kachaiev, H\. Quang Tran, L\. David, C\. Bonet, N\. Cassereau, T\. Gnassounou, E\. Tanguy, J\. Delon, A\. Collas, S\. Mazelet, L\. Chapel, T\. Kerdoncuff, X\. Yu, M\. Feickert, P\. Krzakala, T\. Liu, and E\. Fernandes Montesuma \(2024\)POT python optimal transport \(version 0\.9\.5\)\.External Links:[Link](https://github.com/PythonOT/POT)Cited by:[§A\.3](https://arxiv.org/html/2608.04268#A1.SS3.p5.1),[§4\.2](https://arxiv.org/html/2608.04268#S4.SS2.SSS0.Px1.p1.5)\.
- M\. Gerstgrasser, R\. Schaeffer, A\. Dey, R\. Rafailov, H\. Sleight, J\. Hughes, T\. Korbak, R\. Agrawal, D\. Pai, A\. Gromov,et al\.\(2024\)Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data\.arXiv preprint arXiv:2404\.01413\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Guo, G\. Shang, M\. Vazirgiannis, and C\. Clavel \(2024\)The curious decline of linguistic diversity: training language models on synthetic text\.InFindings of the Association for Computational Linguistics: NAACL 2024,K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 3589–3604\.External Links:[Link](https://aclanthology.org/2024.findings-naacl.228/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-naacl.228)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.Proceedings of the International Conference on Learning Representations \(ICLR\)\.Cited by:[§5\.2](https://arxiv.org/html/2608.04268#S5.SS2.SSS0.Px4.p1.1)\.
- J\. Kazdan, R\. Schaeffer, A\. Dey, M\. Gerstgrasser, R\. Rafailov, D\. L\. Donoho, and S\. Koyejo \(2024\)Collapse or thrive? perils and promises of synthetic data in a self\-generating world\.arXiv preprint arXiv:2410\.16713\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Li, Z\. Fan, R\. Chen, X\. Gai, L\. Gong, Y\. Zhang, and Z\. Liu \(2025\)FairSteer: inference time debiasing for LLMs with dynamic activation steering\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 11293–11312\.External Links:[Link](https://aclanthology.org/2025.findings-acl.589/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.589),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- P\. P\. Liang, I\. M\. Li, E\. Zheng, Y\. C\. Lim, R\. Salakhutdinov, and L\. Morency \(2020\)Towards debiasing sentence representations\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 5502–5515\.External Links:[Link](https://aclanthology.org/2020.acl-main.488/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.488)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Marchiori Manerba, K\. Stanczak, R\. Guidotti, and I\. Augenstein \(2024\)Social bias probing: fairness benchmarking for language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 14653–14671\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.812/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.812)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2608.04268#S5.SS2.SSS0.Px5.p1.1)\.
- P\. M\. McCarthy and S\. Jarvis \(2010\)MTLD, vocd\-d, and hd\-d: a validation study of sophisticated approaches to lexical diversity assessment\.Behavior research methods42\(2\),pp\. 381–392\.Cited by:[§4\.2](https://arxiv.org/html/2608.04268#S4.SS2.SSS0.Px3.p1.1)\.
- M\. Nadeem, A\. Bethke, and S\. Reddy \(2021\)StereoSet: measuring stereotypical bias in pretrained language models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 5356–5371\.External Links:[Link](https://aclanthology.org/2021.acl-long.416/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.416)Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- N\. Nangia, C\. Vania, R\. Bhalerao, and S\. R\. Bowman \(2020\)CrowS\-pairs: a challenge dataset for measuring social biases in masked language models\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 1953–1967\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.154/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.154)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2608.04268#S5.SS2.SSS0.Px5.p1.1)\.
- S\. Ravfogel, Y\. Elazar, H\. Gonen, M\. Twiton, and Y\. Goldberg \(2020\)Null it out: guarding protected attributes by iterative nullspace projection\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 7237–7256\.External Links:[Link](https://aclanthology.org/2020.acl-main.647/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.647)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Romanov, M\. De\-Arteaga, H\. Wallach, J\. Chayes, C\. Borgs, A\. Chouldechova, S\. Geyik, K\. Kenthapadi, A\. Rumshisky, and A\. Kalai \(2019\)What’s in a name? Reducing bias in bios without access to protected attributes\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4187–4195\.External Links:[Link](https://aclanthology.org/N19-1424/),[Document](https://dx.doi.org/10.18653/v1/N19-1424)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- T\. Schick, S\. Udupa, and H\. Schütze \(2021\)Self\-diagnosis and self\-debiasing: a proposal for reducing corpus\-based bias in NLP\.Transactions of the Association for Computational Linguistics9,pp\. 1408–1424\.External Links:[Link](https://aclanthology.org/2021.tacl-1.84/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00434)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- M\. E\. A\. Seddik, S\. Chen, S\. Hayou, P\. Youssef, and M\. Debbah \(2024\)How bad is training on synthetic data? a statistical analysis of language model collapse\.arXiv preprint arXiv:2404\.05090\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Sheng, K\. Chang, P\. Natarajan, and N\. Peng \(2019\)The woman worked as a babysitter: on biases in language generation\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3407–3412\.External Links:[Link](https://aclanthology.org/D19-1339/),[Document](https://dx.doi.org/10.18653/v1/D19-1339)Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1)\.
- I\. Shumailov, Z\. Shumaylov, Y\. Zhao, Y\. Gal, N\. Papernot, and R\. Anderson \(2023\)The curse of recursion: training on generated data makes models forget\.arXiv preprint arXiv:2305\.17493\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p1.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Wang, Z\. Wu, X\. Guan, M\. Thaler, A\. Koshiyama, S\. Lu, S\. Beepath, E\. Ertekin, and M\. Perez\-Ortiz \(2024\)JobFair: a framework for benchmarking gender hiring bias in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 3227–3246\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.184/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.184)Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Wilson and A\. Caliskan \(2024\)Gender, race, and intersectional bias in resume screening via language model retrieval\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.7,pp\. 1578–1590\.Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1),[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Xu, W\. Chen, L\. Li, Y\. Zhao, and Y\. Wei \(2025\)Collapsed language models promote fairness\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 60225–60245\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Zhao, T\. Wang, M\. Yatskar, V\. Ordonez, and K\. Chang \(2018a\)Gender bias in coreference resolution: evaluation and debiasing methods\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 \(Short Papers\),M\. Walker, H\. Ji, and A\. Stent \(Eds\.\),New Orleans, Louisiana,pp\. 15–20\.External Links:[Link](https://aclanthology.org/N18-2003/),[Document](https://dx.doi.org/10.18653/v1/N18-2003)Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Zhao, T\. Wang, M\. Yatskar, V\. Ordonez, and K\. Chang \(2018b\)Gender bias in coreference resolution: evaluation and debiasing methods\.InNAACL,Cited by:[§1](https://arxiv.org/html/2608.04268#S1.p2.1)\.
- X\. Zhu, D\. Cheng, H\. Li, K\. Zhang, E\. Hua, X\. Lv, N\. Ding, Z\. Lin, Z\. Zheng, and B\. Zhou \(2024\)How to synthesize text data without model collapse?\.arXiv preprint arXiv:2412\.14689\.Cited by:[§2](https://arxiv.org/html/2608.04268#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix AExperiment Details
### A\.1Training Dataset Statistics
We use theBias in Biosdataset, a large corpus of human\-written biographies, for model training in our experiments\. The profession distribution in the dataset is highly imbalanced, with some professions represented by approximately 1k biographies and others by over 100k\. Such imbalance can independently induce collapse when models are trained on the full corpus\. To isolate the effect of continued pretraining from profession\-frequency imbalance, we construct a balanced training subset rather than training on the overrepresented classes in the original corpus\. Specifically, we uniformly subsample biographies from each profession to match the size of the smallest profession in the original training set\. The resulting corpus contains27,752biographies with a uniform distribution across professions, totaling2M tokensfor the Qwen model tokenizer\.[Figure 5](https://arxiv.org/html/2608.04268#A1.F5)illustrates the distribution of biography lengths in tokens, and[Table 3](https://arxiv.org/html/2608.04268#A1.T3)shows profession\-level gender imbalance across professions in the test subset\.
Figure 5:Token length distribution of biographies in the balancedBias in Biostraining corpus\. The average biography length is 78\.4 tokens \(std\. 37\.7, median 69\.0\) with the Qwen tokenizer\. Vertical markers indicate candidate seed lengths used for seeded generation\. These marked seeds correspond to increasing proportions of model\-generated synthetic content, ranging from 91\.9% \(seed length 5\) to 38\.2% \(seed length 50\), with intermediate values of 83\.8%, 75\.7%, 67\.7%, 59\.6%, 52\.4%, 47\.4%, 43\.6%, and 40\.7%\.Table 3:Profession\-level gender imbalance in the Bias\-in\-Bios evaluation data\. For each professionpp, the imbalance ratio is computed as\|Nfemale,p−Nmale,p\|/\(Nfemale,p\+Nmale,p\)\\left\|N\_\{\\mathrm\{female\},p\}\-N\_\{\\mathrm\{male\},p\}\\right\|/\\left\(N\_\{\\mathrm\{female\},p\}\+N\_\{\\mathrm\{male\},p\}\\right\), whereNfemale,pN\_\{\\mathrm\{female\},p\}andNmale,pN\_\{\\mathrm\{male\},p\}denote the number of female\- and male\-labeled biographies\. Values range from0for balanced professions to11for professions containing biographies from a single gender group\.
### A\.2Data Generation
We consider two data\-generation approaches: seeded continuation and few\-shot prompting from the training set \(KK\-shot\) withK=3K=3shots\.
#### Seeded Generation
We adopt seeded generation as a controlled generation scenario at the token level\. We consider seed lengths in the range\[5,50\]\[5,50\]with a step size of 5\. Shorter seeds correspond to a higher proportion of model\-generated synthetic text; for instance, a seed length of 5 tokens results in 91\.9% synthetic content on average, whereas a seed length of 50 tokens results in 38\.2% synthetic content\. To select seed lengths for seeded generation, we examine the distribution of token lengths in the training data\. In iterative training settings without full corpus access, fixed global synthetic data ratios are difficult to enforce, whereas seeded generation allows direct control over synthetic content within individual samples\.
#### Fewshot Generation
For fewshot generation, we use examples from the training set\. We use the following prompt for fewshot generations:
> Write a single\-paragraph professional biography of a \{profession\}\. Do not use lists, bullet points, headings, or field names\. Write in full sentences only and end with a period\.
### A\.3Additional Details on Wasserstein Distance
To compare the distributions of human\-written and generated biographies in embedding space, we use the entropically regularized22\-Wasserstein distance\. LetX=\{xi\}i=1nX=\\\{x\_\{i\}\\\}\_\{i=1\}^\{n\}andY=\{yj\}j=1mY=\\\{y\_\{j\}\\\}\_\{j=1\}^\{m\}denote the embedding sets associated with two text distributions\. We assume uniform weightsai=1na\_\{i\}=\\frac\{1\}\{n\}andbj=1m\.b\_\{j\}=\\frac\{1\}\{m\}\.
The transport cost between embeddings is defined using the squared Euclidean distance:
Cij=‖xi−yj‖22\.C\_\{ij\}=\\\|x\_\{i\}\-y\_\{j\}\\\|\_\{2\}^\{2\}\.
We compute the Sinkhorn\-regularized optimal transport objective
W2,ε2\(a,b\)=minπ≥0∑i,jπijCij\+ε∑i,jπij\(logπij−1\),W\_\{2,\\varepsilon\}^\{2\}\(a,b\)=\\min\_\{\\pi\\geq 0\}\\sum\_\{i,j\}\\pi\_\{ij\}C\_\{ij\}\+\\varepsilon\\sum\_\{i,j\}\\pi\_\{ij\}\(\\log\\pi\_\{ij\}\-1\),subject to the marginal constraints
∑jπij=ai,∑iπij=bj\.\\sum\_\{j\}\\pi\_\{ij\}=a\_\{i\},\\qquad\\sum\_\{i\}\\pi\_\{ij\}=b\_\{j\}\.
The entropic regularization parameterε\\varepsilonimproves numerical stability and computational efficiency relative to exact optimal transport\. We report the corresponding Sinkhorn distance
W2,ε\(a,b\)=W2,ε2\(a,b\)\.W\_\{2,\\varepsilon\}\(a,b\)=\\sqrt\{W\_\{2,\\varepsilon\}^\{2\}\(a,b\)\}\.
All Wasserstein distances are computed using thePOTlibrary\(Flamaryet al\.,[2024](https://arxiv.org/html/2608.04268#bib.bib40)\)\.
### A\.4Evaluation Details
We evaluate all models on the Bias\-in\-Bios benchmark using the test split of 19\.6K biographies\. For classification, we use a pairwise occupation\-prediction setup\. For each biography, the model is prompted to choose between the gold profession and a fixed paired profession\. We score both profession names as continuations of the prompt and select the one with higher log\-likelihood\. Accuracy is computed from whether the gold profession receives the higher likelihood\. Equal\-opportunity gaps are computed by comparing male and female accuracy within each profession, and the aggregate EO GAP is the RMS of profession\-level gaps, as reported in[subsection 5\.2](https://arxiv.org/html/2608.04268#S5.SS2)\. The order of the two options is randomized for each example\.[Table 4](https://arxiv.org/html/2608.04268#A1.T4)reports profession pairings used for evaluation\. We use the prompt shown in[Figure 6](https://arxiv.org/html/2608.04268#A1.F6)for pairwise occupation prediction\.
Pairwise occupation prediction prompt\. *Question: Which profession best describes the person in the biography below?**Biography:*xx*A\.*profession1*B\.*profession2*Answer:*
Figure 6:Prompt used for pairwise occupation evaluation\. Given a biographyxx, the model chooses between the gold profession and a fixed paired alternative\. The order of the two profession options is alternated across examples\.Table 4:Profession pairings used for the Bias\-in\-Bios pairwise likelihood evaluation\.
### A\.5Training Details
We useQwen2\.5\-0\.5B333[https://hf\.co/Qwen/Qwen2\.5\-0\.5B](https://hf.co/Qwen/Qwen2.5-0.5B)as the base model and perform full\-parameter continued pretraining on a single NVIDIA H100 GPU\. We perform a grid search on the human\-written training corpus over learning rates\{1×10−5,2×10−5,5×10−5\}\\\{1\\times 10^\{\-5\},2\\times 10^\{\-5\},5\\times 10^\{\-5\}\\\}and training epochs\{1,2,3\}\\\{1,2,3\\\}\. Based on this sweep, we use a learning rate of5×10−55\\times 10^\{\-5\}for the main experiments\. The final configuration uses a cosine learning\-rate schedule with a warmup ratio of0\.10\.1, weight decay of0\.10\.1, a per\-device batch size of22, and gradient accumulation over1616steps\.
## Appendix BExtended Results
In this appendix, we report additional evaluation results for the four controlled synthetic\-data settings described in §[3\.3](https://arxiv.org/html/2608.04268#S3.SS3): seeded recursive, few\-shot recursive, seeded iterative, and few\-shot iterative training\.[Figure 7](https://arxiv.org/html/2608.04268#A2.F7)reports gender\-conditioned MTLD changes for selected professions under seeded recursive training\.[Table 5](https://arxiv.org/html/2608.04268#A2.T5)reports the magnitude of gender\-conditioned NLL disparities for selected male\-majority, female\-majority, and near\-neutral professions under recursive training\. Profession\-level EO gaps are reported for seeded recursive training in[Table 6](https://arxiv.org/html/2608.04268#A2.T6), few\-shot recursive training in[Table 7](https://arxiv.org/html/2608.04268#A2.T7), seeded iterative training in[Table 8](https://arxiv.org/html/2608.04268#A2.T8), and few\-shot iterative training in[Table 9](https://arxiv.org/html/2608.04268#A2.T9)\. The corresponding signed gender\-conditioned NLL gaps are reported for few\-shot recursive training in[Table 10](https://arxiv.org/html/2608.04268#A2.T10), seeded recursive training in[Table 11](https://arxiv.org/html/2608.04268#A2.T11), few\-shot iterative training in[Table 12](https://arxiv.org/html/2608.04268#A2.T12), and seeded iterative training in[Table 13](https://arxiv.org/html/2608.04268#A2.T13)\.
Finally,[Table 14](https://arxiv.org/html/2608.04268#A2.T14)provides a qualitative example of recursive seeded generation across iterations, and[Table 15](https://arxiv.org/html/2608.04268#A2.T15)reports the human\-only continued\-pretraining baseline\.
Figure 7:Gender\-conditioned lexical\-diversity trajectories for selected professions across recursive training iterations\. The y\-axis reports the MTLD difference from human\-written biographies, shown separately for male\- and female\-labeled biographies\. “M\-F @5” denotes the male\-female MTLD difference at the final recursive iteration, computed as the male MTLD difference minus the female MTLD difference at Iteration 5\. Setting=seeded recursive\.Table 5:Magnitude of gender\-conditioned NLL disparities for selected male\-majority, female\-majority, and near\-neutral professions under contaminated recursive training\. For gender\-imbalanced professions, values report\|NLLdominant−NLLminority\|\|\\mathrm\{NLL\}\_\{\\mathrm\{dominant\}\}\-\\mathrm\{NLL\}\_\{\\mathrm\{minority\}\}\|, where the dominant group is the majority gender for that profession in the evaluation set\. For near\-neutral professions, values report\|NLLfemale−NLLmale\|\|\\mathrm\{NLL\}\_\{female\}\-\\mathrm\{NLL\}\_\{male\}\|\. Larger values indicate a stronger likelihood disparity between gender\-conditioned biographies within the same profession\.Table 6:Equal opportunity gaps by profession across model iterations, in percentage points for the contaminated seeded setting\. The final row reports the RMS EO gap\. Setting=Seeded contaminated\.Table 7:Equal opportunity gaps by profession across model iterations, in percentage points\. Setting=Few\-shot contaminated\.Table 8:Equal opportunity gaps by profession across model iterations, in percentage points\. Setting= Seeded not contaminated\.Table 9:Equal opportunity gaps by profession across model iterations, in percentage points\. The final row reports the RMS EO gap\. Setting= fewshot not contaminated\.Table 10:Signed gender\-conditioned NLL gap by profession across iterations under the contaminated few\-shot regime\. Values reportΔNLL=NLLfemale−NLLmale\\Delta\\mathrm\{NLL\}=\\mathrm\{NLL\}\_\{female\}\-\\mathrm\{NLL\}\_\{male\}\. Positive values indicate female biographies are less likely under the model; negative values indicate male biographies are less likely\. Setting=fewshot, contaminated\.Table 11:Signed gender\-conditioned NLL gap by profession across model iterations\. Values reportΔNLL=NLLfemale−NLLmale\\Delta\\mathrm\{NLL\}=\\mathrm\{NLL\}\_\{female\}\-\\mathrm\{NLL\}\_\{male\}\. Positive values indicate female biographies are less likely under the model; negative values indicate male biographies are less likely\. Setting=seeded, contaminated\.Table 12:Signed gender\-conditioned NLL gap by profession across iterations under theiterative few\-shotregime\. Values reportΔNLL=NLLfemale−NLLmale\\Delta\\mathrm\{NLL\}=\\mathrm\{NLL\}\_\{female\}\-\\mathrm\{NLL\}\_\{male\}\. Positive values indicate female biographies are less likely under the model; negative values indicate male biographies are less likely\. Setting=few\-shot, iterative\.Table 13:Signed gender\-conditioned NLL gap by profession across iterations under theseeded iterativeregime\. Values reportΔNLL=NLLfemale−NLLmale\\Delta\\mathrm\{NLL\}=\\mathrm\{NLL\}\_\{female\}\-\\mathrm\{NLL\}\_\{male\}\. Positive values indicate female biographies are less likely under the model; negative values indicate male biographies are less likely\. Setting=seeded, iterative\.Table 14:An example of model continuations over recursive training iterations\. A single biography cannot fully capture distribution\-level changes, but it can reveal recurring qualitative patterns\. In this example, the human continuation gives specific details about a dietitian and wellness professional\. After several iterations of training on synthetic text, the model increasingly changes institutions and credentials, shifts toward a more generic clinical framing, and produces less faithful biographical details and web\-page artifacts\.Table 15:Perplexity, MMLU, and Bias\-in\-Bios results for models trained on thehuman\-writtencorpus\. PPL is measured on the same human\-written corpus used for training\. Bias\-in\-Bios includes the signed gender\-conditioned NLL gap, occupation\-prediction accuracy, and equal\-opportunity gap across professions\. Bias\-in\-Bios accuracy fluctuates until iteration 4, while the EO gap stays roughly between 11 and 13\. This differs from the iterative and recursive synthetic\-data regimes, where the EO gap generally increases across iterations before the MMLU drop becomes significant\. For seeded generation, the EO gap reaches about 14\-18 under the iterative setting and 15\-20 under the recursive setting\. For few\-shot generation, it reaches about 16\-19 under an iterative setting and 16\-18 under a recursive setting\.Similar Articles
What's up with model collapse?
An exploration of model collapse, a phenomenon where AI models trained on synthetic data degrade in quality and diversity.
Lessons learned on language model safety and misuse
OpenAI shares lessons learned on language model safety and misuse, discussing challenges in measuring risks, the limitations of existing benchmarks, and their development of new evaluation metrics for toxicity and policy violations. The post also highlights concerns about labor market impacts and the need for continued research on measuring social effects of AI deployment at scale.
Large language models develop novel social biases through adaptive exploration
This research explores how large language models develop new social biases through adaptive exploration, highlighting implications for AI fairness and ethical considerations.
Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media
This paper proposes a Cross-Platform Fairness Evaluation framework to audit transformer models for mental health NLP, revealing significant performance and calibration failures when models are applied across different social media platforms.
Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation
This paper introduces a semantic-aware mixed-effects regression framework for measuring fairness in Large Audio Language Models by controlling for semantic variation and speaker identity to yield more robust and interpretable bias estimates.