Tag
AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.
This paper studies model collapse in iterative instruction tuning with synthetic data, revealing that collapse manifests as polarization of competence where strong skills are reinforced while weak ones degrade. It proposes KITE, a two-stage framework combining failure-guided data generation and boundary-aware uncertainty curation to ensure stable improvement across iterations.
The article discusses a consulting firm's argument that model collapse, human cognitive debt (skill atrophy), and competitive pressure form a self-reinforcing feedback loop in AI, and questions whether organizations can resist the race to automate.
An exploration of model collapse, a phenomenon where AI models trained on synthetic data degrade in quality and diversity.
This paper demonstrates that data selection in low-resource verification regimes, where verifiers only have access to fragmented and biased slices of the target distribution, can paradoxically accelerate model collapse by pruning globally relevant tail modes. The authors provide theoretical proof and propose a collaborative proxy reference mechanism as a mitigation strategy.
This paper proposes a bilayer coupled SIR/SIRS framework to model synthetic data contamination and model collapse in AI ecosystems, showing that cross-contamination between models and data corpora leads to supercritical dynamics and identifying detection-based filtering as a key intervention.
This article explores model collapse not as a technical bug but as an epistemic problem: when an AI model's outputs become its own inputs, the model's representation of reality gradually flattens into a self-referential average, raising questions about how we distinguish a model that models the world from one that models only itself.
This paper studies self-consuming training in a multi-model regime, showing that human curation can backfire and degrade long-term alignment due to cross-model interactions.
This paper reframes model collapse in LLMs as a cultural transmission phenomenon, showing that iterated learning theory predicts a non-monotonic trajectory of compositionality under self-training, confirmed across multiple languages and models.
Scientists claim to have found a solution to prevent AI models from cannibalizing themselves when human-generated data runs out, addressing the problem of model collapse where LLMs trained on synthetic data produce gibberish and hallucinations.
This paper presents evidence that self-training on language model outputs does not uniformly flatten language but restructures it, with surface markers (discourse connectives, hedges, em-dashes) increasing while deep syntactic structures (passives, subjunctives, parentheticals) collapse, formalized as the Structural Depth Hypothesis.
AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.
This paper identifies a critical 'model collapse' issue in standard fine-tuning for causal reasoning and proposes a semantic loss function with graph-based logical constraints to prevent it.
This article critiques the sensationalized media coverage of mathematical proofs regarding LLM limitations, specifically highlighting how conditional results about self-improvement are often misrepresented as universal impossibilities.