model-collapse

Tag

Cards List
#model-collapse

RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored

arXiv cs.CL · 2026-08-25 Cached

This paper demonstrates that retrieval-augmented generation (RAG) systems suffer from collapse when retrieving self-authored documents, leading to reduced response diversity and self-bias, with experiments showing collapse in 79.6% of simulations.

0 favorites 0 likes
#model-collapse

Reviewing Model Collapse and Countermeasures

arXiv cs.AI · 2026-08-25 Cached

This paper provides an up-to-date overview of the phenomenon of model collapse in generative AI and reviews countermeasures to mitigate it, highlighting challenges and future research opportunities.

0 favorites 0 likes
#model-collapse

Is (or will) AI learn backwards? (Since most of its training data is now AI-generated data)

Reddit r/ArtificialInteligence · 2026-08-21

The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.

0 favorites 0 likes
#model-collapse

Does pre-generative-AI data become more valuable as the internet fills with synthetic material?

Reddit r/artificial · 2026-08-12

A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.

0 favorites 0 likes
#model-collapse

The Fairness Collapse Phenomenon: Bias Amplification in Language Models Trained on Synthetic Data

arXiv cs.CL · 2026-08-06 Cached

This paper introduces the 'fairness collapse' phenomenon, showing that training language models on synthetic data silently amplifies social biases before standard model collapse metrics degrade, highlighting a critical risk for AI fairness.

0 favorites 0 likes
#model-collapse

How does AI know what's AI generated?

Reddit r/ArtificialInteligence · 2026-07-28

Explores the problem of AI models training on AI-generated content, leading to potential degradation and loss of truth as inaccuracies compound.

0 favorites 0 likes
#model-collapse

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

Reddit r/ArtificialInteligence · 2026-07-21 Cached

AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.

0 favorites 0 likes
#model-collapse

Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning

arXiv cs.CL · 2026-07-21 Cached

This paper studies model collapse in iterative instruction tuning with synthetic data, revealing that collapse manifests as polarization of competence where strong skills are reinforced while weak ones degrade. It proposes KITE, a two-stage framework combining failure-guided data generation and boundary-aware uncertainty curation to ensure stable improvement across iterations.

0 favorites 0 likes
#model-collapse

Model collapse + skill atrophy + competitive pressure = one big feedback loop. Thoughts?

Reddit r/ArtificialInteligence · 2026-07-15

The article discusses a consulting firm's argument that model collapse, human cognitive debt (skill atrophy), and competitive pressure form a self-reinforcing feedback loop in AI, and questions whether organizations can resist the race to automate.

0 favorites 0 likes
#model-collapse

What's up with model collapse?

Reddit r/LocalLLaMA · 2026-07-09

An exploration of model collapse, a phenomenon where AI models trained on synthetic data degrade in quality and diversity.

0 favorites 0 likes
#model-collapse

When Sample Selection Bias Precipitates Model Collapse

arXiv cs.AI · 2026-06-15 Cached

This paper demonstrates that data selection in low-resource verification regimes, where verifiers only have access to fragmented and biased slices of the target distribution, can paradoxically accelerate model collapse by pruning globally relevant tail modes. The authors provide theoretical proof and propose a collaborative proxy reference mechanism as a mitigation strategy.

0 favorites 0 likes
#model-collapse

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

arXiv cs.CL · 2026-06-05 Cached

This paper proposes a bilayer coupled SIR/SIRS framework to model synthetic data contamination and model collapse in AI ecosystems, showing that cross-contamination between models and data corpora leads to supercritical dynamics and identifying detection-based filtering as a key intervention.

0 favorites 0 likes
#model-collapse

The interesting part of model collapse isn't technical, it's epistemic

Reddit r/AI_Agents · 2026-06-01

This article explores model collapse not as a technical bug but as an epistemic problem: when an AI model's outputs become its own inputs, the model's representation of reality gradually flattens into a self-referential average, raising questions about how we distinguish a model that models the world from one that models only itself.

0 favorites 0 likes
#model-collapse

When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

arXiv cs.AI · 2026-05-29 Cached

This paper studies self-consuming training in a multi-model regime, showing that human curation can backfire and degrade long-term alignment due to cross-model interactions.

0 favorites 0 likes
#model-collapse

Model Collapse as Cultural Evolution

arXiv cs.CL · 2026-05-25 Cached

This paper reframes model collapse in LLMs as a cultural transmission phenomenon, showing that iterated learning theory predicts a non-monotonic trajectory of compositionality under self-training, confirmed across multiple languages and models.

0 favorites 0 likes
#model-collapse

How can we prevent AI models from cannibalizing themselves when human-generated data runs out? Scientists say they've found the answer.

Reddit r/artificial · 2026-05-22 Cached

Scientists claim to have found a solution to prevent AI models from cannibalizing themselves when human-generated data runs out, addressing the problem of model collapse where LLMs trained on synthetic data produce gibberish and hallucinations.

0 favorites 0 likes
#model-collapse

Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies

arXiv cs.CL · 2026-05-21 Cached

This paper presents evidence that self-training on language model outputs does not uniformly flatten language but restructures it, with surface markers (discourse connectives, hedges, em-dashes) increasing while deep syntactic structures (passives, subjunctives, parentheticals) collapse, formalized as the Structural Depth Hypothesis.

0 favorites 0 likes
#model-collapse

AI is deteriorating in realtime

Reddit r/ArtificialInteligence · 2026-05-20

AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.

0 favorites 0 likes
#model-collapse

On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning

arXiv cs.LG · 2026-05-08 Cached

This paper identifies a critical 'model collapse' issue in standard fine-tuning for causal reasoning and proposes a semantic loss function with graph-based logical constraints to prevent it.

0 favorites 0 likes
#model-collapse

The Problem with “Mathematically Proven” Claims About LLMs (15 minute read)

TLDR AI · 2026-05-07 Cached

This article critiques the sensationalized media coverage of mathematical proofs regarding LLM limitations, specifically highlighting how conditional results about self-improvement are often misrepresented as universal impossibilities.

0 favorites 0 likes
← Back to home

Submit Feedback