model-collapse

Tag

Cards List
#model-collapse

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

Reddit r/ArtificialInteligence · 11h ago Cached

AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.

0 favorites 0 likes
#model-collapse

Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning

arXiv cs.CL · 21h ago Cached

This paper studies model collapse in iterative instruction tuning with synthetic data, revealing that collapse manifests as polarization of competence where strong skills are reinforced while weak ones degrade. It proposes KITE, a two-stage framework combining failure-guided data generation and boundary-aware uncertainty curation to ensure stable improvement across iterations.

0 favorites 0 likes
#model-collapse

Model collapse + skill atrophy + competitive pressure = one big feedback loop. Thoughts?

Reddit r/ArtificialInteligence · 6d ago

The article discusses a consulting firm's argument that model collapse, human cognitive debt (skill atrophy), and competitive pressure form a self-reinforcing feedback loop in AI, and questions whether organizations can resist the race to automate.

0 favorites 0 likes
#model-collapse

What's up with model collapse?

Reddit r/LocalLLaMA · 2026-07-09

An exploration of model collapse, a phenomenon where AI models trained on synthetic data degrade in quality and diversity.

0 favorites 0 likes
#model-collapse

When Sample Selection Bias Precipitates Model Collapse

arXiv cs.AI · 2026-06-15 Cached

This paper demonstrates that data selection in low-resource verification regimes, where verifiers only have access to fragmented and biased slices of the target distribution, can paradoxically accelerate model collapse by pruning globally relevant tail modes. The authors provide theoretical proof and propose a collaborative proxy reference mechanism as a mitigation strategy.

0 favorites 0 likes
#model-collapse

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

arXiv cs.CL · 2026-06-05 Cached

This paper proposes a bilayer coupled SIR/SIRS framework to model synthetic data contamination and model collapse in AI ecosystems, showing that cross-contamination between models and data corpora leads to supercritical dynamics and identifying detection-based filtering as a key intervention.

0 favorites 0 likes
#model-collapse

The interesting part of model collapse isn't technical, it's epistemic

Reddit r/AI_Agents · 2026-06-01

This article explores model collapse not as a technical bug but as an epistemic problem: when an AI model's outputs become its own inputs, the model's representation of reality gradually flattens into a self-referential average, raising questions about how we distinguish a model that models the world from one that models only itself.

0 favorites 0 likes
#model-collapse

When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

arXiv cs.AI · 2026-05-29 Cached

This paper studies self-consuming training in a multi-model regime, showing that human curation can backfire and degrade long-term alignment due to cross-model interactions.

0 favorites 0 likes
#model-collapse

Model Collapse as Cultural Evolution

arXiv cs.CL · 2026-05-25 Cached

This paper reframes model collapse in LLMs as a cultural transmission phenomenon, showing that iterated learning theory predicts a non-monotonic trajectory of compositionality under self-training, confirmed across multiple languages and models.

0 favorites 0 likes
#model-collapse

How can we prevent AI models from cannibalizing themselves when human-generated data runs out? Scientists say they've found the answer.

Reddit r/artificial · 2026-05-22 Cached

Scientists claim to have found a solution to prevent AI models from cannibalizing themselves when human-generated data runs out, addressing the problem of model collapse where LLMs trained on synthetic data produce gibberish and hallucinations.

0 favorites 0 likes
#model-collapse

Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies

arXiv cs.CL · 2026-05-21 Cached

This paper presents evidence that self-training on language model outputs does not uniformly flatten language but restructures it, with surface markers (discourse connectives, hedges, em-dashes) increasing while deep syntactic structures (passives, subjunctives, parentheticals) collapse, formalized as the Structural Depth Hypothesis.

0 favorites 0 likes
#model-collapse

AI is deteriorating in realtime

Reddit r/ArtificialInteligence · 2026-05-20

AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.

0 favorites 0 likes
#model-collapse

On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning

arXiv cs.LG · 2026-05-08 Cached

This paper identifies a critical 'model collapse' issue in standard fine-tuning for causal reasoning and proposes a semantic loss function with graph-based logical constraints to prevent it.

0 favorites 0 likes
#model-collapse

The Problem with “Mathematically Proven” Claims About LLMs (15 minute read)

TLDR AI · 2026-05-07 Cached

This article critiques the sensationalized media coverage of mathematical proofs regarding LLM limitations, specifically highlighting how conditional results about self-improvement are often misrepresented as universal impossibilities.

0 favorites 0 likes
← Back to home

Submit Feedback