Scientists claim to have found a solution to prevent AI models from cannibalizing themselves when human-generated data runs out, addressing the problem of model collapse where LLMs trained on synthetic data produce gibberish and hallucinations.
# How can we prevent AI models from cannibalizing themselves when human-generated data runs out? Scientists say…
Source: [https://www.livescience.com/technology/artificial-intelligence/how-can-we-prevent-ai-models-from-cannibalizing-themselves-when-human-generated-data-runs-out-scientists-say-theyve-found-the-answer](https://www.livescience.com/technology/artificial-intelligence/how-can-we-prevent-ai-models-from-cannibalizing-themselves-when-human-generated-data-runs-out-scientists-say-theyve-found-the-answer)
While the evolution of[artificial intelligence](https://www.livescience.com/technology/artificial-intelligence/what-is-artificial-intelligence-ai)\(AI\) systems has shown no sign of slowing, there's a growing concern that large language models \(LLMs\) will soon run out of human\-made data to ingest and learn from\.
Once this happens, scientists say, AI models will increasingly rely on synthetic AI\-made information, which will lead to an effect called "[model collapse](https://www.livescience.com/technology/artificial-intelligence/ai-models-trained-on-ai-generated-data-could-spiral-into-unintelligible-nonsense-scientists-warn)\." This is where LLMs spout gibberish and the AI systems they underpin deliver inaccurate answers and hallucinate information to queries far more commonly than they do today\.
Get the world’s most fascinating discoveries delivered straight to your inbox\.
As AI models consume finite human-generated data, future training may rely on synthetic data from other AIs, raising questions about long-term implications.
AI models are deteriorating due to training on recursively generated synthetic data, leading to model collapse; multiple studies highlight the risks of scaling with synthetic data.
This paper provides an up-to-date overview of the phenomenon of model collapse in generative AI and reviews countermeasures to mitigate it, highlighting challenges and future research opportunities.