A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.
Hey hey folks, I’ve been thinking about an odd consequence of the generative AI boom. Especially in light of these doomer stories about Anthropic destroying books (boo bad Anthropic bad). The first major LLMs inherited decades of internet that was overwhelmingly produced by humans. Now those same systems and their descendants are producing articles, code, summaries, books, comments, and other material that ends up back in the information environment. Obviously synthetic data itself isn’t inherently bad. Carefully generated and filtered synthetic data can be extremely useful. What interests me is provenance. A book printed in 1980 has a very obvious property: whatever else is wrong with it, it wasn’t written with an LLM. The same applies to old forums, archived websites, academic work, old documentation and other pre-generative material. Does that historical corpus become unusually useful precisely because we know something about its origin? I wrote a longer piece exploring this through Anthropic’s physical book scanning, recursive training/model collapse, old internet archives and human-authorship certification. Full disclosure, it’s mine: https://www.gonzocapital.net/the-internet-ouroboros/ But I’m more interested in the underlying question: does provenance become materially more important for training data, or are filtering and verification techniques good enough that the age/origin of the corpus becomes mostly irrelevant?
This article argues that Reddit's messy, authentic human conversations are becoming increasingly valuable for training AI as the web fills with synthetic content, highlighting the economic shift toward scarce human behavioral data.
As AI models consume finite human-generated data, future training may rely on synthetic data from other AIs, raising questions about long-term implications.
AI is causing online content to be optimized for machines, leading to data scraping and a decline in original website traffic, replacing human content with synthetic alternatives.
The article questions whether the proliferation of AI-generated content on the internet could diminish the quality of human-created information and hinder future AI learning by creating a feedback loop of AI training on AI-generated data.