The thread discusses the harm of derived data in AI training to creatives, where AI rewrites content to evade detection, and calls for mandatory training data transparency laws.
Derived data is creative content that is rewritten by AI before it is trained on.
# Thread by @ednewtonrex on Thread Reader App
Source: [https://threadreaderapp.com/thread/2102311135441547405.html](https://threadreaderapp.com/thread/2102311135441547405.html)
People need to start paying attention to the issue of derived data in AI training\.
There’s rightfully been a big focus on the pirated / web\-scraped data AI companies train on \(resulting in well over 100 lawsuits\) \- and, increasingly, on synthetic data made by ‘tainted’ models\. Derived data gets talked about less, but does the same harm to creatives\.
Derived data is creative content that is rewritten by AI before it is trained on\. So instead of training directly on sentences from your book, a company will first use an AI model to rewrite those sentences, then train on the results\.
It is harder to detect, because the resulting model is less likely to spit out your work verbatim\. But it does the same harm: your work is used without permission to build a product that competes with you\.
This gets easier to do the better AI models get, so it’s happening more and more\. And it makes it that much harder for creatives to know when their work is being exploited, and therefore that much less likely they can defend themselves\.
This is also a huge issue because it may be a way for AI companies to try to get around opt\-outs\. Perhaps you have put your work on some platform, and you have opted out of the company that runs the platform training on it\. But you likely haven’t opted out of them rewriting it and training on that\. They will create derived data from it, train on it, and later say they did nothing wrong\.
The only real solution to this is training data transparency requirements\. Governments must introduce laws that force AI companies to reveal their training data\. Until that happens, creatives are forever playing catch\-up, at best finding out their work has been exploited long after the fact, at worst never finding out and therefore unable to defend against it\.
The article explores ethical questions about transparency in AI-generated content, such as novels and websites, and whether consumers should be informed when AI is used in creative or commercial work.
This article argues that Reddit's messy, authentic human conversations are becoming increasingly valuable for training AI as the web fills with synthetic content, highlighting the economic shift toward scarce human behavioral data.
The article discusses how personal data is often used in AI training without consent, highlighting privacy concerns, and introduces a tool called 'Don’t Train Me' that automates opt-out requests to protect user data.
The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.
Discussion of whether AI companies should face copyright lawsuits to force ethical training data sourcing, citing Suno's loss in Germany and arguing laws should require consent for training on people's data.