People need to start paying attention to the issue of derived data in AI training (2 minute read)

TLDR AI News

Summary

The thread discusses the harm of derived data in AI training to creatives, where AI rewrites content to evade detection, and calls for mandatory training data transparency laws.

Derived data is creative content that is rewritten by AI before it is trained on.
Original Article
View Cached Full Text

Cached at: 09/23/26, 02:41 PM

# Thread by @ednewtonrex on Thread Reader App Source: [https://threadreaderapp.com/thread/2102311135441547405.html](https://threadreaderapp.com/thread/2102311135441547405.html) People need to start paying attention to the issue of derived data in AI training\. There’s rightfully been a big focus on the pirated / web\-scraped data AI companies train on \(resulting in well over 100 lawsuits\) \- and, increasingly, on synthetic data made by ‘tainted’ models\. Derived data gets talked about less, but does the same harm to creatives\. Derived data is creative content that is rewritten by AI before it is trained on\. So instead of training directly on sentences from your book, a company will first use an AI model to rewrite those sentences, then train on the results\. It is harder to detect, because the resulting model is less likely to spit out your work verbatim\. But it does the same harm: your work is used without permission to build a product that competes with you\. This gets easier to do the better AI models get, so it’s happening more and more\. And it makes it that much harder for creatives to know when their work is being exploited, and therefore that much less likely they can defend themselves\. This is also a huge issue because it may be a way for AI companies to try to get around opt\-outs\. Perhaps you have put your work on some platform, and you have opted out of the company that runs the platform training on it\. But you likely haven’t opted out of them rewriting it and training on that\. They will create derived data from it, train on it, and later say they did nothing wrong\. The only real solution to this is training data transparency requirements\. Governments must introduce laws that force AI companies to reveal their training data\. Until that happens, creatives are forever playing catch\-up, at best finding out their work has been exploited long after the fact, at worst never finding out and therefore unable to defend against it\.

Similar Articles