Your personal data is probably being used to train AI — and most people have no idea

Reddit r/ArtificialInteligence Tools

Summary

The article discusses how personal data is often used in AI training without consent, highlighting privacy concerns, and introduces a tool called 'Don’t Train Me' that automates opt-out requests to protect user data.

I’ve been digging into AI training/privacy recently and some of the numbers are pretty wild. The UK’s ICO says generative AI training involves “vast amounts of personal data”, often processed without people knowing it’s happening. It says web-scraped training datasets can contain information relating to millions, if not billions, of people. And “public” doesn’t necessarily mean harmless. Research has demonstrated neural networks memorising unique information like names and IDs even when it appeared in just ONE training sample. The ICO gives a good example: someone posting about a doctor’s visit in 2020 probably wasn’t expecting that post to be scraped years later to train an AI model. Obviously this doesn’t mean ChatGPT has memorised everyone’s private information — newer research actually suggests some claims around PII memorisation have been overstated. But it made me wonder: how many people actually know they can object/opt out with some AI companies? The problem is every company has a different process, form or email address, and policies change. So I’ve built Don’t Train Me to automate the process and periodically resubmit opt-out requests: https://donttrainme.com I’m still very early with it, so genuinely interested in feedback — particularly whether people here actually care about opting out of training, or whether you consider public internet data fair game for AI.
Original Article

Similar Articles