Tag
Forbes reports that US AI data vendors Surge AI, Mercor, AfterQuery, and Turing are selling ~$500M/year in training data to Chinese labs Tencent, Alibaba, and ByteDance, despite also serving OpenAI and Anthropic.
This Forbes article reveals that U.S. data-labeling startups like Surge AI and Mercor are secretly selling high-quality training data to Chinese AI labs, helping China close the gap with American AI despite U.S. chip restrictions. The piece highlights the overlooked role of data infrastructure in the AI arms race and the regulatory blind spot around exporting proprietary data.
AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.
As AI models consume finite human-generated data, future training may rely on synthetic data from other AIs, raising questions about long-term implications.
Sam Altman, CEO of OpenAI, holds nearly 9% of Reddit stock, making him the third largest shareholder, while Reddit bans users for AI-generated content but sold user data to Google for AI training, highlighting conflicts of interest.
Stripe is suggested to be a major source of data feeding frontier AI models, highlighting its role in AI development.
A significant portion of remaining AI training data is on undigitized magnetic tapes stored in warehouses, highlighting a potential data source as internet-based data runs out.
Tech companies like Shift and Pronto are offering free services or payment to record people doing chores in order to gather training data for physical AI and robotics, raising privacy and ethical concerns.
The author presents LQS v3.1, an open methodology for rating AI training data using multi-oracle consensus and signed certificates, with a published paper and public index. The approach aims to solve the bottleneck of independent quality evaluation in the AI training data market.
A New York federal judge granted a $19.5 million default judgment against shadow library Anna's Archive and ordered a global domain takedown, after major publishers sued over copyright infringement and the site's use as a training data hub for AI companies.
Meta employees are protesting the company's new mouse-tracking software, the Model Capability Initiative, arguing it constitutes invasive surveillance just before a major round of layoffs.
Meta is accused of illegally downloading over 80 TB of books from LibGen, Anna's Archive, and Z-Library to train AI models. The article contrasts the case of Aaron Swartz, who faced severe charges for downloading a much smaller amount of papers, highlighting the double standard in copyright enforcement.