ai-training-data

Tag

Cards List
#ai-training-data

Forbes: Surge AI, Mercor, AfterQuery and Turing Sold ~$500M/Year of Training Data to Tencent, Alibaba and ByteDance

Reddit r/ArtificialInteligence · 2d ago

Forbes reports that US AI data vendors Surge AI, Mercor, AfterQuery, and Turing are selling ~$500M/year in training data to Chinese labs Tencent, Alibaba, and ByteDance, despite also serving OpenAI and Anthropic.

0 favorites 0 likes
#ai-training-data

Silicon Valley’s Other China Problem: It’s Training Their AI

Reddit r/ArtificialInteligence · 5d ago Cached

This Forbes article reveals that U.S. data-labeling startups like Surge AI and Mercor are secretly selling high-quality training data to Chinese AI labs, helping China close the gap with American AI despite U.S. chip restrictions. The piece highlights the overlooked role of data infrastructure in the AI arms race and the regulatory blind spot around exporting proprietary data.

0 favorites 0 likes
#ai-training-data

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

Reddit r/ArtificialInteligence · 2026-07-21 Cached

AI companies are buying pre-2022 printed books to avoid AI-generated text in training data, as old books are guaranteed free of AI slop and poisoning. ISBNdb offers bulk book acquisition services to AI labs under NDAs.

0 favorites 0 likes
#ai-training-data

What happens when AI runs out of human-made data?

Reddit r/ArtificialInteligence · 2026-07-16

As AI models consume finite human-generated data, future training may rely on synthetic data from other AIs, raising questions about long-term implications.

0 favorites 0 likes
#ai-training-data

Did you know the CEO of OpenAI owns nearly 9% of Reddit while Reddit bans users for AI generated content?

Reddit r/artificial · 2026-07-14

Sam Altman, CEO of OpenAI, holds nearly 9% of Reddit stock, making him the third largest shareholder, while Reddit bans users for AI-generated content but sold user data to Google for AI training, highlighting conflicts of interest.

0 favorites 0 likes
#ai-training-data

Stripe might be one of the strongest feeders into frontier AI

Reddit r/ArtificialInteligence · 2026-07-02

Stripe is suggested to be a major source of data feeding frontier AI models, highlighting its role in AI development.

0 favorites 0 likes
#ai-training-data

A significant portion of the remaining training data for AI is located on magnetic tapes stored in warehouses.

Reddit r/artificial · 2026-06-24

A significant portion of remaining AI training data is on undigitized magnetic tapes stored in warehouses, highlighting a potential data source as internet-based data runs out.

0 favorites 0 likes
#ai-training-data

Tech companies desperately want to film you doing chores

The Verge · 2026-05-29 Cached

Tech companies like Shift and Pronto are offering free services or payment to record people doing chores in order to gather training data for physical AI and robotics, raising privacy and ethical concerns.

0 favorites 0 likes
#ai-training-data

LQS v3.1 — an open methodology for rating AI training data (multi-oracle consensus + signed certificates) [P]

Reddit r/MachineLearning · 2026-05-23

The author presents LQS v3.1, an open methodology for rating AI training data using multi-oracle consensus and signed certificates, with a published paper and public index. The approach aims to solve the bottleneck of independent quality evaluation in the AI training data market.

0 favorites 0 likes
#ai-training-data

Anna's Archive Hit with $19.5M Default Judgment and Global Domain Takedown Order

Hacker News Top · 2026-05-20 Cached

A New York federal judge granted a $19.5 million default judgment against shadow library Anna's Archive and ordered a global domain takedown, after major publishers sued over copyright infringement and the site's use as a training data hub for AI companies.

0 favorites 0 likes
#ai-training-data

Meta employees protest new mouse-tracking software days before mass layoffs

Reddit r/ArtificialInteligence · 2026-05-13 Cached

Meta employees are protesting the company's new mouse-tracking software, the Model Capability Initiative, arguing it constitutes invasive surveillance just before a major round of layoffs.

0 favorites 0 likes
#ai-training-data

@FinanceYF5: Meta illegally downloaded over 80 TB of books from LibGen, Anna's Archive, and Z-Library to train its AI models. Aaron Swartz downloaded 70 GB of papers from JSTOR in 2010 (only equivalent to...

X AI KOLs Following · 2026-05-08 Cached

Meta is accused of illegally downloading over 80 TB of books from LibGen, Anna's Archive, and Z-Library to train AI models. The article contrasts the case of Aaron Swartz, who faced severe charges for downloading a much smaller amount of papers, highlighting the double standard in copyright enforcement.

0 favorites 0 likes
← Back to home

Submit Feedback