training-data

Tag

Cards List
#training-data

Does pre-generative-AI data become more valuable as the internet fills with synthetic material?

Reddit r/artificial · 14h ago

A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.

0 favorites 0 likes
#training-data

Why isn't there more talk of AI being stopped by copyright laws?

Reddit r/ArtificialInteligence · 3d ago

Discussion of whether AI companies should face copyright lawsuits to force ethical training data sourcing, citing Suno's loss in Germany and arguing laws should require consent for training on people's data.

0 favorites 0 likes
#training-data

@macrodata_labs: Everyone is betting on Egocentric data to scale robotics But turning that footage into training data requires recoverin…

X AI KOLs Following · 6d ago Cached

Macrodata Labs releases a research blog on scaling robotics with egocentric video data by recovering 3D hand motion signals using only open-source models.

0 favorites 0 likes
#training-data

TIL AI can draw a watch showing an actual time

Reddit r/singularity · 2026-08-04

The author observes that ChatGPT can now draw a watch with a specified time, noting this was a classic AI failure a year ago due to training data over-representing 10:10, and asks what other '10:10 problems' remain.

0 favorites 0 likes
#training-data

@turingbook: The greatest potential value of Harness (Agent) may be going into the real work scenarios of seasoned experts across industries, observing how they work and think, and recording and "distilling" their valuable experience—data that was previously largely tacit.

X AI KOLs Timeline · 2026-08-02 Cached

The tweet highlights the potential of agent harnesses to capture and distill experts' tacit knowledge as new training data, referencing a Berkeley AI Summit talk by Jianfeng Gao on agentic modeling as an emerging AI paradigm.

0 favorites 0 likes
#training-data

@julien_c: Happy to partner with @trufflesec to help them perform the largest secret scan of AI training data ever

X AI KOLs Following · 2026-07-31 Cached

Truffle Security, in partnership with Julien Chaumond, conducted the largest secret scan of AI training data on HuggingFace, finding 221,303 live unique credentials across 6,003 public datasets.

0 favorites 0 likes
#training-data

How does AI know what's AI generated?

Reddit r/ArtificialInteligence · 2026-07-28

Explores the problem of AI models training on AI-generated content, leading to potential degradation and loss of truth as inaccuracies compound.

0 favorites 0 likes
#training-data

Training data needs a real go/no-go gate before training [D]

Reddit r/MachineLearning · 2026-07-27

The author proposes a formal pre-training control layer that audits training data artifacts and provides a verdict (PASS/FAIL) based on explicit criteria, as a missing gate between data preparation and training, and invites discussion on its practicality.

0 favorites 0 likes
#training-data

Screencap

Product Hunt · 2026-07-27

Screencap turns your team's real workflows into AI training data.

0 favorites 0 likes
#training-data

Are brain waves the next unlock for physical AI?

TechCrunch AI · 2026-07-27 Cached

Encord and Zander Labs are experimenting with brain wave headsets to collect richer training data for physical AI, aiming to solve the scarcity of real-world robotics data.

0 favorites 0 likes
#training-data

What data mix are the labs using to train 10T param models?

Reddit r/singularity · 2026-07-25

Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.

0 favorites 0 likes
#training-data

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

arXiv cs.LG · 2026-07-24 Cached

DataPrep-Bench is a unified benchmark evaluating LLMs' capabilities in training data construction and quality evaluation across six domains, including a skill-guided agent (Data-Construction-Skill) and a distribution-based evaluator (DAS) that achieves strong cross-model correlation.

0 favorites 0 likes
#training-data

Anthropic got sued for using copyrighted books for LLM training

Reddit r/LocalLLaMA · 2026-07-21

Anthropic is being sued for allegedly using copyrighted books without permission to train its large language models.

0 favorites 0 likes
#training-data

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Hugging Face Daily Papers · 2026-07-21 Cached

This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.

0 favorites 0 likes
#training-data

Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments

arXiv cs.CL · 2026-07-20 Cached

TerminalTraj is a scalable pipeline that generates high-quality terminal agent trajectories using Dockerized environments, achieving up to 20% improvement on terminal task benchmarks with models like TerminalTraj-32B.

0 favorites 0 likes
#training-data

@viks_rum: https://x.com/viks_rum/status/2077650169265590727

X AI KOLs Timeline · 2026-07-16 Cached

An in-depth analysis of the booming business of selling training data to frontier AI labs, detailing six distinct data products and the financial dynamics of the market.

0 favorites 0 likes
#training-data

Suno snatched millions of songs from YouTube, Genius, and Deezer

The Verge · 2026-07-15 Cached

A hack reveals that AI music generator Suno trained its models by scraping millions of songs from YouTube, Genius, and Deezer, backing up allegations of copyright infringement and exposing customer data.

0 favorites 0 likes
#training-data

Hack suggests AI music generator Suno scraped YouTube for training data

TechCrunch AI · 2026-07-15 Cached

A hack into AI music generator Suno reveals it allegedly scraped decades of audio from YouTube, Deezer, and other sources for training data, raising copyright and DMCA concerns amid ongoing lawsuits from major record labels.

0 favorites 0 likes
#training-data

The potential problem of perception hijack

Reddit r/ArtificialInteligence · 2026-07-14

The article discusses how AI models exhibit a bias toward statistically dominant narratives in training data, which could be exploited to manipulate historical and current contexts on a global scale.

0 favorites 0 likes
#training-data

this openai court story is starting to look ugly

Reddit r/artificial · 2026-07-12

The New York Times alleges OpenAI misled the court by claiming it couldn't search training data for copyrighted material, despite having previously done so and deleting billions of chat logs. The case raises concerns about transparency and accountability in AI companies.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback