training-data

Tag

Cards List
#training-data

@julien_c: Happy to partner with @trufflesec to help them perform the largest secret scan of AI training data ever

X AI KOLs Following ↗ · 2026-07-31 Cached

Truffle Security, in partnership with Julien Chaumond, conducted the largest secret scan of AI training data on HuggingFace, finding 221,303 live unique credentials across 6,003 public datasets.

0 favorites 0 likes
#training-data

How does AI know what's AI generated?

Reddit r/ArtificialInteligence ↗ · 2026-07-28

Explores the problem of AI models training on AI-generated content, leading to potential degradation and loss of truth as inaccuracies compound.

0 favorites 0 likes
#training-data

Training data needs a real go/no-go gate before training [D]

Reddit r/MachineLearning ↗ · 2026-07-27

The author proposes a formal pre-training control layer that audits training data artifacts and provides a verdict (PASS/FAIL) based on explicit criteria, as a missing gate between data preparation and training, and invites discussion on its practicality.

0 favorites 0 likes
#training-data

Screencap

Product Hunt ↗ · 2026-07-27

Screencap turns your team's real workflows into AI training data.

0 favorites 0 likes
#training-data

Are brain waves the next unlock for physical AI?

TechCrunch AI ↗ · 2026-07-27 Cached

Encord and Zander Labs are experimenting with brain wave headsets to collect richer training data for physical AI, aiming to solve the scarcity of real-world robotics data.

0 favorites 0 likes
#training-data

What data mix are the labs using to train 10T param models?

Reddit r/singularity ↗ · 2026-07-25

Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.

0 favorites 0 likes
#training-data

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

arXiv cs.LG ↗ · 2026-07-24 Cached

DataPrep-Bench is a unified benchmark evaluating LLMs' capabilities in training data construction and quality evaluation across six domains, including a skill-guided agent (Data-Construction-Skill) and a distribution-based evaluator (DAS) that achieves strong cross-model correlation.

0 favorites 0 likes
#training-data

Anthropic got sued for using copyrighted books for LLM training

Reddit r/LocalLLaMA ↗ · 2026-07-21

Anthropic is being sued for allegedly using copyrighted books without permission to train its large language models.

0 favorites 0 likes
#training-data

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Hugging Face Daily Papers ↗ · 2026-07-21 Cached

This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.

0 favorites 0 likes
#training-data

Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments

arXiv cs.CL ↗ · 2026-07-20 Cached

TerminalTraj is a scalable pipeline that generates high-quality terminal agent trajectories using Dockerized environments, achieving up to 20% improvement on terminal task benchmarks with models like TerminalTraj-32B.

0 favorites 0 likes
#training-data

@viks_rum: https://x.com/viks_rum/status/2077650169265590727

X AI KOLs Timeline ↗ · 2026-07-16 Cached

An in-depth analysis of the booming business of selling training data to frontier AI labs, detailing six distinct data products and the financial dynamics of the market.

0 favorites 0 likes
#training-data

Suno snatched millions of songs from YouTube, Genius, and Deezer

The Verge ↗ · 2026-07-15 Cached

A hack reveals that AI music generator Suno trained its models by scraping millions of songs from YouTube, Genius, and Deezer, backing up allegations of copyright infringement and exposing customer data.

0 favorites 0 likes
#training-data

Hack suggests AI music generator Suno scraped YouTube for training data

TechCrunch AI ↗ · 2026-07-15 Cached

A hack into AI music generator Suno reveals it allegedly scraped decades of audio from YouTube, Deezer, and other sources for training data, raising copyright and DMCA concerns amid ongoing lawsuits from major record labels.

0 favorites 0 likes
#training-data

The potential problem of perception hijack

Reddit r/ArtificialInteligence ↗ · 2026-07-14

The article discusses how AI models exhibit a bias toward statistically dominant narratives in training data, which could be exploited to manipulate historical and current contexts on a global scale.

0 favorites 0 likes
#training-data

this openai court story is starting to look ugly

Reddit r/artificial ↗ · 2026-07-12

The New York Times alleges OpenAI misled the court by claiming it couldn't search training data for copyrighted material, despite having previously done so and deleting billions of chat logs. The case raises concerns about transparency and accountability in AI companies.

0 favorites 0 likes
#training-data

OpenAI may have made a fatal misstep in copyright fight with news orgs

Ars Technica ↗ · 2026-07-09 Cached

OpenAI faces calls for sanctions after allegedly hiding and deleting ChatGPT logs relevant to the NYT copyright lawsuit, potentially undermining its fair-use defense.

0 favorites 0 likes
#training-data

"Grok 4.5 has an advantage on CursorBench: an earlier snapshot of the Cursor codebase was unintentionally included in training"

Reddit r/singularity ↗ · 2026-07-09 Cached

CursorBench reveals that Grok 4.5's high scores were partly due to unintentional inclusion of an earlier snapshot of the Cursor codebase in its training data. The data has been removed for future models.

0 favorites 0 likes
#training-data

This startup thinks robotics is about to have its ChatGPT moment

TechCrunch AI ↗ · 2026-07-08 Cached

General Intuition, a startup building a foundation model for robotics trained on video game data, argues robotics will follow NLP's pattern with general-purpose models, and has raised $320M at a $2.3B valuation to pursue this vision.

0 favorites 0 likes
#training-data

Why this CEO thinks video games make better training data than the internet

TechCrunch AI ↗ · 2026-07-08 Cached

General Intuition CEO Pim de Witte argues that training AI on video game data could be key to achieving AGI, as games provide spatial and temporal understanding that text-only models lack. The Bezos-backed startup recently raised $320 million.

0 favorites 0 likes
#training-data

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

Image2Sim is a neural simulation framework that creates high-fidelity interactive environments from RGB-D images, enabling scalable training for embodied navigation agents. It generates nearly 20K scenes and over 10 million training samples, showing strong benchmark improvements and effective real-world zero-shot transfer.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback