Tag
Truffle Security, in partnership with Julien Chaumond, conducted the largest secret scan of AI training data on HuggingFace, finding 221,303 live unique credentials across 6,003 public datasets.
Explores the problem of AI models training on AI-generated content, leading to potential degradation and loss of truth as inaccuracies compound.
The author proposes a formal pre-training control layer that audits training data artifacts and provides a verdict (PASS/FAIL) based on explicit criteria, as a missing gate between data preparation and training, and invites discussion on its practicality.
Screencap turns your team's real workflows into AI training data.
Encord and Zander Labs are experimenting with brain wave headsets to collect richer training data for physical AI, aiming to solve the scarcity of real-world robotics data.
Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.
DataPrep-Bench is a unified benchmark evaluating LLMs' capabilities in training data construction and quality evaluation across six domains, including a skill-guided agent (Data-Construction-Skill) and a distribution-based evaluator (DAS) that achieves strong cross-model correlation.
Anthropic is being sued for allegedly using copyrighted books without permission to train its large language models.
This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.
TerminalTraj is a scalable pipeline that generates high-quality terminal agent trajectories using Dockerized environments, achieving up to 20% improvement on terminal task benchmarks with models like TerminalTraj-32B.
An in-depth analysis of the booming business of selling training data to frontier AI labs, detailing six distinct data products and the financial dynamics of the market.
A hack reveals that AI music generator Suno trained its models by scraping millions of songs from YouTube, Genius, and Deezer, backing up allegations of copyright infringement and exposing customer data.
A hack into AI music generator Suno reveals it allegedly scraped decades of audio from YouTube, Deezer, and other sources for training data, raising copyright and DMCA concerns amid ongoing lawsuits from major record labels.
The article discusses how AI models exhibit a bias toward statistically dominant narratives in training data, which could be exploited to manipulate historical and current contexts on a global scale.
The New York Times alleges OpenAI misled the court by claiming it couldn't search training data for copyrighted material, despite having previously done so and deleting billions of chat logs. The case raises concerns about transparency and accountability in AI companies.
OpenAI faces calls for sanctions after allegedly hiding and deleting ChatGPT logs relevant to the NYT copyright lawsuit, potentially undermining its fair-use defense.
CursorBench reveals that Grok 4.5's high scores were partly due to unintentional inclusion of an earlier snapshot of the Cursor codebase in its training data. The data has been removed for future models.
General Intuition, a startup building a foundation model for robotics trained on video game data, argues robotics will follow NLP's pattern with general-purpose models, and has raised $320M at a $2.3B valuation to pursue this vision.
General Intuition CEO Pim de Witte argues that training AI on video game data could be key to achieving AGI, as games provide spatial and temporal understanding that text-only models lack. The Bezos-backed startup recently raised $320 million.
Image2Sim is a neural simulation framework that creates high-fidelity interactive environments from RGB-D images, enabling scalable training for embodied navigation agents. It generates nearly 20K scenes and over 10 million training samples, showing strong benchmark improvements and effective real-world zero-shot transfer.