Tag
A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.
Discussion of whether AI companies should face copyright lawsuits to force ethical training data sourcing, citing Suno's loss in Germany and arguing laws should require consent for training on people's data.
Macrodata Labs releases a research blog on scaling robotics with egocentric video data by recovering 3D hand motion signals using only open-source models.
The author observes that ChatGPT can now draw a watch with a specified time, noting this was a classic AI failure a year ago due to training data over-representing 10:10, and asks what other '10:10 problems' remain.
The tweet highlights the potential of agent harnesses to capture and distill experts' tacit knowledge as new training data, referencing a Berkeley AI Summit talk by Jianfeng Gao on agentic modeling as an emerging AI paradigm.
Truffle Security, in partnership with Julien Chaumond, conducted the largest secret scan of AI training data on HuggingFace, finding 221,303 live unique credentials across 6,003 public datasets.
Explores the problem of AI models training on AI-generated content, leading to potential degradation and loss of truth as inaccuracies compound.
The author proposes a formal pre-training control layer that audits training data artifacts and provides a verdict (PASS/FAIL) based on explicit criteria, as a missing gate between data preparation and training, and invites discussion on its practicality.
Screencap turns your team's real workflows into AI training data.
Encord and Zander Labs are experimenting with brain wave headsets to collect richer training data for physical AI, aiming to solve the scarcity of real-world robotics data.
Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.
DataPrep-Bench is a unified benchmark evaluating LLMs' capabilities in training data construction and quality evaluation across six domains, including a skill-guided agent (Data-Construction-Skill) and a distribution-based evaluator (DAS) that achieves strong cross-model correlation.
Anthropic is being sued for allegedly using copyrighted books without permission to train its large language models.
This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.
TerminalTraj is a scalable pipeline that generates high-quality terminal agent trajectories using Dockerized environments, achieving up to 20% improvement on terminal task benchmarks with models like TerminalTraj-32B.
An in-depth analysis of the booming business of selling training data to frontier AI labs, detailing six distinct data products and the financial dynamics of the market.
A hack reveals that AI music generator Suno trained its models by scraping millions of songs from YouTube, Genius, and Deezer, backing up allegations of copyright infringement and exposing customer data.
A hack into AI music generator Suno reveals it allegedly scraped decades of audio from YouTube, Deezer, and other sources for training data, raising copyright and DMCA concerns amid ongoing lawsuits from major record labels.
The article discusses how AI models exhibit a bias toward statistically dominant narratives in training data, which could be exploited to manipulate historical and current contexts on a global scale.
The New York Times alleges OpenAI misled the court by claiming it couldn't search training data for copyrighted material, despite having previously done so and deleting billions of chat logs. The case raises concerns about transparency and accountability in AI companies.