Tag
The paper investigates the conditional diversity gap in large language models by comparing the entropy of generated outputs with training data and proposes an information-theoretic framework to measure and mitigate this gap.
An individual built a Minecraft clone using the Qwen3.8-27B Q4 model and added custom features to address criticism that Minecraft data influenced the training.
AI crawlers are causing significant CPU load on kernel.org by inefficiently scraping git commits for training data, consuming more resources than all legitimate access combined.
The article audits major AI providers like OpenAI and Anthropic, revealing that they train on user chats by default on consumer plans, with opt-out options often hidden in settings.
Google is reportedly in advanced talks for a $1.5 billion deal with AI coding startup Mechanize, which develops virtual environments, benchmarks, and training data to help AI agents handle complex tasks.
The paper introduces Game2World Engine, a framework for removing UI from gameplay videos to create cleaner training data for video world models, with GameCleaner model achieving state-of-the-art results in UI removal.
Napster has pivoted to an AI agent platform, illustrating how AI training data quickly becomes outdated, leading models to provide confident but incorrect information about company identities.
The paper presents Task Model Induction, a method that decomposes screen recordings of real computer work into separate tasks and goal trees to better train AI agents, improving performance by 30% on unseen tasks compared to current summarizers.
The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.
Micro1, an AI data startup, has reached a $500M gross annual run rate in eight months due to high demand for AI training data, though it remains smaller than competitors like Mercor and Handshake.
Y Combinator's Paper Club focuses on data challenges in AI research, with domain experts discussing training data and benchmarks.
Humyn Labs is tackling the robotics data problem by converting human experience into synchronized training data with multiple sensors, such as IMU, depth cameras, and hand tracking, to make it useful for robot training.
Zhipu Founder Tang Jie discusses how AI scaling is evolving beyond parameter count to include factors like training data, compute per forward pass, and post-training, with GLM-5.3 as an example.
A Meta/Oxford study finds that multimodal models need surprisingly little image-generation data if language and visual understanding are trained together, suggesting an optimal 70/25/5 split of language, image understanding, and image generation tokens. It also warns that delaying vision training causes 'vision laziness' where models ignore images.
A blog post and accompanying tweet explore whether pre-generative-AI data becomes more valuable as the internet fills with synthetic content, discussing provenance, model collapse, and Anthropic's book scanning.
Discussion of whether AI companies should face copyright lawsuits to force ethical training data sourcing, citing Suno's loss in Germany and arguing laws should require consent for training on people's data.
Macrodata Labs releases a research blog on scaling robotics with egocentric video data by recovering 3D hand motion signals using only open-source models.
The author observes that ChatGPT can now draw a watch with a specified time, noting this was a classic AI failure a year ago due to training data over-representing 10:10, and asks what other '10:10 problems' remain.
The tweet highlights the potential of agent harnesses to capture and distill experts' tacit knowledge as new training data, referencing a Berkeley AI Summit talk by Jianfeng Gao on agentic modeling as an emerging AI paradigm.
Truffle Security, in partnership with Julien Chaumond, conducted the largest secret scan of AI training data on HuggingFace, finding 221,303 live unique credentials across 6,003 public datasets.