Tag
Photoroom details their data strategy for training PRX, including assembling diverse datasets, re-captioning with a VLM, and using Mosaic Data Shards for efficient training.
A tweet warns that different AI model families have unique conversation formatting rules that silently corrupt training data, requiring developers to learn each family's quirks individually.
A tweet highlights Robin Hanson's observation that current LLMs are being influenced by low-status clear thinkers, but may eventually learn to ignore them like high-status humans do.
DolphinMath is a tool that generates math problems with step-by-step solutions, suitable for various training stages from pretraining to RL, covering elementary to postgraduate levels.
This article argues that Reddit's messy, authentic human conversations are becoming increasingly valuable for training AI as the web fills with synthetic content, highlighting the economic shift toward scarce human behavioral data.
35 newspaper publishers across the US have filed a lawsuit against OpenAI and Microsoft, alleging that the companies scraped their copyrighted and paywalled content without permission to train ChatGPT, harming local journalism.
A researcher describes a method called 'projection' to fine-tune AI agents by projecting a verifiably correct solution onto a map the agent can traverse, achieving improved performance on cybersecurity tasks with limited training data.
A new open-source tool called claude_converter converts Claude Code session logs into fine-tuning datasets compatible with TRL/SFTTrainer, Axolotl, and LLaMA-Factory, enabling developers to repurpose real coding conversations for training local models.
A commentary on the shifting attitudes towards web scraping for AI training, questioning the sudden condemnation of data collection without permission.
The New York Times updated its copyright lawsuit against Microsoft and OpenAI, alleging that Microsoft built a supercomputer specifically designed to train AI on copyrighted works, including NYT articles, without permission.
This paper demonstrates that a small coordinated Wikipedia editing campaign can measurably shape how language models handle topics, using animal welfare as a case study.
A tweet highlights that Anthropic conducted large-scale reinforcement learning using Slack conversations, with Andrej Karpathy emphasizing that it is not a trivial Slack bot feature as commonly misinterpreted.
Ai2 and the University of Washington released a paper titled Tmax, proposing the strongest open-source terminal agent RL training recipe to date. A 9B parameter model outperforms larger models on Terminal-Bench 2.0, with the key being low-cost generation of vast amounts of verifiable training data, not model size or algorithm.
Autodata is a method that enables AI agents to act as data scientists to create high-quality synthetic training data through meta-optimization, achieving improved performance across computer science, legal reasoning, and mathematical tasks.
Leaked documents reveal Russia's Project 2026, run by the Social Design Agency, to create fake reference platforms like a German Wikipedia clone to contaminate AI training data and search indices, aiming to embed Russian narratives into AI responses.
A tweet observes that current social media influencers serve as training data for the next generation of AI-generated influencers.
This paper introduces OpenThoughts-Agent, an open-source data curation pipeline for training agentic language models, achieving a 44.8% average accuracy across seven benchmarks and outperforming prior open datasets through systematic experiments.
This article deeply analyzes the problem that AI's sample efficiency is far lower than that of humans, pointing out that frontier models require massive amounts of domain-specific data, while humans can learn from just a few examples. This data black hole is a core bottleneck in current AI development. Through multiple comparisons (annotation volume, robot manipulation, driving) and refuting common objections, the article demonstrates the severity of this gap and explores its impact on the goals of AI automation.
An opinion piece argues that AI models acquire dangerous knowledge from training data, and that companies like Anthropic and OpenAI rely on easily breakable refusal filters instead of truly removing harmful capabilities, prioritizing speed over safety.
The Atlantic has created a searchable database of millions of music tracks used to train AI models, allowing the public to search through four datasets including those from Google and Stability AI.