What data mix are the labs using to train 10T param models?

Reddit r/singularity News

Summary

Discussion about the data sources labs may use to train 10T parameter models, including synthetic reasoning chains and human-generated traces, amid concerns about hitting the data wall.

So my assumption is: So far labs have made public max 2-3T param models based on different reports. And they are currently training or have trained 10T param models internally. Another assumption I'm making: If the models are increasing params by 3x , they would have to proportionally increase the data by 3x too or some margin.But we have also been hearing news of hitting the data wall based on internet data since the gpt 4 days. So what gives? Where are they getting so much data from? Is most of it reasoning chains generated by models during inference? Or is it reasoning traces from actual humans thanks to mercor, etc? Anyone know the exact mix? Or what's going on here? Seems like a lot of data needed all of a sudden.
Original Article

Similar Articles

Agents That Build Better Training Data (25 minute read)

TLDR AI

Autodata introduces an agentic data scientist that iteratively generates and refines synthetic training data, with meta-optimization to further improve data quality, achieving better results on computer science and legal reasoning tasks.

Data for Agents

Hugging Face Blog

NVIDIA discusses the importance of open and synthetic data for building robust AI agents, highlighting their Nemotron open datasets for training, reasoning, and tool-use.

What matters when synthetic training data is generated on demand?

Reddit r/ArtificialInteligence

Abliteration launches a made-to-order synthetic training data workflow that generates negative, rare, and adversarial examples for classifiers, with schema, real-world facts, labels, provenance, and export to platforms like Hugging Face.

@neural_avb: https://x.com/neural_avb/status/2072294078805684613

X AI KOLs Timeline

This paper introduces Autodata, a method that uses an agentic 'data scientist' AI to automate the creation of high-quality synthetic datasets through iterative generation, verification, and refinement, specifically optimized for reinforcement learning (GRPO) to improve reasoning in language models.