100 Trillion+ Pretraining data??? This is the largest data I've see a model being trained on.
Summary
A new AI model is being trained on over 100 trillion tokens, doubling the typical pretraining data size of 27-50 trillion tokens used by other models like Kimi, Mimo, and DeepSeek.
Similar Articles
Retell vs Vapi vs Plura ai for a production voice agent, which one held up?
A comparison of three voice AI agents — Retell, Vapi, and Plura AI — evaluating their performance for production use cases.
Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification
This study empirically compares Kolmogorov-Arnold Networks (KANs) and Multi-Layer Perceptrons (MLPs) on structured tabular classification tasks, finding that KANs statistically outperform MLPs but with higher computational cost.
Kimi K3 leaks: on par with Fable
Leaked details about Kimi K3 suggest it performs on par with Fable, marking a competitive development in the AI model space.
@jxmnop: ok sorry everyone apparently they did distill lol. but only a tiny bit
Jack Morris corrects his earlier claim about an open-weight model being trained without distillation from OpenAI or Anthropic, acknowledging that it actually did use a small amount of distillation.
Comparing Obelisk with Temporal and Restate
A technical comparison of three workflow systems (Obelisk, Temporal, Restate) implementing a weather forecast workflow, highlighting differences in architecture, Rust APIs, determinism boundaries, and deployment models.