Tag
The article highlights how wet lab data enables a specialized model to outperform GPT-6 Astra in scientific tasks, underscoring the growing importance of domain-specific data for advancing AI at the frontiers of science.
The paper proposes BaguanHR, a framework that uses variable-wise super-resolution to synthesize high-resolution weather data from coarse-resolution sources, overcoming data limitations for ML-based forecasting and demonstrating power-law scaling effects for improved performance.
This paper from the Kathleen series shows that an attention-free, byte-level model with ~0.5M parameters can beat a parameter-matched transformer on WikiText-103 language modeling and generation, introduces a non-parametric 'Form Distance' metric for evaluating text realism, and demonstrates that retrieval-augmented decoding from the model's own training corpus improves generation quality.
Advocates using production traces as data for AI post-training, highlighting the growing scale of data spending.
The article argues that AI scaling is hitting data limits, requiring a civilizational-scale data effort similar to compute projects, and predicts over $100B/year in data spending by 2030.
A comprehensive overview of scaling laws in deep learning, tracing their theoretical roots and empirical findings, and explaining how loss decreases predictably with model size, data, and compute.
VeriEvol is a novel framework for scaling reinforcement learning in visual mathematical reasoning by ensuring reliable reward labels through a two-axis approach separating prompt difficulty from answer reliability, using evolutionary operators and hypothesis-testing verification. It achieves significant accuracy gains on a five-benchmark visual-math suite.
This paper finds that egocentric human video, when processed with a filtering and labeling pipeline, can outperform teleoperated real-robot data for pretraining embodied foundation models, achieving lower validation loss and higher success rates on real-robot tasks.
Article questions why frontier AI labs like OpenAI and Anthropic do not disclose the size of their training data, suggesting that improvements may come from data volume rather than genuine intelligence.
This paper proposes that real-data scaling laws are governed by progressive coverage of a latent predictive contribution spectrum rather than token-frequency tails alone, and provides empirical evidence using a suffix-automaton representation of text corpora.
FrontierSmith is a system that synthesizes open-ended coding problems at scale from closed-ended tasks. It generates, filters, and builds training environments; models trained on its data outperform those trained on human-curated open-ended data.