data-scaling

Tag

Cards List
#data-scaling

@_jasonwei: very cool result showing how wet lab data enables a specialized model to beat gpt-6 astra at a task at the frontier of …

X AI KOLs Timeline ↗ · 2026-09-15 Cached

The article highlights how wet lab data enables a specialized model to outperform GPT-6 Astra in scientific tasks, underscoring the growing importance of domain-specific data for advancing AI at the frontiers of science.

0 favorites 0 likes
#data-scaling

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

arXiv cs.LG ↗ · 2026-08-18 Cached

The paper proposes BaguanHR, a framework that uses variable-wise super-resolution to synthesize high-resolution weather data from coarse-resolution sources, overcoming data limitations for ML-based forecasting and demonstrating power-law scaling effects for improved performance.

0 favorites 0 likes
#data-scaling

Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper from the Kathleen series shows that an attention-free, byte-level model with ~0.5M parameters can beat a parameter-matched transformer on WikiText-103 language modeling and generation, introduces a non-parametric 'Form Distance' metric for evaluating text realism, and demonstrates that retrieval-augmented decoding from the model's own training corpus improves generation quality.

0 favorites 0 likes
#data-scaling

@samsja19: do not delete your production trace, turn them into fuel for your next post training

X AI KOLs Following ↗ · 2026-07-06 Cached

Advocates using production traces as data for AI post-training, highlighting the growing scale of data spending.

0 favorites 0 likes
#data-scaling

@willdepue: A Stargate for Data Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar…

X AI KOLs Following ↗ · 2026-07-06 Cached

The article argues that AI scaling is hitting data limits, requiring a civilizational-scale data effort similar to compute projects, and predicts over $100B/year in data spending by 2030.

0 favorites 0 likes
#data-scaling

Scaling Laws, Carefully (25 minute read)

TLDR AI ↗ · 2026-06-26 Cached

A comprehensive overview of scaling laws in deep learning, tracing their theoretical roots and empirical findings, and explaining how loss decreases predictably with model size, data, and compute.

0 favorites 0 likes
#data-scaling

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

Hugging Face Daily Papers ↗ · 2026-06-22 Cached

VeriEvol is a novel framework for scaling reinforcement learning in visual mathematical reasoning by ensuring reliable reward labels through a two-axis approach separating prompt difficulty from answer reliability, using evolutionary operators and hypothesis-testing verification. It achieves significant accuracy gains on a five-benchmark visual-math suite.

0 favorites 0 likes
#data-scaling

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Hugging Face Daily Papers ↗ · 2026-06-18 Cached

This paper finds that egocentric human video, when processed with a filtering and labeling pipeline, can outperform teleoperated real-robot data for pretraining embodied foundation models, achieving lower validation loss and higher success rates on real-robot tasks.

0 favorites 0 likes
#data-scaling

Why don't frontier labs say how much data they are training on?

Reddit r/ArtificialInteligence ↗ · 2026-06-17

Article questions why frontier AI labs like OpenAI and Anthropic do not disclose the size of their training data, suggesting that improvements may come from data volume rather than genuine intelligence.

0 favorites 0 likes
#data-scaling

Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum

arXiv cs.CL ↗ · 2026-05-21 Cached

This paper proposes that real-data scaling laws are governed by progressive coverage of a latent predictive contribution spectrum rather than token-frequency tails alone, and provides empirical evidence using a suffix-automaton representation of text corpora.

0 favorites 0 likes
#data-scaling

@MangQiuyang: Open-ended coding training data may no longer be the bottleneck: AI can scale open-ended tasks—and even outperform huma…

X AI KOLs Timeline ↗ · 2026-05-15 Cached

FrontierSmith is a system that synthesizes open-ended coding problems at scale from closed-ended tasks. It generates, filters, and builds training environments; models trained on its data outperform those trained on human-curated open-ended data.

0 favorites 0 likes
← Back to home

Submit Feedback