@maxrumpf: Humans Are a Low Ceiling Can Magnus Carlson give feedback to AlphaZero? Obviously not. Even the very best chess players…

X AI KOLs Timeline News

Summary

Max Rumpf argues that human feedback is becoming obsolete for training advanced AI models, citing examples like chess, math, and search. He advocates for human-free methods like self-play and synthetic data, while a quoted tweet from Will Depue calls for a large-scale data infrastructure parallel to compute scaling.

Humans Are a Low Ceiling Can Magnus Carlson give feedback to AlphaZero? Obviously not. Even the very best chess players are so far outclassed that their input is of no value in training future AIs. While not universal, we're already there in other domains too. It's why the modal labeler has become educated and experienced. How many humans can still give good feedback on time-bounded math problems like IMO? 1,000 globally? In 2022, >1B people could have helped the models get better at math. It is most likely easier to make self-play/synthetic data/online learning work than it is to turn the planet into a click-farm. We have seen it in games (chess), math (lean), and search (our work). Human-free methods are most scalable. Most bitter lesson.
Original Article
View Cached Full Text

Cached at: 07/07/26, 09:27 AM

Humans Are a Low Ceiling

Can Magnus Carlson give feedback to AlphaZero? Obviously not. Even the very best chess players are so far outclassed that their input is of no value in training future AIs.

While not universal, we’re already there in other domains too. It’s why the modal labeler has become educated and experienced. How many humans can still give good feedback on time-bounded math problems like IMO? 1,000 globally? In 2022, >1B people could have helped the models get better at math.

It is most likely easier to make self-play/synthetic data/online learning work than it is to turn the planet into a click-farm. We have seen it in games (chess), math (lean), and search (our work).

Human-free methods are most scalable. Most bitter lesson.

will depue (@willdepue): A Stargate for Data

Labs are on a trajectory towards >$100B/year of data spend by 2030. As we begin the trillion-dollar compute project, we need to think about the equivalent civilizational-scale effort for the other core ingredient: data.

At the foundation of the scaling

Similar Articles

What to expect from AlphaZero's value predictions [D]

Reddit r/MachineLearning

The article analyzes how AlphaZero's value predictions are shaped by self-play training data and noise, questioning whether they reliably estimate win chances against opponents with different play styles despite AlphaZero's strong empirical performance.