@geoffreyhinton: I just watched an amazingly good lecture by Adam Brown about the future impact of AI on physics
Summary
Geoffrey Hinton highlights a lecture by Adam Brown on how LLMs have advanced from preschool to PhD-level competence in physics, with scaling laws and benchmarks showing rapid progress.
View Cached Full Text
Cached at: 06/28/26, 10:08 PM
I just watched an amazingly good lecture by Adam Brown about the future impact of AI on physics https://t.co/EjiKpoiKZr
TL;DR
A theoretical physicist turned AI researcher describes how large language models have progressed from preschool-level competence to PhD-level research in just a few years, and what that means for the future of physics.
The Landscape: From Sand to Silicon to Thought
“We live at an extraordinary moment in our civilization’s history,” the speaker opens. “We have collectively figured out how to refine sand into silicon, turn that silicon into chips, assemble those chips into neural networks, and now train those networks to think.”
A theoretical physicist who has written about 40 papers, he stopped because “it felt like too much of a guilty pleasure to handwrite theoretical physics papers one by one when what I should be doing is contributing to the production of a machine that is going to spew out knowledge on an industrial scale.”
Unlike earlier computer aids (pocket calculators, abacuses), a large language model (LLM) is not a special-purpose tool. “A large language model has the capability to do every single part of my job as a theoretical physicist. It is a general intelligence.”
How Large Language Models Are Made: Grown, Not Programmed
LLMs are neural networks inspired by the arrangement of neurons in the human brain. At the start of the decade the largest had about a billion parameters; now they have a few trillion (still short of the human brain’s ~100 trillion synapses).
“They are grown, not programmed.” You start with random weights, train the network to predict the next word given a block of text. Each correct prediction strengthens a synaptic pathway; each wrong one punishes it. After seeing a million words the output is still gibberish; after tens of billions it can converse intelligently on almost any topic.
This is called pre‑training. Then comes post‑training (“finishing school”) where the model is taught to be polite and helpful rather than merely predicting the next word.
Scaling Laws: The Physicist’s Contribution
“Physicists love scaling laws – that’s our bread and butter.” Empirical scaling laws (e.g., animal metabolic rate vs. mass) are often straight lines on a log‑log plot. The same turned out to be true for LLMs: if you spend more compute and scale appropriately, performance improves linearly on a log‑log plot.
“This plot is so simple that even a venture capitalist can understand it.” It told investors that pouring in compute would yield better performance. The original scaling law (discovered by physicists in 2020) spanned eight orders of magnitude; it has now been extended eight more and still holds.
The resources devoted to training have grown exponentially: FLOPS up 4× per year since 2010, cost up 2.7× per year. Yet the biggest driver of progress has been algorithmic improvement – human ingenuity in shearing away inefficiencies.
The Benchmark Rush: From Preschool to PhD
Progress is measured by benchmarks that are rapidly killed: they go from “too hard to be useful” to “too easy to be useful” in about 18 months.
MATH: High School Mathematics
In 2020, the MATH benchmark (high‑school math problems) saw LLMs score 6% – barely better than random guessing. A prediction market said models would reach 50% by 2025. The creators were incredulous: “If I imagine a system getting more than half of these right, I’ll be pretty impressed.”
The next system (Minerva) hit 50% almost immediately. By mid‑2024, a system called Max Math achieved 90%, beating the human expert level. Then, six months later, off‑the‑shelf models got nearly perfect scores.
Over the hardest 20% of problems (level five), the same pattern repeated: from near‑random to saturation in 2.5 years.
GPQA: Graduate Science
GPQA simulates first‑year PhD exams. Example: a problem about the cosmic microwave background. “If you’re in an adjacent field, you don’t know how to answer that.”
From 2024 to 2025, models went from random guessing past expert human level (≈70%) to near perfect. “GPQA is dead. It has suffered the fate of all benchmarks.”
Skeptics might argue the models simply memorized answers. But held‑out tests – including the speaker’s own private graduate exams in general relativity and quantum mechanics – show the same performance. “From late 2023 over the following 18 months, models got 100% accuracy. My benchmark sadly dead.”
The International Math Olympiad
Just over a year ago, a Turing Award winner told the speaker that LLMs would never do something creative like solve an IMO problem. Last summer, a system earned a gold medal (5 out of 6 problems correct). The president of the IMO said: “Their solutions were astonishing … clear, precise, and most of them easy to follow.”
“There are only a very small number of humans in the world now who are better than the AIs at doing the IMO.”
Novel Research: The Centaur Approach
All benchmarks up to this point test known problems. The next step is generating new knowledge. The speaker’s group produced a paper using a “centaur” style: half human, half LLM.
“The output was … the most impressive thing that had yet been done with LLMs in maths.” One co‑author, a Stanford professor and president of the American Mathematical Society, said: “We found that Gemini’s argument was no mere repackaging of existing proofs. It was the kind of insight I would have been proud to have produced myself.”
The proof was assembled under human guidance, but the LLM produced the key novel arguments.
Remaining Challenges and the Road Ahead
Despite this progress, current LLMs have four major weaknesses:
- Low agency
- Slow learning
- Poor planning
- Poor error correction
Every one of these problems is actively being worked on and has improved over the past year, but none is fully solved.
“What doesn’t work is just taking your favorite LLM and saying, ‘Please invent a novel theory of quantum gravity for me.’ It will output AI slop – not worth your time.”
The future is deeply uncertain. The speaker notes a Financial Times plot showing extremely high variance in predicted AI‑driven GDP growth over the next decade. One possibility is that we plateau; another is that we continue the rapid progression. Given the pattern of benchmarks being destroyed in months, the latter seems plausible.
“A good rule of thumb: we’re moving about four times as fast as a human student. For every year that passes, we advance four years into the future.”
Similar Articles
@RohOnChain: just talked to an MIT CS grad who is building the next frontier LLM. he told me this lecture by an OpenAI researcher on…
A tweet recommends a lecture by an OpenAI researcher on how LLMs are built, claiming it taught an MIT CS grad more than his entire degree.
@rohanpaul_ai: Theoretical Physics just crossed the same threshold mathematics did. AI will make a lot of junior-level theoretical phy…
A tweet discusses how AI, specifically Claude, is crossing a threshold in theoretical physics similar to mathematics, potentially making junior-level research much cheaper.
@ProfBuehlerMIT: For science, AI sovereignty and physics-grounded reasoning are non-negotiable. But how can we teach a small LLM like Ge…
mistral.rs now natively supports Agent Skills, enabling locally-run small LLMs to perform complex agentic workflows for scientific tasks, with full control over models, data, and execution.
Interesting post from Adam Majmudar (research at OpenAI) — full text in body
An analysis by Adam Majmudar on the rapid progression of AI capabilities, highlighting how new scaling laws are driving leaps beyond external expectations.
@waynoir: Ex-Google's Jeff Dean just dropped a free 1 hour lecture covering the full path of AI engineering, from LLMs built from…
Jeff Dean, a former Google engineer, has released a free one-hour lecture covering the full path of AI engineering, from building LLMs to prompt engineering and agent coordination.