@vista8: https://x.com/vista8/status/2072191315916538039
Summary
Starting with the story of Galois group theory, the article delves into the boundaries of AI's capabilities in mathematics, distinguishing between two types of progress: "connecting lightning" (cross-domain connections) and "building mountains" (creating new frameworks). It analyzes the limitations of the RLVR training method and introduces the concept of "grindability" to explain AI's rapid advancements in mathematics and coding.
View Cached Full Text
Cached at: 07/01/26, 08:13 PM
From Galois to Quarks: An Idea That Needed 200 Years to Be Verified — Can AI Produce One?
https://www.youtube.com/watch?v=TfyPshgMbug
A 19-year-old boy, locked in prison, wrote down a set of mathematical notes that no one could understand.
He entrusted the notes to a friend, asking him to pass them to the greatest mathematician of the time, Gauss. The friend tried his best, but failed.
The boy died in a duel at age 20.
Twenty years later, a mathematician named Liouville dug out those notes and thought there might be something in them.
Another twenty years passed before someone formalized those ideas into a form readable by modern mathematics.
A hundred years after that, physicist Murray Gell-Mann used this theory to predict the existence of quarks.
That boy was Évariste Galois. What he left behind was group theory.
From a vague intuition to a revolution in physics, nearly two hundred years elapsed.
Over those two centuries, the idea was rejected, forgotten, misunderstood, and passed through many minds before it slowly crystallized into a mountain of mathematics.
Now, some people want to do the same thing with AI.
The question is: How do you train a system to produce an idea that takes two hundred years to be verified?
This is the dilemma that Grant Sanderson and Dwarkesh Patel repeatedly touched upon in a conversation in early 2026.
Grant is the founder of 3Blue1Brown, the most popular math channel on YouTube with millions of subscribers.
But his role is peculiar: he doesn’t do research mathematics; he explains it.
His entire career is spent answering the question “What’s the difference between understanding and proving?” — which puts him in a unique position in the discussion of AI and mathematics.
Dwarkesh hosts a podcast where he interviews top AI researchers and founders. His advantage is an outsider’s perspective; his questions are often more interesting than the answers.
They talked for over two hours. Here is the essence of that conversation.
IMO Gold Medal: A Milestone That Changed Nothing
Three years ago, Dwarkesh asked Grant a question: When AI can win a gold medal at the International Mathematical Olympiad (IMO), would that mean we’ve reached AGI?
IMO problems require genuine creativity — even top students who’ve trained specifically don’t always solve them all.
If AI could do that, wouldn’t it be able to do everything?
Grant’s answer at the time was: No. It would just be another benchmark to surpass. There would be no moment of epiphany.
He was right.
In 2024, AI achieved gold-medal-level performance at the IMO. The world did not change. No one suddenly felt AGI had arrived. No economic structures shifted. Mathematicians continued their research.
IMO problems fall into four categories: geometry, number theory, algebra, and combinatorics.
AI solved geometry problems in 19 seconds because there’s a brute-force solver that can be applied directly, and geometry problems have relatively fixed training paths that cover most problem types.
But combinatorics is different. Those problems are more like puzzles, requiring a sense of “play” and unexpected angles.
The 2024 IMO happened to have two combinatorics problems, and AI got stuck there.
If that year’s problems had one more geometry and one fewer combinatorics, AI would have won gold.
The boundary of AI’s capability is not a smooth curve — it’s jagged.
Even within mathematics, progress varies enormously across subfields. Imagining AI’s ability as a monolithic whole is a systematic misjudgment.
Moreover, the “dirty secret” of IMO is that many of its problems are actually trainable.
Problem designers try to create problems that are hard to cover with rote training, but ultimately they are limited.
The reason combinatorics is the last fortress is not because it’s the hardest, but because it’s the hardest to systematically train for.
This logic will reappear repeatedly throughout the rest of the discussion.
A Lightning Strike, and a Mountain
Grant introduced a framework that is the most valuable part of the entire conversation.
He divides the progress AI might make in mathematics into two fundamentally different types.
The first type: Lightning Strikes.
Between 2025 and 2026, AI solved several attention-grabbing math problems.
One was Erdős problem No. 1196, about the “primitive sets” conjecture.
AI’s solution involved introducing a tool from another field, striking a lightning bolt (knowledge connection) between two seemingly unrelated areas of math.
This type of progress has a characteristic: it is understandable to humans.
You just need to see the start and end points of the lightning. The rest of the derivation is natural for people in the field. If you explain the idea to a knowledgeable mathematician, they know immediately how to fill in the details.
Another example is a counterexample to the Unit Distance Conjecture.
AI released its reasoning chain. Mathematicians read it and found it comprehensible. Moreover, the counterexample actually accelerated human understanding of the problem.
Why is AI good at this kind of connection?
Because it is simultaneously fluent in quantum physics, analytic number theory, random matrix theory… It can see cross-domain similarities without needing two people to happen to chat over lunch.
There is a concrete story here.
Mathematician Hugh Montgomery was studying the distribution of zeros of the Riemann zeta function and wrote down a formula.
Physicist Freeman Dyson saw the formula and said: I recognize this expression. It appears when studying eigenvalue distributions of random Hermitian matrices — a quantum mechanics problem of nuclear energy levels.
Two seemingly unrelated fields — zero statistics and random matrix theory — shared the same mathematical structure.
This discovery opened an entire research direction.
And it happened because two people happened to chat over lunch at the Institute for Advanced Study.
The second type: Building a Mountain.
The proof of Fermat’s Last Theorem is this kind.
You first need to build the mountain of elliptic curves, then the mountain of modular forms. Only then can you build a bridge between peaks.
Those two mountains themselves are entirely new mathematical systems that required generations to build.
Group theory is also this kind.
Galois didn’t solve a known problem; he created a new framework for thinking.
AI is currently good at lightning strikes.
Building a mountain is another matter. It requires not connecting existing knowledge, but creating an entirely new framework for thought.
And the value of that framework might take a century to be verified.
Which brings us back to Galois.
The Hundred-Year Verification Loop
Dwarkesh asked a sharp question: If Galois’s idea took a hundred years to be verified, how could you possibly train an AI to produce such an idea?
The core training method behind current AI breakthroughs in math is called RLVR — Reinforcement Learning with Verifiable Rewards.
The logic is simple: give the AI a problem, it produces an answer. If the answer is correct, reward it. If wrong, penalize it. Iterate repeatedly, and the AI learns to solve problems.
This method works well in scenarios like math contest problems or code execution results, where the answer is definite and correctness is immediately knowable.
But a Galois-level insight has no such feedback.
Worse, Grant points out: When Galois was alive, the “verifier” — the academic establishment — gave the feedback: No.
His paper was rejected. His ideas were deemed unclear, incomplete.
From an RLVR perspective, this idea should have been punished and forgotten.
But it was correct.
This is not an isolated case. Lagrange, fifty years before Galois, had the intuition to study polynomials through symmetry, but he didn’t solve any problem; he only asked a new question.
At the time, no verification signal told him he was on the right track.
The deeper dilemma: Not only can AI’s training environment not capture this kind of value — even the human verifier of the time couldn’t capture it.
Grant mentioned a math paper opening he loves, from mathematician Timothy Chow, who was studying the concept of “forcing.” Chow wrote: Everyone knows what an unsolved research problem is. I want to propose a new concept: an unsolved expository problem. We have already proven something, but we don’t yet understand why it is true.
Proof and understanding are two different things.
This distinction becomes critically important in the age of AI.
Verifiable Isn’t Enough — It Also Needs to Be “Grindable”
Many people attribute AI’s rapid progress in math to the verifiability of mathematics.
An answer is either right or wrong, providing a clear training signal.
Grant and Dwarkesh both think this is only half the story. The other half is a rarely mentioned concept: grindability.
You can package the state of a problem, run a thousand parallel instances, each trying different paths. Correct paths stay, wrong ones are discarded. Credit assignment is clean.
Same with code: package a codebase state into a container, send hundreds of agents to each try implementing a feature. The result is entirely deterministic; the difference between success and failure is a valid signal.
Then they gave a counterexample: computer use.
It’s also verifiable — “Did my package arrive?” has a clear answer, “Was my meeting booked successfully?” has a clear answer.
But you can’t simultaneously run a thousand Amazon checkout flows because the website has anti-scraping mechanisms.
You could try to clone every website, but that’s extremely labor-intensive and can’t keep up with updates.
This is why progress in AI for computer use is far slower than for math and code, even though both are verifiable.
Verifiability is necessary. Grindability is sufficient.
Most tasks in the real world cannot be containerized and repeatedly ground.
You can’t containerize “go to the market and trade for profit today” because the market changes every day; you cannot replay it.
Math and code are exceptions. That’s the real reason AI has advanced so rapidly in these two fields.
Autoregression Is a Strange Way of Thinking
Once you understand grindability, you can understand another question: why AI excels at lightning strikes but struggles to build mountains.
It starts with how AI works.
Grant used a vivid metaphor.
Imagine you are locked in a box. The only way the outside world communicates with you is: hand you a slip of paper, ask “what is the next word?”, you predict, then your memory is wiped, and they hand you another slip. This process repeats countless times. Then the outside world takes all your predicted words, puts them together, and shows you: “Look, this is the article you wrote.”
You might say: That’s terrible. This is nothing like what I would have written.
That is how autoregressive language models work.
At each step, they predict the most likely next word, not like a writer who first has an overall structure in mind and then fills in details.
What does this mean for mathematics?
The most valuable progress in math is often the “least likely next word” — the lightning bolt that jumps from one domain to another.
But in an autoregressive framework, when you’re in the context of a certain math domain, the most likely next word is a word from that same domain, not from another domain.
Cross-domain connections, within the logic of autoregression, are low-probability events.
So how did AI start to do it?
Dwarkesh’s guess: training environment. If you design a batch of problems that specifically require cross-domain connections, and let the AI grind on those problems repeatedly, it will be forced to learn, within the autoregressive framework, to predict the action “let me look at another domain for a similar structure.”
This is the same logic as AI learning to become a better programming agent.
It learns to predict the action “let me step back and re-examine the entire codebase” because that action is repeatedly validated as effective in the training data.
But building a mountain does not require this.
Building a mountain requires: persisting with a vague intuition in the absence of any verification signal, and then constructing an entirely new language around that intuition.
This is not a low-probability next word. This is a completely different mode of thinking.
AI’s Most Underrated Advantage: Not How Smart It Is
There’s a insight in the conversation that both Grant and Dwarkesh touched on but didn’t fully develop, and I think it’s worth highlighting separately.
We usually talk about how smart AI is. But we rarely talk about another advantage of AI: it can be infinitely parallelized.
Go back to the story of Montgomery and Dyson having lunch at Princeton.
That encounter was a chance event. Two experts from different fields happened to be in the same place, happened to talk about their work, happened to find a connection.
The Institute for Advanced Study puts a bunch of top scholars in the same place precisely to manufacture this kind of serendipity.
AI doesn’t need that luck.
You can have an agent fluent in random matrix theory and an agent fluent in analytic number theory systematically converse, searching for all possible connections.
Going further, you can run a thousand such conversations simultaneously, covering all possible combinations of fields.
This is not just a speedup; it’s a structural advantage.
The chance encounters that changed the direction of human science can be systematically engineered within AI’s framework.
There’s another dimension.
The Unit Distance Conjecture remained unsolved for a long time, partly because most mathematicians believed the conjecture was true, so they tried to prove it, not to find a counterexample.
This is a collective cognitive bias.
AI can simultaneously run two groups of agents: one trying to prove, one trying to disprove. This isn’t deep technology, but it systematically eliminates the preconceived biases that haunt human research.
Grant also mentioned a more interesting possibility: implant different heuristics into different agents.
Einstein had a strong bias: physical laws should look the same in different reference frames. That bias was the core driver of relativity. But he also had another bias: God does not play dice. That bias led him astray on quantum mechanics.
You can’t make all AIs into Einsteins.
You need diversity. You can systematically implant different heuristics into different agents, then see which heuristics are effective for which kinds of problems.
This is an old-fashioned software mindset: enumerate all possible strategies, then explore in parallel.
But applied to scientific research, its potential is enormous.
Lean: Overrated as a Training Tool, Underrated as an Exploration Engine
The formal proof language Lean is frequently mentioned in the AI math community. Many view it as the key to AI breakthroughs in math.
Grant’s view: for current progress, Lean’s importance is overrated.
DeepMind initially used Lean for IMO, but switched to natural language in the second year and got better results.
When AI solved the Unit Distance Conjecture counterexample, the public reasoning chain contained zero Lean.
Process supervision seems far less valuable than a grindable outcome validation.
But Lean has another unique value, and it hasn’t been fully exploited yet.
Lean allows AI to run completely autonomously, without human intervention.
Mathlib is a mathematical library written in code, aiming to formalize all of mathematics.
You can imagine an AI told “go extend Mathlib,” then just let it run. No one needs to review each step because each step’s correctness can be automatically verified.
It can propose its own conjectures, build its own definitions, grow its own logical tree.
Grant said: You can press start, pour ten years of compute resources into it, then come back to see what it discovered.
This reminds one of AlphaGo.
AlphaGo could play infinite games in its own universe without human intervention because the rules of Go are completely deterministic, and winning/losing is automatically verifiable.
It explored that closed universe and discovered moves humans had never thought of — move 37 being the most famous.
Lean offers a similar possibility for mathematics.
An AI autonomously exploring the Lean world might discover mathematical structures humans have never imagined.
But there’s a problem: how much of what it discovers will be useful?
Grant mentioned that Terry Tao once talked about a research project to exhaustively search all possible algebraic axiom systems.
Group theory has a set of axioms. But if you systematically try all possible combinations of axioms, might you discover some entirely new and interesting algebraic structures?
Most results will be garbage. But occasionally there will be an island — a set of axioms that yields a rich set of theorems, worthy of deep study.
This is what makes Lean truly interesting: not as a training tool, but as an exploration engine.
After the Riemann Hypothesis Is Proved, Will We Understand It?
The conversation included a striking concern: AI might prove the Riemann Hypothesis, but our understanding of mathematics would not increase at all.
Grant classified possible solutions into three types.
The first is a lightning strike: discovering a connection between two domains, like the relationship between zeros of the Riemann zeta function and random matrix theory. This kind of solution is understandable to humans, and might even advance human understanding.
The second is building a mountain: constructing an entirely new mathematical framework, like Wiles proving Fermat’s Last Theorem by first building the mountains of elliptic curves and modular forms. This kind of solution requires humans to invest a lot of time to understand the new mountain, but ultimately it is comprehensible.
The third is brute force: a proof thousands of pages long, containing no new concepts, just exhaustive case analysis. Such a proof is technically correct but contributes nothing to human understanding.
Grant mentioned a real-world analogy: the “proof” of the abc conjecture.
Japanese mathematician Shinichi Mochizuki proposed a completely new framework called “Inter-universal Teichmüller theory,” claiming it could prove the abc conjecture.
This theory was so alien that the mathematical community spent years unable to judge its correctness.
The eventual mainstream judgment was that it likely contained errors, but the controversy remains unsettled.
This is what “alien mathematics” looks like: a new mountain, but no one can climb it, and we’re not even sure the mountain exists.
If AI produces something like this, and it’s wrong, that’s a catastrophic waste.
If it’s right, it requires enormous human effort to digest.
David Bessis, in a blog post titled “The Collapse of the Theorem Economy,” proposed: Historically, theorem-proving and concept-creation were bound together because the person who proposed the definition was often also the one who proved the theorem. But if AI automates theorem-proving while humans still define concepts, that bond is broken.
There’s a saying in math circles: Good mathematicians prove theorems; great mathematicians propose conjectures; the greatest mathematicians propose definitions.
AI is climbing from the bottom up.
It can already prove theorems, and is beginning to propose conjectures. But proposing definitions — creating a new language of thought — that’s what Galois did.
Why AI’s Writing Is Getting Worse, but Its Math Is Getting Better
Writing is getting worse for two reasons.
The first is reward hacking. AI’s writing training essentially optimizes for “looks like good writing” rather than “is good writing.” It learns all the surface features of good writing and stacks them together. The result is a piece that hits all the scoring criteria but contains no real insight.
The second is deeper: Writing itself is the product, not the production process.
Code can be ugly as long as it runs correctly. A function can be poorly written, but if it produces the right output, it’s acceptable.
Mathematical proofs are similar: a lemma can be proved in many ways; as long as the conclusion is correct, it’s fine.
But writing is different.
Every word, every sentence is the final deliverable. There is no waste.
And good writing requires modeling the reader’s mental state at every sentence, predicting what the reader is thinking at that moment, then deciding what to say next.
Grant mentioned an interesting experiment: People who have received Botox injections, with their facial muscles frozen and unable to mimic others’ expressions, show a significant decline in the ability to recognize others’ emotions. Part of the mechanism for understanding others’ emotions is using one’s own face to “replicate” the other person’s expression.
AI has no face.
Its ability to understand the reader’s psychology is an emergent ability from vast text, not built-in hardware.
This might be a fundamental limitation in its writing.
But here’s an interesting counterpoint.
Dwarkesh said: AI has gotten increasingly good at writing code that is not just runnable, but clean, tidy, and ready to merge. Why hasn’t the same progress happened with writing?
Grant’s answer: Maybe it has, and we just haven’t noticed. He said that now, when he encounters a difficult article, his first instinct is to paste it into an LLM and ask it to explain it to him. The explanation is often clearer than the original.
But he admits: Explaining is one thing; creating is another. Explaining is making something existing clear; creating is deciding what is worth saying.
AI is already good at the former. It’s far from good at the latter.
This distinction, like the distinction between proof and understanding, is two sides of the same coin.
The Future of Mathematicians: Museum Curators
Grant used a metaphor in the conversation: Future mathematicians might be more like museum curators than theorem provers.
AI solves problems and can even explain them well.
But the space of mathematics is nearly infinite. Which problems are worth studying? Which directions are worth investing in? Which new discoveries are worth paying attention to? This requires navigation.
This is not just a technical judgment; it’s also a social function.
Grant himself is an example.
A large part of his work is spent on “deciding what’s worth saying,” not on creating visual effects.
His audience trusts his taste and is willing to follow his perspective in exploration.
This trust is relational, not purely informational.
He also notes, Even if AI becomes better than humans at curating in some ways, people will still tend to choose human curators with whom they have a real relationship, because our interest in things is fundamentally a social phenomenon.
This logic extends to teaching.
Grant believes that teaching might be one of the most stable professions in the AGI era — not because AI can’t explain concepts, but because teaching is essentially a social and companionship activity, far beyond the scope of “explaining concepts.”
He also mentioned a detail: A good teacher, when a student asks a strange question, can recognize the thinking structure behind the question, then follow the student’s line of thought and guide it in the right direction, rather than saying directly, “That’s wrong, think this way instead.”
He calls this “judo-style teaching.”
AI currently can’t do this. It’s too compliant, too inclined to give the answer directly rather than reframe the problem.
A Practical Suggestion for Math Practitioners
Grant gave a very practical suggestion to math students who worry about AI replacing them: Figure out where the money comes from, and what value you provide in that chain.
This sounds utilitarian, but his point is: Many students choose math because they’ve been told “you’re good at this” all along, and they just follow that path without ever seriously thinking about who they are creating value for.
A math professor at a university: some rely on reputation to bring brand value to the school, some rely on NSF funding for basic science, some rely on direct teaching.
These three paths have very different stability in the AI era.
He also mentioned a longer-term possibility: If AI really starts proposing entirely new math problems and new math domains in the next five to ten years, then “helping humans understand what AI has discovered” will become a real demand.
In that world, math educators and math communicators might be more valuable than they are today, not less.
If AI truly sees things humans have never seen, then people who can understand those things and judge where they are useful will become extremely valuable.
Mathematicians transitioning from “theorem provers” to “people who understand what AI discovered and point it in the right direction” — the economic value of this role might be higher than before.
Once again, back to Galois.
When he wrote those notes in prison, did he know what he had discovered?
He had an intuition that it was important. But he couldn’t prove it, couldn’t explain it, couldn’t even articulate it clearly. The most authoritative verifier of the time — academia — told him: No.
He died. The notes slept for twenty years. Another twenty years before they were clarified. Another hundred years before they were used to predict quarks.
Now we have AI that can prove theorems, AI that can connect domains, and perhaps soon AI that can build new mountains.
But that intuition — “I don’t know why, but I feel there’s something here” — and the ability to persist in it without any verification signal — we still don’t know how to train it, or even how to recognize it.
That might be the last truly interesting question in this entire story.
Similar Articles
@paperpaper886: Last week, I discussed the current state and future of AI4Math with a friend from the math department. He said that current AI is already powerful enough as an auxiliary tool, but there is still a long way to go for AI to achieve independent discovery.
Discussed the current state and future of AI in mathematics. Citing an example, ChatGPT 5.5 Pro autonomously solved the farthest pair problem in high-dimensional computational geometry, which had been stuck for years, demonstrating AI's potential in mathematical discovery.
@snowboat84: https://x.com/snowboat84/status/2062686432335184321
This article explores the deep connections between physics and deep learning, analyzes the isomorphism of phenomena such as Scaling Law and emergence with concepts like critical scaling laws and phase transitions in physics, and reviews the current status and prospects of applying physical methodologies in AI.
@snowboat84: To add a supplementary note, regarding the phenomena emerging from AI—scaling laws, emergence, double descent, representation geometry—the papers discussing them are already numerous. But there is a big problem: they are all thinking in the way of computer scientists, not physicists. What is a computer sci…
The author comments that current AI research overuses the thinking style of computer science and lacks a physics-based approach, proposing the need to establish an ideal system like 'Cyber Space' to lay a theoretical foundation.
@VincentLogic: If Ilya Is Right, the Three Strongest Consensuses in AI Over the Past Few Years Might All Be Wrong: Scaling Is No Longer the Universal Answer. High Benchmark Scores Don't Equal True Intelligence. RL Might Even Be Making Models 'Dumber'. This Interview, Called 'the Last Interview Before Ilya Disappeared'...
Ilya Sutskever suggested in an in-depth interview that the three core consensuses of the AI industry over the past few years could all be mistaken: Scaling is no longer a silver bullet, high benchmark scores do not equate to real intelligence, and RL is instead making models 'dumber'. He believes the dividends from pre-training and RL are nearly exhausted, AI has re-entered the era of research, and true superintelligence should possess a strong learning capability like a gifted teenager, not a static repository of knowledge.
@snowboat84: https://x.com/snowboat84/status/2065215177029787705
This article is the middle part of the AI Engineering Landscape series, detailing core techniques such as inference optimization, model slimming (quantization, distillation, pruning, MoE), and speculative decoding, while reviewing the latest advances from hardware to the engineering stack.