@berryxia: Honestly, only truly brilliant people dare to say such things! An undergraduate student can handle the math training of LLMs! In a recent interview, Terence Tao laid out the core mystery of LLMs directly. The Fields Medal winner, the highest honor in mathematics — often called the Nobel Prize of math — and one of the most top contemporary…
Summary
Terence Tao pointed out that the math behind current LLMs is actually very simple, but the real puzzle lies in the intermediate zone of natural language data, which leads to unpredictable model behavior.
View Cached Full Text
Cached at: 05/17/26, 05:29 AM
To be honest, only truly brilliant people would dare to say something like this!
Undergrads can handle the math training for LLMs!
Terence Tao recently laid out the core mystery of LLMs plainly in an interview.
The Fields Medalist, winner of the highest honor in mathematics — often called the Nobel Prize of math — one of the most accomplished mathematicians of our time, said:
The math behind today’s large models is actually quite simple.
Linear algebra, matrix multiplication, plus a bit of calculus — an undergraduate can fully grasp it.
We know exactly how to train them and how to run them.
But what’s truly puzzling is: why do they perform amazingly on some tasks, yet suddenly fail on others, and we have no way to predict that in advance?
The core reason lies in real-world data — natural language text.
It’s neither pure noise nor fully structured data; it sits in a “middle ground”: partially ordered, partially random. Current mathematical theory for this intermediate zone is still very weak.
So we can build powerful models, but we cannot reliably predict their capability boundaries.
This contradiction of “simple mechanisms vs. unpredictable behavior” is the core puzzle of AI today.
Full interview video here (Dr. Brian Keating’s channel):
Rohan Paul (@rohanpaul_ai): Terence Tao says the math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models.
The real mystery is
Full video:
Similar Articles
@rohanpaul_ai: Terence Tao says the math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra,…
Terence Tao states that the mathematics underlying modern LLMs is simple, using basic linear algebra and calculus, but the unpredictability of model performance across tasks remains a mystery due to the complex nature of natural language data.
@haider1: Yann LeCun says LLMs are strongest in domains where language itself is the substrate of reasoning, like math and code T…
Yann LeCun states that LLMs are strongest in domains where language is the substrate of reasoning, like math and code, but they are not creative mathematicians, software architects, or computer scientists.
@FinanceYF5: Google new paper: Let LLM solve math competition problems, accuracy jumps from 10% to 70%. [LEAP framework] Instead of having the model write a complete proof at once, it breaks down the problem into a goal tree, learns step by step from Lean verifier feedback, and reuses proven lemmas. Result: All 12 problems of Putnam 2025 solved, IMO style…
Google new paper proposes the LEAP framework, which decomposes math problems into goal trees, learns from Lean verifier feedback, and improves LLM accuracy on math competition problems from 10% to 70%. It solves all 12 problems of Putnam 2025 and surpasses dedicated gold-medal-level systems on IMO-style benchmarks.
Tim Gowers: What sort of maths are LLMs good at?
Tim Gowers reflects on what kinds of mathematical problems LLMs are good at, noting that the most famous solved problems involve counterexamples and discussing potential explanations.
@freeman1266: You don't need math to understand most AI papers—just understand this chain: token → embedding → position encoding → attention → FFN → residual stream → next-token prediction. LLMs essentially stack Transf…
A Chinese science tweet that intuitively explains the core chain of LLMs (Large Language Models): from token, embedding, position encoding, attention, FFN to residual stream and next-token prediction, helping readers without a math background understand AI papers.