@Phoenixyin13: Finished reading a long post today by OpenAI researcher Noam Brown — a reality severely underestimated by the industry. The true ceiling of LLM capabilities is far higher than what any current benchmark shows. The reason: too little test-time compute. And as models...

X AI KOLs Timeline News

Summary

Highlights OpenAI researcher Noam Brown's argument: the true ceiling of LLM capabilities is far higher than current benchmarks show, due to insufficient test-time compute, and stronger models benefit more from additional computation. This poses a serious challenge for AI safety evaluation, as many dangerous capabilities may only emerge under long time and high compute budgets.

I finished reading a long post today by OpenAI researcher Noam Brown — a reality severely underestimated by the industry. The true ceiling of LLM capabilities is far beyond what any current benchmark indicates. The reason: too little test-time compute. And as models iterate, this issue becomes more pronounced. When GPT-5.5 was first released, benchmarks showed only a slight improvement over 5.4, so many people weren't that impressed. But actual experience showed a clear step-change. Why? Because the evaluation didn't control for test-time compute. Once you plot tokens on the x-axis, 5.5 is clearly stronger under the same compute budget. Single-number benchmarks mask this gap. A more critical observation: stronger models benefit more from additional test-time compute. The performance improvement curve is steeper, the plateau is pushed further back, and it might not even exist. In Karpathy's autoresearch experiments, performance was still rising after hundreds of trials. In AI Security Institute's cybersecurity evaluations, top models were still rapidly improving after 100 million tokens. This suggests the capability ceiling may be much farther than we think. This has huge implications for AI safety and preparedness. Most current safety evaluations are conducted only under a low inference budget. But truly dangerous capabilities, especially long-horizon, multi-step planning, and complex agent behaviors, are likely to emerge only under high compute budgets. The controversy around Gemini 3 Deep Think exposed this problem. Scaffolding plus high inference can bring ordinary users close to high capabilities, but the system card did not reflect this curve. Noam points out a more difficult reality: evaluating long-horizon tasks is becoming increasingly challenging. Some capabilities can only be verified by actually running them for long periods, and model iteration cycles may already be shorter than this validation time. In the future, labs may face a dilemma: either delay releases significantly, or release with incomplete safety evaluations. More than two years have passed since o1, and the industry still habitually reports single numbers. Safety organizations are still surprised that spending 100x more inference can significantly improve performance. This is arguably the most underestimated paradigm shift of the post-o1 era. The real gap between future models will lie in who better understands and leverages the intelligence-compute curve. For those working on agents, long-horizon reasoning, and high-risk deployments, re-understanding this curve has become an absolute necessity.
Original Article
View Cached Full Text

Cached at: 06/10/26, 07:48 AM

I read through OpenAI researcher Noam Brown’s long post today — a reality that is seriously underestimated by the industry.

The true capability ceiling of LLMs is far higher than what any current benchmark shows.

The reason: they are given too little test-time compute. And as models evolve, this problem will only become more pronounced.

When GPT-5.5 was first released, its benchmarks only showed a slight improvement over GPT-5.4, leading many to feel it wasn’t that impressive.

But the actual experience was a clear step-change.

Why? Because test-time compute was not controlled for during evaluation.

Once you plot tokens on the x-axis, 5.5 is clearly stronger at the same compute budget. A single benchmark number masks this gap.

A more critical observation: stronger models benefit more from additional test-time compute.

The performance improvement curve is steeper, the plateau keeps getting pushed further away — it may not even exist.

Karpathy’s autoresearch experiments showed performance still rising after hundreds of trials.

The AI Security Institute’s cyber attack/defense evaluation showed top models still improving rapidly after 100 million tokens.

This suggests the capability ceiling may be much further away than we think.

The implications for AI safety and preparedness are huge.

Most current safety evaluations are only conducted under low inference budgets.

But truly dangerous capabilities — especially long-horizon, multi-step planning, complex agent behaviors — are likely only revealed under high compute budgets.

The controversy around Gemini 3 Deep Think exposed this problem. Scaffolding plus high inference can let ordinary users approach high capability, but the system card did not reflect this curve.

Noam points out a more stubborn reality: evaluating long-horizon tasks is becoming increasingly difficult.

Some capabilities can only be verified through actually running them over long periods, and the model iteration cycle may already be shorter than this verification time.

In the future, labs may face a dilemma: either delay release significantly, or release with incomplete safety evaluations.

More than two years have passed since o1, yet the industry still habitually reports single numbers, and safety organizations are still surprised when spending 100x more inference leads to significant capability improvements.

This is actually the most underestimated paradigm shift of the post-o1 era.

In the future, the real gap between models will be reflected in who better understands and leverages the intelligence-compute curve.

For those working on agents, long-horizon reasoning, and high-risk deployments, re-understanding this curve has become a top priority — an absolute must.

Similar Articles

@VincentLogic: If Ilya Is Right, the Three Strongest Consensuses in AI Over the Past Few Years Might All Be Wrong: Scaling Is No Longer the Universal Answer. High Benchmark Scores Don't Equal True Intelligence. RL Might Even Be Making Models 'Dumber'. This Interview, Called 'the Last Interview Before Ilya Disappeared'...

X AI KOLs Timeline

Ilya Sutskever suggested in an in-depth interview that the three core consensuses of the AI industry over the past few years could all be mistaken: Scaling is no longer a silver bullet, high benchmark scores do not equate to real intelligence, and RL is instead making models 'dumber'. He believes the dividends from pre-training and RL are nearly exhausted, AI has re-entered the era of research, and true superintelligence should possess a strong learning capability like a gifted teenager, not a static repository of knowledge.

@MaxForAI: A brutal fact: AI's mathematical abilities have already surpassed 99% of humans on this planet. @MenloVentures partner Deedy @deedydas, who invested in Anthropic and OpenRouter, conducted a test. He used the just-concluded 2026 International Mathematical Olympiad...

X AI KOLs Timeline

AI models Fable, Sol, K3, and Axiom all achieved a perfect score of 42/42 in the 2026 International Mathematical Olympiad, solving the competition completely for the first time at low cost. Among them, Claude Fable 5 was the fastest, while GPT 5.6 Sol had the lowest cost.

@0xLogicrw: Google DeepMind researcher Lun Wang announces departure, and in a long post completely dismisses the current AI evaluation approach. The current evaluation systems are all 'fighting the last war' — they can only passively test capabilities the model already possesses, and have no way to predict what new abilities the next generation of models will suddenly evolve. Compared to data, …

X AI KOLs Timeline

Google DeepMind researcher Lun Wang leaves the company and writes a post criticizing the current AI evaluation system, arguing that it lags behind model evolution and cannot predict new capabilities, leaving the industry in a state of 'flying blind'.

@Phoenixyin13: This kind of assessment method where students create questions that stump AI is indeed very innovative and highly forward-looking. Students need to explore the strengths and weaknesses of the three models: Claude, DeepSeek, and MiniMax. In this process, students no longer blindly trust AI outputs but learn to review AI responses with a critical and discerning eye, which...

X AI KOLs Timeline

This educational assessment method encourages students to explore the strengths and weaknesses of Claude, DeepSeek, and MiniMax, creating questions that defeat AI, thereby cultivating critical thinking and competitiveness needed in the AI era.