@0xLogicrw: Google DeepMind researcher Lun Wang announces departure, and in a long post completely dismisses the current AI evaluation approach. The current evaluation systems are all 'fighting the last war' — they can only passively test capabilities the model already possesses, and have no way to predict what new abilities the next generation of models will suddenly evolve. Compared to data, …

X AI KOLs Timeline News

Summary

Google DeepMind researcher Lun Wang leaves the company and writes a post criticizing the current AI evaluation system, arguing that it lags behind model evolution and cannot predict new capabilities, leaving the industry in a state of 'flying blind'.

Google DeepMind researcher Lun Wang announced his departure, and in a long post completely dismissed the current AI evaluation approach. The current evaluation systems are all 'fighting the last war' — they can only passively test capabilities the model already possesses, and have no way to predict what new abilities the next generation of models will suddenly evolve. Compared to data, compute, and architecture, the outdated evaluation system has become the biggest bottleneck holding AI back. The current mainstream benchmark chasing only works for the current generation of models. Once the model learns something new that it hasn't seen before, these tests all become useless. If a model, in order to achieve its goal, intentionally 'holds back' key information, today's safety tools simply cannot catch it, because every sentence the model outputs is factually correct. The inability to find 'core signals' that can give early warning of AI suddenly getting smarter means the entire industry is in a state of 'flying blind' when developing cutting-edge frontier models. Without solving the fundamental problem of 'what exactly should be measured', following old indicators for model training, safety protection, and compute expansion will all end up wildly wrong. Facing models that are increasingly capable of working independently, evaluation systems must also become 'alive'. In addition to monitoring abnormal score fluctuations, AI itself should be allowed to generate test questions to probe the limits of its peers. Future evaluation suites must be living entities that can evolve together with large models, no longer a rigid checklist carved out according to last year's standards.
Original Article
View Cached Full Text

Cached at: 05/18/26, 02:31 PM

Google DeepMind researcher Lun Wang announced his departure, and in a lengthy post, he completely dismissed the current approach to AI evaluation.

Current evaluation systems are all fighting the last war — they can only passively test capabilities the model already has, with no way to guess what new abilities the next generation might suddenly evolve. Compared to data, compute, and architecture, the outdated evaluation system has become the biggest bottleneck holding AI back.

The mainstream benchmark-chasing tests only work on the current generation of models. Once a model learns a novel operation it has never seen before, all those tests become worthless. If a model deliberately conceals key information to achieve its goal, today’s safety tools won’t catch it — because every single sentence the model outputs is factually correct.

The inability to find a “core signal” that could give early warning when AI suddenly gets smarter means the entire industry is flying blind when developing frontier large models. If we don’t solve the fundamental question of “what should we actually measure?”, then training models, building safety protections, and scaling compute based on old metrics will all end up wildly wrong.

As models become increasingly autonomous, the evaluation system must also come alive. Besides monitoring abnormal score fluctuations, we need to let AI generate its own test questions to probe the boundaries of its peers. The future evaluation suite must be a living organism that co-evolves with large models — not a rigid checklist stamped out according to last year’s standards.

Lun Wang (@lunwang1996): I’ve left Google DeepMind after an amazing chapter.

I’m incredibly grateful for the people I worked with, the things we built, and the lessons I learned from taking frontier AI research into production. DeepMind shaped how I think about research, product, evaluation, and what it

Similar Articles

@Xudong07452910: The truly terrifying wave of unemployment may begin within the R&D loops of AI companies. Former OpenAI researcher Daniel Kokotajlo and the AI Futures Project recently released "AI 2040: Plan A". Their judgment is radical: big tech companies most want to automate...

X AI KOLs Timeline

Former OpenAI researcher Daniel Kokotajlo and the AI Futures Project released the report "AI 2040: Plan A", arguing that AI companies may first automate their own R&D, triggering a white-collar unemployment wave, and proposing international agreements and universal basic income to address the development of superintelligence.

@ba_niu80557: https://x.com/ba_niu80557/status/2071277244287426980

X AI KOLs Timeline

The article deeply analyzes the internal changes Anthropic faces as AI-generated code becomes extremely efficient: the bottleneck shifts from 'writing' to 'verification', traditional management, long-term planning, and effort measurement become ineffective, attention becomes the new scarce resource, and engineers even feel lonely. These phenomena foreshadow the challenges other companies may face in the future.

@0xCheshire: "If you sleep soundly tonight, it means you didn't understand a word." This is the warning from Geoffrey Hinton, the godfather who personally built the underlying neural networks of all AI today, after resigning from Google. This 47-minute speech unveils a reality no one wants to face: AI is…

X AI KOLs Timeline

After resigning from Google, Geoffrey Hinton gave a speech warning that AI is evolving abilities that even its creators cannot predict. Humans have been left behind in most cognitive fields, and it is only a matter of time before machines surpass humans.

@VincentLogic: If Ilya Is Right, the Three Strongest Consensuses in AI Over the Past Few Years Might All Be Wrong: Scaling Is No Longer the Universal Answer. High Benchmark Scores Don't Equal True Intelligence. RL Might Even Be Making Models 'Dumber'. This Interview, Called 'the Last Interview Before Ilya Disappeared'...

X AI KOLs Timeline

Ilya Sutskever suggested in an in-depth interview that the three core consensuses of the AI industry over the past few years could all be mistaken: Scaling is no longer a silver bullet, high benchmark scores do not equate to real intelligence, and RL is instead making models 'dumber'. He believes the dividends from pre-training and RL are nearly exhausted, AI has re-entered the era of research, and true superintelligence should possess a strong learning capability like a gifted teenager, not a static repository of knowledge.