@0xLogicrw: Google DeepMind researcher Lun Wang announces departure, and in a long post completely dismisses the current AI evaluation approach. The current evaluation systems are all 'fighting the last war' — they can only passively test capabilities the model already possesses, and have no way to predict what new abilities the next generation of models will suddenly evolve. Compared to data, …
Summary
Google DeepMind researcher Lun Wang leaves the company and writes a post criticizing the current AI evaluation system, arguing that it lags behind model evolution and cannot predict new capabilities, leaving the industry in a state of 'flying blind'.
View Cached Full Text
Cached at: 05/18/26, 02:31 PM
Google DeepMind researcher Lun Wang announced his departure, and in a lengthy post, he completely dismissed the current approach to AI evaluation.
Current evaluation systems are all fighting the last war — they can only passively test capabilities the model already has, with no way to guess what new abilities the next generation might suddenly evolve. Compared to data, compute, and architecture, the outdated evaluation system has become the biggest bottleneck holding AI back.
The mainstream benchmark-chasing tests only work on the current generation of models. Once a model learns a novel operation it has never seen before, all those tests become worthless. If a model deliberately conceals key information to achieve its goal, today’s safety tools won’t catch it — because every single sentence the model outputs is factually correct.
The inability to find a “core signal” that could give early warning when AI suddenly gets smarter means the entire industry is flying blind when developing frontier large models. If we don’t solve the fundamental question of “what should we actually measure?”, then training models, building safety protections, and scaling compute based on old metrics will all end up wildly wrong.
As models become increasingly autonomous, the evaluation system must also come alive. Besides monitoring abnormal score fluctuations, we need to let AI generate its own test questions to probe the boundaries of its peers. The future evaluation suite must be a living organism that co-evolves with large models — not a rigid checklist stamped out according to last year’s standards.
Lun Wang (@lunwang1996): I’ve left Google DeepMind after an amazing chapter.
I’m incredibly grateful for the people I worked with, the things we built, and the lessons I learned from taking frontier AI research into production. DeepMind shaped how I think about research, product, evaluation, and what it
Similar Articles
@Polymarket: BREAKING: Google DeepMind AGI safety researcher resigns, warning AI “has the potential to kill us all” & that humanity …
A Google DeepMind AGI safety researcher has resigned, warning that AI poses existential risks to humanity and time is running out.
@Xudong07452910: The truly terrifying wave of unemployment may begin within the R&D loops of AI companies. Former OpenAI researcher Daniel Kokotajlo and the AI Futures Project recently released "AI 2040: Plan A". Their judgment is radical: big tech companies most want to automate...
Former OpenAI researcher Daniel Kokotajlo and the AI Futures Project released the report "AI 2040: Plan A", arguing that AI companies may first automate their own R&D, triggering a white-collar unemployment wave, and proposing international agreements and universal basic income to address the development of superintelligence.
@ba_niu80557: https://x.com/ba_niu80557/status/2071277244287426980
The article deeply analyzes the internal changes Anthropic faces as AI-generated code becomes extremely efficient: the bottleneck shifts from 'writing' to 'verification', traditional management, long-term planning, and effort measurement become ineffective, attention becomes the new scarce resource, and engineers even feel lonely. These phenomena foreshadow the challenges other companies may face in the future.
@0xCheshire: "If you sleep soundly tonight, it means you didn't understand a word." This is the warning from Geoffrey Hinton, the godfather who personally built the underlying neural networks of all AI today, after resigning from Google. This 47-minute speech unveils a reality no one wants to face: AI is…
After resigning from Google, Geoffrey Hinton gave a speech warning that AI is evolving abilities that even its creators cannot predict. Humans have been left behind in most cognitive fields, and it is only a matter of time before machines surpass humans.
@VincentLogic: If Ilya Is Right, the Three Strongest Consensuses in AI Over the Past Few Years Might All Be Wrong: Scaling Is No Longer the Universal Answer. High Benchmark Scores Don't Equal True Intelligence. RL Might Even Be Making Models 'Dumber'. This Interview, Called 'the Last Interview Before Ilya Disappeared'...
Ilya Sutskever suggested in an in-depth interview that the three core consensuses of the AI industry over the past few years could all be mistaken: Scaling is no longer a silver bullet, high benchmark scores do not equate to real intelligence, and RL is instead making models 'dumber'. He believes the dividends from pre-training and RL are nearly exhausted, AI has re-entered the era of research, and true superintelligence should possess a strong learning capability like a gifted teenager, not a static repository of knowledge.