Is anyone else surprised Google DeepMind isn't leading these mathematical benchmarks?
Summary
The author expresses surprise that Google DeepMind is not leading mathematical AI benchmarks despite its past foundational work in the area, while noting OpenAI's recent progress.
Similar Articles
[Google DeepMind] the AI co-mathematician also achieves state of the art results on hard problemsolving benchmarks, including scoring 48% on FrontierMath Tier 4, a new high score among all AI systems evaluated.
Google DeepMind's AI co-mathematician achieves state-of-the-art results on hard problem-solving benchmarks, scoring 48% on FrontierMath Tier 4, the highest among all AI systems evaluated.
The new benchmarks like DeepSWE now show a very big gap in proprietary models and open source
New benchmarks like DeepSWE reveal a significant performance gap between proprietary and open-source AI models, causing disappointment in the open-source community.
DeepMind is now reportedly struggling to compete with Anthropic and OpenAI while 3.5 Pro is not the step change they'd need to be competitive
DeepMind reportedly struggles to keep up with Anthropic and OpenAI, with its 3.5 Pro model not providing the significant advancement needed to be competitive.
Bench Maxing
Discusses slipping market share for OpenAI while Meta and Google gain, questioning whether high benchmark scores matter to average users and suggesting AI's true value is as a feature within existing product ecosystems.
"OpenAI has overcome their pre-training issues, and a much larger model code named “Doug” is actively in the works"
SemiAnalysis reports that Google DeepMind is no longer a frontier lab following a leadership overhaul and key departures, while noting OpenAI has overcome pre-training issues and is working on a larger model code-named 'Doug'.