Tag
OpenAI 的 GPT-6.1 sol (max) 在 Epoch AI 的 FrontierMath Tier 4 v2 基准测试中取得 100% 的成绩,登上该基准的榜首。
FrontierMath has successfully solved its first 'Major Advance' problem, marking a significant milestone in mathematical research and AI collaboration.
David Turturean solved a 40-year-old open problem in p-adic Galois theory using voice input, in collaboration with problem proposer David Roe, under EpochAI Research's FrontierMath initiative.
Epoch AI released a v2 update to the FrontierMath benchmark, correcting errors in 42% of problems and increasing scores across all models, though rankings remained largely unchanged; Tiers 1-4 are approaching saturation.
GPT-5.5 was used by Epoch to identify fatal errors in approximately one-third of the FrontierMath benchmark problems, demonstrating the model's capability to sanity-check evaluation standards.
Google DeepMind's AI co-mathematician achieves state-of-the-art results on hard problem-solving benchmarks, scoring 48% on FrontierMath Tier 4, the highest among all AI systems evaluated.
This paper introduces the AI Co-Mathematician, a workbench that uses agentic AI to support mathematicians in open-ended research tasks like ideation and theorem proving. Early tests show the system achieving state-of-the-art results on hard problem-solving benchmarks, including a 48% score on FrontierMath Tier 4.