GPT-6.1 sol (max) scores 100% in frontier math 4
Summary
OpenAI 的 GPT-6.1 sol (max) 在 Epoch AI 的 FrontierMath Tier 4 v2 基准测试中取得 100% 的成绩,登上该基准的榜首。
Similar Articles
GPT-6.1 Sol reaches third place in Humanity's Last Exam — Diamond
GPT-6.1 Sol has reached third place on the Humanity's Last Exam (Diamond) leaderboard, marking a significant milestone in frontier AI benchmark performance.
GPT-6.1 Sol is now #1 on MathArena: 86.3% accuracy for $0.94, beating Astra’s 81.9% at $2.26
GPT-6.1 Sol claims the top spot on MathArena with 86.3% accuracy at a cost of just $0.94, outperforming Astra's 81.9% achieved at $2.26. The result highlights a significant advance in math reasoning performance and cost-efficiency among frontier models.
GPT-6 Sol & Astra dominate the ARC-AGI-3 leaderboard
GPT-6 Sol and Astra have taken the top spots on the ARC-AGI-3 leaderboard, marking a notable advance in abstract reasoning benchmarks for frontier AI models.
GPT 5.6 Sol benchmarks
GPT 5.6 Sol achieves new benchmark results, showcasing performance improvements in AI language modeling.
Evaluating AI’s ability to perform scientific research tasks
OpenAI introduces FrontierScience, a new benchmark for measuring expert-level AI scientific capabilities across physics, chemistry, and biology, with GPT-5.2 achieving 77% on olympiad-style tasks and 25% on research-style tasks. The paper presents early evidence that GPT-5 meaningfully accelerates real scientific workflows, shortening work from weeks to hours while establishing metrics for tracking progress toward AI-accelerated science.