GPT-6.1 sol (max) scores 100% in frontier math 4

Reddit r/singularity Models

Summary

OpenAI 的 GPT-6.1 sol (max) 在 Epoch AI 的 FrontierMath Tier 4 v2 基准测试中取得 100% 的成绩,登上该基准的榜首。

https://epoch.ai/benchmarks/frontiermath-tier-4-v2?view=graph&tab=leaderboard
Original Article

Similar Articles

GPT 5.6 Sol benchmarks

Reddit r/singularity

GPT 5.6 Sol achieves new benchmark results, showcasing performance improvements in AI language modeling.

Evaluating AI’s ability to perform scientific research tasks

OpenAI Blog

OpenAI introduces FrontierScience, a new benchmark for measuring expert-level AI scientific capabilities across physics, chemistry, and biology, with GPT-5.2 achieving 77% on olympiad-style tasks and 25% on research-style tasks. The paper presents early evidence that GPT-5 meaningfully accelerates real scientific workflows, shortening work from weeks to hours while establishing metrics for tracking progress toward AI-accelerated science.