Tag
This study evaluates frontier and open-weight LLMs on IMO 2026 problems, demonstrating that specialized harnesses like AutoFyn significantly improve performance of sub-frontier models, though hallucination issues persist on the hardest problem.
A comparison shows that LLM performance on the 2026 IMO varies dramatically depending on the evaluation harness, with structured multi-agent setups achieving far higher scores than simple web UI, indicating that current gains are absorbed at the frontier by better orchestration.
AI models Fable, Sol, K3, and Axiom all achieved a perfect score of 42/42 in the 2026 International Mathematical Olympiad, solving the competition completely for the first time at low cost. Among them, Claude Fable 5 was the fastest, while GPT 5.6 Sol had the lowest cost.
This article compares two AI approaches for mathematical problem solving: DeepMind's AlphaProof, which uses reinforcement learning in the Lean proof language, and OpenAI's raw LLM that achieved a gold medal at the 2025 International Math Olympiad without formal methods.
Starting with the story of Galois group theory, the article delves into the boundaries of AI's capabilities in mathematics, distinguishing between two types of progress: "connecting lightning" (cross-domain connections) and "building mountains" (creating new frameworks). It analyzes the limitations of the RLVR training method and introduces the concept of "grindability" to explain AI's rapid advancements in mathematics and coding.
MaxProof introduces a test-time scaling framework that combines proof generation, verification, and repair using generative-verifier RL, enabling the M3 model to exceed human gold-medal thresholds on IMO 2025 and USAMO 2026.
Gemini 3.2 Flash can solve IMO 2025 P6, but only GPT-5.5-Pro can do so without any scaffolding or harness engineering.
MIT and the IMO release MathNet, a massive dataset of International Math Olympiad problems and solutions spanning 40 years and 40+ countries, 5x larger than prior datasets.
Google DeepMind's advanced Gemini with Deep Think achieved gold-medal standard at the International Mathematical Olympiad 2025, solving 5 out of 6 problems for 35 points—a significant advance over last year's silver-medal performance, operating end-to-end in natural language within competition time limits.