imo

Tag

Cards List
#imo

We compared different LLMs on IMO 2026 [R]

Reddit r/MachineLearning · 6d ago

This study evaluates frontier and open-weight LLMs on IMO 2026 problems, demonstrating that specialized harnesses like AutoFyn significantly improve performance of sub-frontier models, though hallucination issues persist on the hardest problem.

0 favorites 0 likes
#imo

@enginenerdx: same model, three medals: sonnet 5 scored 14/42 on the IMO in a web ui, 21 in claude code, 35 in a structured multi-age…

X AI KOLs Timeline · 2026-07-25 Cached

A comparison shows that LLM performance on the 2026 IMO varies dramatically depending on the evaluation harness, with structured multi-agent setups achieving far higher scores than simple web UI, indicating that current gains are absorbed at the frontier by better orchestration.

0 favorites 0 likes
#imo

@MaxForAI: A brutal fact: AI's mathematical abilities have already surpassed 99% of humans on this planet. @MenloVentures partner Deedy @deedydas, who invested in Anthropic and OpenRouter, conducted a test. He used the just-concluded 2026 International Mathematical Olympiad...

X AI KOLs Timeline · 2026-07-21 Cached

AI models Fable, Sol, K3, and Axiom all achieved a perfect score of 42/42 in the 2026 International Mathematical Olympiad, solving the competition completely for the first time at low cost. Among them, Claude Fable 5 was the fastest, while GPT 5.6 Sol had the lowest cost.

0 favorites 0 likes
#imo

@ChrisHayduk: https://x.com/ChrisHayduk/status/2076196217109746095

X AI KOLs Timeline · 2026-07-12 Cached

This article compares two AI approaches for mathematical problem solving: DeepMind's AlphaProof, which uses reinforcement learning in the Lean proof language, and OpenAI's raw LLM that achieved a gold medal at the 2025 International Math Olympiad without formal methods.

0 favorites 0 likes
#imo

@vista8: https://x.com/vista8/status/2072191315916538039

X AI KOLs Timeline · 2026-07-01 Cached

Starting with the story of Galois group theory, the article delves into the boundaries of AI's capabilities in mathematics, distinguishing between two types of progress: "connecting lightning" (cross-domain connections) and "building mountains" (creating new frameworks). It analyzes the limitations of the RLVR training method and introduces the concept of "grindability" to explain AI's rapid advancements in mathematics and coding.

0 favorites 0 likes
#imo

Maxproof

Hacker News Top · 2026-06-12 Cached

MaxProof introduces a test-time scaling framework that combines proof generation, verification, and repair using generative-verifier RL, enabling the M3 model to exceed human gold-medal thresholds on IMO 2025 and USAMO 2026.

0 favorites 0 likes
#imo

Gemini 3.2 Flash is capable of solving IMO 2025 P6. Only GPT-5.5-Pro can solve it currently without any scaffolding / harness engineering.

Reddit r/singularity · 2026-05-18

Gemini 3.2 Flash can solve IMO 2025 P6, but only GPT-5.5-Pro can do so without any scaffolding or harness engineering.

0 favorites 0 likes
#imo

MIT & the IMO released MathNet, the world’s largest dataset of International Math Olympiad problems & solutions. MathNet is 5x larger than previous datasets & is sourced from over 40 countries across 4 decades

Reddit r/LocalLLaMA · 2026-04-22

MIT and the IMO release MathNet, a massive dataset of International Math Olympiad problems and solutions spanning 40 years and 40+ countries, 5x larger than prior datasets.

0 favorites 0 likes
#imo

Advanced Gemini with Deep Think Achieves Gold Medal Standard at International Mathematical Olympiad

Google DeepMind Blog · 2025-10-24 Cached

Google DeepMind's advanced Gemini with Deep Think achieved gold-medal standard at the International Mathematical Olympiad 2025, solving 5 out of 6 problems for 35 points—a significant advance over last year's silver-medal performance, operating end-to-end in natural language within competition time limits.

0 favorites 0 likes
← Back to home

Submit Feedback