MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
Summary
MaxProof is a test-time scaling framework that enhances mathematical proof generation using a generative verifier and population-level search, achieving scores exceeding human gold-medal thresholds on IMO 2025 and USAMO 2026.
View Cached Full Text
Cached at: 06/12/26, 02:52 AM
Paper page - MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
Source: https://huggingface.co/papers/2606.13473 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
MaxProof is a test-time scaling framework that enhances mathematical proof generation by combining multiple proof-oriented capabilities and using population-level search with tournament selection to achieve competitive performance on high-level mathematical competitions.
We present MaxProof, a population-leveltest-time scalingframework for competition-levelmathematical proofin theMiniMax-M3 series. M3 first trains three proof-oriented capabilities --proof generation,proof verification, andcritique-conditioned proof repair-- using a defense-in-depthgenerative verifierengineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof throughtournament selection. With MaxProoftest-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2606\.13473
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.13473 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.13473 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.13473 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Maxproof
MaxProof introduces a test-time scaling framework that combines proof generation, verification, and repair using generative-verifier RL, enabling the M3 model to exceed human gold-medal thresholds on IMO 2025 and USAMO 2026.
@0xLogicrw: MiniMax Developer Relations Lead Ryan Lee announced that MaxProof, a test-time scaling framework for large language model mathematical proofs, has been officially open-sourced, along with a companion technical paper. MaxProof restructures mathematical proof during inference into an evolutionary search system, enabling inference scaling through verification, repair, and elimination mechanisms.
MiniMax open-sourced MaxProof, a test-time scaling framework for LLM mathematical proofs, and released a companion paper. The framework uses an evolutionary search mechanism to enable the M3 model to achieve gold-medal scores on both the IMO 2025 and USAMO 2026 test sets.
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
AdvancedMathBench is a new benchmark suite for evaluating LLMs on advanced mathematical proof generation and verification. It includes ProverBench for generation and VerifierBench for verification, demonstrating that current models like GPT-5.5-xhigh achieve only modest performance.
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
RePro integrates Lean-oriented neural automated theorem provers into benchmark rewriting to ensure problem validity and answer correctness for reliable evaluation of LLMs in mathematical problem solving.
Evaluating Research-Level Math Proofs via Strict Step-Level Verification
This paper introduces a strict step-level verification framework for evaluating research-level mathematical proofs using LLMs, addressing context poisoning and outperforming global evaluation. The approach shifts focus to deductive constraints and reveals that remaining errors are often due to pedantic hyper-rigor, exposing implicit ambiguities in benchmarks.