MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Hugging Face Daily Papers Papers

Summary

MaxProof is a test-time scaling framework that enhances mathematical proof generation using a generative verifier and population-level search, achieving scores exceeding human gold-medal thresholds on IMO 2025 and USAMO 2026.

We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.
Original Article
View Cached Full Text

Cached at: 06/12/26, 02:52 AM

Paper page - MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Source: https://huggingface.co/papers/2606.13473 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

MaxProof is a test-time scaling framework that enhances mathematical proof generation by combining multiple proof-oriented capabilities and using population-level search with tournament selection to achieve competitive performance on high-level mathematical competitions.

We present MaxProof, a population-leveltest-time scalingframework for competition-levelmathematical proofin theMiniMax-M3 series. M3 first trains three proof-oriented capabilities --proof generation,proof verification, andcritique-conditioned proof repair-- using a defense-in-depthgenerative verifierengineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof throughtournament selection. With MaxProoftest-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2606\.13473

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2606.13473 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2606.13473 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.13473 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Maxproof

Hacker News Top

MaxProof introduces a test-time scaling framework that combines proof generation, verification, and repair using generative-verifier RL, enabling the M3 model to exceed human gold-medal thresholds on IMO 2025 and USAMO 2026.

@0xLogicrw: MiniMax Developer Relations Lead Ryan Lee announced that MaxProof, a test-time scaling framework for large language model mathematical proofs, has been officially open-sourced, along with a companion technical paper. MaxProof restructures mathematical proof during inference into an evolutionary search system, enabling inference scaling through verification, repair, and elimination mechanisms.

X AI KOLs Timeline

MiniMax open-sourced MaxProof, a test-time scaling framework for LLM mathematical proofs, and released a companion paper. The framework uses an evolutionary search mechanism to enable the M3 model to achieve gold-medal scores on both the IMO 2025 and USAMO 2026 test sets.

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

arXiv cs.AI

This paper introduces a strict step-level verification framework for evaluating research-level mathematical proofs using LLMs, addressing context poisoning and outperforming global evaluation. The approach shifts focus to deductive constraints and reveals that remaining errors are often due to pedantic hyper-rigor, exposing implicit ambiguities in benchmarks.