imo-gradingbench

Tag

Cards List
#imo-gradingbench

Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

arXiv cs.CL · 18h ago Cached

This paper studies whether cheap open-weight LLMs can judge natural-language mathematical proofs as reliably as frontier models at far lower cost. On IMO-GradingBench, three cheap judges match frontier pass/fail agreement, and the authors recommend an all-three-pass consensus rule for cost-effective deployment.

0 favorites 0 likes
← Back to home

Submit Feedback