formal-math

Tag

Cards List
#formal-math

TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations

arXiv cs.AI · 2026-08-11 Cached

Introduces TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving transformations of mathematical formulas. The best tested model achieves only 60.73% accuracy, showing that theorem knowledge is fragile under representation changes.

0 favorites 0 likes
#formal-math

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

arXiv cs.AI · 2026-06-30 Cached

This paper audits five widely used Lean theorem-proving benchmarks, uncovering 398 mechanically certified issues such as counterexamples, vacuous theorems, and unsound axioms. It proposes a fault taxonomy, automated checkers, and release standards to improve evaluation reliability and trustworthiness.

0 favorites 0 likes
#formal-math

@rohanpaul_ai: Another great paper from Google. Shows general LLMs can solve formal math by planning proofs and checking each step. Ra…

X AI KOLs Following · 2026-06-04 Cached

A new Google paper introduces LEAP, an agentic framework that enables general LLMs to solve formal math problems by planning proofs and checking each step, raising performance from under 10% to 70% on the Lean IMO benchmark and solving all 2025 Putnam problems.

0 favorites 0 likes
← Back to home

Submit Feedback