theorem-proving

Tag

Cards List
#theorem-proving

Learning to Discover Interesting Mathematics

arXiv cs.LG ↗ · 23h ago Cached

This paper introduces a framework for LLMs to discover and prove interesting mathematical theorems by optimizing for a metric based on proof difficulty, leading to more novel and useful mathematical knowledge with reduced overlap with existing libraries.

0 favorites 0 likes
#theorem-proving

Learning to Discover Interesting Mathematics

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper defines intrinsic interestingness for mathematical theorems using proof-to-statement length ratio and trains a 27B model to predict proof difficulty, enabling the generation of more interesting theorems and self-expanding mathematical libraries with reduced overlap with existing knowledge like Mathlib.

0 favorites 0 likes
#theorem-proving

@ElliotGlazer: Wtf CIC + LEM proves Con(ZF) ?!? Which then gives a proof of Con(ZF) in axiom-free Lean + LEM. Found by Mario Carneiro …

X AI KOLs Following ↗ · 4d ago Cached

Mario Carneiro discovered that CIC plus LEM proves the consistency of ZF set theory, enabling a proof in axiom-free Lean with LEM, advancing type theoretic metamathematics.

0 favorites 0 likes
#theorem-proving

Deep theorems were scarce. AI has broken this system (15 minute read)

TLDR AI ↗ · 2026-09-14 Cached

The article discusses how artificial intelligence is revolutionizing the production of deep theorems in mathematics, challenging traditional measures of success, with a focus on the Nivat conjecture.

0 favorites 0 likes
#theorem-proving

@dhuynh95: So @OpenAI just spent $10M worth of tokens to win a $1M math prize by teaching AI to do a rasengan in Lean? What a time…

X AI KOLs Following ↗ · 2026-09-13 Cached

OpenAI reportedly spent $10M worth of tokens to solve the Navier-Stokes Millennium Prize Problem, winning a $1M math prize by using an advanced AI model.

0 favorites 0 likes
#theorem-proving

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

arXiv cs.CL ↗ · 2026-09-10 Cached

StochBench introduces a domain-specific benchmark of 450 graduate stochastic processes problems in Lean 4, evaluated with an AI agent achieving a 34.9% proof rate, to advance formal theorem proving in applied mathematics.

0 favorites 0 likes
#theorem-proving

OpenAI Says It Has Cracked One of Math’s ‘Millennium Problems’

Reddit r/singularity ↗ · 2026-09-08

OpenAI claims to have solved one of the Millennium Prize Problems in mathematics, marking a potential major breakthrough in AI and mathematical research.

0 favorites 0 likes
#theorem-proving

C*: Unifying Programming and Verification in C

Hacker News Top ↗ · 2026-09-08 Cached

This paper introduces C*, a proof-integrated language that unifies C programming with formal verification, enabling real-time verification through embedded proof-code blocks.

0 favorites 0 likes
#theorem-proving

New lean proof repos by Openai ahead of Astra release

Reddit r/singularity ↗ · 2026-09-03

OpenAI has released new Lean theorem proving repositories on GitHub, ahead of their upcoming Astra release.

0 favorites 0 likes
#theorem-proving

ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving

arXiv cs.AI ↗ · 2026-08-28 Cached

ProofEvolve is a neuro-symbolic framework that evolves formally verified proof structures using neural models to enhance automated theorem proving, achieving high solve rates on Lean benchmarks by preserving verified knowledge from incomplete attempts.

0 favorites 0 likes
#theorem-proving

FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence

arXiv cs.AI ↗ · 2026-08-28 Cached

FaithSieve is a Lean-assisted framework for fine-grained evaluation of mathematical proofs that uses semantic alignment scoring to ensure faithful formal evidence, achieving higher accuracy than baseline methods on expert-verified datasets.

0 favorites 0 likes
#theorem-proving

MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

arXiv cs.CL ↗ · 2026-08-27 Cached

MathAdv is a diagnostic benchmark for formal theorem proving in mathematics, covering 13 domains with auxiliary tasks to evaluate knowledge, reasoning, and robustness. The study reveals formalization bottlenecks and performance variations across models.

0 favorites 0 likes
#theorem-proving

Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving

arXiv cs.CL ↗ · 2026-08-20 Cached

This paper introduces a compiler-guided adaptive proof search framework for context-dependent theorem proving in Lean 4, using cross-model synergy to improve proof success rates while reducing computational cost.

0 favorites 0 likes
#theorem-proving

MathCode, Mathematical Coding Agent

Hacker News Top ↗ · 2026-08-16 Cached

MathCode is an AI-powered coding assistant that converts mathematical problems into Lean 4 theorems and attempts formal proofs, featuring a persistent REPL, theorem libraries, and agent-mode proving.

0 favorites 0 likes
#theorem-proving

PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper presents Prove-RT, an LLM-assisted framework for generating Prosa/Rocq mechanized theorem prover scripts for schedulability analysis in real-time systems, achieving a 44.7% success rate on a curated evaluation set.

0 favorites 0 likes
#theorem-proving

@rohanpaul_ai: $2K of tokens costs less than sending 2 people to a conference, and that gap is what changes the calculation for open p…

X AI KOLs Following ↗ · 2026-08-03 Cached

OpenAI reports that an internal version of its next major model (Astra) solved 10 long-standing open problems in math and theoretical computer science for roughly $2,000 in tokens, with formal Lean certificates.

0 favorites 0 likes
#theorem-proving

Ten advances in mathematics and theoretical computer science

OpenAI Blog ↗ · 2026-08-01 Cached

OpenAI announces ten results on long-standing open problems in mathematics and theoretical computer science, achieved by an internal version of its next model Astra, with proofs formalized in Lean.

0 favorites 0 likes
#theorem-proving

Why Rocq is better than Lean for program verification

Lobsters Hottest ↗ · 2026-07-28 Cached

A technical blog post argues that Rocq (Coq) is better than Lean for program verification due to Rocq's native support for coinductive types and cofixpoints, contrasting with Lean's less mature, library-based approach.

0 favorites 0 likes
#theorem-proving

Learned Interventions in Lean 4 grind

arXiv cs.LG ↗ · 2026-07-28 Cached

A research paper introducing a failure-triggered cascade approach to safely integrate machine learning into Lean 4's grind tactic, achieving improved efficiency and solving previously unsolvable proofs without regressions.

0 favorites 0 likes
#theorem-proving

LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

arXiv cs.AI ↗ · 2026-07-24 Cached

LeanFlow is an LLM agent system for translating mathematical papers into formalized Lean projects, evaluated through case studies and benchmarks with Kimi-K2.6 and GPT-5.5, achieving high completion rates within budget constraints.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback