autoformalization

Tag

Cards List
#autoformalization

@mweber_PU: Moving from individual proofs to large-scale autoformalization requires new tools. We introduce Choir, an open protocol…

X AI KOLs Timeline ↗ · yesterday Cached

The authors introduce Choir, an open, modular protocol for distributed multi-agent autoformalization that decomposes projects into GitHub-based tasks with deterministic trust-gate verification, supporting Lean 4, Isabelle, and Rocq; it includes a preprint and an open-source repo funded by DARPA's expMath program.

0 favorites 0 likes
#autoformalization

When Scientific Contradictions Are Lost in Translation

arXiv cs.AI ↗ · 3d ago Cached

This arXiv paper studies how language models decide whether two scientific findings are comparable before resolving contradictions. In a controlled task with unsatisfiable XOR constraints translated into lab reports, GPT-5.6 Sol and Claude Opo 5 recover the best-supported assignment well when constraints are explicit, but models often defer to biological expectations when findings are presented in scientific prose.

0 favorites 0 likes
#autoformalization

Sage: Formalization with Semantic Correction

arXiv cs.LG ↗ · 4d ago Cached

华为 Lagrange 数学计算研究中心提出 Sage,一个四阶段分解生成管线加双重信号语义校正循环的自动形式化框架,将答案泄漏率从70.9%降至2.7%,在 Omni-MATH NP 上达到73.3% pass@4,并在新提出的 IMO-Unformalized 基准上零样本达到87.4%验证保真度。

0 favorites 0 likes
#autoformalization

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

arXiv cs.LG ↗ · 2026-09-11 Cached

This paper proposes Generative Verification (GenV), a method using generative reward models to detect reference-equivalence failures in autoformalization, addressing vulnerabilities in neurosymbolic systems and improving verification accuracy.

0 favorites 0 likes
#autoformalization

SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper introduces SA-Pass, a method for evaluating semantic alignment in autoformalization, and presents ShadowBench, a Lean 4 benchmark with 178 problems, demonstrating high agreement with expert judgments.

0 favorites 0 likes
#autoformalization

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

arXiv cs.CL ↗ · 2026-08-21 Cached

FormalTCS is a benchmark for evaluating large language models on end-to-end theoretical computer science research, revealing significant limitations, especially in autoformalization.

0 favorites 0 likes
#autoformalization

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

Hugging Face Daily Papers ↗ · 2026-08-14 Cached

MathForm introduces a framework for mathematical autoformalization using knowledge retrieval and verification-guided refinement, yielding the FormalVerse dataset and an 8B model that outperforms specialized baselines.

0 favorites 0 likes
#autoformalization

LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

arXiv cs.AI ↗ · 2026-07-24 Cached

LeanFlow is an LLM agent system for translating mathematical papers into formalized Lean projects, evaluated through case studies and benchmarks with Kimi-K2.6 and GPT-5.5, achieving high completion rates within budget constraints.

0 favorites 0 likes
#autoformalization

Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases

arXiv cs.AI ↗ · 2026-07-16 Cached

This position paper argues for theory-level autoformalization, which formalizes entire theories including axioms, definitions, and lemmas as coherent libraries, rather than isolated statements. It discusses the significance, alternative views, open challenges, and proposes paths forward for this shift in formalization research.

0 favorites 0 likes
#autoformalization

Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics

arXiv cs.AI ↗ · 2026-07-01 Cached

Presents an agentic framework using general coding LLMs to autoformalize research-level mathematics into Lean 4 code, evaluated on Putnam problems and STOC conference papers.

0 favorites 0 likes
#autoformalization

Leanstral 1.5

Hacker News Top ↗ · 2026-06-30 Cached

Mistral AI releases Leanstral 1.5, an updated Lean 4 formal proof engineering model optimized for automated theorem proving and autoformalization, with 119B total parameters and 6.5B active parameters.

0 favorites 0 likes
#autoformalization

Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper introduces SD-GPS, a solver-driven framework for geometry problem solving that uses autoformalization guided by solver feedback and verified theorem proposing to overcome bottlenecks in neuro-symbolic systems.

0 favorites 0 likes
#autoformalization

The Signal-Coverage Matrix: Stratifying Type and Semantic Errors in Statement Autoformalization

arXiv cs.CL ↗ · 2026-06-29 Cached

This paper introduces a signal-coverage matrix that decomposes type-correctness gains in autoformalization into four strata, revealing the mechanisms behind LLM refinements and showing that headline metrics can obscure which errors are actually resolved.

0 favorites 0 likes
#autoformalization

Autoformalization of Agent Instructions into Policy-as-Code

arXiv cs.AI ↗ · 2026-06-26 Cached

This paper presents an autoformalization pipeline that translates agent prompts, MCP tool descriptions, and natural language policy documents into formally verified policies using an LLM-based generator-critic loop, achieving better coverage than hand-coded enforcement on MedAgentBench.

0 favorites 0 likes
#autoformalization

Closing the Loop: Formally Verified Law as a Reward Signal for Self-Improving Legal AI

arXiv cs.LG ↗ · 2026-06-24 Cached

This paper presents an architecture that uses formally verified law as a reward signal for training legal AI, adaptively autoformalizing legal rules into a formal calculus and employing a verifier to ensure provable correctness, demonstrated on German and US law examples.

0 favorites 0 likes
#autoformalization

Evaluating the Robustness of Proof Autoformalization in Lean 4

arXiv cs.CL ↗ · 2026-06-16 Cached

This paper evaluates the robustness of proof autoformalization models in Lean 4 under global and local perturbations, finding that current LLM-based models are sensitive to perturbations and often fail to faithfully reflect local changes.

0 favorites 0 likes
#autoformalization

PrologMCP: A Standardized Prolog Tool Interface for LLM Agents

arXiv cs.AI ↗ · 2026-06-16 Cached

Introduces PrologMCP, an open-source server that exposes Prolog as a stateful tool via the Model Context Protocol, enabling LLM agents to delegate reasoning to a symbolic solver. Evaluation shows competitive or superior accuracy on deductive reasoning tasks compared to frontier reasoning LLMs.

0 favorites 0 likes
#autoformalization

Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptance

arXiv cs.AI ↗ · 2026-06-15 Cached

This paper presents an agent pipeline for formalizing a numerical analysis textbook in Lean 4 and introduces a quality audit framework that evaluates semantic correctness and library reuse beyond kernel acceptance, revealing common unfaithful formalization patterns.

0 favorites 0 likes
#autoformalization

Sorries Are Not the Hard Part: An Expert-Review Case Study of a Semi-Autonomous Formalization

arXiv cs.AI ↗ · 2026-06-15 Cached

This paper presents a case study of using a large language model (Claude Code) to formalize Grothendieck's vanishing theorem in the Lean theorem prover. It finds that while agents can produce verified code, they struggle with definitions and API design, emphasizing the need for expert review beyond mere compilation.

0 favorites 0 likes
#autoformalization

Characterizing initial human-AI proof formalization workflows

arXiv cs.AI ↗ · 2026-06-04 Cached

Researchers from Oxford, Cambridge, MIT, CMU and other institutions conduct a mixed-methods study examining how people integrate AI tools into mathematical proof formalization workflows, finding that participants generally achieve higher formalization accuracy with AI assistance while preferring to maintain high-level human control over the proof discovery process.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback