formal-verification

Tag

Cards List
#formal-verification

Learning to Discover Interesting Mathematics

arXiv cs.LG ↗ · 23h ago Cached

This paper introduces a framework for LLMs to discover and prove interesting mathematical theorems by optimizing for a metric based on proof difficulty, leading to more novel and useful mathematical knowledge with reduced overlap with existing libraries.

0 favorites 0 likes
#formal-verification

Provably Complete Generalized Planning with LLMs

arXiv cs.AI ↗ · yesterday Cached

This paper introduces an approach for automatically generating generalized plans in Lean with formal proofs of completeness using LLMs, evaluated on benchmark domains showing significant advancements in automatic plan verification.

0 favorites 0 likes
#formal-verification

@jiayq: TLA+ is also (one of the) method that we at Intent Lab used to build mission critical infra software like Agent FS. Key…

X AI KOLs Following ↗ · 2d ago Cached

The tweet discusses using formal verification methods like TLA+ and Lean to ensure the correctness and scalability of mission-critical AI infrastructure software, with references to Intent Lab and Boris Cherny's work on the Claude Agent SDK.

0 favorites 0 likes
#formal-verification

@bcherny: I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs…

X AI KOLs Timeline ↗ · 3d ago Cached

The author used Claude's Opus 5.5 model to formally verify the Claude Agent SDK using Lean, generating 16 PRs to fix bugs and race conditions, and recommends combining Lean with TLA+ for enhanced bug finding.

0 favorites 0 likes
#formal-verification

Security auditing in the age of (good enough) AI

Hacker News Top ↗ · 4d ago Cached

Trail of Bits employed AI agents to develop custom security auditing tools and formal models for the Miden zero-knowledge VM, uncovering critical vulnerabilities and generating machine-checked correctness proofs.

0 favorites 0 likes
#formal-verification

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

arXiv cs.LG ↗ · 4d ago Cached

This paper introduces SWE-Proof, a benchmark of formally verified code patches for real-world software issues, demonstrating that formal verification improves error detection in LLM-generated code and identifies specification synthesis as a key open problem.

0 favorites 0 likes
#formal-verification

@poteto: very bullish on Bend, this is exactly what ive been trying to do in other languages, but you can only get so far with w…

X AI KOLs Timeline ↗ · 6d ago Cached

@poteto expresses enthusiasm for Bend, a language enabling formal verification to accelerate software development and solve code review, with @VictorTaelin agreeing on its potential for building complex software without mistakes.

0 favorites 0 likes
#formal-verification

A Lean-verified proof can still prove the wrong version of a problem

Reddit r/ArtificialInteligence ↗ · 6d ago

The article discusses the verification of OpenAI's Lean proof for the Navier-Stokes problem, highlighting the importance of ensuring formal proofs align with intended mathematical problems and the need for further scrutiny by mathematicians.

0 favorites 0 likes
#formal-verification

Bend 2 and the Vibe-Coding Trap

Hacker News Top ↗ · 2026-09-18 Cached

The article critiques Bend 2, a programming language designed for the AI coding era, for falling into a 'vibe-coding trap' and compares it unfavorably to formal verification approaches like SPARK.

0 favorites 0 likes
#formal-verification

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

arXiv cs.AI ↗ · 2026-09-18 Cached

MAGS introduces a multi-agent framework that uses formal verification with Dafny to generate executable programs with safety guarantees from LLM coding agents, achieving 100% success in producing verified code across domains like CUDA kernels and robotic tasks.

0 favorites 0 likes
#formal-verification

RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines

arXiv cs.CL ↗ · 2026-09-15 Cached

The paper evaluates LLMs' ability to interpret network protocol specifications and map them to formal finite state machines, assessing their reasoning through designed tasks and queries for various protocols.

0 favorites 0 likes
#formal-verification

Developing provably correct Rust code with Verus

Hacker News Top ↗ · 2026-09-14 Cached

Verus is an open-source automated program verifier for Rust that helps ensure code correctness through mathematical proofs, and is used at Amazon for projects like the Nitro Isolation Engine.

0 favorites 0 likes
#formal-verification

Is truth futureproof? On the possible futures of mechanized proofs

Lobsters Hottest ↗ · 2026-09-12 Cached

This paper explores the future possibilities of mechanized proofs, questioning whether truth can be futureproof in the context of automated theorem proving.

0 favorites 0 likes
#formal-verification

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

arXiv cs.CL ↗ · 2026-09-10 Cached

StochBench introduces a domain-specific benchmark of 450 graduate stochastic processes problems in Lean 4, evaluated with an AI agent achieving a 34.9% proof rate, to advance formal theorem proving in applied mathematics.

0 favorites 0 likes
#formal-verification

A timeline of recent events leading to the solution of Navier-Stokes

Reddit r/singularity ↗ · 2026-09-08

A timeline of events from 2025 to 2026 detailing the collaborative and competitive efforts to solve the Navier-Stokes equations, culminating in OpenAI's announcement of a solution using AI models.

0 favorites 0 likes
#formal-verification

Finding a bug in Dummit and Foote's Abstract Algebra

Hacker News Top ↗ · 2026-09-04 Cached

A blog post detailing how the author found a bug in Dummit and Foote's Abstract Algebra textbook while formalizing it in Rocq, specifically that a proposition about injective functions and left inverses is false for empty sets.

0 favorites 0 likes
#formal-verification

RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving

arXiv cs.CL ↗ · 2026-09-02 Cached

RePro integrates Lean-oriented neural automated theorem provers into benchmark rewriting to ensure problem validity and answer correctness for reliable evaluation of LLMs in mathematical problem solving.

0 favorites 0 likes
#formal-verification

SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper introduces SA-Pass, a method for evaluating semantic alignment in autoformalization, and presents ShadowBench, a Lean 4 benchmark with 178 problems, demonstrating high agreement with expert judgments.

0 favorites 0 likes
#formal-verification

I'm building an independent verification layer for Ai generated-claims and I'm lokking for researchers and partners to build with us.

Reddit r/artificial ↗ · 2026-08-29

The author is developing a deterministic verification engine for AI-generated claims, focusing on formal verification methods and seeking researchers and partners for collaboration.

0 favorites 0 likes
#formal-verification

ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving

arXiv cs.AI ↗ · 2026-08-28 Cached

ProofEvolve is a neuro-symbolic framework that evolves formally verified proof structures using neural models to enhance automated theorem proving, achieving high solve rates on Lean benchmarks by preserving verified knowledge from incomplete attempts.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback