Tag
The author describes using AI to assist in proving Conway's refinement conjecture on omnific integers, claiming to have obtained a Lean proof after extensive token use. The proof has passed mechanical checks but awaits independent verification.
This paper introduces PRISMA-LLM, an empirically grounded framework for reporting AI-assisted systematic reviews, addressing inconsistencies in how LLMs and AI software are documented in review workflows.
Microsoft issued over 974 security patches, including critical zero-day vulnerabilities, marking its largest single patch batch to date, with AI aiding in faster vulnerability discovery.
The article describes the development of a new economic theory on wage determination using AI tools like Opus and Fable, formalized in a paper that extends task-based models from Acemoglu and Restrepo.
Paper Pilot is a human-in-the-loop expert system designed for evidence-traceable scientific manuscript generation, using approval gates and audit logs to ensure accountability and prevent fabrication in AI-assisted research workflows.
The article reflects on how frontier LLMs and Lean automation have drastically reduced the effort required for formal proofs in programming language research, leading to doubled conference submissions and shifting publication norms.
This paper presents randomized algorithms for the Shortest Vector Problem (SVP), improving the best-known time complexity to 2^{0.6039n} classically and 2^{0.5411n} quantumly using the Hessian of the periodic Gaussian function at mid-points.
This paper evaluates how well LLMs (ChatGPT, Claude, DeepSeek) can generate one-page project plans in physics, astrophysics, and cosmology, and how human and AI reviewers assess them. Results show that human reviewers rate AI and human proposals similarly, while AI reviewers prefer AI-written proposals and can perfectly distinguish them from human-written ones.
This paper presents a case study of human-AI co-discovery in mathematics, where AI assisted in expanding an intuition about sign-embedding quantum algorithms into a formal framework and proofs, with human judgment guiding route selection.
Shared an open collaborative repository Awesome Vibe Research maintained by ModelScope. This repository collects and curates reusable, verifiable, and evolvable AI-assisted components across the full research workflow, including agents, skills, workflows, tools, and best practices. It aims to help researchers and developers leverage AI to improve research efficiency.
This paper proposes that reliability in AI-assisted social science research depends on decision architecture—how cognitive labor is divided between humans and machines. Through a pre-specified factorial experiment, the authors show that an unconstrained multi-agent baseline fails in 72% of runs, while one organized with three architectural commitments (LLMs restricted to reasoning, deterministic data/estimation, and three human decision gates) fails in only 16%.
Researchers from Charles University introduce Bolzano, an open-source multi-agent LLM system that orchestrates prover and verifier agents to assist with mathematical research, reporting new results on six problems where four reached publishable quality and three were produced essentially autonomously.
GPT-5.2 assisted in deriving a new theoretical physics result showing that single-minus gluon tree amplitudes can be nonzero under specific half-collinear momentum conditions, challenging decades of assumptions in particle physics. The AI model identified patterns in complex Feynman diagram expressions and conjectured a general formula that was subsequently verified through formal proofs.