Tag
This paper introduces the Hypothesis Evolution Protocol (HEP) for LLM agents, which makes hypothesis generation, testing, and belief updates explicit and auditable. Experiments on materials-science tasks show that HEP-equipped agents generalize across research questions and become more effective with stronger base LLMs.
Introduces SDABench, a benchmark evaluating LLMs on six scientific analysis capabilities across five domains, finding models struggle with tasks requiring assumption selection and mechanistic reasoning.
An analysis of over 40 million papers finds that AI tools boost individual researchers' productivity and career advancement but narrow the scope of scientific inquiry, leading to less diverse and original discoveries, as published in Nature.
Noam Brown announces that GPT-5.6 Sol Ultra proved a 50-year-old math conjecture, demonstrating impressive prompt engineering and agent prompting.
Introduces ZendoWorld, a controlled interactive environment for evaluating AI agents on active visual concept induction, where agents must perceive scenes, infer hidden logical rules, and design informative experiments. Experiments with various agent classes reveal that high prediction accuracy does not guarantee rule recovery, and VLM-based agents struggle with informative experimentation, highlighting gaps compared to human inductive reasoning.
This paper introduces Statistically Meaningful Geometry (SMG), a geometric framework for modeling over-parameterized learning systems as infinite-dimensional non-parametric Orlicz fiber bundles. It proposes that under out-of-distribution stimuli, the system undergoes a gauge symmetry break, leading to the emergence of new causal axes that can distinguish genuine scientific discovery from hallucinations.
FirstResearch introduces a structured framework for LLM scientific discovery agents that generates a Research Question Certificate containing primitive definitions, assumptions, mechanism, falsifiable hypothesis, and failure update rules, making the proposed research question inspectable before execution. Preliminary evaluations using LLM judges show that the certificate-centered approach outperforms baseline systems in audibility and score.
This paper presents SCION, an agentic scientific operating system that integrates AI tools for scientific discovery through a Research Execution Plan (REP) and hierarchical multi-agent execution. It demonstrates applications in materials analysis, molecule design, and protein screening, outperforming existing autonomous research-agent baselines.
This paper introduces LLM-PySR, a method where language models guide symbolic equation discovery by controlling search parameters while using numerical symbolic regression for fitting. The approach achieves strong balance of accuracy and complexity across benchmark tasks.
Damo Academy (Alibaba) introduces Elements Claw, an AI agent that discovered four new superconducting materials by screening millions of crystal structures, potentially accelerating materials research.
EO-Agents presents a three-agent LLM pipeline for generating Earth observation hypotheses, leveraging a NASA knowledge graph and graph neural network to rank candidate dataset pairings, with LLM agents filtering, generating, and evaluating structured research hypotheses.
Sakana AI has released AI Scientist-v2, an end-to-end automated research system that can autonomously go from generating research hypotheses to writing papers, and has been accepted by the ICLR2025 Workshop after peer review.
This paper addresses the limitation of static surprisal in LLM-based scientific discovery by introducing evidence-informed non-stationary beliefs, and proposes belief-update filtering and diversity maximization to improve discovery, achieving 30.62% higher non-stationary surprisal across five domains.
Discusses the challenge of verifying AI-generated hypotheses in scientific discovery where no ground truth exists, and presents Apodex's multi-agent approach with independent verifier agents as a solution.
This paper proposes formalizing scientific discovery as a meta-optimization problem where LLMs generate and aggregate objective functions via correlation-weighted voting, applied to 3-SAT algorithm discovery using digital MemComputing, achieving a 67x speedup on large instances.
This paper examines Ray Kurzweil's thesis of accelerating returns and argues that while quantitative capabilities may accelerate, genuine scientific discovery requires a different capacity: qualitative reasoning about conceptual frameworks. It proposes the Qualitative Engine for Science (QES) as a response to this gap.
PSI is building a vertically integrated factory for physical superintelligence to accelerate physics breakthroughs with artificial superintelligence, and has open-sourced an AI copilot for physicists called Get Physics Done (GPD).
Introduces NatureBench, a cross-disciplinary benchmark of 90 tasks from Nature papers to test AI coding agents, finding the best agent (Claude Opus 4.7) surpasses SOTA on only 17.8% of tasks and often succeeds by reducing science to supervised ML rather than genuine discovery.
OpenAI's GPT-5 Pro helped immunologist Derya Unutmaz solve a three-year-old mystery about how glucose affects T cell specialization by suggesting that deoxyglucose interferes with IL-2 protein construction, leading to increased inflammatory Th17 cells.
NatureBench is a cross-disciplinary benchmark of 90 scientific tasks from Nature publications, designed to evaluate AI coding agents' ability to achieve genuine discovery. Current agents succeed mainly through methodological translation, not scientific innovation.