scientific-discovery

Tag

Cards List
#scientific-discovery

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

arXiv cs.AI ↗ · 2026-07-13 Cached

This paper introduces the Hypothesis Evolution Protocol (HEP) for LLM agents, which makes hypothesis generation, testing, and belief updates explicit and auditable. Experiments on materials-science tasks show that HEP-equipped agents generalize across research questions and become more effective with stronger base LLMs.

0 favorites 0 likes
#scientific-discovery

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Hugging Face Daily Papers ↗ · 2026-07-13 Cached

Introduces SDABench, a benchmark evaluating LLMs on six scientific analysis capabilities across five domains, finding models struggle with tasks requiring assumption selection and mechanistic reasoning.

0 favorites 0 likes
#scientific-discovery

AI Boosts Research Careers but Flattens Scientific Discovery

Hacker News Top ↗ · 2026-07-12 Cached

An analysis of over 40 million papers finds that AI tools boost individual researchers' productivity and career advancement but narrow the scope of scientific inquiry, leading to less diverse and original discoveries, as published in Nature.

0 favorites 0 likes
#scientific-discovery

@muratcan: The prompt engineering here is super impressive! Such a great example of agent prompting:

X AI KOLs Timeline ↗ · 2026-07-10 Cached

Noam Brown announces that GPT-5.6 Sol Ultra proved a 50-year-old math conjecture, demonstrating impressive prompt engineering and agent prompting.

0 favorites 0 likes
#scientific-discovery

Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

arXiv cs.AI ↗ · 2026-07-10 Cached

Introduces ZendoWorld, a controlled interactive environment for evaluating AI agents on active visual concept induction, where agents must perceive scenes, infer hidden logical rules, and design informative experiments. Experiments with various agent classes reveal that high prediction accuracy does not guarantee rule recovery, and VLM-based agents struggle with informative experimentation, highlighting gaps compared to human inductive reasoning.

0 favorites 0 likes
#scientific-discovery

Statistically Meaningful Geometry and Gauge Symmetry Breaking: A Geometric Foundation for Scientific Discovery and Intelligence Emergence

arXiv cs.LG ↗ · 2026-07-08 Cached

This paper introduces Statistically Meaningful Geometry (SMG), a geometric framework for modeling over-parameterized learning systems as infinite-dimensional non-parametric Orlicz fiber bundles. It proposes that under out-of-distribution stimuli, the system undergoes a gauge symmetry break, leading to the emergence of new causal axes that can distinguish genuine scientific discovery from hallucinations.

0 favorites 0 likes
#scientific-discovery

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

arXiv cs.AI ↗ · 2026-07-08 Cached

FirstResearch introduces a structured framework for LLM scientific discovery agents that generates a Research Question Certificate containing primitive definitions, assumptions, mechanism, falsifiable hypothesis, and failure update rules, making the proposed research question inspectable before execution. Preliminary evaluations using LLM judges show that the certificate-centered approach outperforms baseline systems in audibility and score.

0 favorites 0 likes
#scientific-discovery

Rethinking Scientific Discovery in an Agentic Era

arXiv cs.CL ↗ · 2026-07-07 Cached

This paper presents SCION, an agentic scientific operating system that integrates AI tools for scientific discovery through a Research Execution Plan (REP) and hierarchical multi-agent execution. It demonstrates applications in materials analysis, molecule design, and protein screening, outperforming existing autonomous research-agent baselines.

0 favorites 0 likes
#scientific-discovery

Language models guide symbolic equation discovery by controlling search

arXiv cs.AI ↗ · 2026-07-07 Cached

This paper introduces LLM-PySR, a method where language models guide symbolic equation discovery by controlling search parameters while using numerical symbolic regression for fitting. The approach achieves strong balance of accuracy and complexity across benchmark tasks.

0 favorites 0 likes
#scientific-discovery

Damo Academy unveils an AI agent able to discover superconductors, which could revolutionise scientific materials research and innovation

Reddit r/singularity ↗ · 2026-07-04 Cached

Damo Academy (Alibaba) introduces Elements Claw, an AI agent that discovered four new superconducting materials by screening millions of crystal structures, potentially accelerating materials research.

0 favorites 0 likes
#scientific-discovery

EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation

arXiv cs.AI ↗ · 2026-07-03 Cached

EO-Agents presents a three-agent LLM pipeline for generating Earth observation hypotheses, leveraging a NASA knowledge graph and graph neural network to rank candidate dataset pairings, with LLM agents filtering, generating, and evaluating structured research hypotheses.

0 favorites 0 likes
#scientific-discovery

@ai_suxiaole: Many AI tools can save researchers time, but most only solve one step: reading papers, writing code, polishing papers, or summarizing abstracts. Sakana AI's AI Scientist-v2 aims to be an AI system that runs the full research workflow, from generating research hypotheses to designing experiments...

X AI KOLs Timeline ↗ · 2026-07-02 Cached

Sakana AI has released AI Scientist-v2, an end-to-end automated research system that can autonomously go from generating research hypotheses to writing papers, and has been accepted by the ICLR2025 Workshop after peer review.

0 favorites 0 likes
#scientific-discovery

Evidence-Informed LLM Beliefs for Continual Scientific Discovery

arXiv cs.AI ↗ · 2026-06-30 Cached

This paper addresses the limitation of static surprisal in LLM-based scientific discovery by introducing evidence-informed non-stationary beliefs, and proposes belief-update filtering and diversity maximization to improve discovery, achieving 30.62% higher non-stationary surprisal across five domains.

0 favorites 0 likes
#scientific-discovery

When there is no answer key for scientific discovery how do we verify an ai hypothesis

Reddit r/artificial ↗ · 2026-06-26

Discusses the challenge of verifying AI-generated hypotheses in scientific discovery where no ground truth exists, and presents Apodex's multi-agent approach with independent verifier agents as a solution.

0 favorites 0 likes
#scientific-discovery

Scientific discovery as meta-optimization: a combinatorial optimization case study

arXiv cs.AI ↗ · 2026-06-26 Cached

This paper proposes formalizing scientific discovery as a meta-optimization problem where LLMs generate and aggregate objective functions via correlation-weighted voting, applied to 3-SAT algorithm discovery using digital MemComputing, achieving a 67x speedup on large instances.

0 favorites 0 likes
#scientific-discovery

Accelerating Returns and the Qualitative Engine for Science

arXiv cs.AI ↗ · 2026-06-26 Cached

This paper examines Ray Kurzweil's thesis of accelerating returns and argues that while quantitative capabilities may accelerate, genuine scientific discovery requires a different capacity: qualitative reasoning about conceptual frameworks. It proposes the Qualitative Engine for Science (QES) as a response to this gap.

0 favorites 0 likes
#scientific-discovery

Physical Superintelligence

Reddit r/singularity ↗ · 2026-06-24 Cached

PSI is building a vertically integrated factory for physical superintelligence to accelerate physics breakthroughs with artificial superintelligence, and has open-sourced an AI copilot for physicists called Get Physics Done (GPD).

0 favorites 0 likes
#scientific-discovery

@OkhayIea: Everyone's racing to build "AI scientists." So we asked a blunt question: Can today's best coding agents beat the publi…

X AI KOLs Timeline ↗ · 2026-06-24 Cached

Introduces NatureBench, a cross-disciplinary benchmark of 90 tasks from Nature papers to test AI coding agents, finding the best agent (Claude Opus 4.7) surpasses SOTA on only 17.8% of tasks and often succeeds by reducing science to supervised ML rather than genuine discovery.

0 favorites 0 likes
#scientific-discovery

How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery

OpenAI Blog ↗ · 2026-06-23 Cached

OpenAI's GPT-5 Pro helped immunologist Derya Unutmaz solve a three-year-old mystery about how glucose affects T cell specialization by suggesting that deoxyglucose interferes with IL-2 protein construction, leading to increased inflammatory Th17 cells.

0 favorites 0 likes
#scientific-discovery

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

NatureBench is a cross-disciplinary benchmark of 90 scientific tasks from Nature publications, designed to evaluate AI coding agents' ability to achieve genuine discovery. Current agents succeed mainly through methodological translation, not scientific innovation.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback