Tag
This paper presents AIMC, a visual analytics framework for human oversight of autonomous scientific discovery, enabling monitoring and understanding of AI-generated research artifacts.
This podcast episode explores the concept of AI scientists, which combine large language models with automated laboratories to design and conduct experiments. Ant Rowstron from ARIA discusses the emerging architecture and ongoing research in the field.
Jeff Dean, former Google Chief Scientist, tearfully discusses the decision to leave Google after 27 years and start an independent startup, expressing deep feelings for his time at Google.
This article recommends paying attention to AI Scientist and discusses how research agents can learn from failures by analogizing to fuzz testing, thereby mapping the unknown and guiding subsequent experiments.
OmniScientist is an end-to-end omni-modal AI scientist that performs multidisciplinary research directly from heterogeneous raw evidence using autonomous agents and lifecycle-wide perception. Evaluated on 36 real-data cases, it improves evidence-grounded discovery across diverse scientific modalities.
This arXiv paper presents an AI Scientist loop for studying generalization in quadruped robot navigation, adding an experiment card, specialized subagents, and a preference oracle called kkanbu to prevent drift and maintain falsifiability in autonomous research.
This paper proposes a benchmarking protocol using automated multi-model LLM review to evaluate AI Scientist systems, comparing frameworks like Sakana AI, CycleResearcher, and Data-to-Paper, and finds that FARS benchmark papers significantly outperform other systems.
Sakana AI has released AI Scientist-v2, an end-to-end automated research system that can autonomously go from generating research hypotheses to writing papers, and has been accepted by the ICLR2025 Workshop after peer review.
This paper introduces Xcientist, a research harness that externalizes AI-driven scientific research synthesis and validation into inspectable, contract-governed processes to ensure accountability and traceability.
Researchers at MIT present a paper on self-evolving AI scientists that can discover and adapt their own scientific vocabulary, using a categorical framework to mathematically quantify genuine novelty and separate discovery from mere search or retrieval.
Harvard University's AutoScientists proposes a decentralized multi-agent team approach, allowing multiple agents to share experimental status, automatically form teams, and review research plans, significantly outperforming existing methods on multiple benchmarks.
A new paper from Meta, Stanford, and Google introduces AutoResearchClaw, which improves automated research by integrating failure recovery, debate, and selective human input. It outperforms AI Scientist v2 by 54.7% on ARC-Bench and reveals that autonomy is enhanced when constrained by process rather than given unlimited freedom.
A study evaluates frontier models' ability to forecast scientific progress across 4,760 events, finding they can identify plausible directions but cannot reliably predict outcomes or timelines, with systematic overconfidence.
A comprehensive open-source collection of 138 scientific agent skills that transform AI coding assistants like Claude Code and Codex into AI scientists, covering biology, chemistry, medicine, and more, with integration of over 100 scientific databases and specialized Python packages.
This paper presents AI CFD Scientist, an open-source AI agent for computational fluid dynamics that autonomously discovers physics corrections using vision-language verification and code modification, outperforming general AI scientists on CFD tasks.
EvoScientist is an adaptive multi-agent framework for end-to-end scientific discovery that continuously improves through persistent memory modules, comprising three specialized agents for idea generation, experiment execution, and knowledge distillation. It outperforms 7 state-of-the-art systems in scientific idea generation and improves code execution success rates through multi-agent evolution.