Tag
The largest dark matter detector has identified a single anomalous particle, marking a potential breakthrough in the search for dark matter.
Physical Superintelligence (PSI) has raised $58M to build an AI-driven lab for discovering physics breakthroughs, aiming to restart the golden age of physics by using AI to generate and rigorously test hypotheses.
This paper presents AIMC, a visual analytics framework for human oversight of autonomous scientific discovery, enabling monitoring and understanding of AI-generated research artifacts.
The paper explores how language-model agents in SwarmWorld self-organize into technological societies without predefined roles, using stigmergy and cultural mechanisms to develop specialized behaviors and resilient technologies across generations.
InsightSR is a framework that leverages Large Language Models to refine the search space for symbolic regression, improving accuracy and physical consistency through iterative semantic and structural guidance.
The article introduces ActFlow, a continued pre-training scheme that expands the valid design space for flow and diffusion models, enabling out-of-distribution generative modeling and evolvable search spaces in scientific discovery.
The paper presents theStation, an open-world multi-agent environment where AI agents autonomously collaborate on mathematical research, achieving novel results on several open problems and releasing all dialogues, proofs, and code for transparency.
This paper introduces Certification-Driven Reinforcement Learning (CDRL), a framework that leverages symbolic reasoning to generate reusable constraints for improving reinforcement learning in combinatorial search spaces, demonstrated in neutrino flavor model discovery with higher valid model rates and efficiency.
This article examines how historical volcanic eruptions, such as the 1883 Krakatau eruption, have altered global climate and societies, laying the groundwork for modern volcanology.
This paper proposes a logit-based energy scoring method for evaluating scientific hypotheses using large language models, which outperforms prompted LLM-as-judge methods in hypothesis ranking tasks.
dig.bench is a benchmark for evaluating AI models' ability to discover unknown rules in text-based games, measuring scientific discovery capabilities with 70 interactive games and a leaderboard comparing frontier models.
Scientists have demonstrated that by exciting a strontium atom in a Bose-Einstein condensate to form a Rydberg atom, they can contain over 170 other atoms within its orbital, creating a Rydberg polaron.
Recent scientific discoveries, including the identification of 'ghost' ancestors and ancient tools, are radically revising the understanding of human origins through DNA analysis and multidisciplinary research.
The paper introduces MDA, a framework that uses LLMs for hypothesis generation and Bayesian inference for mechanism scoring, significantly reducing experiment needs while improving accuracy on scientific benchmarks like FORCEBENCH.
The paper introduces Large Discovery Model (LDM), a recurrent architecture that couples generative models with Bayesian non-parametric surrogates to guide uncertainty-aware search in scientific domains like molecules and proteins, achieving significant performance gains over existing methods.
TTT-Discover is a framework that trains LLMs at test time using reinforcement learning on individual problems, setting new records in GPU kernel engineering and improving mathematical bounds.
This paper proposes a self-evolving memory framework for lifelong AI partners in materials science, storing scientific experience as facts and skills to improve agent performance across models. Evaluations show memory nearly doubles GPT-5.2 task success in materials tool-use questions and reduces repeated errors in simulations.
MIT Technology Review's daily newsletter covers an op-ed on AI agents for scientific discovery, an investigation into the 'censorship-industrial complex' influencing US policy, and news about an Amazon data center's potential pollution.
This MIT Technology Review article argues that AlphaFold-style deep learning on massive datasets is not the ideal template for accelerating science, and that AI agents capable of reasoning and experimentation will drive future breakthroughs.
This paper introduces SEE, a multimodal benchmark of expert-curated questions for scientific discovery in chemistry, biology, and materials science. Evaluation of 19 MLLMs shows the best model reaches only 48.7% accuracy, and even with tool use only 52.7%, revealing that current models lack reliable evidence-bounded scientific reasoning.