Tag
Anima Anandkumar announces four Lean-related papers from their group at ICML workshops, covering verified ML systems, functional program synthesis, proof assistant interoperability, and scientific reasoning, positioning Lean as infrastructure for AI.
This paper presents SCION, an agentic scientific operating system that integrates AI tools for scientific discovery through a Research Execution Plan (REP) and hierarchical multi-agent execution. It demonstrates applications in materials analysis, molecule design, and protein screening, outperforming existing autonomous research-agent baselines.
Proposes DDIAgents, a mechanism-conditioned multi-agent framework for drug-drug interaction prediction that dynamically routes relevant biomedical knowledge to specialized expert agents and aggregates their analyses, outperforming existing feature-based, graph-based, and LLM-based methods.
This paper introduces SciRisk-Bench, a benchmark for evaluating the safety of large language models in AI4Science contexts, covering 7 disciplines, 31 subdisciplines, and 10 risk dimensions to assess both scientific competence and risk awareness.
LakeFM is a foundation model for aquatic systems, pre-trained on large-scale ecological datasets to forecast lake dynamics using irregular multivariate multi-depth time series data, achieving competitive performance compared to existing models.
Introduces SciPaths, a benchmark for forecasting the enabling contributions required to realize a target scientific discovery, and evaluates frontier and open-weight language models, finding significant room for improvement in reasoning backward from contributions to enabling building blocks.
MeasHalu is a novel framework for mitigating scientific measurement hallucinations in LLMs through a two-stage reasoning-aware fine-tuning strategy and progressive reward curriculum. It introduces a fine-grained taxonomy of measurement-specific hallucinations and demonstrates improved accuracy on the MeasEval benchmark.