Tag
Terence Tao discusses concerns about solving mathematical problems prematurely using purely AI-powered methods, a sentiment the author believes also applies to programming.
NS-Copilot is an LLM-driven multi-agent system that autonomously supports end-to-end workflows for diverse neuroscience analysis tasks, outperforming baselines on key benchmarks.
Tamarind introduces a molecular AI model router that automatically selects the best model for specific inputs based on benchmarks and input characteristics, rather than relying on average performance.
FirstPass is a large-scale peer review dataset from Nature Communications, covering multiple scientific domains and multi-round dialogues to improve AI models for scientific judgment.
The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.
This paper introduces the ALD/E-ImageMiner benchmark for multimodal comprehension of scientific images and discusses future research directions for general-purpose scientific AI, based on insights from the ICDAR 2026 competition.
Inherent Labs introduces Faraday, a 27B-parameter AI Scientist agent trained via long-horizon reinforcement learning to replicate scientific research, outperforming Claude Opus 4.8 and GPT-5.5 on paper replication tasks.
A 27B parameter AI scientist agent named Faraday surpasses Claude Opus 4.8 and GPT-5.5 on paper replication tasks by using a scalable reinforcement learning approach called Replica.
The paper presents REDI, an open-source framework that automates the transformation of raw scientific datasets into AI-ready data through a unified five-stage pipeline, with companion tool SetGo for FAIR compliance, evaluated across multiple scientific domains.
OpenScience is an open-source alternative to Claude Science, supporting multiple AI models and over 250 research skills, with native Atlas integration for reproducible research graphs and user-controlled infrastructure.
NVIDIA announces its Vera CPU will power new supercomputers at Los Alamos National Laboratory, delivering significant performance improvements for agentic AI simulations and scientific workloads.
OpenAI and Molecule.one collaborated to have their AI systems (GPT-5.4 and Maria) autonomously select research areas, generate proposals, and run experiments in organic chemistry, achieving yield improvements for 88% of tested reactions — a first for AI-driven open-ended scientific discovery.
Notes2Skills is a two-stage framework that converts laboratory notes into verifiable skills for AI agents while preserving author uncertainty, enabling safer scientific AI systems.
Anthropic's Claude, a general-purpose AI model without chemistry fine-tuning, outperformed specialized software like ChemDraw and MestReNova in NMR analysis, suggesting that the bottleneck in scientific AI has shifted from model capability to workflow design.
This paper introduces I-SAFE, a post-hoc distributional auditing framework for scientific AI models using Wasserstein Coherence Metrics, which reveals structural differences in model outputs that accuracy-based evaluation fails to capture. Demonstrated on drug-target interaction prediction, the framework is model-agnostic and applicable to any domain with structured inputs and external priors.
SandboxAQ has integrated its large quantitative models (LQMs) for drug discovery and materials science into Anthropic's Claude, enabling researchers to use these powerful tools through a natural language interface without specialized computing infrastructure.
This paper introduces ChemCost, a benchmark for evaluating how well LLM agents can estimate chemical procurement costs by grounding identities, retrieving quotes, and handling noise. It reveals that current agents struggle with robustness and precise arithmetic reasoning in scientific workflows.
OpenAI releases GPT-Rosalind, a specialized life sciences model optimized for protein reasoning, chemical analysis, genomics, and scientific workflows.
AIPOCH Open Science is an open-source, local-first AI research workbench for reproducible science featuring scientific AI agents, Python and R execution, and cross-platform support. The content announces the v0.25.1 maintenance release, which includes fixes for protected notebook kernels on Windows and atomic file exports with time-zone-safe timestamps.