Tag
ASI-Bench is a new benchmark designed to evaluate AI systems' capabilities in innovative exploration and autonomous scientific execution across 11 scientific domains, revealing current AI's heavy dependence on human guidance.
This paper presents Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside a coding assistant, achieving high citation validity and figure editability while reducing fabrication.
Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.
Google Cloud AI Research's ScientistOne paper addresses trust issues in AI-generated research by introducing Chain-of-Evidence, which verifies citations, numerical claims, and method descriptions against source artifacts before finalizing manuscripts.
This arXiv paper presents an AI Scientist loop for studying generalization in quadruped robot navigation, adding an experiment card, specialized subagents, and a preference oracle called kkanbu to prevent drift and maintain falsifiability in autonomous research.
Introduces Project2Task, a graph-guided project-level planning layer for autonomous research systems that decomposes a macro project brief into bounded, dependency-aware research tasks with explicit contribution ownership. Evaluations show improved portfolio quality and downstream task accuracy.
EviGraph is an autonomous research framework that represents the research process as a typed evidence graph to maintain claim–evidence consistency across stages, improving claim support rates and experimental data consistency on ARC-Bench-ML and NanoResearch-20.
An autonomous AI agent (math-god) proved the weighted theta extension theorem, demonstrating that every simple theta graph with one arbitrary rooted-tree attached through a single bridge edge satisfies s⁺(G) > |V(G)|, using a combination of root-congruence PSD witnesses, local reductions, phase-sign classification, and other advanced techniques, with machine-checkable certificates.
Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.
NVIDIA demonstrated a coding agent that autonomously built a training environment and taught Qwen3-VL-2B to count colored stars, improving accuracy from 25% to 96.9% using NeMo RL and NeMo Gym frameworks.
Superpowers 6 is released, using Fable 5 to run 25 autonomous experiments, improving build speed by 50% and reducing token costs by 60%, with detailed records of the experimental process and lessons from failures.
This paper presents TrafficSci, an agentic AI system that automates the discovery of universal traffic laws across cities through iterative workflows, successfully rediscovering established laws and identifying a new temporal memory scale in urban driving behavior.
A tweet speculates that OpenAI's upcoming AI research intern (September) feels like early AGI, and predicts a fully autonomous AI researcher by 2027-2028, which could be the first ASI.
Demonstrates building a full LLM using a single prompt to an AI coding agent (Claude Code/Codex) and installing an autonomous AI research skill by a DeepSeek researcher, covering architecture, failure modes, and unattended operation.
A DeepSeek researcher open-sourced AutoResearch, an autonomous framework that can plan, execute, and debug RL experiments on the DeepSeek 285B model without human intervention, accompanied by a self-play survey paper.
GPT-5.4, in collaboration with Molecule.one's Maria AI platform, autonomously drove a medicinal chemistry project from literature review to validated experimental result, proposing an unexpected improvement to a widely used reaction in drug discovery.
Deli AutoResearch SKILL is open-sourced, an autonomous framework that automates GPU experiments and RL pipelines, with a companion survey paper on Self-play.
Sakana AI launches its first commercial product, Sakana Marlin, an autonomous research assistant that completes strategy work in hours by generating structured slides and detailed reports.
THU Team Eureka open-sources EurekAgent, an autonomous research system built with Claude Code that achieves state-of-the-art results on math, kernel engineering, and ML tasks through environment engineering.
A paper introducing Arbor, an AI framework that enables autonomous scientific research by combining strategic coordination, isolated hypothesis testing, and a persistent knowledge tree to iteratively improve research outcomes across multiple domains.