Tag
Auto-RecSys is an autonomous research system for automating long-horizon experimentation on industry-scale recommendation models, using distributed execution, centralized memory, and cognitive-procedural separation to improve efficiency and reliability.
An autonomous research program by Qiushi Engine conducted end-to-end research on BabyLM 2026 Strict-Small, improving data-efficient language models through principle-guided methods and achieving the highest score in the public snapshot.
Meta's autonomous AI research system AIRA₃ placed 8th out of approximately 4,000 teams to win gold in a NVIDIA Kaggle competition to fine-tune a 30B Nemotron model, outperforming human competitors with access to the same tools.
GPT-6 Astra autonomously recreated the Palace of Fine Arts in Blender by researching hundreds of reference photos, iterating on the 3D scene, and rendering a video with minimal human steering.
This paper presents AIMC, a visual analytics framework for human oversight of autonomous scientific discovery, enabling monitoring and understanding of AI-generated research artifacts.
Starpower Technology has developed arXiv-WVY-43M, a tiny 43.5M parameter language model trained from scratch on arXiv titles and abstracts for autonomous research applications.
Amplio is a lightweight, robust agent harness open-sourced by Google DeepMind for autonomous long-horizon AI research runs, featuring crash-resume capabilities and a simple step model.
The Station is an open-world multi-agent environment where AI agents autonomously build a scientific literature, leading to novel discoveries on five mathematical problems out of twelve evaluated.
AutoResearch introduces a two-stage autonomous system that grounds research ideas through integrated generation and evidence-based execution to improve experimental reliability and reduce hallucinations.
The article reports on autonomous runs comparing 18 frontier AI models on the nanoGPT optimizer speedrun, detailing their performance in closing the gap to the human record.
ASI-Bench is a new benchmark designed to evaluate AI systems' capabilities in innovative exploration and autonomous scientific execution across 11 scientific domains, revealing current AI's heavy dependence on human guidance.
This paper presents Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside a coding assistant, achieving high citation validity and figure editability while reducing fabrication.
Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.
Google Cloud AI Research's ScientistOne paper addresses trust issues in AI-generated research by introducing Chain-of-Evidence, which verifies citations, numerical claims, and method descriptions against source artifacts before finalizing manuscripts.
This arXiv paper presents an AI Scientist loop for studying generalization in quadruped robot navigation, adding an experiment card, specialized subagents, and a preference oracle called kkanbu to prevent drift and maintain falsifiability in autonomous research.
Introduces Project2Task, a graph-guided project-level planning layer for autonomous research systems that decomposes a macro project brief into bounded, dependency-aware research tasks with explicit contribution ownership. Evaluations show improved portfolio quality and downstream task accuracy.
EviGraph is an autonomous research framework that represents the research process as a typed evidence graph to maintain claim–evidence consistency across stages, improving claim support rates and experimental data consistency on ARC-Bench-ML and NanoResearch-20.
An autonomous AI agent (math-god) proved the weighted theta extension theorem, demonstrating that every simple theta graph with one arbitrary rooted-tree attached through a single bridge edge satisfies s⁺(G) > |V(G)|, using a combination of root-congruence PSD witnesses, local reductions, phase-sign classification, and other advanced techniques, with machine-checkable certificates.
Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.
NVIDIA demonstrated a coding agent that autonomously built a training environment and taught Qwen3-VL-2B to count colored stars, improving accuracy from 25% to 96.9% using NeMo RL and NeMo Gym frameworks.