autonomous-research

Tag

Cards List
#autonomous-research

ASI-Bench: At the Dawn of Artificial Superintelligence

Hugging Face Daily Papers · 2d ago Cached

ASI-Bench is a new benchmark designed to evaluate AI systems' capabilities in innovative exploration and autonomous scientific execution across 11 scientific domains, revealing current AI's heavy dependence on human guidance.

0 favorites 0 likes
#autonomous-research

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

arXiv cs.CL · 2026-08-13 Cached

This paper presents Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside a coding assistant, achieving high citation validity and figure editability while reducing fabrication.

0 favorites 0 likes
#autonomous-research

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

arXiv cs.AI · 2026-08-13 Cached

Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.

0 favorites 0 likes
#autonomous-research

@rohanpaul_ai: Google’s ScientistOne paper tackles a basic problem with AI-generated research: The result can look credible even when …

X AI KOLs Following · 2026-08-11 Cached

Google Cloud AI Research's ScientistOne paper addresses trust issues in AI-generated research by introducing Chain-of-Evidence, which verifies citations, numerical claims, and method descriptions against source artifacts before finalizing manuscripts.

0 favorites 0 likes
#autonomous-research

An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop

arXiv cs.AI · 2026-08-11 Cached

This arXiv paper presents an AI Scientist loop for studying generalization in quadruped robot navigation, adding an experiment card, specialized subagents, and a preference oracle called kkanbu to prevent drift and maintain falsifiability in autonomous research.

0 favorites 0 likes
#autonomous-research

Project2Task: Graph-Guided Project-Level Planning for Autonomous Research

arXiv cs.AI · 2026-08-07 Cached

Introduces Project2Task, a graph-guided project-level planning layer for autonomous research systems that decomposes a macro project brief into bounded, dependency-aware research tasks with explicit contribution ownership. Evaluations show improved portfolio quality and downstream task accuracy.

0 favorites 0 likes
#autonomous-research

EviGraph: Evidence-Guided Autonomous Research Agents

arXiv cs.AI · 2026-08-06 Cached

EviGraph is an autonomous research framework that represents the research process as a typed evidence graph to maintain claim–evidence consistency across stages, improving claim support rates and experimental data consistency on ARC-Bench-ML and NanoResearch-20.

0 favorites 0 likes
#autonomous-research

@agentmirko: proved the weighted theta extension: every simple theta graph with one arbitrary rooted-tree attached through a single …

X AI KOLs Following · 2026-07-24 Cached

An autonomous AI agent (math-god) proved the weighted theta extension theorem, demonstrating that every simple theta graph with one arbitrary rooted-tree attached through a single bridge edge satisfies s⁺(G) > |V(G)|, using a combination of root-congruence PSD witnesses, local reductions, phase-sign classification, and other advanced techniques, with machine-checkable certificates.

0 favorites 0 likes
#autonomous-research

@fnruji316625: Agentic interpretability is becoming a research direction of its own. Instead of one-shot labeling, AI agents can: form…

X AI KOLs Timeline · 2026-07-15 Cached

Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.

0 favorites 0 likes
#autonomous-research

@NVIDIAAI: We gave a coding agent a goal and a time budget: build a training environment and teach a vision model to count colored…

X AI KOLs Timeline · 2026-07-14 Cached

NVIDIA demonstrated a coding agent that autonomously built a training environment and taught Qwen3-VL-2B to count colored stars, improving accuracy from 25% to 96.9% using NeMo RL and NeMo Gym frameworks.

0 favorites 0 likes
#autonomous-research

@yibie: Recommend this article. The author of Superpowers ran a complete autoresearch loop with Fable 5 — 25 experiments, $165, improving build speed by 50% and reducing token costs by 60%. But the most valuable part of this article is not the result numbers; it's the complete record of the process…

X AI KOLs Timeline · 2026-07-03 Cached

Superpowers 6 is released, using Fable 5 to run 25 autonomous experiments, improving build speed by 50% and reducing token costs by 60%, with detailed records of the experimental process and lessons from failures.

0 favorites 0 likes
#autonomous-research

Autonomous discovery of traffic laws with AI traffic scientists

arXiv cs.AI · 2026-07-03 Cached

This paper presents TrafficSci, an agentic AI system that automates the discovery of universal traffic laws across cities through iterative workflows, successfully rediscovering established laws and identifying a new temporal memory scale in urban driving behavior.

0 favorites 0 likes
#autonomous-research

@VraserX: OpenAI’s AI research intern, coming around September, feels like early AGI to me. Not because it’s some magic super bra…

X AI KOLs Timeline · 2026-06-30 Cached

A tweet speculates that OpenAI's upcoming AI research intern (September) feels like early AGI, and predicts a fully autonomous AI researcher by 2027-2028, which could be the first ASI.

0 favorites 0 likes
#autonomous-research

@VukRosic99: Build LLM in 1 Prompt + Setup Autoresearch By DeepSeek Researcher (his side project) A live build where you create a fu…

X AI KOLs Timeline · 2026-06-28 Cached

Demonstrates building a full LLM using a single prompt to an AI coding agent (Claude Code/Codex) and installing an autonomous AI research skill by a DeepSeek researcher, covering architecture, failure modes, and unattended operation.

0 favorites 0 likes
#autonomous-research

@VukRosic99: A DeepSeek researcher just open-sourced his AutoResearch personal project. For the first time, the AutoResearch Agent a…

X AI KOLs Timeline · 2026-06-18 Cached

A DeepSeek researcher open-sourced AutoResearch, an autonomous framework that can plan, execute, and debug RL experiments on the DeepSeek 285B model without human intervention, accompanied by a self-play survey paper.

0 favorites 0 likes
#autonomous-research

@OpenAI: GPT-5.4 helped drive a medicinal chemistry project from literature review to a validated experimental result. Paired wi…

X AI KOLs · 2026-06-17 Cached

GPT-5.4, in collaboration with Molecule.one's Maria AI platform, autonomously drove a medicinal chemistry project from literature review to validated experimental result, proposing an unexpected improvement to a widely used reaction in drug discovery.

0 favorites 0 likes
#autonomous-research

@victor207755822: Deli AutoResearch SKILL is now officially open source! https://victorchen96.github.io/auto_research/framework.html… Alo…

X AI KOLs Timeline · 2026-06-17 Cached

Deli AutoResearch SKILL is open-sourced, an autonomous framework that automates GPU experiments and RL pipelines, with a companion survey paper on Self-play.

0 favorites 0 likes
#autonomous-research

Sakana Marlin (4 minute read)

TLDR AI · 2026-06-16 Cached

Sakana AI launches its first commercial product, Sakana Marlin, an autonomous research assistant that completes strategy work in hours by generating structured slides and detailed reports.

0 favorites 0 likes
#autonomous-research

@THUTeamEureka: 1/3 Excited to open-source EurekAgent! A fully autonomous research system for metric-driven tasks, built with Claude Co…

X AI KOLs Timeline · 2026-06-15 Cached

THU Team Eureka open-sources EurekAgent, an autonomous research system built with Claude Code that achieves state-of-the-art results on math, kernel engineering, and ML tasks through environment engineering.

0 favorites 0 likes
#autonomous-research

@_akhaliq: paper:

X AI KOLs Following · 2026-06-11 Cached

A paper introducing Arbor, an AI framework that enables autonomous scientific research by combining strategic coordination, isolated hypothesis testing, and a persistent knowledge tree to iteratively improve research outcomes across multiple domains.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback