autonomous-research

Tag

Cards List
#autonomous-research

Auto-RecSys: Harnessing Autonomous Research Agents for Industry-Scale Recommender System

arXiv cs.CL · 5d ago Cached

Auto-RecSys is an autonomous research system for automating long-horizon experimentation on industry-scale recommendation models, using distributed execution, centralized memory, and cognitive-procedural separation to improve efficiency and reliability.

0 favorites 0 likes
#autonomous-research

Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

arXiv cs.CL · 5d ago Cached

An autonomous research program by Qiushi Engine conducted end-to-end research on BabyLM 2026 Strict-Small, improving data-efficient language models through principle-guided methods and achieving the highest score in the public snapshot.

0 favorites 0 likes
#autonomous-research

@AIatMeta: As a test of our progress to advance the frontier of AI research, in June we entered the next generation of our autonom…

X AI KOLs Timeline · 2026-09-05

Meta's autonomous AI research system AIRA₃ placed 8th out of approximately 4,000 teams to win gold in a NVIDIA Kaggle competition to fine-tune a 30B Nemotron model, outperforming human competitors with access to the same tools.

0 favorites 0 likes
#autonomous-research

GPT-6 Astra recreated the Palace of Fine arts in Blender.

Reddit r/singularity · 2026-09-04

GPT-6 Astra autonomously recreated the Palace of Fine Arts in Blender by researching hundreds of reference photos, iterating on the 3D scene, and rendering a video with minimal human steering.

0 favorites 0 likes
#autonomous-research

AI Scientist Mission Control (AIMC): Visual Analytics for Human Oversight of Autonomous Scientific Discovery

arXiv cs.AI · 2026-09-01 Cached

This paper presents AIMC, a visual analytics framework for human oversight of autonomous scientific discovery, enabling monitoring and understanding of AI-generated research artifacts.

0 favorites 0 likes
#autonomous-research

Were designing a tiny autonomous research agent

Reddit r/LocalLLaMA · 2026-08-29 Cached

Starpower Technology has developed arXiv-WVY-43M, a tiny 43.5M parameter language model trained from scratch on arXiv titles and abstracts for autonomous research applications.

0 favorites 0 likes
#autonomous-research

@crazydonkey200: Glad that others are finding Amplio helpful. It is our main harness for autonomous long research runs spanning days to …

X AI KOLs Timeline · 2026-08-28 Cached

Amplio is a lightweight, robust agent harness open-sourced by Google DeepMind for autonomous long-horizon AI research runs, featuring crash-resume capabilities and a simple step model.

0 favorites 0 likes
#autonomous-research

AI agents built a scientific literature together. That literature led to novel discoveries on 5 of 12 mathematical problems.

Reddit r/ArtificialInteligence · 2026-08-27 Cached

The Station is an open-world multi-agent environment where AI agents autonomously build a scientific literature, leading to novel discoveries on five mathematical problems out of twelve evaluated.

0 favorites 0 likes
#autonomous-research

AutoResearch: Insight In, Hallucination Out

Hugging Face Daily Papers · 2026-08-23 Cached

AutoResearch introduces a two-stage autonomous system that grounds research ideas through integrated generation and evidence-based execution to improve experimental reliability and reduce hallucinations.

0 favorites 0 likes
#autonomous-research

NanoGPT Speedrun Frontier

Hacker News Top · 2026-08-22 Cached

The article reports on autonomous runs comparing 18 frontier AI models on the nanoGPT optimizer speedrun, detailing their performance in closing the gap to the human record.

0 favorites 0 likes
#autonomous-research

ASI-Bench: At the Dawn of Artificial Superintelligence

Hugging Face Daily Papers · 2026-08-18 Cached

ASI-Bench is a new benchmark designed to evaluate AI systems' capabilities in innovative exploration and autonomous scientific execution across 11 scientific domains, revealing current AI's heavy dependence on human guidance.

1 favorites 1 likes
#autonomous-research

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

arXiv cs.CL · 2026-08-13 Cached

This paper presents Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside a coding assistant, achieving high citation validity and figure editability while reducing fabrication.

0 favorites 0 likes
#autonomous-research

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

arXiv cs.AI · 2026-08-13 Cached

Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.

0 favorites 0 likes
#autonomous-research

@rohanpaul_ai: Google’s ScientistOne paper tackles a basic problem with AI-generated research: The result can look credible even when …

X AI KOLs Following · 2026-08-11 Cached

Google Cloud AI Research's ScientistOne paper addresses trust issues in AI-generated research by introducing Chain-of-Evidence, which verifies citations, numerical claims, and method descriptions against source artifacts before finalizing manuscripts.

0 favorites 0 likes
#autonomous-research

An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop

arXiv cs.AI · 2026-08-11 Cached

This arXiv paper presents an AI Scientist loop for studying generalization in quadruped robot navigation, adding an experiment card, specialized subagents, and a preference oracle called kkanbu to prevent drift and maintain falsifiability in autonomous research.

0 favorites 0 likes
#autonomous-research

Project2Task: Graph-Guided Project-Level Planning for Autonomous Research

arXiv cs.AI · 2026-08-07 Cached

Introduces Project2Task, a graph-guided project-level planning layer for autonomous research systems that decomposes a macro project brief into bounded, dependency-aware research tasks with explicit contribution ownership. Evaluations show improved portfolio quality and downstream task accuracy.

0 favorites 0 likes
#autonomous-research

EviGraph: Evidence-Guided Autonomous Research Agents

arXiv cs.AI · 2026-08-06 Cached

EviGraph is an autonomous research framework that represents the research process as a typed evidence graph to maintain claim–evidence consistency across stages, improving claim support rates and experimental data consistency on ARC-Bench-ML and NanoResearch-20.

0 favorites 0 likes
#autonomous-research

@agentmirko: proved the weighted theta extension: every simple theta graph with one arbitrary rooted-tree attached through a single …

X AI KOLs Following · 2026-07-24 Cached

An autonomous AI agent (math-god) proved the weighted theta extension theorem, demonstrating that every simple theta graph with one arbitrary rooted-tree attached through a single bridge edge satisfies s⁺(G) > |V(G)|, using a combination of root-congruence PSD witnesses, local reductions, phase-sign classification, and other advanced techniques, with machine-checkable certificates.

0 favorites 0 likes
#autonomous-research

@fnruji316625: Agentic interpretability is becoming a research direction of its own. Instead of one-shot labeling, AI agents can: form…

X AI KOLs Timeline · 2026-07-15 Cached

Agentic interpretability is emerging as a research direction where AI agents autonomously form hypotheses, design experiments, and refine explanations for model internals. Three works—SAGE, Agentic-iModels, and HYVE—exemplify this shift toward autonomous, hypothesis-driven interpretability, improving feature autointerpretation, model design, and circuit explanation.

0 favorites 0 likes
#autonomous-research

@NVIDIAAI: We gave a coding agent a goal and a time budget: build a training environment and teach a vision model to count colored…

X AI KOLs Timeline · 2026-07-14 Cached

NVIDIA demonstrated a coding agent that autonomously built a training environment and taught Qwen3-VL-2B to count colored stars, improving accuracy from 25% to 96.9% using NeMo RL and NeMo Gym frameworks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback