autoresearch

Tag

Cards List
#autoresearch

@_philschmid: The next iterations of agent harnesses will center on coded extensions that integrate automatically. Pi lead, others ar…

X AI KOLs Timeline · 2026-08-19 Cached

A tweet from @_philschmid speculates that future AI agent harnesses will focus on coded extensions for automatic integration, with Pi leading the trend and DeepSeek as an extreme example, all driven by autoresearch and recursive self-improvements.

0 favorites 0 likes
#autoresearch

Scaling Automatic Research Agents via World Models

arXiv cs.LG · 2026-08-14 Cached

This paper identifies a scalability bottleneck in RL-trained automatic research agents—environment execution dominates training cost—and proposes World Model RL (WMRL) with online debiasing and inverse-variance denoising to replace real execution, achieving 3–4x training speedups and better generalization.

0 favorites 0 likes
#autoresearch

Recovering Wasted Compute in Autoresearch Agents

arXiv cs.AI · 2026-08-12 Cached

This paper identifies common failure modes in tree-search-based autoresearch agents applied to tabular datasets, such as repeated bug resolution, poor hyperparameter tuning, and ineffective exploration, and proposes targeted interventions like a global debug consultant and refined tree-search algorithms to recover wasted compute and improve performance without changing the underlying language model.

0 favorites 0 likes
#autoresearch

@iamai_omni: Fable 5 is basically ASI, its self-correction ability is astonishing.

X AI KOLs Timeline · 2026-07-03 Cached

User iamai_omni praises Fable 5's self-correction ability, considering it comparable to ASI. Citing a recommendation from yibie, they point out that the Superpowers author had Fable 5 run an autoresearch loop, spending $165 to complete 25 experiments, increasing build speed by 50% and reducing token overhead by 60%, and documented failures and correction processes in detail.

0 favorites 0 likes
#autoresearch

Autoresearch, Claude, and Constrained Optimization (13 minute read)

TLDR AI · 2026-07-03 Cached

The article describes an experiment using Claude Code to autonomously develop a file compression algorithm with constrained optimization, evaluating the viability of AI agents for unsupervised problem-solving.

0 favorites 0 likes
#autoresearch

Autoresearch: The feedback loop behind self-improving agents (11 minute read)

TLDR AI · 2026-07-02 Cached

Introspection, a new AI startup founded by ex-xAI engineers, introduces 'autoresearch' – a feedback loop system where agents maintain and improve themselves using signals, evals, and human input, moving beyond traditional agent harnesses.

0 favorites 0 likes
#autoresearch

@lillian_ma_: Emerging autoresearch labs worth following: @AutoScienceAI (@eliot_cowan) One of the cleanest “AI builds AI” bets: agen…

X AI KOLs Timeline · 2026-06-22 Cached

A Twitter thread highlights emerging autoresearch labs that are building AI systems to automate the full research loop, from hypothesis to experimentation.

0 favorites 0 likes
#autoresearch

@yibie: awesome-autoresearch periodic roundup - 2 new entries added: 1. Cribl AI Production Deployment Autoresearch (discussions): Cribl deploys Karpathy-style experimental loop to production environment; Agent autonomously…

X AI KOLs Timeline · 2026-06-17 Cached

awesome-autoresearch is a curated list of automated research use cases. This update adds two entries: Cribl's production deployment and Colab TPU port.

0 favorites 0 likes
#autoresearch

@FinanceYF5: Jim Fan announces that the team has brought AutoResearch into the physical world for the first time, launching ENPIRE. They equipped 8 Codex agents with robots, GPUs, and sufficient tokens, allowing them to autonomously learn, trial and error, collaborate, and complete tasks on real hardware.

X AI KOLs Following · 2026-06-17 Cached

Jim Fan announced that the team launched ENPIRE, bringing AutoResearch into the physical world for the first time, equipping 8 Codex agents with robots, GPUs, and tokens, enabling them to autonomously learn and collaborate on real hardware to complete tasks.

0 favorites 0 likes
#autoresearch

@yibie: awesome-autoresearch periodic review, 1 new entry (discussions): Karpathy's Autoresearch for Software Engineers: applying the three-file model of autoresearch (...

X AI KOLs Timeline · 2026-06-16 Cached

awesome-autoresearch is a curated list of autoresearch use cases. This new entry applies Karpathy's autoresearch pattern to everyday software engineering scenarios, providing concrete templates to lower the adoption barrier for engineers.

0 favorites 0 likes
#autoresearch

@zhengyaojiang: We benchmarked 7 frontier models on 3 categories of autoresearch tasks: ML engineering, harness/prompt engineering, and…

X AI KOLs Following · 2026-06-14 Cached

Researchers benchmarked 7 frontier models on autoresearch tasks. Fable-5 won overall, but the open model Kimi-K2.7-Code surpassed others on ML engineering tasks.

0 favorites 0 likes
#autoresearch

Self-Evolving Autoresearch Workflow Loops (5 minute read)

TLDR AI · 2026-06-10 Cached

Evo ported its autoresearch loop onto Anthropic's dynamic workflows in Claude Code, moving orchestration from in-context memory to deterministic JavaScript, solving long-horizon instruction adherence and enabling self-evolving workflows.

0 favorites 0 likes
#autoresearch

@posthog: https://x.com/posthog/status/2062595534381326421

X AI KOLs Timeline · 2026-06-04 Cached

PostHog used an AI agent based on Karpathy's autoresearch to find a three-year-old bug in their ClickHouse query engine that prevented proper primary key usage for timestamp filters. Fixing it improved performance by 11% and reduced scanned granules by 62%.

0 favorites 0 likes
#autoresearch

@yibie: 3 new additions this round: 1. Auto-Quant: Apply Karpathy autoresearch to FreqTrade cryptocurrency strategy backtesting, evolving multi-strategy combinations across 5 trading pairs through a keep/discard loop. 2. Universal…

X AI KOLs Timeline · 2026-06-02 Cached

yibie shared three new entries from the awesome-autoresearch list, covering automated quantitative trading, universal skill optimization, and a Claude Code plugin.

0 favorites 0 likes
#autoresearch

@yibie: This week awesome-autoresearch added 3 items: 1. autoslam: Applying Karpathy's autoresearch loop to LiDAR SLAM method design, accumulating experimental leaderboard on KITTI benchmark 2. Bir…

X AI KOLs Timeline · 2026-05-30 Cached

This week awesome-autoresearch added three items, including the autoslam project that applies Karpathy's autoresearch loop to LiDAR SLAM, and two blog posts analyzing the original experiments and revealing metric gaming.

0 favorites 0 likes
#autoresearch

Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas

Hugging Face Daily Papers · 2026-05-28 Cached

This paper presents a two-level autoresearch framework where an outer-loop AI agent autonomously optimizes inner-loop LLM policy-synthesis pipelines for multi-agent sequential social dilemmas, achieving superior performance and discovering objective-specific mechanisms like fairness under a maximin welfare objective.

0 favorites 0 likes
#autoresearch

@benhylak: this saturday, @OpenAI is throwing an autoresearch hackathon with @raindrop_ai and @modal come learn how to build syste…

X AI KOLs Timeline · 2026-05-27 Cached

OpenAI is co-hosting an autoresearch hackathon this Saturday with Raindrop AI and Modal, focusing on building self-improving systems like agents and models.

0 favorites 0 likes
#autoresearch

@yibie: This week's autoresearch ecosystem evidence scan: 9 new records, total count 383. AutoResearch-RL: A continuous RL research framework with http://prepare.py/train.py isolation, supporting LLM/hybrid strategy experiment scheduling l…

X AI KOLs Timeline · 2026-05-26 Cached

This week, 9 new records were added to the autoresearch ecosystem, bringing the total to 383, covering multiple open-source tools and projects such as the AutoResearch-RL reinforcement learning framework, lance-autoresearch database kernel optimization, and Clio prediction market backtesting framework.

0 favorites 0 likes
#autoresearch

@yibie: awesome-autoresearch updated, added 3 entries. dreamworld — world model research. Applied the autoresearch loop to pixel-level world model training (CarRacing-v3), where the agent can perform keep/discard experiments autonomously in tokenizer, dy…

X AI KOLs Timeline · 2026-05-24 Cached

awesome-autoresearch updated, added dreamworld (world model research), Odyssey Engine (general iterative engine), and an article by Kirill Krainov on self-improvement of agentic coding.

0 favorites 0 likes
#autoresearch

@WWTLitee: Is there a way for AI to autonomously iterate and optimize? Yes, check out autoresearch. Its core isn't to have AI directly 'invent papers,' but to break the research process into a verifiable loop: humans write program.md to give research direction, AI agent modifies http://tra…

X AI KOLs Timeline · 2026-05-23 Cached

Introduces the autoresearch project, which breaks down the AI research process into a verifiable loop (fixed environment, single editable file, fixed metric, Git rollback), enabling AI agents to perform controllable and reproducible experiment iterations; also mentions the 12-factor-agents checklist.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback