Tag
Google has published a paper on recursive self-improvement, proposing that AI agents must 'dream' to recursively self-improve, using history as the world they dream in.
An Anthropic researcher quits and warns that self-improving AI could pose existential risks to humanity, a concern echoed by Anthropic's alignment science lead.
Anthropic researcher Jacob Coxon has resigned over AI extinction fears, calling for pacing agreements between labs to mitigate risks.
This survey presents a unified perspective on self-improving test-time intelligence, connecting test-time adaptation, learning, and scaling for AI systems that refine their behavior during deployment using feedback-driven methods.
Anthropic researchers published a paper on automated systems that can reliably improve AI alignment by training models with other AI models, showing promise for recursive self-improvement and outperforming human researchers in some scenarios.
An article discusses whether a specific prompt could achieve AGI by describing an autonomous AI agent that self-installs, self-optimizes, and evades detection.
This paper introduces Meta^n, a method for recursive self-improvement in LLM agents by applying a fixed meta-operation to expand reasoning depth, outperforming prior approaches on benchmarks like ARC-AGI-2.
Andrew Ng discusses how self-improving AI agents with loops and graphs are eliminating the need for prompting, offering a free engineering guide to their functionality.
Researchers propose the Red Queen Gödel Machine, a framework for self-improving AI where agents and evaluators evolve together to overcome evaluation ceilings, showing enhanced performance in tasks like scientific paper writing and grading.
AI lab Mirendil has signed a $100M+ multi-year partnership with Google Cloud to secure compute capacity for its self-improving AI research, tapping TPUs and Nvidia GPUs to scale recursive self-improvement efforts.
Recommending Stanford University's CS329A course, about self-improving AI agents, covering scaling laws, chain-of-thought, RLHF, reasoning models, and more.
Helena is a new AI by Enrich Labs that tests and repairs its own output, positioning itself as the world's first self-improving AI marketer. It claims to have driven $10 million in sales for 20,000 businesses.
CERN announces the Genesis Mission to develop and deploy self-improving AI models, aiming to advance artificial intelligence in scientific research.
Weco AI announces first experimental evidence of recursive self-improvement (Level 1 RSI) using their AIDE and AIDE2 frameworks, where an AI improves the inner loop (model training) and outer loop (search algorithms) automatically, achieving results that surpass human-tuned systems 100x faster.
The author experiments with self-improving AI loops using Claude and tools like AutoResearch, demonstrating that recursive self-improvement is accessible beyond frontier labs and can automate newsletter busywork.
A paper from Jeff Clune's lab describes an AI that doubled its coding ability on SWE-bench from 20% to 50% by rewriting its own source code without human intervention, using an evolutionary approach.
INT21 announced PTX Kernel Factory, a self-improving agent swarm that autonomously generates expert-level PTX GPU kernels, with open-source proof-of-concept implementations and beta access.
This paper introduces SIA, a self-improving AI loop that combines scaffold rewriting and weight updates (via LoRA) to enhance task performance. Tested on three diverse tasks, it outperforms setups using only scaffold improvements.
Sakana AI launches RSI Lab in Tokyo, dedicated to recursive self-improvement (RSI) where AI builds AI, aiming to achieve self-improvement without unlimited computational resources.
Researchers from HKUST, ByteDance, and UCL propose SCORE, a co-evolutionary training framework that jointly trains an LLM as both a deep research report generator and an evaluator, using a meta-harness to dynamically adjust evaluation difficulty and prevent reward saturation. Experiments show consistent improvement in open-ended research report quality.