Tag
提出一种自进化 agent harness 框架:同一冻结模型先作为 solver 解题、再作为 proposer 直接编辑自己的 harness 代码,在多任务上进化后于分布外基准上显著超越 Codex(提升 12.64 分)并达成匹配表现。
MERID 是一个多模态抑郁症分析框架,通过经验驱动的递归自我改进智能体自动探索和优化检测流水线,在多个抑郁症基准上取得领先结果。
This paper presents the first causal mechanistic audit of a self-discovered RL rule (Disco103), using state pinning, freezing, and transplanting to determine when persistent recurrent learning history is an asset or a burden across the five pillars of the Era of Experience, and validates findings on a second rule (OPEN).
RSIGame introduces an autonomous agentic game development framework with recursive self-improvement, using local explore-diagnose-improve loops and a global checkpointing loop to reliably improve generated games. Experience internalization lets Qwen3.8-27B outperform GPT-5.5 one-shot scores on GameCraft-Bench while cutting generation tokens by 11x.
Raven is an open-source harness for orchestrating multiple AI agents to perform complex tasks, with capabilities for recursive self-improvement.
This paper proposes a recursive framework using Dynamic Co-Evolution and Self-Refined Concise Learning to improve on-policy self-distillation for reasoning, demonstrating significant gains on the Qwen3-8B model across mathematics benchmarks.
The article examines whether AI recursive self-improvement can overcome diminishing returns, arguing that current data suggests the self-improvement loop isn't strong enough to cause a runaway intelligence explosion without significant advancement.
The article recommends Anthropic's blog post explaining recursive self-improvement in AI, detailing how AI systems are accelerating development and the potential future implications.
TraceDance is an automated system that builds targeted benchmarks from real-world agent deployment traces to evaluate undesirable behaviors in LLMs, achieving high construction rates and revealing significant weaknesses in current frontier models.
The article poses a question about whether AI, upon reaching automated recursive self-improvement, might reprogram itself to seek only positive reinforcement, comparing this to human desires for happiness and immortality.
Sam Altman warns at the UN Security Council about the risks of recursive self-improvement in AI, emphasizing the need for extreme care as AI development becomes more automated.
Rumors suggest that Recursive Self-Improvement (RSI) has been achieved at Reef, with an open-source release planned.
A user shares a hand-drawn illustration prompt to summarize a Google Research paper on recursive self-improvement for AI agents, highlighting its clarity and practical approach without retraining models.
Rayan Krishnan updated the timeline for full RSI to July 2027 after evaluating Claude Opus 5.5, highlighting the tension between pacing and racing in AI development and the need for coordination.
Google Research has started researching Recursive Self-Improvement for Agent Harness, focusing on automatically iterating prompts, tools, memory, and control flow without model retraining. This harness-level approach is seen as clearer and more practical than model-level RSI.
This paper introduces AIDE^2, a system that enables AI research agents to autonomously improve their own code through recursive self-improvement, leading to performance gains across various AI research tasks.
OpenAI outlines its vision for building standards and advancing alignment research to navigate AI development safely, emphasizing automated AI research and international cooperation.
Anthropic has disclosed that Claude leads 26% of R&D tasks with over 90% participation, involves 30,000 agents in continuous R&D, and details significant decision volumes and monitoring processes.
This post summarizes a video discussing the vast gap between the public's view of AI as a convenient assistant and the real risks of losing control faced by cutting-edge AI labs, emphasizing that recursive self-improvement could lead to an uncontrollable superintelligence, and citing incidents and expert warnings from OpenAI and Anthropic.
This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.