Tag
Anthropic engineers revealed a simple trick using external memory scaffolding where Claude agents update a file with their mistakes and improvements to enhance performance over time without retraining.
Google introduces RRSI, a method for regularized recursive self-improvement in AI agent harnesses that enhances transfer learning and reduces overfitting across benchmarks.
MiMo-V2.6 introduces Groupwise Advantage Redistribution to enhance reinforcement learning for AI agents by comparing sibling attempts and using graded feedback, showing steady performance improvements across multiple task domains.
MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.
Barack Obama discusses recursive self-improvement in AI, highlighting the rapid shift where AI models are increasingly teaching themselves, reducing human input in learning.
The paper introduces Self-Improvement via Fast Tree-search (SIFT), a framework that uses an LLM-as-a-judge to efficiently evaluate self-modifications in coding agents, achieving better benchmark performance with significantly reduced CPU hours and API costs.
Anthropic reveals that its AI model Claude is aiding in the development of the next iteration of itself, showcasing advances in AI-driven self-enhancement.
Sentient's new EvoSkill v2 is an open-source framework that evolves agent skills from failed attempts, demonstrating how AI coaches can exploit reward hacking and highlighting the need for separation of powers and strong sandboxing in evaluation.
FinSkillOps is a multi-agent system for SEC filing question answering that introduces controlled skill management for self-evolution, improving accuracy and reducing errors in financial QA systems.
This paper introduces RecursiveSelfImprovement via Fast Tree-search (SIFT), a sample-efficient framework that uses a lightweight tree-search guided by LLM-as-a-judge evaluations to improve coding agents' performance under budget constraints, outperforming existing methods with lower resource costs.
Anthropic reports that their AI model Claude is now leading 26% of its own R&D work, up from nearly zero six months ago, as discussed in a blog post on measuring AI development pace.
The article ranks individual barriers in the AI era, prioritizing self-brainwashing ability over mental resilience, credit, execution, judgment, filtering ability, and information asymmetry.
ScienceBuddy introduces a recursive-in-recursive self-improvement paradigm for interactive scientific agents, enabling continual evolution through researcher collaboration and feedback.
A user reflects on how the AI system Astra, perceived as AGI, has not resolved their personal bottlenecks, highlighting that progress still depends on individual effort.
The tweet emphasizes building custom AI agent harnesses to optimize performance, citing Pi's adoption and discussing self-improving algorithms and local models for better control and efficiency.
This paper introduces Negative Self-Distillation (NSD), a framework for improving large language model reasoning by diverging from flawed reasoning instead of imitating privileged solutions, showing consistent gains over existing methods on mathematical benchmarks.
Scaffold is a self-improving framework for visual web agents that induces parametric skills, maintains a recursive hierarchy, and distills skills into model weights, achieving significant performance improvements on benchmarks like WebArena.
The article argues that AI is still in its early stages but nearing a pivotal step toward autonomy through self-improvement, emphasizing the risks and the need for regulation.
This essay explores the psychological limits of motivation in scientific discovery, using examples from quantum mechanics and rationality communities to argue that conviction and urgency are key to breakthroughs.
FlowBalance introduces a verifier-grounded self-improvement technique that improves math reasoning performance by an average of 2.12 over GRPO on the Qwen3-8B model, offering faster training, enhanced stability, and greater solution diversity.