Tag
This essay explores the psychological limits of motivation in scientific discovery, using examples from quantum mechanics and rationality communities to argue that conviction and urgency are key to breakthroughs.
FlowBalance introduces a verifier-grounded self-improvement technique that improves math reasoning performance by an average of 2.12 over GRPO on the Qwen3-8B model, offering faster training, enhanced stability, and greater solution diversity.
This article explores the importance of the Harness (framework) in AI, demonstrates how improving the Harness can significantly enhance model performance, and introduces cutting-edge exploration of self-improving Harnesses.
FlowBalance is a verifier-grounded self-improvement method that calibrates on-policy reasoning experience to enhance AI model performance and stability in mathematical reasoning tasks.
Meta has built an AI agent that acts as a secondary expert in domains, using a knowledge architecture and self-improvement loop to capture and share institutional knowledge, saving time for subject matter experts.
This blog post reflects on the romanticized notion of solitary grinding, using wrestling as a metaphor, and cautions against isolation in technological bubbles that reinforce personal assumptions.
S3Gym is an interactive benchmark that evaluates large language models' ability to self-test, self-judge, and self-improve through text-based games, revealing that self-improvement effectiveness varies by task structure and experience representation.
The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.
A paper from Google discusses separating skill-evolution systems into components like raw execution traces and a persistent wiki of knowledge for AI agents, related to self-improvement concepts.
This article introduces how Warp uses Claude to build self-improving AI agents, automatically optimizing skills through human feedback, and summarizes best practices for effective agent development.
This survey paper categorizes 80 unsupervised post-training methods for foundation models, organizing them by the internal update signals and presenting a unified framework for selection and evaluation.
PILOT is a supervisor-worker harness for live self-improvement in long-horizon agents, enabling real-time redirection and experience distillation to improve accuracy and efficiency.
Tim Ferriss questions whether people have outgrown their systems and beliefs, referencing Jerry Colonna's coaching perspective on personal complicity in creating unwanted conditions, especially in Silicon Valley tech circles.
The article presents AI4AI-Bench, a benchmark evaluating AI agents' ability to improve training algorithms, showing low performance scores and high exploration costs across ten research repositories.
The author discusses AI agent frameworks that minimize human intervention, highlighting projects like GitHub Agentic Workflows, OpenClaw, Hermes, and Aeon, and seeks community input on reliable options for long-running agent work.
An Anthropic engineer discusses the shift from prompting to AI engineering, emphasizing agents and self-improvement systems, with a live demonstration of setting up Claude Code.
The paper presents EvalCEGAR, a method for automatically evolving evaluation metrics using a pool of Python operators that flag specific defects in AI outputs, improving accuracy over hand-written operators and LLM judges.
Ornith-1.5 has launched a family of open AI models in 397B, 35B, and 9B sizes, featuring a self-improvement loop and achieving competitive benchmarks against top models like Claude Opus 4.8.
Ornith-1.5 is an AI model that uses self-generated tasks and scaffolds for continuous self-improvement, achieving state-of-the-art performance on benchmarks like Terminal-Bench and SWE-Bench compared to other models.
Ornith-1.5 is a family of open-source large language models with 9B, 35B, and 397B parameters, achieving state-of-the-art performance in reasoning, agentic, and coding tasks through self-improvement strategies.