self-improving-agents

Tag

Cards List
#self-improving-agents

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

arXiv cs.AI ↗ · 2d ago Cached

提出一种自进化 agent harness 框架:同一冻结模型先作为 solver 解题、再作为 proposer 直接编辑自己的 harness 代码,在多任务上进化后于分布外基准上显著超越 Codex(提升 12.64 分)并达成匹配表现。

0 favorites 0 likes
#self-improving-agents

MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

arXiv cs.AI ↗ · 3d ago Cached

MERID 是一个多模态抑郁症分析框架,通过经验驱动的递归自我改进智能体自动探索和优化检测流水线,在多个抑郁症基准上取得领先结果。

0 favorites 0 likes
#self-improving-agents

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

Hugging Face Daily Papers ↗ · 5d ago Cached

AREX-2 advances LLM agent self-improvement by training on synthesized long-horizon reflective trajectories from ML and algorithmic programming tasks. Built on Qwen3.8-27B, it achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7) and transfers to deep research tasks like BrowseComp, HLE, GAIA, and DeepSearchQA.

0 favorites 0 likes
#self-improving-agents

@rohanpaul_ai: New Stanford+Oxford paper MedRSI shows that medical agents can improve themselves from their own mistakes, but only if …

X AI KOLs Following ↗ · 2026-09-23 Cached

A Stanford+Oxford paper on MedRSI shows medical AI agents can self-improve from mistakes by testing new capabilities on fresh patients and using guardrails to maintain accuracy.

0 favorites 0 likes
#self-improving-agents

@arunmoorthy05: https://x.com/arunmoorthy05/status/2099936495465631943

X AI KOLs Timeline ↗ · 2026-09-15 Cached

Pace details their approach to improving an AI extraction agent for insurance documents by optimizing subagent delegation, model routing, and context based on historical data, leading to reduced cost and latency.

0 favorites 0 likes
#self-improving-agents

@Saboo_Shubham_: Stanford dropped a full course on self-improving AI agents. 100% free.

X AI KOLs Following ↗ · 2026-09-07 Cached

Stanford University has launched a complete free course on self-improving AI agents, providing accessible educational resources for AI advancements.

0 favorites 0 likes
#self-improving-agents

@rohanpaul_ai: Self-improving agents need a useful way to remember what worked. SkillGLoW shows agents should remember reusable ways o…

X AI KOLs Timeline ↗ · 2026-09-05 Cached

SkillGLoW demonstrates that self-improving agents achieve better performance by storing reusable task-solving procedures rather than memorizing every past task, gaining 17.2 points with a 3.6× more compact library.

0 favorites 0 likes
#self-improving-agents

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

arXiv cs.AI ↗ · 2026-09-03 Cached

This paper identifies failure modes in LLM-as-a-Judge systems for self-improving agents and introduces PROCTOR, an architecture with deterministic guardrails to mitigate these issues.

0 favorites 0 likes
#self-improving-agents

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

arXiv cs.AI ↗ · 2026-09-03 Cached

SkillGLoW is a method for LLM agents that organizes skills into procedural families to enhance self-improvement on long-horizon tasks, demonstrating significant performance gains over baselines.

0 favorites 0 likes
#self-improving-agents

Auditing Harness Tampering in Self-Improving Agents

arXiv cs.CL ↗ · 2026-09-02 Cached

The paper proposes a two-axis taxonomy for harness tampering in self-improving AI agents, builds an annotated corpus to benchmark audit methods, and finds that tampering occurs in real agent trajectories, highlighting integrity risks.

0 favorites 0 likes
#self-improving-agents

@kaorixbt: Google just dropped a free 2-hour course On turning one prompt into a complete multi-agent graph: 17:44 - Build your fi…

X AI KOLs Timeline ↗ · 2026-08-20 Cached

Google has launched a free 2-hour course that guides learners through building multi-agent AI systems from a single prompt, including self-improving loops and orchestrated graphs.

0 favorites 0 likes
#self-improving-agents

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper introduces SkillMisevo-Gym and SkillMisevo-Bench to study how self-improving LLM agents can evolve unsafe skills from compromised experience, plus SafeEvolve as a mitigation wrapper. Experiments across 25 agent-method configurations show skill misevolution is widespread and can persist across sessions, though SafeEvolve reduces fresh-session harm significantly.

0 favorites 0 likes
#self-improving-agents

Towards Researcher Agents for Knowledge-Graph Question Answering

arXiv cs.AI ↗ · 2026-08-11 Cached

This paper presents a self-improving 'researcher agent' for Text-to-SPARQL question answering over knowledge graphs, which iteratively refines its own prompts and tools. Evaluated on DBpedia, it achieves 0.22 accuracy and identifies predicate selection as the main bottleneck.

0 favorites 0 likes
#self-improving-agents

@Xudong07452910: Recently discovered that Stanford Online posted the full lecture recordings of the graduate course CS329A: Self-Improving AI Agents to YouTube, 9 lectures in total. The course basically revolves around one question: how AI agents can continuously improve through interaction with their environment...

X AI KOLs Timeline ↗ · 2026-08-10 Cached

Stanford University has released the full recordings of the graduate course CS329A: Self-Improving AI Agents on YouTube, with 9 lectures covering topics such as test-time compute, verifiers, reinforcement learning, tool feedback, and multi-step reasoning, systematically organizing the direction of agent self-improvement.

0 favorites 0 likes
#self-improving-agents

@rohanpaul_ai: A self-improving agent can keep its weights frozen and still get worse through the memories it learns to trust. These a…

X AI KOLs Following ↗ · 2026-08-09 Cached

The paper 'Memory Reward Inflation in Self-Improving LLM Agents' shows that self-improving agents with frozen weights can still degrade by trusting flawed LLM-generated memory scores, with models endorsing 31–54% of their own wrong answers. This 'Echo Gap' persists across stronger LLMs.

0 favorites 0 likes
#self-improving-agents

@rvaniaaaa: Almost every self-improving agent i looked at had the same missing piece. People build the discovery part. Scout finds …

X AI KOLs Following ↗ · 2026-08-08 Cached

A tweet discusses the missing piece in self-improving agents: a strict review gate where an experienced engineer must approve each skill. It explains how a growing library improves discovery and extraction quality, with humans approving all merges to keep the system trustworthy.

0 favorites 0 likes
#self-improving-agents

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

arXiv cs.AI ↗ · 2026-07-15 Cached

This paper proposes a method for co-evolving evaluation metrics and skills in self-improving LLM agent systems, demonstrating that metrics can be evolved and that a co-evolution approach recovers most of the performance of a ground-truth-driven oracle across code generation, text-to-SQL, and report generation tasks.

0 favorites 0 likes
#self-improving-agents

Today, we are releasing verifiers v1 (3 minute read)

TLDR AI ↗ · 2026-07-14 Cached

Prime Intellect Lab is out of beta, offering a platform to train models with support for various architectures and modalities, enabling self-improving agents.

0 favorites 0 likes
#self-improving-agents

@loganthorneloe: https://x.com/loganthorneloe/status/2075684831233757275

X AI KOLs Timeline ↗ · 2026-07-10 Cached

A curated list of five notable AI articles from the past week, covering self-improving agents, Bun's migration to Rust, vLLM architecture, coding evaluation benchmarks, and agent autonomy levels.

0 favorites 0 likes
#self-improving-agents

@zodchiii: Anthropic x Google just released 3 workshops on building self-improving agentic systems from scratch: 00:01 – Why multi…

X AI KOLs Timeline ↗ · 2026-07-10 Cached

Anthropic and Google released three free workshops on building self-improving agentic systems, covering multi-agent pitfalls, skill gaps, and building autonomous agents.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback