Tag
This paper proposes a dual-loop self-evolution framework for multi-turn empathetic dialogue, using verifiable emotion feedback to optimize policy and adapt training distribution. On SAGE, it improves Qwen3-8B Overall from 53.87 to 79.24, outperforming uniform emotion-reward RL by 7.23 points.
A survey paper proposing a three-stage taxonomy of co-evolution in agentic systems, covering agent-agent, agent-environment, and meta co-evolution to enable open-ended improvement beyond fixed human-designed paths.
This paper proposes Circuit-Anchored Evolution (CAE), a method that uses mechanistic interpretability to identify and anchor a tiny safety circuit in LLMs during self-evolution, preventing models from misevolving into capable but dangerous systems while preserving capability.
SkillHEX proposes a closed-loop framework for autonomous skill evolution in LLM agents, using hypothesis-driven self-verification and evidence-guided tree search to overcome sparse reward challenges. It outperforms existing self-evolving methods on SkillsBench with limited interaction budgets.
Argus is a persistent, self-evolving agentic runtime designed for long-horizon reasoning, using Manager, Planner, Engineer, and Reviewer roles with verification-gated persistence and pivoting. It demonstrates strong results across seven benchmark arenas, including ~78% on SWE-Bench Pro, while reducing token usage after runtime self-evolution.
This paper proposes A/B Agent, a closed-loop agent framework that organizes historical A/B testing knowledge into a hierarchical experience tree, retrieves transferable strategies via multi-path Tree-RAG, and self-evolves through online experiment feedback, achieving a 4.829% GMV improvement in a short-video e-commerce recommendation system.
GDPevo is a benchmark for evaluating agent self-evolution on real business tasks, covering CRM, ERP, finance, healthcare, legal, and data-centric workflows. The authors release an automated pipeline and find that self-evolution improves held-out accuracy by up to 16.44 percentage points, though agents remain well below an oracle ceiling.
The Qwen team proposes the Skill Self-Play framework, which significantly improves model capabilities on tool-calling and reasoning tasks through the collaboration of Proposer, Solver, and a dynamic skill controller in self-play.
This paper proposes SkillBoost, a three-stage constrained exploration-exploitation framework to mitigate skill overfitting in LLM agent self-evolution. It achieves state-of-the-art performance across 23 model-benchmark configurations and demonstrates that optimized skills transfer to other agents.
The paper introduces Operation-wise TableQA, a new task with fine-grained question taxonomy, and proposes SkillTGR, a skill-augmented table graph reasoning framework that uses graph traversal and a hierarchical SkillBank for self-evolving reasoning, achieving superior performance and efficiency.
This paper presents MetaEvolve, a framework that uses reinforcement learning to train LLMs in self-evolution meta-skills for iterative refinement, achieving significant improvements on coding benchmarks.
The paper introduces Skill Self-Play (Skill-SP), a co-evolutionary framework that uses a proposer, solver, and skill controller to bridge structured verification and open-ended exploration, improving LLM performance on tool-use and reasoning benchmarks.
This paper introduces a method for self-evolution of open-ended dialogue skills using future-feedback prediction, converting conversational feedback into a fixed offline objective to enable reproducible skill optimization without live traffic. The approach achieves over 75% prediction accuracy on a privacy-preserving sales-assistant dataset.
Cura 1T is a healthcare-specialized LLM trained via a human-gated self-evolution loop that iteratively improves on patient consultation, clinical reasoning, and agentic healthcare tasks, achieving top performance on medical benchmarks while maintaining general reasoning ability.
The paper presents ABot-AgentOS, a general robotic agent operating system with lifelong multi-modal memory, and introduces EmbodiedWorldBench for evaluating long-horizon embodied tasks. It demonstrates significant improvements in task success and memory benchmarks, suggesting that a dedicated agent OS layer enhances execution and persistent memory.
Qiao Bangzhu open-sourced the design Skill qiaomu-design, offering features like Style Fitting Room, anti-AI-style design, self-evolution mechanism, and references from mature websites, supporting GLM or Claude models.
The paper introduces Sealed Joint Search (SJS) and the Agora system, where five specialized LLM agent classes collaborate to evolve alpha factors. On a 91-day CSI 1000 holdout, Agora achieves a portfolio Sharpe of +1.87, significantly outperforming baselines, and the discovered metrics appear as emergent properties of the system.
Introduces RSEA, a method for recursive self-evolution of LLM agents using a three-layer natural-language state and a held-out selection gate to prevent regression. Evaluated across four benchmarks, it shows that context evolution is benchmark-dependent and that a strict selection gate is crucial for reliability.
This paper proposes the EDV framework, which uses multiple heterogeneous agents in execute-distill-verify stages to build reliable experiences for LLM agents, preventing self-confirmatory errors and improving performance on long-horizon benchmarks.
Introduces Autogenesis Protocol (AGP), a self-evolving agent protocol that decouples components from their evolution, enabling lifecycle management, version tracking, and safe rollback for prompts, agents, tools, environments, and memory in LLM-based multi-agent systems.