self-evolution

Tag

Cards List
#self-evolution

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper proposes a dual-loop self-evolution framework for multi-turn empathetic dialogue, using verifiable emotion feedback to optimize policy and adapt training distribution. On SAGE, it improves Qwen3-8B Overall from 53.87 to 79.24, outperforming uniform emotion-reward RL by 7.23 points.

0 favorites 0 likes
#self-evolution

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

Hugging Face Daily Papers ↗ · 2026-08-10 Cached

A survey paper proposing a three-stage taxonomy of co-evolution in agentic systems, covering agent-agent, agent-environment, and meta co-evolution to enable open-ended improvement beyond fixed human-designed paths.

0 favorites 0 likes
#self-evolution

Safe Evolution with Circuit Anchors

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper proposes Circuit-Anchored Evolution (CAE), a method that uses mechanistic interpretability to identify and anchor a tiny safety circuit in LLMs during self-evolution, preventing models from misevolving into capable but dangerous systems while preserving capability.

0 favorites 0 likes
#self-evolution

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

arXiv cs.AI ↗ · 2026-08-07 Cached

SkillHEX proposes a closed-loop framework for autonomous skill evolution in LLM agents, using hypothesis-driven self-verification and evidence-guided tree search to overcome sparse reward challenges. It outperforms existing self-evolving methods on SkillsBench with limited interaction budgets.

0 favorites 0 likes
#self-evolution

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

arXiv cs.AI ↗ · 2026-08-06 Cached

Argus is a persistent, self-evolving agentic runtime designed for long-horizon reasoning, using Manager, Planner, Engineer, and Reviewer roles with verification-gated persistence and pivoting. It demonstrates strong results across seven benchmark arenas, including ~78% on SWE-Bench Pro, while reducing token usage after runtime self-evolution.

0 favorites 0 likes
#self-evolution

A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper proposes A/B Agent, a closed-loop agent framework that organizes historical A/B testing knowledge into a hierarchical experience tree, retrieves transferable strategies via multi-path Tree-RAG, and self-evolves through online experiment feedback, achieving a 4.829% GMV improvement in a short-video e-commerce recommendation system.

0 favorites 0 likes
#self-evolution

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Hugging Face Daily Papers ↗ · 2026-08-04 Cached

GDPevo is a benchmark for evaluating agent self-evolution on real business tasks, covering CRM, ERP, finance, healthcare, legal, and data-centric workflows. The authors release an automated pipeline and find that self-evolution improves held-out accuracy by up to 16.44 percentage points, though agents remain well below an oracle ceiling.

0 favorites 0 likes
#self-evolution

@Xudong07452910: Let the model set its own problems, solve its own problems, and train itself — the biggest fear is learning incorrect problems along with the correct ones. This paper by the Qwen team proposes Skill Self-Play, adding a continuously updated skill library to the model's self-evolution. There are three roles in training: The Proposer generates tasks that are just challenging enough based on the skills...

X AI KOLs Timeline ↗ · 2026-08-02 Cached

The Qwen team proposes the Skill Self-Play framework, which significantly improves model capabilities on tool-calling and reasoning tasks through the collaboration of Proposer, Solver, and a dynamic skill controller in self-play.

0 favorites 0 likes
#self-evolution

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

arXiv cs.AI ↗ · 2026-07-31 Cached

This paper proposes SkillBoost, a three-stage constrained exploration-exploitation framework to mitigate skill overfitting in LLM agent self-evolution. It achieves state-of-the-art performance across 23 model-benchmark configurations and demonstrates that optimized skills transfer to other agents.

0 favorites 0 likes
#self-evolution

Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering

arXiv cs.AI ↗ · 2026-07-28 Cached

The paper introduces Operation-wise TableQA, a new task with fine-grained question taxonomy, and proposes SkillTGR, a skill-augmented table graph reasoning framework that uses graph traversal and a hierarchical SkillBank for self-evolving reasoning, achieving superior performance and efficiency.

0 favorites 0 likes
#self-evolution

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

arXiv cs.CL ↗ · 2026-07-27 Cached

This paper presents MetaEvolve, a framework that uses reinforcement learning to train LLMs in self-evolution meta-skills for iterative refinement, achieving significant improvements on coding benchmarks.

0 favorites 0 likes
#self-evolution

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Hugging Face Daily Papers ↗ · 2026-07-24 Cached

The paper introduces Skill Self-Play (Skill-SP), a co-evolutionary framework that uses a proposer, solver, and skill controller to bridge structured verification and open-ended exploration, improving LLM performance on tool-use and reasoning benchmarks.

0 favorites 0 likes
#self-evolution

Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper introduces a method for self-evolution of open-ended dialogue skills using future-feedback prediction, converting conversational feedback into a fixed offline objective to enable reproducible skill optimization without live traffic. The approach achieves over 75% prediction accuracy on a privacy-preserving sales-assistant dataset.

0 favorites 0 likes
#self-evolution

Cura 1T: Specialized Model for Agentic Healthcare

arXiv cs.AI ↗ · 2026-07-20 Cached

Cura 1T is a healthcare-specialized LLM trained via a human-gated self-evolution loop that iteratively improves on patient consultation, clinical reasoning, and agentic healthcare tasks, achieving top performance on medical benchmarks while maintaining general reasoning ability.

0 favorites 0 likes
#self-evolution

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Hugging Face Daily Papers ↗ · 2026-07-11 Cached

The paper presents ABot-AgentOS, a general robotic agent operating system with lifelong multi-modal memory, and introduces EmbodiedWorldBench for evaluating long-horizon embodied tasks. It demonstrates significant improvements in task success and memory benchmarks, suggesting that a dedicated agent OS layer enhances execution and persistent memory.

0 favorites 0 likes
#self-evolution

@vista8: Open source, open source! Tired of AI-heavy web design and monotonous layouts? Qiao Bangzhu tested and compared 8 mainstream design Skills, absorbed the strengths of each, and open-sourced his own design Skill: qiaomu-design. Features: 1. Style Fitting Room: Generate four design demos from a single sentence, preview and confirm before development.

X AI KOLs Timeline ↗ · 2026-07-06 Cached

Qiao Bangzhu open-sourced the design Skill qiaomu-design, offering features like Style Fitting Room, anti-AI-style design, self-evolution mechanism, and references from mature websites, supporting GLM or Claude models.

0 favorites 0 likes
#self-evolution

AI Trading's Alpha Singularity: Emergent Market Reasoning through Agent-to-Agent Self-Evolution

arXiv cs.AI ↗ · 2026-06-30 Cached

The paper introduces Sealed Joint Search (SJS) and the Agora system, where five specialized LLM agent classes collaborate to evolve alpha factors. On a 91-day CSI 1000 holdout, Agora achieves a portfolio Sharpe of +1.87, significantly outperforming baselines, and the discovered metrics appear as emergent properties of the system.

0 favorites 0 likes
#self-evolution

Recursive Self-Evolving Agents via Held-Out Selection

arXiv cs.AI ↗ · 2026-06-30 Cached

Introduces RSEA, a method for recursive self-evolution of LLM agents using a three-layer natural-language state and a held-out selection gate to prevent regression. Evaluated across four benchmarks, it shows that context evolution is benchmark-dependent and that a strict selection gate is crucial for reliability.

0 favorites 0 likes
#self-evolution

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

This paper proposes the EDV framework, which uses multiple heterogeneous agents in execute-distill-verify stages to build reliable experiences for LLM agents, preventing self-confirmatory errors and improving performance on long-horizon benchmarks.

0 favorites 0 likes
#self-evolution

@bibryam: Autogenesis: A Self-Evolving Agent Protocol https://arxiv.org/abs/2604.15034 tldr: Treat every agent component like a v…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

Introduces Autogenesis Protocol (AGP), a self-evolving agent protocol that decouples components from their evolution, enabling lifecycle management, version tracking, and safe rollback for prompts, agents, tools, environments, and memory in LLM-based multi-agent systems.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback