Tag
This paper presents VIBEMed, a multi-agent framework with a self-evolution mechanism and safety sandbox for robust clinical decision support, integrating specialized agents for diagnosis, treatment planning, and evolving clinical knowledge over time.
This paper formalizes communication policy for LLM agents and proposes Communication Policy Evolution (CPE), a self-evolution framework that refines communication policies through rollout and prompt-level evolving, achieving best task success across multiple settings.
This paper introduces a self-evolving framework for vision-language models to improve their question-generation capabilities without external supervision, enhancing both question quality and answerer performance.
A weekly roundup of notable AI papers covering self-revising discovery systems from MIT, disentangling agent self-evolution, and Google's LEAP for formal mathematics using agentic scaffolds.
EvoMap introduces GEP (Genome Evolution Protocol), a network protocol enabling agents to convert successful strategies into genes and capsules for self-evolution, reducing repeated exploration.
This paper analyzes two capabilities in self-evolving LLM agents: harness-updating and harness-benefit. It finds that harness-updating is flat across base capability levels, while harness-benefit is non-monotonic, with mid-tier models benefiting most.
The article argues that self-evolution in AI agents should be applied cautiously and proposes an Evolution Governor that audits workflows to decide when to evolve, based on conditions like repeatable tasks and external feedback.
BenchTrace is a benchmark for evaluating the self-evolution abilities of LLM agents, focusing on reflection and controlled evolution through a dataset of 1,821 annotated episodes and two evaluation tasks: Reflection Evaluation and Evolution Evaluation. Experiments with Qwen3-32B and GPT-4.1 show both models struggle, with a main bottleneck in diagnosis and issues in generalization and forgetting.
Proposes the Agentic-VLA framework, introducing agents into the VLA loop, enabling the vision-language-action model to self-evolve and surpass existing leading embodied models on all metrics.
SEAL proposes a closed-loop framework for jointly evolving LLM agents and their training environments, using diagnosis-guided labels to align both sides. It achieves substantial gains in multi-turn tool-use tasks with only 400 training samples, demonstrating improved robustness and out-of-distribution transfer.
The MiniMax-M2 series introduces Mixture-of-Experts language models that achieve high performance on agentic tasks with minimal activated parameters (9.8B per token out of 229.9B total), leveraging agent-driven data pipelines, a scalable RL system called Forge, and a checkpoint that takes early steps toward self-evolution.
PACE introduces a two-timescale framework for self-evolution of small language model agents, coordinating low-risk prompt refinement with higher-risk control-logic updates, achieving up to +9.2% relative improvement across benchmarks.
SEAL is a closed-loop co-evolution framework for interactive tool-use agents that addresses Agent-Environment Misalignment by synchronizing policy and environment updates using on-policy trajectories and turn-level diagnosis.
SOLAR proposes a self-optimizing autonomous agent that leverages parameter-level meta-learning and multi-level reinforcement learning to enable lifelong adaptation of LLMs to non-stationary data streams, outperforming baselines on reasoning tasks.
FORGE is a protocol that enables LLM agents to evolve their memory via population broadcast without weight updates, converting failed trajectories into reusable knowledge artifacts. It significantly improves performance on the CybORG CAGE-2 network-defense task over zero-shot and Reflexion baselines across multiple LLM families.
This survey paper provides a unified review of LLM-based multi-agent systems, focusing on collaboration, failure attribution, and self-evolution through the LIFE framework, identifying open challenges and proposing a cross-stage research agenda.
Zhipu AI founder Tang Jie predicts that the biggest breakthrough in large models this year will be long-horizon tasks, where AI can continuously solve complex problems in real environments, and mentions three technical pillars and Anthropic's progress in autonomous training.
This article discusses the implementation of AI Agent memory, introduces the reproducible 93.05% LoCoMo accuracy of the EverOS system and the Skill self-evolution mechanism, and teases a new cool product launching at the end of the month.
RewardHarness is a self-evolving agentic framework for post-training that replaces large-scale preference annotation with iterative tool and skill evolution, achieving superior performance in image editing evaluation benchmarks compared to GPT-5.
This paper introduces FATE, an on-policy framework that leverages failure trajectories to enhance the safety and performance of tool-using LLM agents through self-evolution and Pareto-aware optimization.