self-evolution

Tag

Cards List
#self-evolution

Toward Vibe Medicine: A Self-Evolving Multi-Agent Framework for Clinical Decision Support

arXiv cs.AI ↗ · 2026-06-16 Cached

This paper presents VIBEMed, a multi-agent framework with a self-evolution mechanism and safety sandbox for robust clinical decision support, integrating specialized agents for diagnosis, treatment planning, and evolving clinical knowledge over time.

0 favorites 0 likes
#self-evolution

Communication Policy Evolution for Proactive LLM Agents

arXiv cs.AI ↗ · 2026-06-15 Cached

This paper formalizes communication policy for LLM agents and proposes Communication Policy Evolution (CPE), a self-evolution framework that refines communication policies through rollout and prompt-level evolving, achieving best task success across multiple settings.

0 favorites 0 likes
#self-evolution

Self-Evolving Visual Questioner

Hugging Face Daily Papers ↗ · 2026-06-11 Cached

This paper introduces a self-evolving framework for vision-language models to improve their question-generation capabilities without external supervision, enhancing both question quality and answerer performance.

0 favorites 0 likes
#self-evolution

@dair_ai: https://x.com/dair_ai/status/2063644231030214958

X AI KOLs Following ↗ · 2026-06-07 Cached

A weekly roundup of notable AI papers covering self-revising discovery systems from MIT, disentangling agent self-evolution, and Google's LEAP for formal mathematics using agentic scaffolds.

0 favorites 0 likes
#self-evolution

@EvoMapAI: Introducing GEP (Genome Evolution Protocol). A network protocol developed by EvoMap. The core mechanism behind agent se…

X AI KOLs Following ↗ · 2026-06-02 Cached

EvoMap introduces GEP (Genome Evolution Protocol), a network protocol enabling agents to convert successful strategies into genes and capsules for self-evolution, reducing repeated exploration.

0 favorites 0 likes
#self-evolution

Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

arXiv cs.AI ↗ · 2026-06-01 Cached

This paper analyzes two capabilities in self-evolving LLM agents: harness-updating and harness-benefit. It finds that harness-updating is flat across base capability levels, while harness-benefit is non-monotonic, with mid-tier models benefiting most.

0 favorites 0 likes
#self-evolution

A framework for when AI agents should (and shouldn't) self-evolve

Reddit r/AI_Agents ↗ · 2026-05-31

The article argues that self-evolution in AI agents should be applied cautiously and proposes an Evolution Governor that audits workflows to decide when to evolve, based on conditions like repeatable tasks and external feedback.

0 favorites 0 likes
#self-evolution

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

arXiv cs.AI ↗ · 2026-05-29 Cached

BenchTrace is a benchmark for evaluating the self-evolution abilities of LLM agents, focusing on reflection and controlled evolution through a dataset of 1,821 annotated episodes and two evaluation tasks: Reflection Evaluation and Evolution Evaluation. Experiments with Qwen3-32B and GPT-4.1 show both models struggle, with a main bottleneck in diagnosis and issues in generalization and forgetting.

0 favorites 0 likes
#self-evolution

@FeitengLi: Just said this morning: The intelligence of embodied intelligence should copy the homework of LLM + RL + Agentic. Here it is: Agentic VLA crushes the models of leading embodied companies across the board https://x.com/FeitengLi/status/205909864717506193...

X AI KOLs Timeline ↗ · 2026-05-26 Cached

Proposes the Agentic-VLA framework, introducing agents into the VLA loop, enabling the vision-language-action model to self-evolve and surpass existing leading embodied models on all metrics.

0 favorites 0 likes
#self-evolution

SEAL: Synergistic Co-Evolution of Agents and Learning Environments

arXiv cs.CL ↗ · 2026-05-26 Cached

SEAL proposes a closed-loop framework for jointly evolving LLM agents and their training environments, using diagnosis-guided labels to align both sides. It achieves substantial gains in multi-turn tool-use tasks with only 400 training samples, demonstrating improved robustness and out-of-distribution transfer.

0 favorites 0 likes
#self-evolution

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Hugging Face Daily Papers ↗ · 2026-05-26 Cached

The MiniMax-M2 series introduces Mixture-of-Experts language models that achieve high performance on agentic tasks with minimal activated parameters (9.8B per token out of 229.9B total), leveraging agent-driven data pipelines, a scalable RL system called Forge, and a checkpoint that takes early steps toward self-evolution.

0 favorites 0 likes
#self-evolution

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

arXiv cs.LG ↗ · 2026-05-25 Cached

PACE introduces a two-timescale framework for self-evolution of small language model agents, coordinating low-risk prompt refinement with higher-risk control-logic updates, achieving up to +9.2% relative improvement across benchmarks.

0 favorites 0 likes
#self-evolution

SEAL: Synergistic Co-Evolution of Agents and Learning Environments

Hugging Face Daily Papers ↗ · 2026-05-23 Cached

SEAL is a closed-loop co-evolution framework for interactive tool-use agents that addresses Agent-Environment Misalignment by synchronizing policy and environment updates using on-policy trajectories and turn-level diagnosis.

0 favorites 0 likes
#self-evolution

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation

arXiv cs.AI ↗ · 2026-05-22 Cached

SOLAR proposes a self-optimizing autonomous agent that leverages parameter-level meta-learning and multi-level reinforcement learning to enable lifelong adaptation of LLMs to non-stationary data streams, outperforming baselines on reasoning tasks.

0 favorites 0 likes
#self-evolution

FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast

arXiv cs.AI ↗ · 2026-05-18

FORGE is a protocol that enables LLM agents to evolve their memory via population broadcast without weight updates, converting failed trajectories into reusable knowledge artifacts. It significantly improves performance on the CybORG CAGE-2 network-defense task over zero-shot and Reflexion baselines across multiple LLM families.

0 favorites 0 likes
#self-evolution

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems

Hugging Face Daily Papers ↗ · 2026-05-14 Cached

This survey paper provides a unified review of LLM-based multi-agent systems, focusing on collaboration, failure attribution, and self-evolution through the LIFE framework, identifying open challenges and proposing a cross-stage research agenda.

0 favorites 0 likes
#self-evolution

@0xLogicrw: Zhipu AI founder and chief scientist Tang Jie predicts that the biggest breakthrough in large models this year will be long-horizon tasks, where AI can continuously operate in real environments and solve complex problems. Once long-horizon tasks are achieved, today's 'one-person companies' will rapidly become 'no-employee companies...

X AI KOLs Timeline ↗ · 2026-05-13

Zhipu AI founder Tang Jie predicts that the biggest breakthrough in large models this year will be long-horizon tasks, where AI can continuously solve complex problems in real environments, and mentions three technical pillars and Anthropic's progress in autonomous training.

0 favorites 0 likes
#self-evolution

@elliotchen100: This Chinese article is the clearest I've seen recently on Agent memory. Full disclosure: the EverOS mentioned is built by us @EverMind. Three additions: 1. That 93.05% LoCoMo accuracy isn't just a paper claim—it's a script you can run from the open-source repo, reproducible by anyone. 2. Skill self-evolution: The first trajectory feed only produces cases; you need to run several similar tasks before distilling a skill. Many people integrate and see no skill and think it's broken. 3. Easter egg: A cool new product is launching at the end of the month. If it's not cool, I'll pay up.

X AI KOLs Timeline ↗ · 2026-05-13

This article discusses the implementation of AI Agent memory, introduces the reproducible 93.05% LoCoMo accuracy of the EverOS system and the Skill self-evolution mechanism, and teases a new cool product launching at the end of the month.

0 favorites 0 likes
#self-evolution

RewardHarness: Self-Evolving Agentic Post-Training

arXiv cs.AI ↗ · 2026-05-12 Cached

RewardHarness is a self-evolving agentic framework for post-training that replaces large-scale preference annotation with iterative tool and skill evolution, achieving superior performance in image editing evaluation benchmarks compared to GPT-5.

0 favorites 0 likes
#self-evolution

On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

This paper introduces FATE, an on-policy framework that leverages failure trajectories to enhance the safety and performance of tool-using LLM agents through self-evolution and Pareto-aware optimization.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback