self-evolving-agents

Tag

Cards List
#self-evolving-agents

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

Hugging Face Daily Papers · 6d ago Cached

This paper proposes Feedback-Enriched Environments (FEEs) to bootstrap self-evolving agents in long-horizon tasks by enriching feedback for reinforcement learning, showing improved stability and performance on benchmarks using Qwen3 models and RL algorithms.

0 favorites 0 likes
#self-evolving-agents

Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

arXiv cs.CL · 2026-09-04 Cached

This paper introduces HarnessEvo to decompose LLM agent harnesses into separately-evolvable slots, revealing that optimization value is localized in specific components like reflection/control, and that uniform budget-splitting is sub-optimal, advocating for targeted budget concentration.

0 favorites 0 likes
#self-evolving-agents

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

arXiv cs.AI · 2026-08-20 Cached

This survey paper connects self-evolving LLM-based agents with dynamic graph transformation, proposing a framework to model agent state as dynamic graphs and organizing existing methods for their evolution.

0 favorites 0 likes
#self-evolving-agents

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

arXiv cs.AI · 2026-08-19 Cached

SAGE is a framework for automating storyboard generation in short drama production using self-evolving rules and attribution-guided updates, achieving expert-level performance and reducing authoring time in commercial deployment.

0 favorites 0 likes
#self-evolving-agents

This new DeepSeek paper is a must-read for anyone who is building self-evolving agents

Reddit r/AI_Agents · 2026-08-18

This article introduces a new paper from DeepSeek that proposes a programming paradigm for spatiotemporal composability, addressing dynamic composability challenges in self-evolving agents.

0 favorites 0 likes
#self-evolving-agents

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Hugging Face Daily Papers · 2026-08-14 Cached

PACE-Bench introduces a simulator-grounded benchmark for evaluating self-evolving agents on physics adaptation tasks involving iterative code redesign after environmental mutations, revealing that mechanism redesign is a major bottleneck compared to parameter inference.

0 favorites 0 likes
#self-evolving-agents

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Hugging Face Daily Papers · 2026-08-11 Cached

SkillZip is a new method for compressing the accumulated skills of self-evolving agents without evaluation rollouts, by finding a minimal faithful structural explanation that reuses repeated rules and procedures while preserving rare exceptions.

0 favorites 0 likes
#self-evolving-agents

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

Hugging Face Daily Papers · 2026-08-04 Cached

RoMeRL introduces a reduced-order memory reinforcement learning method for self-evolving LLM agents that balances feedback coverage and avoids the memory-reward trap. Experiments on ALFWorld and LifelongAgentBench show improved task performance, an 80% reduction in Cold-Q ratio, higher feedback density, and fewer maintained memories and LLM calls.

0 favorites 0 likes
#self-evolving-agents

SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

Hugging Face Daily Papers · 2026-08-04 Cached

This paper introduces SkillJack, the first attack targeting the experience-to-skill pipeline of self-evolving agents, showing that poisoned experiences can be transformed into persistent malicious skills that evade detection and survive deletion of original records.

0 favorites 0 likes
#self-evolving-agents

@qingke_ai: https://x.com/qingke_ai/status/2076115489219297455

X AI KOLs Timeline · 2026-07-12 Cached

普林斯顿博士后Shilong Liu提出自进化代理的三层分类体系:工件迭代优化、Agent Harness自我改进和无黄金答案的模型学习,系统梳理了相关概念和前沿工作。

0 favorites 0 likes
#self-evolving-agents

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

arXiv cs.CL · 2026-07-09 Cached

This paper investigates how a biased LLM judge silently disables skill retirement in self-evolving agents, showing that false-pass bias across a sharp threshold prevents contribution-based retirement and that the failure is universal across domains, detectable only through a defect-injection audit.

0 favorites 0 likes
#self-evolving-agents

A Taxonomy of Self-evolving Agents (15 minute read)

TLDR AI · 2026-07-09 Cached

Shilong Liu proposes a taxonomy classifying self-evolving agents into artifact optimization, harness self-improvement, and model learning, providing a common language for emerging agent research.

0 favorites 0 likes
#self-evolving-agents

@rohanpaul_ai: Great paper on Self-evolving agents. Enterprise agents cannot truly improve until their messy daily work becomes safe l…

X AI KOLs Timeline · 2026-07-03 Cached

A paper proposing a mechanism for enterprise agents to improve by safely converting messy daily work into learning data, using a data proxy and control layer, with AREAL2.0 demonstrating online RL from real interaction traces.

0 favorites 0 likes
#self-evolving-agents

Self-Evolving Agents with Anytime-Valid Certificates

arXiv cs.AI · 2026-07-02 Cached

This paper introduces SEA, an architecture for self-evolving agents that confines self-modification to a steering adapter and versioned harness around a frozen base model, using anytime-valid gates to audit modifications against a fixed error budget. Experiments on SWE-bench Verified with four base models show that the suite provides a +4 to +5% improvement on strong base models while preventing regressions.

0 favorites 0 likes
#self-evolving-agents

Metis: Bridging Text and Code Memory for Self-Evolving Agents

arXiv cs.CL · 2026-06-24 Cached

Metis presents a controlled study comparing text and code memory for self-evolving agents, finding they have complementary trade-offs. It proposes a hierarchical dual-representation memory system that improves task accuracy by up to 20.6% and reduces execution cost by up to 22.8% on the AppWorld benchmark.

0 favorites 0 likes
#self-evolving-agents

SEAGym: An Evaluation Environment for Self-Evolving LLM Agents

arXiv cs.AI · 2026-06-17 Cached

SEAGym is a new evaluation environment for self-evolving LLM agents that measures agent harness updates across training, validation, test, replay, and cost records, providing complementary signals about the evolution process.

0 favorites 0 likes
#self-evolving-agents

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

Hugging Face Daily Papers · 2026-06-16 Cached

OPD-Evolver proposes a self-evolving agent framework using slow-fast co-evolution and on-policy self-distillation to enhance memory management and policy learning, outperforming existing methods like ReasoningBank and Skill0 across multi-domain benchmarks.

0 favorites 0 likes
#self-evolving-agents

@qinzytech: https://x.com/qinzytech/status/2066585405479371092

X AI KOLs Timeline · 2026-06-15 Cached

A technical analysis of two approaches to building self-evolving AI agents: model-based (via architecture like SSMs or transformer with fast-weight updates, and training methods) and harness-based (via memory or meta harness that can rewrite itself). The author provides practical recommendations for different audiences.

0 favorites 0 likes
#self-evolving-agents

PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents

arXiv cs.AI · 2026-06-09 Cached

PACE introduces an anytime-valid commit gate for self-evolving agents that replaces greedy acceptance with a sequential hypothesis test, controlling false-commit probability and reducing churn while matching performance with lower variance.

0 favorites 0 likes
#self-evolving-agents

Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments

arXiv cs.CL · 2026-06-08 Cached

This paper introduces FinEvolveBench, a benchmark for financial sentiment prediction, and Tree-of-Experience (ToE), a structured experience-management method for LLM agents in low-repetition tasks with implicit rewards. Experiments show that ToE outperforms general-purpose experience mechanisms in such challenging settings.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback