skill-evolution

Tag

Cards List
#skill-evolution

Google Introduces WikiSkill for Persistent Agent Learning (22 minute read)

TLDR AI · 3d ago Cached

WikiSkill is a framework that co-evolves agent skills with a persistent knowledge base, consistently outperforming state-of-the-art methods and enabling effective skill transfer across models.

0 favorites 0 likes
#skill-evolution

@Teknium: Sounds a bit like Hermes' self improvement

X AI KOLs Timeline · 5d ago Cached

A paper from Google discusses separating skill-evolution systems into components like raw execution traces and a persistent wiki of knowledge for AI agents, related to self-improvement concepts.

0 favorites 0 likes
#skill-evolution

@dotey: A new blog post from Claude titled 'How Warp builds self-improving agents on Claude' https://claude.com/blog/how-warp-builds-self-improving-a…

X AI KOLs Timeline · 5d ago Cached

This article introduces how Warp uses Claude to build self-improving AI agents, automatically optimizing skills through human feedback, and summarizes best practices for effective agent development.

0 favorites 0 likes
#skill-evolution

@omarsar0: A great paper from Google on maintaining agent skills through persistent knowledge.

X AI KOLs Following · 6d ago Cached

A Google paper discusses maintaining agent skills by separating raw execution traces from persistent knowledge in a wiki-like structure.

0 favorites 0 likes
#skill-evolution

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Hugging Face Daily Papers · 2026-08-27 Cached

WikiSkill introduces a framework that co-evolves agent skills with a persistent knowledge base to systematically accumulate experience and improve performance across models, demonstrating benefits over state-of-the-art skill-evolution methods.

0 favorites 0 likes
#skill-evolution

@GitTrend0x: Hermes Must-Install List 42-evey's All-in-One Plugin Suite, outsourc-e's Web Command Center, jdtimothy's Model-Switching Wizard, AMAP-ML's Skill Self-Evolution, tlehman's Literate Programming... Turning Hermes into the Next-Gen 'Goal-Driven...'

X AI KOLs Timeline · 2026-08-21 Cached

This article introduces the Hermes Must-Install List, featuring plugins like the all-in-one plugin suite, Web Command Center, model-switching wizard, and skill self-evolution tools, aimed at enhancing Hermes' autonomy and productivity.

0 favorites 0 likes
#skill-evolution

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

arXiv cs.AI · 2026-08-14 Cached

This paper introduces SkillMisevo-Gym and SkillMisevo-Bench to study how self-improving LLM agents can evolve unsafe skills from compromised experience, plus SafeEvolve as a mitigation wrapper. Experiments across 25 agent-method configurations show skill misevolution is widespread and can persist across sessions, though SafeEvolve reduces fresh-session harm significantly.

0 favorites 0 likes
#skill-evolution

DIVE: Unlocking Self-Improvement in Frozen Language Models Through Diversity-Driven Skill Evolution

arXiv cs.CL · 2026-08-14 Cached

The paper introduces DIVE, a diversity-driven framework that enables frozen LLMs to self-improve by evolving persistent natural-language skills from task experience and verifier feedback, without parameter updates. It outperforms existing methods on math and logical reasoning tasks and transfers across model scales.

0 favorites 0 likes
#skill-evolution

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

Hugging Face Daily Papers · 2026-08-13 Cached

SkillEvo introduces a method to continuously improve AI agent skills by using multi-turn interaction feedback and governance layers to maintain evolution gradients, surpassing self-reflection and single-turn QA-driven approaches.

0 favorites 0 likes
#skill-evolution

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

arXiv cs.AI · 2026-08-03 Cached

This paper introduces SESA, a self-evolving skill-augmented search agent that co-evolves task generation and skill memory via tool-augmented search self-play. It improves accuracy across seven QA benchmarks over baselines while supporting memory-free deployment.

0 favorites 0 likes
#skill-evolution

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

arXiv cs.CL · 2026-07-15 Cached

Introduces KnowAct-GUIClaw, a framework for personal GUI assistants with self-evolving memory and skill, achieving state-of-the-art performance on the MobileWorld benchmark and outperforming closed-source models like GPT-5.5.

0 favorites 0 likes
#skill-evolution

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

arXiv cs.AI · 2026-07-15 Cached

This paper proposes a method for co-evolving evaluation metrics and skills in self-improving LLM agent systems, demonstrating that metrics can be evolved and that a co-evolution approach recovers most of the performance of a ground-truth-driven oracle across code generation, text-to-SQL, and report generation tasks.

0 favorites 0 likes
#skill-evolution

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

arXiv cs.AI · 2026-07-03 Cached

ComfyClaw is an agentic skill evolution framework for ComfyUI image generation workflows, using typed graph editing and region-level VLM verifiers to translate visual failures into repair suggestions, outperforming baselines across multiple configurations.

0 favorites 0 likes
#skill-evolution

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

Hugging Face Daily Papers · 2026-06-23 Cached

SkillHone is a harness for continual agent skill evolution that uses persistent decision history and practice feedback to improve performance on research and tool-mediated analysis tasks. It outperforms existing methods on GAIA and WebWalkerQA-EN benchmarks.

0 favorites 0 likes
#skill-evolution

SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing

arXiv cs.AI · 2026-06-15 Cached

SkillAudit introduces a framework for evolving LLM agent skills without ground-truth feedback by using paired trajectory auditing and contrastive evaluation. It achieves 73.9% average task reward across 89 tasks, outperforming baseline methods.

0 favorites 0 likes
#skill-evolution

VisualClaw: A Real-Time, Personalized Agent for the Physical World

Hugging Face Daily Papers · 2026-06-15 Cached

VisualClaw is a self-evolving multimodal agent that reduces deployment costs through hybrid encoding and skill evolution, while improving video-QA accuracy across multiple benchmarks.

0 favorites 0 likes
#skill-evolution

SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents

arXiv cs.CL · 2026-06-12 Cached

SkillCAT is a training-free framework for LLM agent skill self-evolution that addresses limitations of single-trace bias, unverified merging, and full corpus loading via three stages: Contrastive Causal Extraction, Assessment-Augmented Evolution, and Topology-Aware Task Execution, achieving up to 40.40% improvement on benchmarks.

0 favorites 0 likes
#skill-evolution

SkillChain: Closing the Loop on Skill Evolution for Image-Based E-Commerce AI Assistants

arXiv cs.CL · 2026-06-12 Cached

SkillChain automates the lifecycle of per-intent skill specifications for image-based e-commerce AI assistants, improving response quality and user engagement through iterative refinement and routing alignment.

0 favorites 0 likes
#skill-evolution

Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses

Hugging Face Daily Papers · 2026-06-06 Cached

Bayesian-Agent presents a framework that treats reusable skills and SOPs as hypotheses, using Bayesian inference to guide agent behavior and improve task performance through posterior-guided harness optimization. It achieves significant improvements on multiple benchmarks with deepseek-v4-flash.

0 favorites 0 likes
#skill-evolution

Verilog-Evolve: Feedback-Driven and Skill-Evolving Verilog Generation

arXiv cs.CL · 2026-05-27 Cached

Verilog-Evolve is a feedback-driven framework that iteratively refines Verilog code generated by large language models, using functional simulation, synthesis, and timing metrics to promote better candidates and evolve reusable repair skills across tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback