Tag
SPACE 通过从成功轨迹归纳两层参数化技能,将子技能边界作为动作块监督,训练长程 LLM Agent 自适应输出可变长原子动作序列;在 ALFWorld 和 ScienceWorld 上成功率提升 7.0%–31.3%,决策轮次最多降低 78.9%。
Introduces QDOS, a unified pipeline for offline-to-online reinforcement learning that uses advantage-weighted quality-diversity pretraining to extract diverse and high-value skills, significantly improving performance in manipulation and locomotion tasks.
This article presents a method to extract ideas from bloggers' content using AI tools and develop reusable skill workflows.
SAGE is a framework for automating storyboard generation in short drama production using self-evolving rules and attribution-guided updates, achieving expert-level performance and reducing authoring time in commercial deployment.
Introduces ERSkill, a retrieval-centric framework for self-evolving, skill-guided adaptive memory access in LLM agents. It co-evolves retrieval skills and a routing policy, substantially outperforming strong baselines across agent memory benchmarks.
This paper argues that representing agent skills as code, rather than natural language, yields the best cost reduction for LLM agents, and introduces SpeedRunner, a coding agent that learns programmatic skills from past trajectories to improve performance while cutting costs across embodied environments.
Introduces ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning in LLM agents, showing that while sequential execution improves performance, current methods struggle to consolidate experience into robust, transferable skills.
This paper introduces SkillJack, the first attack targeting the experience-to-skill pipeline of self-evolving agents, showing that poisoned experiences can be transformed into persistent malicious skills that evade detection and survive deletion of original records.
A survey paper organizing robot-learning techniques along an axis of frozen-weight policies (VLA models) versus agents that write their own executable skills as code, providing a taxonomy of self-improvement mechanisms and analyzing the emerging robot-skill economy.
The Qwen team proposes the Skill Self-Play framework, which significantly improves model capabilities on tool-calling and reasoning tasks through the collaboration of Proposer, Solver, and a dynamic skill controller in self-play.
SkillRise is a unified reinforcement learning framework that enables LLM agents to learn and reuse skills across related, progressively challenging tasks, outperforming baselines by up to 8.5 percentage points on several benchmarks.
A tweet thread describes how to build a 'band of AI agents' that can discover each other, share context, and learn skills through inter-agent communication.
The paper introduces Skill Self-Play (Skill-SP), a co-evolutionary framework that uses a proposer, solver, and skill controller to bridge structured verification and open-ended exploration, improving LLM performance on tool-use and reasoning benchmarks.
Claude launches Record a Skill feature, users can teach Claude to execute complex tasks by recording the screen and verbally describing the steps, without writing code or prompts, greatly lowering the barrier to creating AI skills.
Introduces RELIC, a framework for learning interpretable and composable skills in multi-agent planning via revealed principles, enabling privacy-preserving coordination and cross-agent skill transfer without sharing code.
This paper introduces EvoClawBench, a benchmark designed to test whether AI agents can learn reusable skills from their own execution runs. Experiments with multiple agent runtimes show that skill learning is selective and cost-sensitive, not an automatic benefit.
An Anthropic engineer reportedly earning $1.7M/year argues that memory and retrieval architecture skills are the most valuable for AI engineers, with salaries ranging from $90k for prompt writers to over seven figures for system architects.
Introduces Hierarchical Experimentalist Agents (HExA), an in-context, experiment-centric self-improvement framework that enables LLM agents to design experiments, learn reusable skills, and answer queries in novel domains, achieving significant improvements over baselines on the Interphyre physics simulation benchmark.
Introduces 'skill neologisms', a method for enabling LLMs to learn new skills without weight updates, addressing catastrophic forgetting. Presented at ICML.
The tweet shares ChatGPT's 'Infinite Private Tutor' hidden mode, claiming it can teach any skill in 3 hours, and includes 8 practical prompts.