Aspire: Can Models Self-Evolve from Vague Goals?
Summary
ASPIRE introduces a benchmark for self-evolving LLM agents from vague natural-language goals, revealing challenges in goal interpretation and stable weight-level improvement.
View Cached Full Text
Cached at: 09/03/26, 03:49 AM
Paper page - Aspire: Can Models Self-Evolve from Vague Goals?
Source: https://huggingface.co/papers/2608.31111 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
ASPIRE introduces a benchmark for self-evolving LLM agents from vague natural-language goals, revealing challenges in goal interpretation, data selection, and stable weight-level improvement.
Many important forms of human learning begin with a vague goal, such as “become a better physicist” or “improve at research.” Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work onLLM self-evolutiontypically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduceASPIRE, a benchmark forvague-goal-driven self-evolution.ASPIREprovides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate.ASPIREsupports both model-weight andagent-harness evolutionin a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrowself-evaluations, so local gains fail to transfer tohidden evaluationand continued search and training can erase earlier improvements.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2608\.31111
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.31111 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.31111 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.31111 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
This paper proposes a method to train LLM agents with intrinsic meta-evolution capabilities, enabling spontaneous self-improvement without external rewards at inference time. Applied to Qwen3-30B and Seed-OSS-36B, the approach yields a 20% performance boost on web navigation benchmarks, with a 14B model outperforming Gemini-2.5-Flash.
SAGE: A Statistical Acceptance Gate for Self-Evolving Agents
The paper introduces SAGE, a statistical acceptance gate for LLM agents that self-evolve by editing persistent skill documents, using per-item paired comparisons and a one-sided paired test to prevent regressions and avoid the Optimizer's Curse, achieving lower regression rates and higher scores across five benchmarks and four backbone LLMs.
@dair_ai: // MetaSkill-Evolve // Great paper on self-improving agents. Most self-improving agents rewrite what the agent does and…
MetaSkill-Evolve introduces a recursive two-timescale framework for LLM agents to evolve both task skills and the improvement procedure itself, achieving notable accuracy gains on OfficeQA, SealQA, and ALFWorld benchmarks.
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
This paper proposes MASA, a framework that adapts skills to each LLM backbone without modifying weights, using hierarchical evolution and a model-conditioned rewriter, achieving gains of up to 25.8 points over baselines.
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
AgentStream introduces a unified framework to evaluate self-evolving LLM agents under streaming task scenarios, showing that self-evolution reliability varies across scenarios and is gated by model capability.