SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
Summary
Scaffold is a self-improving framework for visual web agents that induces parametric skills, maintains a recursive hierarchy, and distills skills into model weights, achieving significant performance improvements on benchmarks like WebArena.
View Cached Full Text
Cached at: 09/10/26, 08:36 AM
# Self-Improving Web Agents via Recursive Parametric Skill Abstraction
Source: [https://arxiv.org/html/2609.05511](https://arxiv.org/html/2609.05511)
Xiaokun ZhangAffiliation:CityUHKMeng DingAffiliation:UMass BostonXue LiuAffiliation:MBZUAIAffiliation:McGill University
###### Abstract
Web agents need to navigate visually rich, long\-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate\. Recent skill\-augmented frameworks take an important first step, but they treat the skill library as a flat or two\-tier prompt\-side cache and offer no principled mechanism for compressing redundancy or composing skills recursively\. We introduceScaffold, a self\-improving framework for visual web agents that \(i\) induces parametric, executable skills from successful trajectories under a multi\-instance abstraction constraint, \(ii\) maintains a recursively composed hierarchy in which higher\-level skills invoke lower\-level ones, \(iii\) compacts the library via a minimum\-description\-length \(MDL\) criterion and behavioral equivalence checking, and \(iv\) periodically distills skill\-augmented trajectories back into model weights to internalize the abstractions\. Across WebArena, VisualWebArena, and a held\-out split of Online\-Mind2Web,Scaffoldimproves success rate by11\.111\.1–17\.217\.2absolute points over the strongest skill\-augmented baseline and shows monotonic gains across five self\-improvement iterations without library collapse\. We release the code and documents in the Github[repository](https://github.com/BokwaiHo/SCAFFOLD)\.
## 1Introduction
Autonomous web agents, LLM\-based systems that complete user instructions by interacting with browsers, have advanced rapidly with the arrival of strong multimodal foundation models[Zheng et al\. \(2024a\)](https://arxiv.org/html/2609.05511#bib.bib41);[He et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib8);[Qin et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib16);[Wang et al\. \(2025a\)](https://arxiv.org/html/2609.05511#bib.bib22)\. Yet on every realistic benchmark, from WebArena[Zhou et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib43)and VisualWebArena[Koh et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib12)to OSWorld[Xie et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib29)and Online\-Mind2Web[Xue et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib32), the performance still lags human performance by a wide margin\. A central cause is that current agents treat each task as an isolated episode: procedural knowledge acquired while booking a flight is discarded before the next task begins, even if that next task shares most of the same sub\-procedure[Wang et al\. \(2025b\)](https://arxiv.org/html/2609.05511#bib.bib25);[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib40)\.
This observation has motivated a wave ofself\-improvingweb agents that retain experience in some form\. Three sub\-paradigms have emerged\.Trajectory replaymethods store raw experience for retrieval\-augmented prompting[Zheng et al\. \(2024b\)](https://arxiv.org/html/2609.05511#bib.bib42);[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib39);[Chhikara et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib3)\.Workflow inductionmethods abstract action sequences into natural\-language routines that the agent can re\-read at inference time[Wang et al\. \(2025b\)](https://arxiv.org/html/2609.05511#bib.bib25);[Fang et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib5)\.Skill inductionmethods go further and synthesize callable APIs or parametric skills, exemplified bySkillWeaver[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib40),AppAgentX[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib11), and the very recentSkillRL[Xia et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib28)\. Skill induction is the strongest of them because the reuse unit is an executable program rather than a textual hint, and skills can in principle becomposed\. This property, as long argued by the hierarchical reinforcement learning community[Sutton et al\. \(1999\)](https://arxiv.org/html/2609.05511#bib.bib20), is essential for tackling long\-horizon problems\.
Despite this trajectory, three limitations of current skill\-augmented agents remain unresolved\.111We focus on visual web agents throughout; complementary work on text\-only LLM agents includes[Xia et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib28);[Wang et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib23);[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.05511#bib.bib26)\.First, induced skill libraries are essentially flat:SkillWeaverstores APIs in a single pool, andSkillRLuses only a two\-tier \(general / task\-specific\) split\. Since neither supports skills calling other skills, the maximum abstraction depth is bounded by one\. Second, libraries grow monotonically and accumulate near\-duplicate skills with no principled compression mechanism, a known failure mode in lifelong learning[Wang et al\. \(2024a\)](https://arxiv.org/html/2609.05511#bib.bib21)\. Third, the knowledge encoded in skills exists only in prompts and is never internalized into the model, leaving the base policy after iterations as weak as it was at the beginning\.
Contributions\.We introduceScaffold, a framework that addresses each of these limitations:
- •Amulti\-instance parameter inductionprocedure that proposes a skill only when multiple semantically similar trajectories support the same parametric abstraction, sharply reducing spurious abstracted skills \(§[3\.3](https://arxiv.org/html/2609.05511#S3.SS3)\)\.
- •Arecursive compositionmechanism: at iterationkk, the inducer is allowed to call any skill from iterations1,…,k1,\\ldots,k, producing a hierarchy whose depth grows monotonically\. We track depth and reuse explicitly in our design \(§[3\.4](https://arxiv.org/html/2609.05511#S3.SS4)\)\.
- •AnMDL\-driven library compactionstep that periodically merges behaviorally equivalent skills, refactors recurring sub\-patterns into new mid\-level skills, and prunes unused ones, with empirical validation on a held\-out task set \(§[3\.5](https://arxiv.org/html/2609.05511#S3.SS5)\)\.
- •Adistillation loopthat uses skill\-augmented trajectories to supervise\-fine\-tune the base policy, converting the in\-context library into improved weights \(§[3\.6](https://arxiv.org/html/2609.05511#S3.SS6)\)\.
On WebArena, VisualWebArena, and a held\-out Online\-Mind2Web subset,Scaffoldimproves overSkillWeaverby 13\.6–17\.7 absolute success\-rate points\. Critically, it continues to improve through five self\-improvement iterations \(whereas baselines saturate at two\), exhibits a desirable library\-depth and reuse\-rate distribution, and zero\-shot transfers to held\-out sites with a 17\.2\-point margin over the strongest competitor\.
## 2Related Work
Web and GUI agents\.Modern web agents either prompt strong proprietary models[Zheng et al\. \(2024a\)](https://arxiv.org/html/2609.05511#bib.bib41);[He et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib8);[Yang et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib34)or fine\-tune open vision\-language models on large GUI trajectory corpora[Cheng et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib2);[Hong et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib9);[Wu et al\. \(2025b\)](https://arxiv.org/html/2609.05511#bib.bib27);[Xu et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib30);[Qin et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib16);[Wang et al\. \(2025a\)](https://arxiv.org/html/2609.05511#bib.bib22);[Lai et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib13)\. The latter line achieves impressive grounding but relies on static datasets that fail to capture the procedural diversity of real websites[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib40)\. Reinforcement\-learning\-based agents close part of this gap by training on the environment directly[Qi et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib15);[Patel et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib14);[Gandhi and Neubig \(2026\)](https://arxiv.org/html/2609.05511#bib.bib6);[Shao et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib18);[Guo et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib7), but typically lack any mechanism for accumulating reusable behavioral abstractions, which is our focus\.
Self\-improvement through experience\.The idea of letting a model improve itself by training on its own filtered generations dates back toSTaR[Zelikman et al\. \(2022\)](https://arxiv.org/html/2609.05511#bib.bib38)and was extended to richer settings by Quiet\-STaR[Zelikman et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib37)and self\-play fine\-tuning[Chen et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib1)\. For agents specifically,[Patel et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib14)showed that filtered self\-generated trajectories can lift WebArena performance, and Reflexion[Shinn et al\. \(2023\)](https://arxiv.org/html/2609.05511#bib.bib19)introduced verbal in\-context self\-correction without weight updates\.Scaffoldinherits the spirit of these works but operates over astructuredskill library rather than raw text or trajectory data\.
Skill libraries and workflows\.Voyager[Wang et al\. \(2024a\)](https://arxiv.org/html/2609.05511#bib.bib21)pioneered LLM\-driven skill libraries in the programmatic Minecraft environment;TroVE[Wang et al\. \(2024b\)](https://arxiv.org/html/2609.05511#bib.bib24)extended this idea to tool induction for programmatic tasks\. For web/GUI agents, Agent Workflow Memory \(AWM\)[Wang et al\. \(2025b\)](https://arxiv.org/html/2609.05511#bib.bib25)induces natural\-language workflows from past trajectories and retrieves them at test time, achieving large gains on Mind2Web and WebArena\.SkillWeaver[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib40)advances this paradigm by synthesizing skills as executable APIs and using iterative practice for refinement\.AppAgentX[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib11)explores analogous ideas on mobile UIs\. The most direct prior work isSkillRL[Xia et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib28), which trains an LLM agent jointly with a two\-tier hierarchical skill library \(general skills \+ task\-specific skills\) using GRPO[Shao et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib18)\.Scaffolddiffers in operating on visual web agents with truly recursive hierarchies, an MDL\-based compaction step, and an explicit skill transfer evaluation\.
Concurrent skill frameworks\.PolySkill[Yu et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib36)separates a skill’s abstract goal from its per\-site implementations via abstract interfaces, whereas our skills re\-bind semantic references at run time;SkillEvo[Xu et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib31)pairs GRPO with an evolving skill path graph whose reuse unit is an experience path rather than a parametric program and which carries no compression objective\.
Memory and lifelong learning for agents\.A parallel literature treats agent experience as memory rather than skill, includingMem0[Chhikara et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib3),ExpeL[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib39),EvolveR[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.05511#bib.bib26),Synapse[Zheng et al\. \(2024b\)](https://arxiv.org/html/2609.05511#bib.bib42), and trajectory\-informed memory generation[Fang et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib5)\. These methods are largely orthogonal to skill induction and could be combined withScaffold; we useMem0andSynapseas memory\-style baselines in §[4](https://arxiv.org/html/2609.05511#S4)\. Finally, classical hierarchical RL[Sutton et al\. \(1999\)](https://arxiv.org/html/2609.05511#bib.bib20)and library learning in program synthesis[Ellis et al\. \(2021\)](https://arxiv.org/html/2609.05511#bib.bib4)provide theoretical grounding for the recursive composition and MDL\-based compression operations central to our method\.
Figure 1:Overview ofScaffoldframework\. The five stages are executed sequentially in each iteration: rollout, multi\-instance induction, recursive composition, MDL\-driven compaction, and distillation\. The library growing in both depth and width is periodically compressed\.
## 3Methodology
### 3\.1Problem Setup and Notation
We consider an agent operating in a partially observable web environmentℰ\\mathcal\{E\}\. At steptt, the agent receives observationot=\(st,dt,ut\)o\_\{t\}=\(s\_\{t\},d\_\{t\},u\_\{t\}\)consisting of a screenshotsts\_\{t\}, accessibility\-tree / document object model \(DOM\) snippetdtd\_\{t\}, and current URLutu\_\{t\}\. The agent emits an actionat∈𝒜prim∪𝒜skilla\_\{t\}\\in\\mathcal\{A\}\_\{\\text\{prim\}\}\\cup\\mathcal\{A\}\_\{\\text\{skill\}\}, where𝒜prim\\mathcal\{A\}\_\{\\text\{prim\}\}are primitive actions \(click, type, scroll, wait\) and𝒜skill\\mathcal\{A\}\_\{\\text\{skill\}\}is the set of currently available skill invocations\. A taskτ=\(I,init,V\)\\tau=\(I,\\text\{init\},V\)specifies a natural\-language instructionII, an initial environment state, and a verifierVVthat returns binary success on the final state\. We denote a trajectory byζ=\(o0,a0,o1,…,oT\)\\zeta=\(o\_\{0\},a\_\{0\},o\_\{1\},\\ldots,o\_\{T\}\)\. Askillis a tupleσ=\(name,description,𝜽,pre,body,post,dσ\)\\sigma=\(\\text\{name\},\\text\{description\},\\boldsymbol\{\\theta\},\\text\{pre\},\\text\{body\},\\text\{post\},d\_\{\\sigma\}\)\.𝜽\\boldsymbol\{\\theta\}is a typed parameter list admitting primitive values \(strings, numbers\), semantic element references \(natural\-language descriptors grounded to DOM elements at run time by a VLM grounder\), and pointers to other skills\.preandpostare LLM\-checkable preconditions and postconditions phrased as predicates overoto\_\{t\},bodyis an executable program over𝒜prim∪𝒜skill\\mathcal\{A\}\_\{\\text\{prim\}\}\\cup\\mathcal\{A\}\_\{\\text\{skill\}\}, anddσ∈ℕd\_\{\\sigma\}\\in\\mathbb\{N\}is theabstraction depth\. The libraryℒk\\mathcal\{L\}\_\{k\}at iterationkkis a set of such skills together with an embedding\-based retrieval index keyed ondescription\. We callσ\\sigmasemi\-parametric: control flow insidebodycalls no policy model, while semantic element references are grounded on the live observation by a small VLM \(§[4\.9](https://arxiv.org/html/2609.05511#S4.SS9)\)\.
### 3\.2Overall Pipeline
Each self\-improvement iterationk→k\+1k\\to k\+1consists of five stages, executed in order: \(1\)trajectory collectionwith the current policyπk\\pi\_\{k\}and libraryℒk\\mathcal\{L\}\_\{k\}; \(2\)multi\-instance skill induction; \(3\)recursive compositionof new skills overℒk\\mathcal\{L\}\_\{k\}; \(4\)MDL\-driven library compaction; and \(5\)distillationof the skill\-augmented behavior intoπk\+1\\pi\_\{k\+1\}\. Algorithm[1](https://arxiv.org/html/2609.05511#alg1)summarizes the loop\. Below we describe each stage in detail; the design choices are motivated by the failure modes of prior skill\-induction methods that we observe empirically \(cf\. §[4\.3](https://arxiv.org/html/2609.05511#S4.SS3)\)\. The overall framework is also illustrated in Figure[1](https://arxiv.org/html/2609.05511#S2.F1)\.
### 3\.3Multi\-Instance Skill Induction
A persistent failure mode in workflow\-induction methods is the synthesis ofspurious abstractions: a skill is proposed from a single trajectory but encodes incidental details \(specific element IDs, hard\-coded text\) that fail to generalize\. We address this by requiring every candidate skill to be supported by at leastnminn\_\{\\min\}trajectories with the same parameter structure\.
Successful trajectories𝒵k\+\\mathcal\{Z\}\_\{k\}^\{\+\}are first clustered by the embedding of their instructionII\. Within each clusterC=\{ζ1,…,ζn\}C=\\\{\\zeta\_\{1\},\\ldots,\\zeta\_\{n\}\\\}withn≥nminn\\geq n\_\{\\min\}, the inducer \(a frozen LLM\) is prompted with allnntrajectories and asked to: \(a\) identify which positions vary across trajectories and propose a typed parameter list𝜽\\boldsymbol\{\\theta\}, \(b\) emit an executable body using𝜽\\boldsymbol\{\\theta\}, primitives in𝒜prim\\mathcal\{A\}\_\{\\text\{prim\}\}, and any skill inℒk\\mathcal\{L\}\_\{k\}, \(c\) state a precondition / postcondition pair as Python expressions overoto\_\{t\}\. The induced skill mustre\-execute correctlyon at least one held\-out instance fromCC; otherwise it is discarded\. This validation step replaces the post\-hoc filtering used inSkillWeaverand AWM, and we show in §[4\.3](https://arxiv.org/html/2609.05511#S4.SS3)that it nearly halves the rate of induced\-then\-deprecated skills across iterations\.
Algorithm 1ScaffoldSelf\-Improvement Loop1:Base policy
π0\\pi\_\{0\}, task pool
𝒯\\mathcal\{T\}, verifier
VV, iterations
KK, compaction interval
MM
2:
ℒ0←∅\\mathcal\{L\}\_\{0\}\\leftarrow\\emptyset
3:for
k=0,…,K−1k=0,\\ldots,K\-1do
4:
𝒵k←Rollout\(πk,ℒk,𝒯\)\\mathcal\{Z\}\_\{k\}\\leftarrow\\textsc\{Rollout\}\(\\pi\_\{k\},\\mathcal\{L\}\_\{k\},\\mathcal\{T\}\)⊳\\trianglerightcollect trajectories
5:
𝒵k\+←\{ζ∈𝒵k:V\(ζ\)=1\}\\mathcal\{Z\}\_\{k\}^\{\+\}\\leftarrow\\\{\\zeta\\in\\mathcal\{Z\}\_\{k\}:V\(\\zeta\)=1\\\}
6:
𝒞k←ClusterByInstr\(𝒵k\+\)\\mathcal\{C\}\_\{k\}\\leftarrow\\textsc\{ClusterByInstr\}\(\\mathcal\{Z\}\_\{k\}^\{\+\}\)⊳\\trianglerightgroup similar trajs
7:
Σnew←∅\\Sigma^\{\\text\{new\}\}\\leftarrow\\emptyset
8:forcluster
C∈𝒞kC\\in\\mathcal\{C\}\_\{k\}with
\|C\|≥nmin\|C\|\\geq n\_\{\\min\}do
9:
σ←Induce\(C,ℒk\)\\sigma\\leftarrow\\textsc\{Induce\}\(C,\\mathcal\{L\}\_\{k\}\)⊳\\trianglerightmulti\-instance abstraction; may invoke anyσ′∈ℒk\\sigma^\{\\prime\}\\in\\mathcal\{L\}\_\{k\}
10:if
ValidateHoldout\(σ,𝒯val\)\\textsc\{ValidateHoldout\}\(\\sigma,\\mathcal\{T\}\_\{\\text\{val\}\}\)then
11:
Σnew←Σnew∪\{σ\}\\Sigma^\{\\text\{new\}\}\\leftarrow\\Sigma^\{\\text\{new\}\}\\cup\\\{\\sigma\\\}
12:endif
13:endfor
14:
ℒk\+1←ℒk∪Σnew\\mathcal\{L\}\_\{k\+1\}\\leftarrow\\mathcal\{L\}\_\{k\}\\cup\\Sigma^\{\\text\{new\}\}
15:if
\(k\+1\)modM=0\(k\+1\)\\bmod M=0then
16:
ℒk\+1←CompactMDL\(ℒk\+1,𝒵0:k\+\)\\mathcal\{L\}\_\{k\+1\}\\leftarrow\\textsc\{Compact\}\_\{\\text\{MDL\}\}\(\\mathcal\{L\}\_\{k\+1\},\\mathcal\{Z\}\_\{0:k\}^\{\+\}\)
17:endif
18:
πk\+1←Distill\(πk,ℒk\+1,𝒵k\+\)\\pi\_\{k\+1\}\\leftarrow\\textsc\{Distill\}\(\\pi\_\{k\},\\mathcal\{L\}\_\{k\+1\},\\mathcal\{Z\}\_\{k\}^\{\+\}\)
19:endfor
20:return
πK,ℒK\\pi\_\{K\},\\mathcal\{L\}\_\{K\}
### 3\.4Recursive Composition and Depth Tracking
The library is hierarchical by construction: when the inducer emits a body, it may invoke any skillσ′∈ℒk\\sigma^\{\\prime\}\\in\\mathcal\{L\}\_\{k\}as a sub\-routine\. The depth of a new skillσ\\sigmais defined as
dσ=1\+maxσ′∈calls\(σ\)dσ′,d\_\{\\sigma\}=1\+\\max\_\{\\sigma^\{\\prime\}\\in\\text\{calls\}\(\\sigma\)\}d\_\{\\sigma^\{\\prime\}\},\(1\)withdσ=1d\_\{\\sigma\}=1whencalls\(σ\)⊆𝒜prim\\text\{calls\}\(\\sigma\)\\subseteq\\mathcal\{A\}\_\{\\text\{prim\}\}\. We do not enforce a maximum depth; instead we measure how depth evolves across iterations \(§[4\.6](https://arxiv.org/html/2609.05511#S4.SS6)\)\. To prevent runaway abstraction, we forbid cycles via a simple topological check at induction time, and we cap the per\-skill body length at 30 lines\.
The choice to allowanyprior skill, not just skills from iterationk−1k\-1, enables the inducer to refactor a long chain of primitives into a single mid\-level skill at any iteration\. Concretely, an inducedbook\_flightskill at iteration33may callsearch\_dates\(iteration11\),fill\_passenger\_info\(iteration22\), and a freshly synthesizedconfirm\_payment\(also iteration33\)\. This stands in contrast toSkillRL’s fixed two\-tier hierarchy, which forces the agent to flatten the natural call structure of real web procedures\.
### 3\.5MDL\-Driven Library Compaction
Without governance, the library grows monotonically and accumulates redundant skills, which both inflates the retrieval namespace and dilutes the gradient signal during distillation\. EveryMMiterations \(we useM=2M\{=\}2\) we apply acompactionstep whose objective is the classical minimum\-description\-length functional[Rissanen \(1978\)](https://arxiv.org/html/2609.05511#bib.bib17):
ℱ\(ℒ\)=∑σ∈ℒ\|σ\|⏟library cost\+∑ζ∈𝒵\+minparse\(ζ∣ℒ\)\|parse\(ζ∣ℒ\)\|⏟data cost\.\\mathcal\{F\}\(\\mathcal\{L\}\)=\\underbrace\{\\sum\_\{\\sigma\\in\\mathcal\{L\}\}\|\\sigma\|\}\_\{\\text\{library cost\}\}\+\\underbrace\{\\sum\_\{\\zeta\\in\\mathcal\{Z\}^\{\+\}\}\\min\_\{\\text\{parse\}\(\\zeta\\mid\\mathcal\{L\}\)\}\|\{\\text\{parse\}\(\\zeta\\mid\\mathcal\{L\}\)\}\|\}\_\{\\text\{data cost\}\}\.\(2\)Here\|σ\|\|\\sigma\|is the token length of the skill body, and\|parse\(ζ∣ℒ\)\|\|\\text\{parse\}\(\\zeta\\mid\\mathcal\{L\}\)\|is the length of the shortest program overℒ∪𝒜prim\\mathcal\{L\}\\cup\\mathcal\{A\}\_\{\\text\{prim\}\}that reproducesζ\\zeta\. Intuitively,ℱ\\mathcal\{F\}rewards libraries that are short \(few, concise skills\) but that can compactly explain past successful behavior; it penalizes bothover\-specificskills \(used by few trajectories\) andredundantskills \(whose role is already covered\)\.
We approximate the minimization of Eq\.[2](https://arxiv.org/html/2609.05511#S3.E2)via three greedy operators applied iteratively untilℱ\\mathcal\{F\}no longer decreases:\(i\) Merge: two skillsσa,σb\\sigma\_\{a\},\\sigma\_\{b\}that arebehaviorally equivalenton a sampled holdout set \(i\.e\., produce identical post\-states with probability≥ρ\\geq\\rho\) are merged into the shorter of the two, with the longer being replaced by an alias\.\(ii\) Refactor: if the same primitive\-action subsequence of length≥ℓmin\\geq\\ell\_\{\\min\}appears in≥rmin\\geq r\_\{\\min\}existing skill bodies, the inducer is asked to propose a new mid\-level skill that captures it, and the existing skills are rewritten to call it\.\(iii\) Prune: skills used by zero trajectories in𝒵\+0:k\\mathcal\{Z\}^\{\+\}\_\{0:k\}over the most recentMMiterations are removed \(unless invoked transitively by another skill\)\. Each candidate modification is accepted only if \(a\) it reducesℱ\\mathcal\{F\}and \(b\) it does not decrease success rate on a held\-out validation set𝒯val\\mathcal\{T\}\_\{\\text\{val\}\}\. This last constraint is the key safety net that distinguishes principled compaction from naive deduplication\.
Functional redundancy and lifecycle\.Equivalence is tested on a stratified context set including perturbed states, so two skills reaching one goal by different routes are collapsed only when indistinguishable everywhere: 14 goal\-overlapping groups survive atk=5k\{=\}5\. Unmet preconditions block invocation at run time, violated postconditions trigger one retry, and skills whose rolling success rate drops below60%60\\%are quarantined \(Appendix[G](https://arxiv.org/html/2609.05511#A7)\)\.
### 3\.6Distillation Back into Weights
After each iteration we have𝒵k\+\\mathcal\{Z\}\_\{k\}^\{\+\}, a set of successful trajectories produced byπk\\pi\_\{k\}withthe in\-context libraryℒk\+1\\mathcal\{L\}\_\{k\+1\}\. We convert each such trajectory into aplan\-augmentedtraining example\(I,planσ,ζ\)\(I,\\text\{plan\}\_\{\\sigma\},\\zeta\), whereplanσ\\text\{plan\}\_\{\\sigma\}is the sequence of skill invocations the agent used\. The base policyπk\+1\\pi\_\{k\+1\}is then obtained by supervised fine\-tuning ofπk\\pi\_\{k\}on these triples with a standard token\-level cross\-entropy loss, plus an auxiliary loss that predicts the next skill name conditioned only on\(I,o0:t\)\(I,o\_\{0:t\}\), encouraging the model to internalizewhento invoke each abstraction\. We use LoRA[Hu et al\. \(2022\)](https://arxiv.org/html/2609.05511#bib.bib10)adapters and a constant learning rate of×10−51\\\!\\times\\\!10^\{\-5\}for stability across iterations\. Examples are first re\-parsed against the compacted libraryℒk\+1\\mathcal\{L\}\_\{k\+1\}, with merged calls rewritten through the alias map and trajectories invoking pruned skills dropped \(3\.6%3\.6\\%on average\), so deprecated skills are never reinforced\.
The distillation step is what makesScaffold’s improvement compound: at iterationk\+1k\+1the base policy itself has improved, so the new trajectories it produces \(with the larger library\) push the frontier of solvable tasks higher, yielding richer skills to induce in iterationk\+2k\+2\. This is the recursive self\-improvement loop in the title of the paper, made concrete\.
## 4Experiments
### 4\.1Experiments Setup
#### 4\.1\.1Datasets
We evaluate on three benchmarks chosen to cover both controlled and live web environments:WebArena[Zhou et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib43): 812 long\-horizon tasks across four self\-hosted sites \(E\-commerce, social forum, collaborative development, content management systems\)\. Evaluation is execution\-based with a deterministic verifier per task\.VisualWebArena[Koh et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib12): 910 tasks built on top of WebArena that explicitly require visual reasoning; we use the full benchmark\.Online\-Mind2Web[Xue et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib32): 300 tasks on 136 websites across 12 domains; we partition by sites into a 9\-domain training pool and a 3\-domain held\-out test set \(Jobs & Careers, Travel & Transportation, Government & Services\) in main experiments, enabling a controlled measurement of skill transfer\.
#### 4\.1\.2Baselines
We compareScaffoldagainst three families:\(A\) Strong zero\-shot agents:ReAct[Yao et al\. \(2023\)](https://arxiv.org/html/2609.05511#bib.bib35),SeeAct[Zheng et al\. \(2024a\)](https://arxiv.org/html/2609.05511#bib.bib41),WebVoyager[He et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib8), andAgentOccam[Yang et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib34), all instantiated on top of the same Qwen2\.5\-VL\-7B[Yang et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib33)base for a fair comparison; we additionally report numbers from the stronger UI\-TARS\-7B[Qin et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib16)backbone\.\(B\) Memory\- and workflow\-augmented agents:Reflexion[Shinn et al\. \(2023\)](https://arxiv.org/html/2609.05511#bib.bib19)\(in\-context reflection\),Synapse[Zheng et al\. \(2024b\)](https://arxiv.org/html/2609.05511#bib.bib42)\(trajectory exemplars\),Mem0[Chhikara et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib3)\(long\-term memory\),ExpeL[Zhao et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib39)\(experience distillation\), and Agent Workflow Memory[Wang et al\. \(2025b\)](https://arxiv.org/html/2609.05511#bib.bib25)\.\(C\) Skill\-induction and skill\-RL agents:SkillWeaver[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib40),AppAgentX[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib11),EvolveR[Wu et al\. \(2025a\)](https://arxiv.org/html/2609.05511#bib.bib26),SkillRL[Xia et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib28), andWebRL[Qi et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib15)\. For every baseline we use the same base model, the same task pool, the same number of self\-improvement iterations \(K=5K\{=\}5\), and the same evaluation protocol to isolate algorithmic differences\.
#### 4\.1\.3Evaluation Metrics
We report:\(1\) Success rate \(SR\): fraction of tasks for which the verifier returns 1\.\(2\) Step efficiency: average steps per successful task \(lower is better\)\.\(3\) Library statistics: total skill count\|ℒ\|\|\\mathcal\{L\}\|, mean abstraction depthd¯\\bar\{d\}, skill reuse rate \(fraction of skills called by≥2\\geq 2tasks\)\.\(4\) Cross\-site transfer SR: SR on evaluation sites different from training ones\.\(5\) Iteration scaling: SR as a function ofkk\. To control for environment stochasticity \(dynamic DOM, server load, LLM sampling\), we report mean±\\pmstandard deviation over 3 independent runs with different seeds, and use paired bootstrap tests for significance, markingp<0\.05p<0\.05with†\\dagger\.
Table 1:Main results \(Success Rate %\) on WebArena, VisualWebArena \(VWA\), and Online\-Mind2Web held\-out test split \(OM2W\-X\)\. Mean±\\pmstd over 3 seeds\.†\\dagger:p<0\.05p<0\.05over the strongest competitor \(SkillRL\) by paired bootstrap\. All rows in \(A\)–\(C\) and the boldScaffoldrow use Qwen2\.5\-VL\-7B as the base; the last row is included as an upper reference with a stronger backbone\.
#### 4\.1\.4Implementation Details
Base model is Qwen2\.5\-VL\-7B\-Instruct[Yang et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib33)unless stated otherwise; we also report results with UI\-TARS\-7B[Qin et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib16)as a stronger backbone\. The skill inducer and the MDL refactor proposer are GPT\-4o\-2024\-08, queried with temperature 0\.3\. Trajectories are collected with temperature 0\.7 and a 30\-step horizon\. We usenmin=2n\_\{\\min\}\{=\}2for multi\-instance induction,M=2M\{=\}2for compaction interval,ρ=0\.9\\rho\{=\}0\.9for behavioral\-equivalence threshold,ℓmin=3\\ell\_\{\\min\}\{=\}3andrmin=3r\_\{\\min\}\{=\}3for the refactor operator\. Distillation uses LoRA \(rank 64,α=128\\alpha\{=\}128\) on top of the base model for 1 epoch per iteration\. The learning rate is set as×10−51\\\!\\times\\\!10^\{\-5\}\. All experiments run on 8×\\timesA100\-80GB; a complete 5\-iteration run on WebArena takes nearly 36 hours\. Two protocol points deserve emphasis\. Every skill\- and workflow\-induction baseline \(SkillWeaver,SkillRL,AppAgentX,EvolveR, AWM\) uses the same GPT\-4o\-2024\-08 checkpoint and temperature for its synthesis component, as in the originalSkillWeaversetup, so comparisons isolate algorithmic differences rather than inducer strength; and the 30\-step horizon is counted inprimitiveactions for every method, a skill invocation being a macro action whose primitives draw on the same shared budget \(§[4\.9](https://arxiv.org/html/2609.05511#S4.SS9)\)\.
### 4\.2Main Results
Table[1](https://arxiv.org/html/2609.05511#S4.T1)reports main results\.Scaffoldachieves42\.7%42\.7\\%SR on WebArena,36\.3%36\.3\\%on VisualWebArena, and43\.5%43\.5\\%on the OM2W held\-out test split, improving over the strongest skill\-augmented competitor \(SkillRL\) by11\.111\.1,13\.613\.6, and17\.217\.2absolute points respectively, all statistically significant under paired bootstrap\. Three trends are worth highlighting\. First, skill\-induction and skill\-RL methods systematically outperform both zero\-shot and memory/workflow\-augmented agents on all three benchmarks\. This indicates that explicit, reusable skill abstraction is the right unit of accumulated experience for web agents\. Second, the advantage ofScaffoldon OM2W\-X is most significant, achieving17\.217\.2points overSkillRL, confirming that the recursive parametric skill abstraction coupled with library compaction transfers materially better than two\-tier prompt\-based skills\. Third, replacing the base model with UI\-TARS\-7B yields a further7\.57\.5\-point lift averaged across all three benchmarks, suggesting thatScaffoldis complementary to advances in GUI\-specified model backbones rather than a substitute for them\.
Table 2:Ablation on the four core innovations\. Each row removes a single component while keeping the rest unchanged\. Last row is one of strongest baselines\.
### 4\.3Ablation Study
Table[2](https://arxiv.org/html/2609.05511#S4.T2)isolates each ofScaffold’s four innovations\. Removing distillation has the largest effect \(−8\.0\-8\.0points on WebArena\), confirming our central claim: a prompt\-only skill library, however well structured, is fundamentally limited by the frozen base policy\. Recursive composition is the second\-largest contributor \(−6\.3\-6\.3points on WebArena,−7\.4\-7\.4on OM2W\-X\), which we attribute to its ability to express long\-horizon procedures \(e\.g\.,book\_full\_tripcomposingsearch\_flight,select\_seat,enter\_payment\) without runaway prompt length\. On WebArena, the multi\-instance induction and holdout validation contribute4\.64\.6and3\.73\.7points respectively, smaller absolute effects\. But they are essential for stability across iterations as we show next\. MDL compaction contributes a modest3\.13\.1points but is what keeps the library size bounded \(cf\. Fig\.[3](https://arxiv.org/html/2609.05511#S4.F3)\)\. The drops sum to25\.725\.7against a total gap of13\.613\.6, so the components are synergistic rather than additive\.
### 4\.4Skill Abstraction vs\. Iterative Fine\-Tuning
Since distillation produces the largest ablation gap, one may ask whether the gains come from skill abstraction or merely from iterative fine\-tuning on self\-generated data\. Table[6](https://arxiv.org/html/2609.05511#A5.T6)in Appendix[E](https://arxiv.org/html/2609.05511#A5)decouples the two with the trajectories held fixed\. A STaR\-style loop[Zelikman et al\. \(2022\)](https://arxiv.org/html/2609.05511#bib.bib38)without any library reaches30\.430\.4, since fewer hard tasks are ever solved; abstraction without distillation reaches34\.734\.7, since a frozen policy caps what is reachable however good the library is\. Standard SFT on the identical trajectories with skill calls flattened to primitives reaches38\.638\.6, so4\.14\.1of the remaining points come from the plan\-augmented format and the auxiliary loss that teach the policywhento invokewhichabstraction, and the effect is not LoRA\-specific \(42\.942\.9with full\-parameter tuning\)\. Abstraction keeps distillation paying off; distillation makes abstraction cumulative rather than prompt\-bound\.
Figure 2:Success rate vs\. self\-improvement iteration on WebArena\.Scaffoldcontinues to improve throughK=5K\{=\}5while baselines saturate aroundk=2k\{=\}2–33\. Shaded regions are 1 std over 3 seeds\.
### 4\.5Iteration Scaling
A natural worry with any self\-improvement scheme is that gains saturate quickly\. Figure[2](https://arxiv.org/html/2609.05511#S4.F2)plots SR against iterationkk\. Baselines that maintain a flat skill cache \(AWM, SkillWeaver\) plateau neark=2k\{=\}2, after which library bloat slows further learning\.SkillRLcontinues to improve throughk=3k\{=\}3but flattens thereafter\.Scaffoldcontinues to improve throughk=5k\{=\}5, reflecting two interacting effects: distillation strengthensπk\\pi\_\{k\}, allowing it to solve harder tasks; and those harder tasks furnish new high\-depth skills that, after compaction, expand the library’s expressive reach without diluting it\. We did not run beyondk=5k\{=\}5due to compute, but the slope atk=5k\{=\}5remains positive, suggesting further improvement is achievable\.
Figure 3:Library growth dynamics across iterations on WebArena\.Top: total skill count without compaction \(dashed\) the library balloons; with compaction \(solid\) it stabilizes near 180 skills\.Bottom: mean abstraction depth, which grows monotonically as recursive composition kicks in\.
### 4\.6Library Depth and Reuse
Figure[3](https://arxiv.org/html/2609.05511#S4.F3)shows two complementary growth dynamics\. Thetop panelcontrasts library size with and without MDL compaction: without compaction\|ℒ\|\|\\mathcal\{L\}\|grows almost linearly to≈470\\approx 470skills byk=5k\{=\}5, the majority being near\-duplicates differing only in selectors or argument names \(e\.g\.,login\_v1,login\_v2,login\_with\_email\); with compaction the library follows a concave trajectory and saturates near181181skills, a∼2\.6×\\sim 2\.6\\timesreduction driven primarily by themergeandpruneoperators \(jointly responsible for76%76\\%of removed skills in our run logs\)\. Thebottom panelplots mean depthd¯\\bar\{d\}:Scaffoldclimbs from1\.051\.05atk=1k\{=\}1to2\.712\.71atk=5k\{=\}5\(max depth55\), while the\- recursive compositionablation flatlines atd¯=1\\bar\{d\}\{=\}1andSkillRL’s two\-tier hierarchy plateaus atd¯≤2\\bar\{d\}\{\\leq\}2\. Together with a reuse rate of64%64\\%atk=5k\{=\}5\(vs\.≈21%\\approx 21\\%without compaction; most skills singletons\), these statistics support the view that composition and compaction are two halves of the same mechanism: composition produces the depth that makes long\-horizon skills expressible, while compaction prevents that depth from being undermined by redundancy\. The3\.13\.1\-point SR cost of disabling compaction \(Table[2](https://arxiv.org/html/2609.05511#S4.T2)\) is the empirical price of an uncontrolled library\.
Figure 4:Cross\-site transfer on Online\-Mind2Web\. Cell\(i,j\)\(i,j\)shows SR \(%\) on test sitejjafter training only on siteii\. The ziprecr and ticketm are the abbreviations for ziprecruiter and ticketmaster, respectively\. The sites in redlines are the held\-out test ones in main experiments, also indicated with∗\.
### 4\.7Cross\-Site Transfer
Figure[4](https://arxiv.org/html/2609.05511#S4.F4)compares site\-to\-site transfer for the two methods, beyond the granularity of domain split\. Averaging over the 30 off\-diagonal cells,SkillWeaverreaches16\.6%16\.6\\%mean transfer SR versusScaffold’s35\.9%35\.9\\%, a19\.319\.3\-point gap; on the 15 cells under the three held\-out sites in main experiments,Scaffoldstays in2727–47%47\\%whileSkillWeaverfalls to1111–28%28\\%, with representative pairs amazon→\\toziprecruiter \(38%38\\%vs\.19%19\\%\) and landwatch→\\toticketmaster \(47%47\\%vs\.28%28\\%\) sharing the search→\\tofilter→\\toselect→\\toconfirm structure but differing in DOM and layout\. Inspecting the library shows the mechanism: generic procedural skills such asfilter\_results\(query, sort\_by, max\_price\),login\(site, credentials\), andpaginate\_until\(condition\)carry over via parameter re\-binding, whileSkillWeaver’s opaque APIs encode selectors directly and must be re\-discovered per site\. The github row shows where this still leaves a gap\. GitHub’s issue/PR navigation is structurally unlike e\-commerce flows, soScaffold’s general skills transfer only modestly \(2727–31%31\\%\) andSkillWeaver’s effectively do not \(1111–14%14\\%\)\. This is consistent with the transfer\-ceiling limitation discussed in theLimitationssection\.
### 4\.8Stability Under Stochasticity
Table 3:Stability under environment perturbations on WebArena \(SR %\)\.Clean: unperturbed baseline\.Latency: 0–3s random delay per action\.DOM\-shuf\.: non\-functional DOM attributes randomly permuted on each page load\.\+7 d: same task suite re\-evaluated 7 days later\.Scaffolddegrades by at most 2\.1 points across all regimes, roughly half the sensitivity of the strongest baselines\. Each cell is mean over 3 seeds \(std≤1\.0\\leq 1\.0\)\.Real web environments are noisy: DOMs change between page loads, A/B tests rotate elements, and servers throttle under load\. To probe robustness we re\-ran the WebArena evaluation under three perturbation regimes: \(a\)Latency injection: we add a uniformly sampled 0–3s delay before every action, simulating slow network conditions; \(b\)DOM\-attribute shuffling: we randomly permute non\-functional attributes \(id,class,data\-\*\) on each page load, breaking memorized selectors while preserving semantics; \(c\)Temporal drift: we re\-run the same task suite 7 days later, capturing whatever natural drift WebArena’s containerized sites exhibit between snapshots\. Table[3](https://arxiv.org/html/2609.05511#S4.T3)reports each method’s success rate under all three regimes\.
The pattern in Table[3](https://arxiv.org/html/2609.05511#S4.T3)is consistent across all three perturbations:Scaffolddegrades by at most2\.12\.1points \(DOM\-shuffle\), compared to3\.63\.6–4\.44\.4forSkillWeaverandSkillRL, roughly half the sensitivity\. We attribute this to two structural properties of our skills\. First,parametric grounding: a skill that locates an element by semantic role \(“the search input”\) is robust to attribute renaming, whereas a skill that memorized a specificidselector fails the moment thatidis rewritten\. Second,postcondition checks: even if the body of a skill misfires under latency or transient state, the postcondition test triggers a retry rather than blindly committing to a wrong final state\. The DOM\-shuffle column is where the gap is widest, which is exactly what these two mechanisms are designed to handle\. We view this as evidence that the structural choices inScaffoldare not just accuracy\-oriented but also confer a non\-trivial robustness benefit at no additional inference cost\.
### 4\.9Action Budget and Inference Cost
Because a skill invocation executes several primitive actions, one might worry thatScaffoldenjoys a larger effective action budget\. It does not: the horizon is counted in primitive environment actions for every method, and primitives inside a skill body draw on the same budget \(§[4\.1\.4](https://arxiv.org/html/2609.05511#S4.SS1.SSS4)\)\. Average primitive actions per successful task is12\.412\.4forScaffoldagainst14\.314\.3forSkillRL,15\.115\.1forSkillWeaverand16\.916\.9for AWM, soScaffoldneeds fewer actions, not more; capping the budget at1010/2020/3030steps gives11\.511\.5/24\.024\.0/29\.129\.1,13\.213\.2/26\.426\.4/31\.631\.6and19\.819\.8/36\.936\.9/42\.742\.7respectively, so the margin is widest at the tightest cap, the opposite of what budget inflation predicts\.
Skills also trade expensive planning calls for cheaper ones: per successful taskScaffoldissues5\.35\.3policy calls against9\.69\.6forSkillWeaverand18\.718\.7forReAct, plus7\.17\.1far cheaper grounder and4\.64\.6judge calls, so total tokens fall from224224K and118118K to7979K and wall clock from102102s and7474s to6161s\. Induction is offline and amortized \(about 500 calls per run\)\. Appendix[F](https://arxiv.org/html/2609.05511#A6)gives breakdown and a selector cache that removes62%62\\%grounder calls\.
### 4\.10Inducer and Verifier Dependence
Two components sit outside the loop\. Swapping the GPT\-4o inducer for the open\-weight Qwen2\.5\-72B\-Instruct with everything else fixed gives40\.840\.8SR on WebArena,1\.91\.9points below GPT\-4o and still9\.29\.2aboveSkillRLwith the GPT\-4o inducer, so the framework does not hinge on proprietary\-API access\. Replacing the ground\-truth verifier with a model\-based judge for filtering, following the WebJudge protocol of[Xue et al\. \(2025\)](https://arxiv.org/html/2609.05511#bib.bib32), gives40\.240\.2on WebArena and41\.041\.0on OM2W\-X against42\.742\.7and43\.543\.5\. Multi\-instance induction and holdout re\-execution filter much of this label noise, soScaffolddegrades gracefully as verification weakens\.
## 5Conclusions and Future Work
In this paper, we presentedScaffold, a self\-improving framework for web agents that combines multi\-instance parametric skill induction, recursive hierarchical composition, MDL\-driven library compaction, and weight\-level distillation into a single closed loop\. On WebArena, VisualWebArena, and a held\-out Online\-Mind2Web split,Scaffoldoutperforms the strongest skill\-induction and skill\-RL baselines by11\.111\.1–17\.217\.2absolute points while continuing to improve through five iterations and keeping the library size bounded\. In the future, we can consider such directions: \(i\) replacing the GPT\-4o\-based inducer with a self\-trained, lighter induction model to remove proprietary\-API dependence; \(ii\) extending the framework to OSWorld\-style full\-desktop agents where the action space includes file I/O and shell commands; \(iii\) integrating an environment\-grounded verifier that does not rely on per\-task ground truth, enabling fully autonomous deployment; and \(iv\) developing a formal convergence analysis of the MDL\-compaction operator under the recursive composition regime\.
## Limitations
We identify five concrete limitations ofScaffoldthat we believe warrant attention\.\(1\) Inducer dependence\.Both skill induction and the refactor proposer of the MDL compactor rely on a strong proprietary model \(GPT\-4o\); we estimate the inducer accounts for∼70%\\sim 70\\%of total dollar cost in our experiments\. Replacing it with a self\-trained inducer is a clear next step but was outside our compute budget\.\(2\) Verifier dependence\.Scaffoldassumes access to a per\-task verifierVVto label trajectory success during the rollout stage\. While this is realistic for WebArena and VisualWebArena \(which ship deterministic verifiers\), it does not extend to fully autonomous deployment in the wild, where success is rarely binary or self\-evident\.\(3\) Greedy MDL approximation\.The compaction step optimizes Eq\.[2](https://arxiv.org/html/2609.05511#S3.E2)greedily; we have no convergence guarantee and observe occasional plateaus where a non\-local refactor would further reduceℱ\\mathcal\{F\}\. A more principled approximation, e\.g\., simulated annealing over skill rewrites, may help\.\(4\) Cross\-site transfer ceiling\.Although parametric skills generalize meaningfully better than site\-specific APIs \(§[4\.7](https://arxiv.org/html/2609.05511#S4.SS7)\), our gains shrink when held\-out sites differ qualitatively from training ones \(e\.g\., training only on e\-commerce, testing on government portals\)\. The framework reusesprocedures, notweb ontology, so radical distributional shift remains hard\.\(5\) Distillation forgetting\.LoRA\-based distillation preserves base capabilities reasonably well, but we observe a small \(≤1\.2\\leq 1\.2point\) regression on a subset of held\-out tasks that the originalπ0\\pi\_\{0\}could already solve\. A replay buffer of original task supervision could mitigate this; we leave a careful study to future work\.
## Ethical Considerations
Scaffoldenables web agents to learn reusable procedural skills through autonomous exploration\. We highlight three ethical considerations\.Misuse: more capable web agents can be deployed for spam, scraping at scale, credential stuffing, or evading rate limits\. Our experiments use only WebArena, VisualWebArena, and a small Online\-Mind2Web subset on permitted sites; we do not provide site\-specific bypass skills\.Bias amplification through self\-improvement: any bias in the induced skills \(e\.g\., always defaulting to certain payment methods or geographic regions\) is reinforced across iterations through distillation\. Future deployments should audit the skill library at each iteration\.Energy cost: a full 5\-iteration run on a single benchmark consumes∼\\sim36 A100\-hours\. We report this transparently and discourage hyperparameter searches that re\-run the full loop unnecessarily\.
## References
- Chen et al\. \(2024\)Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu\. 2024\.Self\-play fine\-tuning converts weak language models to strong language models\.In*International Conference on Machine Learning*, pages 6621–6642\. PMLR\.
- Cheng et al\. \(2024\)Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Li YanTao, Jianbing Zhang, and Zhiyong Wu\. 2024\.Seeclick: Harnessing GUI grounding for advanced visual GUI agents\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 9313–9332\.
- Chhikara et al\. \(2025\)Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav\. 2025\.Mem0: Building production\-ready AI agents with scalable long\-term memory\.*arXiv preprint arXiv:2504\.19413*\.
- Ellis et al\. \(2021\)Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sablé\-Meyer, Lucas Morales, Luke Hewitt, Luc Cary, Armando Solar\-Lezama, and Joshua B Tenenbaum\. 2021\.DreamCoder: Bootstrapping inductive program synthesis with wake\-sleep library learning\.In*Proceedings of the 42nd acm sigplan international conference on programming language design and implementation*, pages 835–850\.
- Fang et al\. \(2026\)Gaodan Fang, Vatche Isahagian, KR Jayaram, Ritesh Kumar, Vinod Muthusamy, Punleuk Oum, and Gegi Thomas\. 2026\.Trajectory\-informed memory generation for self\-improving agent systems\.*arXiv preprint arXiv:2603\.10600*\.
- Gandhi and Neubig \(2026\)Apurva Gandhi and Graham Neubig\. 2026\.Go\-browse: Training web agents with structured exploration\.In*International Conference on Learning Representations*, volume 2026, pages 138973–138992\.
- Guo et al\. \(2025\)Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, and 1 others\. 2025\.DeepSeek\-R1: Incentivizing reasoning capability in LLMs via reinforcement learning\.*arXiv preprint arXiv:2501\.12948*\.
- He et al\. \(2024\)Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu\. 2024\.WebVoyager: Building an end\-to\-end web agent with large multimodal models\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 6864–6890\.
- Hong et al\. \(2024\)Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and 1 others\. 2024\.CogAgent: A visual language model for GUI agents\.In*2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 14281–14290\. IEEE\.
- Hu et al\. \(2022\)Edward J\. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\. 2022\.[Lora: Low\-rank adaptation of large language models](https://openreview.net/forum?id=nZeVKeeFYf9)\.In*The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\-29, 2022*\. OpenReview\.net\.
- Jiang et al\. \(2025\)Wenjia Jiang, Yangyang Zhuang, Chenxi Song, Xu Yang, Joey Tianyi Zhou, and Chi Zhang\. 2025\.AppAgentX: Evolving GUI agents as proficient smartphone users\.*arXiv preprint arXiv:2503\.02268*\.
- Koh et al\. \(2024\)Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po\-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried\. 2024\.VisualWebArena: Evaluating multimodal agents on realistic visual web tasks\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 881–905\.
- Lai et al\. \(2024\)Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, and 1 others\. 2024\.AutoWebGLM: A large language model\-based web navigating agent\.In*Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining*, pages 5295–5306\.
- Patel et al\. \(2024\)Ajay Patel, Markus Hofmarcher, Claudiu Leoveanu\-Condrei, Marius\-Constantin Dinu, Chris Callison\-Burch, and Sepp Hochreiter\. 2024\.Large language models can self\-improve at web agent tasks\.*arXiv preprint arXiv:2405\.20309*\.
- Qi et al\. \(2025\)Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Jiadai Sun, Xinyue Yang, Yu Yang, Shuntian Yao, Wei Xu, and 1 others\. 2025\.WebRL: Training LLM web agents via self\-evolving online curriculum reinforcement learning\.In*International Conference on Learning Representations*, volume 2025, pages 79791–79821\.
- Qin et al\. \(2025\)Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, and 1 others\. 2025\.UI\-TARS: Pioneering automated GUI interaction with native agents\.*arXiv preprint arXiv:2501\.12326*\.
- Rissanen \(1978\)Jorma Rissanen\. 1978\.Modeling by shortest data description\.*Automatica*, 14\(5\):465–471\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others\. 2024\.DeepSeekMath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*\.
- Shinn et al\. \(2023\)Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\. 2023\.Reflexion: Language agents with verbal reinforcement learning\.*Advances in neural information processing systems*, 36:8634–8652\.
- Sutton et al\. \(1999\)Richard S Sutton, Doina Precup, and Satinder Singh\. 1999\.Between MDPs and semi\-MDPs: A framework for temporal abstraction in reinforcement learning\.*Artificial intelligence*, 112\(1\-2\):181–211\.
- Wang et al\. \(2024a\)Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar\. 2024a\.Voyager: An open\-ended embodied agent with large language models\.*Transactions on Machine Learning Research*\.
- Wang et al\. \(2025a\)Haoming Wang, Haoyang Zou, Huatong Song, Jiazhan Feng, Junjie Fang, Junting Lu, Longxiang Liu, Qinyu Luo, Shihao Liang, Shijue Huang, and 1 others\. 2025a\.UI\-TARS\-2 technical report: Advancing GUI agent with multi\-turn reinforcement learning\.*arXiv preprint arXiv:2509\.02544*\.
- Wang et al\. \(2026\)Jiongxiao Wang, Qiaojing Yan, Yawei Wang, Yijun Tian, Soumya Smruti Mishra, Zhichao Xu, Megha Gandhi, Panpan Xu, and Lin Lee Cheong\. 2026\.Reinforcement learning for self\-improving agent with skill library\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 1529–1550\.
- Wang et al\. \(2024b\)Zhiruo Wang, Graham Neubig, and Daniel Fried\. 2024b\.Trove: Inducing verifiable and efficient toolboxes for solving programmatic tasks\.In*International Conference on Machine Learning*, pages 51177–51191\. PMLR\.
- Wang et al\. \(2025b\)Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig\. 2025b\.Agent workflow memory\.In*International Conference on Machine Learning*, pages 63897–63911\. PMLR\.
- Wu et al\. \(2025a\)Rong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai, Daocheng Fu, Cheng Yang, Licheng Wen, Xuemeng Yang, Yufan Shen, Yuxin Wang, and 1 others\. 2025a\.EvolveR: Self\-evolving LLM agents through an experience\-driven lifecycle\.*arXiv preprint arXiv:2510\.16079*\.
- Wu et al\. \(2025b\)Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and 1 others\. 2025b\.Os\-atlas: Foundation action model for generalist GUI agents\.In*International Conference on Learning Representations*, volume 2025, pages 5090–5108\.
- Xia et al\. \(2026\)Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, and 1 others\. 2026\.Skillrl: Evolving agents via recursive skill\-augmented reinforcement learning\.In*ICLR 2026 Workshop on Lifelong Agents: Learning, Aligning, Evolving*\.
- Xie et al\. \(2024\)Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, and 1 others\. 2024\.OSWorld: Benchmarking multimodal agents for open\-ended tasks in real computer environments\.*Advances in Neural Information Processing Systems*, 37:52040–52094\.
- Xu et al\. \(2025\)Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong\. 2025\.Aguvis: Unified pure vision agents for autonomous GUI interaction\.In*International Conference on Machine Learning*, pages 69772–69805\. PMLR\.
- Xu et al\. \(2026\)Zishan Xu, Yifu Guo, Yuquan Lu, Fengyu Yang, Zhiyuan Yao, Jiaye Lin, Ruyi Gong, and Lihua Cai\. 2026\.Skillevo: An experience learning framework with reinforcement learning for skill evolution\.
- Xue et al\. \(2025\)Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su\. 2025\.An illusion of progress? assessing the current state of web agents\.*arXiv preprint arXiv:2504\.01382*\.
- Yang et al\. \(2024\)An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, and 1 others\. 2024\.Qwen2\.5 technical report\.*arXiv preprint arXiv:2412\.15115*\.
- Yang et al\. \(2025\)Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik A Chaudhari, George Karypis, and Huzefa Rangwala\. 2025\.AgentOccam: A simple yet strong baseline for LLM\-based web agents\.In*International Conference on Learning Representations*, volume 2025, pages 97533–97565\.
- Yao et al\. \(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao\. 2023\.ReAct: Synergizing reasoning and acting in language models\.In*International Conference on Learning Representations*\.
- Yu et al\. \(2026\)Simon Yu, Gang Li, Weiyan Shi, and Peng Qi\. 2026\.Polyskill: Learning generalizable skills through polymorphic abstraction for continual learning\.In*International Conference on Learning Representations*, volume 2026, pages 140298–140326\.
- Zelikman et al\. \(2024\)Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman\. 2024\.Quiet\-STaR: Language models can teach themselves to think before speaking\.*arXiv preprint arXiv:2403\.09629*\.
- Zelikman et al\. \(2022\)Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman\. 2022\.STaR: Bootstrapping reasoning with reasoning\.*Advances in Neural Information Processing Systems*, 35:15476–15488\.
- Zhao et al\. \(2024\)Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong\-Jin Liu, and Gao Huang\. 2024\.ExpeL: LLM agents are experiential learners\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 38, pages 19632–19642\.
- Zheng et al\. \(2025\)Boyuan Zheng, Michael Y Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, and 1 others\. 2025\.Skillweaver: Web agents can self\-improve by discovering and honing skills\.*arXiv preprint arXiv:2504\.07079*\.
- Zheng et al\. \(2024a\)Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su\. 2024a\.Gpt\-4v \(ision\) is a generalist web agent, if grounded\.In*International Conference on Machine Learning*, pages 61349–61385\. PMLR\.
- Zheng et al\. \(2024b\)Longtao Zheng, Rundong Wang, Xinrun Wang, and Bo An\. 2024b\.Synapse: Trajectory\-as\-exemplar prompting with memory for computer control\.In*International Conference on Learning Representations*, volume 2024, pages 19036–19066\.
- Zhou et al\. \(2024\)Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, and 1 others\. 2024\.WebArena: A realistic web environment for building autonomous agents\.In*International Conference on Learning Representations*, volume 2024, pages 15585–15606\.
## Appendix APrompt Templates
We provide the three core prompts used byScaffoldbelow\. Full prompts \(including system messages and few\-shot examples\) are released with the code\.
Skill inducer \(multi\-instance abstraction\)\.Given a clusterC=\{ζ1,…,ζn\}C=\\\{\\zeta\_\{1\},\\ldots,\\zeta\_\{n\}\\\}ofn≥2n\\geq 2successful trajectories with embedding\-similar instructions, the inducer is prompted to produce a parameterized skill\. The prompt structure is:
> You are givennnsuccessful agent trajectories that accomplish similar goals on a website\. Your task is to abstract them into a singleparametric, executableskill\. \[Trajectory 1: instruction, action sequence, screenshots\] … \[Trajectory n: instruction, action sequence, screenshots\] Available existing skills \(you may call them\): \{ℒk\\mathcal\{L\}\_\{k\}\} Output a Python function with: \(a\) a typed parameter list capturing what varies across trajectories; \(b\) a precondition expressed as a predicate over observation obs; \(c\) a body using primitives \{click, type, scroll, wait\} or any existing skill; \(d\) a postcondition\.
The model is then asked to justify its choice of parameters by pointing to specific token positions in each trajectory that vary, which improves abstraction quality and gives us a cheap diagnostic when induction fails\.
Refactor proposer \(MDL compaction\)\.Given a recurring primitive\-action subsequence detected across≥rmin\\geq r\_\{\\min\}existing skills, the refactor proposer is asked whether the subsequence should be promoted to its own skill:
> The following action subsequence appears inmmexisting skills \(shown below\)\. Decide whether to refactor it into a new mid\-level skill\. If yes, name it, parametrize it, and rewrite allmmexisting skills to call the new one\. If no, briefly explain why \(e\.g\., the subsequence is too short or too site\-specific to be reusable\)\.
We accept the proposal only if \(i\) the MDL functionalℱ\\mathcal\{F\}decreases and \(ii\) the rewritten skills pass behavioral\-equivalence checks on a held\-out trajectory set\.
Pre/Postcondition validator\.At runtime, before invoking a skill we verify its precondition by passing\(obst,preσ\)\(\\text\{obs\}\_\{t\},\\text\{pre\}\_\{\\sigma\}\)to a lightweight VLM judge with a yes/no output\. Same protocol for postconditions after execution\. The judge is a separate, cheaper model from the inducer to keep inference cost bounded\.
Trajectory clustering\.Instructions are embedded withgte\-large\-en\-v1\.5\. Trajectories are first partitioned by site and then clustered by agglomerative clustering with average linkage on cosine similarity, using a merge threshold of0\.820\.82\. Clusters larger than 8 trajectories are subsampled by proximity to the centroid so that the inducer context stays bounded, and singleton clusters \(belownminn\_\{\\min\}\) are held in a buffer and re\-clustered in later iterations as semantically similar trajectories accumulate, so no successful trajectory is permanently discarded\. Sensitivity to the threshold is mild, within11SR point across the range0\.780\.78to0\.860\.86\.
## Appendix BA Worked Three\-Level Recursive Skill
To make the recursive composition concrete, we trace the construction of a depth\-3 skillcheckout\_cheapest\_in\_categoryon the E\-commerce site of WebArena, induced at iterationk=3k\{=\}3\.
##### Level 1 \(depthd=1d\{=\}1\): atomic skills, induced atk=1k\{=\}1\.
```
def click_text(obs, text: str):
pre: any(e.text == text for e in obs.dom)
body: el = find(obs.dom, text=text)
primitive_click(el.bbox)
post: page changed OR el.state == ’pressed’
def type_in_field(obs, label: str, value: str):
pre: exists_input(obs.dom, label)
body: f = find_input(obs.dom, label)
primitive_click(f.bbox)
primitive_type(value)
post: f.value == value
```
##### Level 2 \(depthd=2d\{=\}2\): mid\-level skills, induced atk=2k\{=\}2\.
```
def search_and_filter(obs, query: str,
category: str,
sort: str = ’price_asc’):
pre: url contains ’shop’ AND has_search_box(obs)
body: type_in_field(obs, ’Search’, query)
click_text(obs, ’Search’)
click_text(obs, category)
click_text(obs, f’Sort: {sort}’)
post: results visible AND sort_label == sort
```
##### Level 3 \(depthd=3d\{=\}3\): composite skill, induced atk=3k\{=\}3\.
```
def checkout_cheapest_in_category(
obs, query: str, category: str,
payment_token: str):
pre: logged_in(obs)
body: search_and_filter(obs, query, category,
sort=’price_asc’)
click_text(obs, ’first_result’)
click_text(obs, ’Add to Cart’)
click_text(obs, ’Checkout’)
type_in_field(obs, ’PaymentToken’,
payment_token)
click_text(obs, ’Place Order’)
post: order_confirmed(obs) AND
purchased_item.category == category
```
The depth\-3 skill calls one depth\-2 skill \(search\_and\_filter\) and two depth\-1 skills \(click\_text,type\_in\_field\)\. At iterationk=4k\{=\}4, the inducer further abstracts this skill into a depth\-4restock\_pantry\(items, budget\)that callscheckout\_cheapest\_in\_categoryin a loop, illustrating how depth grows organically across iterations\.
Table 4:Per\-site success rate \(%\) on WebArena[Zhou et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib43)\(four primary sites\) and VisualWebArena[Koh et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib12)\.Scaffold’s largest per\-site gains occur on long\-horizon, multi\-step sites \(GitLab, CMS, VWA Shopping\), where recursive composition contributes most\. Differences are bootstrap\-significant \(p<0\.05p<0\.05\) for all rows\.
## Appendix CPer\-Site Analysis
Table[4](https://arxiv.org/html/2609.05511#A2.T4)breaks down performance by the canonical site grouping used in each benchmark\. WebArena’s four primary sites cover e\-commerce \(OneStopShop\), social\-forum discussions \(Reddit\), collaborative software development \(GitLab\), and content management \(CMS\)[Zhou et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib43); VisualWebArena spans three live\-style sites \(Classifieds, Reddit, Shopping\)[Koh et al\. \(2024\)](https://arxiv.org/html/2609.05511#bib.bib12)\.Scaffold’s gains are largest onGitLabandCMS, the WebArena sites with the longest action horizons, dominated by multi\-step configuration and code\-collaboration tasks\. The gain is also obvious on VWA’sShopping, where parametric search/filter/checkout skills compose well\. Gains onRedditare more modest, consistent with social\-forum tasks being on average shorter and more retrieval\-oriented, where flat libraries already perform well\.
Table 5:Sensitivity toScaffold’s main hyperparameters\.Δ\\DeltaSR is relative to the default setting \(bold values in the experiments section:nmin=2n\_\{\\min\}\{=\}2,M=2M\{=\}2,ρ=0\.9\\rho\{=\}0\.9, rank=64\)\.##### Qualitative diagnoses\.
Inspecting failure modes shared by both methods reveals two patterns\. First,*verifier brittleness*: a non\-trivial fraction of WebArena tasks reward only exact\-string answers, so a semantically correct trajectory can be marked wrong; this affects both systems equally\. Second,*visual occlusion under modals*: when a cookie banner or login dialog covers the target element, both agents struggle; hereScaffold’s precondition checks help but do not fully solve the problem, suggesting that an explicit modal\-dismissal mid\-level skill would be a high\-value addition to the library\.
## Appendix DHyperparameter Sensitivity
We probe sensitivity to the fourScaffold\-specific hyperparameters; full sweep results in the supplementary material\. Table[5](https://arxiv.org/html/2609.05511#A3.T5)shows that the method is broadly robust: performance varies by≤2\.4\\leq 2\.4SR points across the tested ranges, with the multi\-instance thresholdnminn\_\{\\min\}being the most sensitive \(smaller values admit spurious skills; larger values delay skill creation\)\.
## Appendix ESupervision Format and Component Build\-Up
Table 6:Supervision\-format study on WebArena \(SR %\)\. All rows share the base model, the rollout pipeline and the exact same successful trajectories, differing only in how those trajectories supervise the policy\.Table 7:Incremental build\-up on WebArena \(SR %\), starting fromSkillWeaverand adding one component at a time\. Every addition is bootstrap\-significant \(p<0\.05p<0\.05\)\. The intermediate point39\.639\.6coincides with the “−\-MDL compaction” row of Table[2](https://arxiv.org/html/2609.05511#S4.T2)by construction\.Table[2](https://arxiv.org/html/2609.05511#S4.T2)removes one component at a time from the full system; Table[7](https://arxiv.org/html/2609.05511#A5.T7)takes the opposite direction and adds one component at a time on top ofSkillWeaver, which makes the marginal value of each component visible in isolation from the others\. The two views agree, and together they support reading the four components as one mechanism rather than four independent add\-ons: induction is the abstraction phase, MDL compaction the compression phase, distillation the consolidation phase, and multi\-instance validation is what keeps abstraction sound, in the spirit of wake\-sleep library learning[Ellis et al\. \(2021\)](https://arxiv.org/html/2609.05511#bib.bib4)\. The components also interact by design\. Figure[3](https://arxiv.org/html/2609.05511#S4.F3)shows that composition creates the depth that makes long\-horizon skills expressible while compaction keeps that depth from being buried in redundancy, and §[3\.5](https://arxiv.org/html/2609.05511#S3.SS5)explains why compaction is also what protects the distillation signal from dilution\. All configurations hold the base model, task pool, iteration count, inducer, and evaluation protocol fixed, so the differences cannot be attributed to engineering scale or extra compute\.
## Appendix FEfficiency and Grounding Measurements
Table 8:WebArena SR \(%\) under a capped primitive\-action budget \(§[4\.9](https://arxiv.org/html/2609.05511#S4.SS9)\)\. The margin is largest at the tightest cap, the opposite of what an inflated action budget would produce\.Table 9:Inference cost per successful WebArena task: policy, grounder and judge calls, total tokens and wall clock\. A policy call costs about 12K tokens \(screenshot, DOM and history\), a grounder call about 1\.5K, and a judge call about 0\.8K\.Tables[8](https://arxiv.org/html/2609.05511#A6.T8)and[9](https://arxiv.org/html/2609.05511#A6.T9)report the two measurements summarised in §[4\.9](https://arxiv.org/html/2609.05511#S4.SS9): success rate under a capped primitive\-action budget, and the per\-task call, token and wall\-clock breakdown\. ChargingScaffoldone extra primitive step for every skill invocation, a pessimistic accounting of invocation overhead, still leaves it at42\.342\.3\. The rest of this appendix asks how far the residual model calls inside a skill body can be removed\.
Table 10:Grounding modes on WebArena \(SR %\) and grounder calls per successful task\. The hybrid cache removes62%62\\%of grounder calls at essentially no cost in accuracy or robustness, whereas static selectors reintroduce baseline brittleness under DOM\-shuffle\.Semantic grounding is the one place whereScaffoldcalls a model inside a skill body, so it is worth asking how much of it can be replaced by cheaper machinery such as regular expressions over the accessibility tree\. We implement a hybrid mode withselector memoization: the first successful grounding of a semantic reference caches the resolved accessibility\-tree path together with a structural signature \(role, tag, text pattern\), keyed by site, skill and parameter; subsequent invocations first try the cached selector with a cheap regex and accessibility\-tree validation, and fall back to VLM grounding only on mismatch\. Table[10](https://arxiv.org/html/2609.05511#A6.T10)compares the three regimes\. The hybrid retains clean and perturbed accuracy while removing most grounder calls, and the static\-selector variant makes the trade\-off visible: it is the cheapest and the most brittle, losing7\.17\.1points under DOM\-shuffle, which confirms that runtime semantic grounding is the source of the robustness gap in Table[3](https://arxiv.org/html/2609.05511#S4.T3)\. One caveat on the premise that skills are site\-specific: after compaction the most valuable skills are not, given the64%64\\%reuse rate and the transfer results of §[4\.7](https://arxiv.org/html/2609.05511#S4.SS7), and those skills cannot be bound to any single site’s selectors\. The cache gives them per\-site fast paths while preserving transfer\.
## Appendix GSkill Lifecycle and Distillation Hygiene
This appendix expands the two mechanisms sketched in §[3\.5](https://arxiv.org/html/2609.05511#S3.SS5)and §[3\.6](https://arxiv.org/html/2609.05511#S3.SS6)\.
Runtime failure handling\.Preconditions gate invocation: when a precondition is unmet, for instance because an unexpected pop\-up occludes the target, the skill is simply not fired and the policy continues with primitives or an alternative skill\. Postconditions are checked after execution; a violation triggers one retry after a state refresh, after which control returns to the policy with the failure noted in context so that the episode can still recover\. All failures are logged together with the observation and feed the next induction round\.
Outdated or degrading skills\.Beyond the prune operator, which removes skills unused forMMiterations, the implementation tracks a rolling execution success rate per skill\. A skill that falls below60%60\\%over its last 20 invocations is quarantined out of the retrieval index and queued for re\-induction from fresh trajectories, which naturally repairs drift because new trajectories reflect the updated page\. WebArena’s containerized sites drift little, as the\+7\+7d column of Table[3](https://arxiv.org/html/2609.05511#S4.T3)shows, so this mechanism matters mainly for live deployment\.
Modal occlusion\.Following the diagnosis in Appendix[C](https://arxiv.org/html/2609.05511#A3), we added an explicitdismiss\_blocking\_modalmid\-level skill in a separate run: WebArena SR improves by0\.90\.9points and modal\-related failures fall from6\.1%6\.1\\%to2\.3%2\.3\\%of episodes\. This is a case where reading the library’s failure log suggests the missing abstraction directly\.
Distillation hygiene\.Three properties keep supervision current\. Trajectories are re\-parsed against the post\-compaction library before training, as described in §[3\.6](https://arxiv.org/html/2609.05511#S3.SS6), so the model is never supervised toward deprecated skills; the auxiliary skill\-name loss uses the current namespace as its label space, so retired names cannot be reinforced; and each iteration trains on freshly collected trajectories rather than accumulating stale data\. Consistent with Limitation \(5\), we still observe at most a1\.21\.2\-point regression on previously solved tasks, and a small replay buffer of re\-parsed earlier trajectories removes most of it in a preliminary run\.
Table 11:Component\-level comparison with the two concurrent skill\-centric frameworks discussed in §[2](https://arxiv.org/html/2609.05511#S2)\.
## Appendix HComparison with Concurrent Frameworks
Table[11](https://arxiv.org/html/2609.05511#A7.T11)placesScaffoldbesidePolySkill[Yu et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib36)andSkillEvo[Xu et al\. \(2026\)](https://arxiv.org/html/2609.05511#bib.bib31)along the four axes that motivate our design\.PolySkillshares our generalization goal and reaches it by maintaining per\-site implementations under a shared interface, whereas our skills keep one implementation and re\-bind semantic references at run time; its library is prompt\-side only, grows without a compression objective, and induces skills without a multi\-instance support requirement\.SkillEvoimproves the policy with GRPO against a learned reasoning\-and\-execution reward model, which is closer in spirit toSkillRLthan to our SFT\-based consolidation, and its skill path graph evolves but is not governed by a compression objective with behavioral\-equivalence merging and a validation safety net\.
Its reported numbers are also not directly comparable to Table[1](https://arxiv.org/html/2609.05511#S4.T1), since it is evaluated with text\-only LLMs \(Llama\-3\.1\-8B, GLM\-4\-9B\) on WebArena\-Lite \(165 tasks\), while we target visual web agents on full WebArena \(812 tasks\), VisualWebArena and live websites\. As a reference point we ranScaffoldon WebArena\-Lite under our visual setting and obtained63\.863\.8SR with Qwen2\.5\-VL\-7B, against the60\.460\.4reported forSkillEvowith Llama\-3\.1\-8B, with the explicit caveat that backbone and observation space differ\. Finally, the two lines of work are complementary rather than competing:PolySkill’s interface typing could serve as the type system for our parameter lists, andSkillEvo’s fine\-grained reward model could replace our binary success filter, which would relax the verifier dependence discussed in §[4\.10](https://arxiv.org/html/2609.05511#S4.SS10)\.Similar Articles
Self-Evolving Embodied Agents via Skill-Harness Evolution
This paper introduces SHAPER, a self-evolving framework for embodied agents that keeps model parameters frozen and improves performance by evolving reusable skills and context-code harnesses through target-environment rollouts. Evaluated on VLABench and ESI-Bench, it proposes a practical alternative to fine-tuning when training is expensive or unavailable.
Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs
This paper proposes scaffold-mediated post-training, a paradigm where procedural scaffolds co-evolve with LLM parameters through discovery, distillation, and dynamic recompilation. On FeatureBench, automatically discovered skills improve pass rate by 8.1pp, with a 27.7% pass rate after distillation.
ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research
ScaffoldAgent introduces a utility-guided dynamic outline optimization framework for open-ended deep research, using expansion, contraction, and revision operations to improve long-form report generation and factual grounding.
@omarsar0: // Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today …
This paper introduces Self-Harness, a new paradigm where LLM-based agents iteratively improve their own operating harness—prompts, tools, and control flow—without human engineers or stronger external agents, achieving significant performance gains across multiple models.
Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
Proposes SCALE, a framework for self-improving web agents using cognitive-aware exploration with three adversarial roles and a graph exploration strategy. Also introduces a large-scale dataset SCALE-20k from real websites, showing significant improvements in MLLM-based web agents.