DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

arXiv cs.AI Papers

Summary

DRIVE proposes a dual-level skill modeling framework that separates reasoning knowledge from interaction knowledge for web agents under continual learning, achieving a 52.8% task success rate on WebArena, outperforming the skill-free baseline by 7.3 percentage points.

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different tasks. However, these knowledge types differ fundamentally: reasoning knowledge (e.g., booking a flight requires first searching for routes) is abstract and transferable across websites, while interaction knowledge (e.g., clicking the Search button at a specific coordinate on Site A) depends heavily on page-specific contexts. Existing methods store experiences uniformly. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface-level differences or attempt infeasible actions from outdated page structures. To disentangle them, we propose DRIVE, a dual-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations. A scene-aware coordination mechanism adaptively retrieves and invokes these dual-level skills based on task semantics. DRIVE also uses skill-level reflection to identify hierarchy-specific failure modes, enabling targeted skill library expansion and refinement. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52.8%, exceeding the skill-free baseline by 7.3 percentage points. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page-level operations.
Original Article
View Cached Full Text

Cached at: 05/26/26, 09:01 AM

# DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning
Source: [https://arxiv.org/html/2605.23939](https://arxiv.org/html/2605.23939)
Sihang Zhou[sihangjoe@gmail\.com](https://arxiv.org/html/2605.23939v1/mailto:[email protected])Yanning HouRong ZhouHaoyuan ChenMaolin HeSiwei WangHao ChenJian Huang

###### Abstract

Web agents require both high\-level reasoning \(for task decomposition\) and low\-level interactions \(for page elements manipulation\) to conduct different tasks\. However, these knowledge types differ fundamentally: reasoning knowledge \(e\.g\., booking a flight requires first searching for routes\) is abstract and transferable across websites, while interaction knowledge \(e\.g\., clicking the Search button at a specific coordinate on Site A\) depends heavily on page\-specific contexts\. Existing methods store experiences uniformly\. This creates a dilemma: abstract representations lose executability on concrete pages, while concrete representations fail to generalize across domains\. This entanglement limits capability accumulation: on new websites, agents either fail to recognize reusable task logic due to surface\-level differences or attempt infeasible actions from outdated page structures\. To disentangle them, we propose DRIVE, a dual\-level skill modeling framework separating historical experience into natural language reasoning skills, which capture transferable task logic, and programmatic interaction skills, grounding abstract actions to executable operations\. A scene\-aware coordination mechanism adaptively retrieves and invokes these dual\-level skills based on task semantics\. DRIVE also uses skill\-level reflection to identify hierarchy\-specific failure modes, enabling targeted skill library expansion and refinement\. Experiments across five WebArena domains show DRIVE attains an average task success rate of 52\.8%, exceeding the skill\-free baseline by 7\.3 percentage points\. Further ablations show reasoning and interaction skills provide distinct, complementary benefits, supporting separation of transferable task logic from executable page\-level operations\.

###### keywords:

Web agents; Web automation; Procedural memory; Skill induction; Continual learning; Failure attribution

††journal:Pattern Recognition\\affiliation

\[1\] organization=College of Intelligence Science and Technology, National University of Defense Technology, city=Changsha, postcode=410073, state=Hunan, country=China\\affiliation\[2\] organization=College of Computer Science and Technology, National University of Defense Technology, city=Changsha, postcode=410073, state=Hunan, country=China

## 1Introduction

As large language models improve the ability of web agents to complete tasks[zheng2024gpt](https://arxiv.org/html/2605.23939#bib.bib1),[he2024webvoyager](https://arxiv.org/html/2605.23939#bib.bib2),[lai2024autowebglm](https://arxiv.org/html/2605.23939#bib.bib3), a central question is how these agents can achieve stable, long\-term adaptation in dynamic web environments[liu2024domain](https://arxiv.org/html/2605.23939#bib.bib4)\. Compared with static benchmarks, web tasks present distinct challenges because their states change over time and their solution paths are often diverse[zheng2024webolympus](https://arxiv.org/html/2605.23939#bib.bib5),[pan2406webcanvas](https://arxiv.org/html/2605.23939#bib.bib6),[ye2025realwebassist](https://arxiv.org/html/2605.23939#bib.bib7),[xue2025illusion](https://arxiv.org/html/2605.23939#bib.bib8)\. Even when the goal is similar, action sequences may shift as page components and layouts are updated, and the appropriate next action often depends on the current page state and its surrounding context[lee2025learning](https://arxiv.org/html/2605.23939#bib.bib9),[gou2024navigating](https://arxiv.org/html/2605.23939#bib.bib10)\. As a result, a web agent cannot rely solely on on\-the\-fly reasoning or task\-local exploration, as in many existing agent frameworks[yao2022react](https://arxiv.org/html/2605.23939#bib.bib11),[he2024webvoyager](https://arxiv.org/html/2605.23939#bib.bib2),[lai2024autowebglm](https://arxiv.org/html/2605.23939#bib.bib3)\. It must also accumulate, reuse, and revise experience from past interactions[han2024adaptive](https://arxiv.org/html/2605.23939#bib.bib12)\. Such experience may include task\-level regularities for decomposing recurring goals, as well as interaction\-level cues indicating how particular page states enable or constrain actions\. This longer\-term use of experience is essential for robust adaptation in changing web environments\.

To make web agents more adaptive over long horizons, recent work has studied how to store and reuse past experience\. Representative methods distill experience into natural\-language reflections or rules that guide later decisions[zhao2024expel](https://arxiv.org/html/2605.23939#bib.bib13), or retain interaction trajectories and demonstrations that can be consulted when similar tasks arise[wang2024agent](https://arxiv.org/html/2605.23939#bib.bib14),[liu2025contextual](https://arxiv.org/html/2605.23939#bib.bib15),[zheng2504skillweaver](https://arxiv.org/html/2605.23939#bib.bib16),[zhou2025proposer](https://arxiv.org/html/2605.23939#bib.bib17)\. Reusing such experience can reduce repeated trial and error\. For instance, an agent may learn from earlier failures that a task should be decomposed into search, filtering, and confirmation steps, or that a particular page state requires closing a modal window before the target action becomes available\.

More recent work has studied programmatic skills as executable representations of reusable web experience[prabhu2025walt](https://arxiv.org/html/2605.23939#bib.bib18),[zhong2026actionengine](https://arxiv.org/html/2605.23939#bib.bib19)\. Rather than treating raw trajectories only as demonstrations, these methods abstract recurring interaction patterns into callable functions so that an agent can perform structurally similar operations by providing task\- or page\-specific arguments[wang2024agent](https://arxiv.org/html/2605.23939#bib.bib14),[prabhu2025walt](https://arxiv.org/html/2605.23939#bib.bib18)\. This formulation is well suited to interaction\-level knowledge, since it represents procedural constraints, action order, and executable grounding more directly than natural\-language reflections[zhong2026actionengine](https://arxiv.org/html/2605.23939#bib.bib19)\. For example, a form\-filling trajectory can be converted into a parameterized function that locates the relevant input fields, fills in the required values, and submits the form\. The agent can then reuse this operation without planning each low\-level action again from scratch[prabhu2025walt](https://arxiv.org/html/2605.23939#bib.bib18)\.

These efforts point to a deeper challenge: different forms of experience tend to encode different levels of knowledge[liu2025class](https://arxiv.org/html/2605.23939#bib.bib20)\. Natural\-language reflections can express transferable task\-level lessons, such as recognizing when an ineffective strategy should be revised, but they often provide too little grounding to determine which page element should be operated on\. Trajectories and programmatic skills, by contrast, retain richer executable details, including clicks, inputs, page states, and procedural constraints\. These details, however, usually transfer only when the new task shares similar page layouts, element structures, or interaction contexts\. The main limitation is therefore not simply the choice of representation, but the absence of a clear separation and coordination between reasoning knowledge and interaction knowledge\. When these two levels are stored and reused as a single experience object, agents face a persistent trade\-off between cross\-site transferability and page\-level executability\.

Motivated by this observation, we propose DRIVE, a dual\-level skill modeling framework that separates historical web experience into transferable reasoning skills and executable interaction skills\. Rather than forcing all experience into a single representation, DRIVE uses a form suited to each level of reuse\. Reasoning skills are written in natural language and capture knowledge for task understanding and decision making, while interaction skills are represented as programs that encode executable page\-operation patterns and action constraints\. To support reuse, DRIVE links each skill to its usage scenario, described by task semantics and page\-context conditions\. It then uses a scene\-aware mechanism to retrieve and coordinate dual\-level skills for the current task\-page scene\. DRIVE also uses skill\-level failure feedback to revise, expand, and deduplicate the skill library over time, forming a continual learning pipeline based on layered representation, coordinated reuse, and closed\-loop updating\.

We evaluate DRIVE on five website domains from WebArena[zhou2023webarena](https://arxiv.org/html/2605.23939#bib.bib21)\. DRIVE achieves an average task success rate of 52\.8%, improving over the skill\-free baseline by 7\.3 percentage points\. Its performance also increases as more historical trajectories are used for skill induction and updating\. These results suggest that dual\-level skill modeling provides an effective way to accumulate experience, reuse capabilities, and support continual improvement in web agents\.

The main contributions of this work are as follows\.

- •We motivate DRIVE by identifying the heterogeneity between reasoning knowledge and interaction knowledge in web interaction experience, highlighting the limitation of treating historical experience as a unified representation\.
- •We present DRIVE, a dual\-level skill modeling framework that represents historical experience as natural\-language reasoning skills and programmatic interaction skills\. DRIVE reuses and updates these skills over time by retrieving and invoking them according to the current scene, and by incorporating failure feedback at the skill level\.
- •We evaluate DRIVE across the five website domains in WebArena\. The results show that DRIVE improves task success rates, and its performance continues to increase as experience accumulates\.

## 2Related work

Prior work has improved web agents by reusing past trajectories through episodic memory and experience replay[wang2025continual](https://arxiv.org/html/2605.23939#bib.bib22)\. AWM[wang2024agent](https://arxiv.org/html/2605.23939#bib.bib14)extracts reusable workflows from previous interactions, CER[liu2025contextual](https://arxiv.org/html/2605.23939#bib.bib15)retrieves relevant patterns from a replay buffer, WebCoach[liu2025webcoach](https://arxiv.org/html/2605.23939#bib.bib23)compresses navigation histories into reusable guidance, and ExpSeek[zhang2026expseek](https://arxiv.org/html/2605.23939#bib.bib24)intervenes with prior experience when the agent is uncertain\. Together, these methods show that accumulated trajectories can support decisions in later tasks\. In most cases, however, the experience is reused as a single object, such as a memory entry or workflow\. This design is useful for recalling high\-level strategies, but it leaves low\-level interaction knowledge largely implicit, including the knowledge needed to carry out those strategies on a concrete web page\. For example, a summary may tell an agent to search, filter, and confirm, yet the agent may still fail when it must identify the right interface element or handle page\-specific constraints\.

Executable skill induction\. A related line of work studies how web interaction trajectories can be abstracted into programmatic or callable skills\. SkillWeaver[zheng2504skillweaver](https://arxiv.org/html/2605.23939#bib.bib16)synthesizes reusable web skills as APIs through exploration and practice, and CASCADE[huang2025cascade](https://arxiv.org/html/2605.23939#bib.bib25)and ALITA[qiu2025alita](https://arxiv.org/html/2605.23939#bib.bib26)study how repeated interaction patterns can be exposed as callable tools or skills\. These methods show that interaction data can be converted into reusable operations rather than kept only as contextual memory\. The resulting skills or policies, however, are usually represented as single procedural units\. This improves executability, but offers limited support for capturing the reasoning experience that explains when and why an operation should be used\. Other self\-improving frameworks, including WebRL[qi2025webrltrainingllmweb](https://arxiv.org/html/2605.23939#bib.bib27), WebAgent\-R1[wei2025webagent](https://arxiv.org/html/2605.23939#bib.bib28), and Agent Q[putta2024agent](https://arxiv.org/html/2605.23939#bib.bib29), improve agent behavior through online exploration, reinforcement learning, or search\-based correction\. Their feedback signals are typically used for overall behavior optimization rather than skill\-level revision\. This leaves existing methods limited in how they accumulate reasoning experience and support the continual refinement of reusable skills\.

Our work connects experience reuse with executable skill induction, but argues that both face a shared representational problem: experience derived from trajectories is often treated as a single reusable object, even though it contains knowledge with different levels of abstraction and different grounding requirements\. DRIVE addresses this abstraction–grounding mismatch by representing experience at the level where it can be reused most effectively\. Instead of treating memory retrieval and executable skills as separate components, DRIVE separates the knowledge that can generalize across tasks from the knowledge that must be grounded in a specific web context\. This decomposition provides a stronger basis for continual skill evolution, allowing reusable experience to be refined and expanded according to its level of abstraction rather than updated as one monolithic memory or procedural unit\.

## 3Method

![Refer to caption](https://arxiv.org/html/2605.23939v1/x1.png)Figure 1:Overview of DRIVE\. DRIVE consists of an offline stage for skill abstraction and evolution and an online stage for scene\-aware skill reuse\. The offline stage abstracts historical web interaction trajectories into reasoning and interaction skills, which are organized in hierarchical skill libraries and refined through evolution\. During online execution, the agent retrieves relevant skills according to task semantics and page observations, and feeds execution feedback back to the offline stage to form a closed\-loop refinement process\.Figure[1](https://arxiv.org/html/2605.23939#S3.F1)illustrates the overall framework of DRIVE\. DRIVE follows an offline–online loop for building, reusing, and refining web\-agent skills\. In the offline phase, historical trajectories are abstracted into two skill libraries\. Reasoning skills capture reusable task logic, while interaction skills record page\-grounded operation experience\. DRIVE further analyzes failures at the skill level, so reasoning failures update the reasoning library, whereas interaction failures repair the interaction library\. In the online phase, given a task instruction and the current page observation, DRIVE retrieves skills that match both the task intent and the webpage context\. The reasoning skill provides corrective guidance to avoid recurring reasoning errors, while the interaction skill reuses successful page\-level experience to reduce execution failures\. The resulting feedback is returned to the offline phase to refine the libraries\. By decoupling transferable reasoning from page\-grounded interaction while coordinating them during execution, DRIVE enables web agents to accumulate capabilities across dynamic web environments\.

### 3\.1Problem Formulation

We study a web agent built on a language model backboneℒ\\mathcal\{L\}that interacts with a web environment to solve a natural\-language task instructionqq\. At environment steptt, letsts\_\{t\}denote the underlying environment state, and letoto\_\{t\}denote the corresponding observation available to the agent\. The agent acts in a predefined primitive action space𝒜\\mathcal\{A\}, and the environment evolves according to

st\+1=𝒯​\(st,at\),s\_\{t\+1\}=\\mathcal\{T\}\(s\_\{t\},a\_\{t\}\),whereat∈𝒜a\_\{t\}\\in\\mathcal\{A\}is the action executed at steptt\. The interaction continues until the agent emits a termination action or the environment reaches a terminal condition\.

To adapt continuously in dynamic web environments, we assume that the agent maintains a dual\-level skill library

𝒦=\(𝒦r,𝒦i\),\\mathcal\{K\}=\(\\mathcal\{K\}^\{r\},\\mathcal\{K\}^\{i\}\),where𝒦r\\mathcal\{K\}^\{r\}is a reasoning skill library and𝒦i\\mathcal\{K\}^\{i\}is an interaction skill library\. We also use𝒦all=𝒦r∪𝒦i\\mathcal\{K\}^\{\\mathrm\{all\}\}=\\mathcal\{K\}^\{r\}\\cup\\mathcal\{K\}^\{i\}to denote the union of all skill entries\. Reasoning skills capture transferable knowledge for task understanding and decision making, such as subgoal decomposition, strategy selection, state assessment, and outcome verification\. Interaction skills capture executable web procedures grounded in page context, including element grounding, action sequences, triggering conditions, and execution constraints\.

At each step, the agent retrieves relevant reasoning and interaction skills conditioned on the task instruction and the current observation:

\(ktr,dktr\)=Retriever​\(q,ot,𝒦r\),\(kti,dkti\)=Retrievei​\(q,ot,𝒦i\)\.\(k\_\{t\}^\{r\},d\_\{k\_\{t\}^\{r\}\}\)=\\mathrm\{Retrieve\}^\{r\}\(q,o\_\{t\},\\mathcal\{K\}^\{r\}\),\\qquad\(k\_\{t\}^\{i\},d\_\{k\_\{t\}^\{i\}\}\)=\\mathrm\{Retrieve\}^\{i\}\(q,o\_\{t\},\\mathcal\{K\}^\{i\}\)\.The retrieved reasoning skill provides high\-level guidance for decision making, for example by identifying the next subgoal or action intent\. The retrieved interaction skill is then invoked directly as an executable procedure to carry out the intended web operation under the current page context\. In this way, reasoning skills guide what to do, while interaction skills determine how to do it\.

A task execution produces an interaction episode

τ=\(q,\{\(ot,at\)\}t=1T,y\),\\tau=\\left\(q,\\\{\(o\_\{t\},a\_\{t\}\)\\\}\_\{t=1\}^\{T\},y\\right\),whereTTis the trajectory length andy∈\{0,1\}y\\in\\\{0,1\\\}indicates whether the task is completed successfully\. Over time, the agent accumulates successful episodesΓsucc\\Gamma\_\{\\mathrm\{succ\}\}and failed episodesΓfail\\Gamma\_\{\\mathrm\{fail\}\}\.

Our goal is to continually update the dual\-level skill library from historical interaction experience so that the resulting skills can be reused to improve future task completion in changing web environments\. Letnndenote the skill evolution round, and letΦn\\Phi\_\{n\}denote the skill\-level feedback collected before roundnn\. The self\-evolution process is defined as

\(𝒦r,\(n\+1\),𝒦i,\(n\+1\)\)=Update​\(𝒦r,\(n\),𝒦i,\(n\),Γsucc,Γfail,Φn\),\\left\(\\mathcal\{K\}^\{r,\(n\+1\)\},\\mathcal\{K\}^\{i,\(n\+1\)\}\\right\)=\\mathrm\{Update\}\\\!\\left\(\\mathcal\{K\}^\{r,\(n\)\},\\mathcal\{K\}^\{i,\(n\)\},\\Gamma\_\{\\mathrm\{succ\}\},\\Gamma\_\{\\mathrm\{fail\}\},\\Phi\_\{n\}\\right\),whereUpdate​\(⋅\)\\mathrm\{Update\}\(\\cdot\)denotes the skill evolution operator that revises existing skills, induces new skills, and removes obsolete or redundant ones based on accumulated experience\. During online deployment, invoking retrieved skills also produces skill\-level feedback, especially information about why skill use fails under the current task and page context\. This failure feedback is recorded and incorporated into later updates, allowing both reasoning and interaction skills to be refined over time\.

### 3\.2Skill Representation

To support continual self\-evolution, the agent maintains a dual\-level skill library

𝒦=\(𝒦r,𝒦i\),\\mathcal\{K\}=\(\\mathcal\{K\}^\{r\},\\mathcal\{K\}^\{i\}\),where𝒦r\\mathcal\{K\}^\{r\}and𝒦i\\mathcal\{K\}^\{i\}denote the reasoning skill library and the interaction skill library, respectively\. We use𝒦all=𝒦r∪𝒦i\\mathcal\{K\}^\{\\mathrm\{all\}\}=\\mathcal\{K\}^\{r\}\\cup\\mathcal\{K\}^\{i\}to denote all skill entries\. The two types of skills capture complementary aspects of historical experience\. Reasoning skills encode transferable knowledge about what to do, whereas interaction skills encode executable knowledge about how to do it\. Each skill is also associated with a scenario descriptor that summarizes the task semantics and page\-context conditions under which the skill is applicable\. This descriptor serves as the interface between offline skill induction and online skill retrieval\.

Interaction skills in𝒦i\\mathcal\{K\}^\{i\}encode reusable executable web procedures grounded in page context\. They are primarily induced from successful trajectories and can be further refined using execution\-level failure feedback collected during online deployment\. Rather than functioning as a static open\-loop mapping, each interaction skill is represented as a parameterized executable procedure \(or sub\-policy\)\. Invoking an interaction skillkik^\{i\}under a specific web contextC∈𝒞C\\in\\mathcal\{C\}with input argumentsX∈𝒳X\\in\\mathcal\{X\}triggers a dynamic, closed\-loop execution process against the environment\. This process yields a realized action traceα\\alphaand a final execution outcomeee:

ki​\(C,X\)→\(α,e\),k^\{i\}\(C,X\)\\rightarrow\(\\alpha,e\),where𝒞\\mathcal\{C\}denotes the space of web contexts \(including page structure, available elements, and local state conditions\),𝒳\\mathcal\{X\}denotes the space of input arguments required for execution,α∈𝒜∗\\alpha\\in\\mathcal\{A\}^\{\*\}is the finite sequence of primitive actions executed during the invocation, ande∈𝒴e\\in\\mathcal\{Y\}is the execution outcome\. During execution,kik^\{i\}actively interacts with the page \(e\.g\., iteratively verifying locators and handling dynamic UI changes\) rather than blindly emitting a predefined sequence\. In this sense, an interaction skill is not simply a behavioral hint, but a robust executable module that directly supports action execution on the web page\.

Reasoning skills in𝒦r\\mathcal\{K\}^\{r\}encode reusable guidance for task understanding and decision making\. Unlike interaction skills, they do not execute web actions directly\. Instead, they provide high\-level support for subgoal decomposition, strategy selection, state assessment, and outcome verification\. We represent a reasoning skill as

kr=⟨ℳ,ℬ,𝒱⟩,k^\{r\}=\\langle\\mathcal\{M\},\\mathcal\{B\},\\mathcal\{V\}\\rangle,whereℳ\\mathcal\{M\}summarizes a previously observed reasoning error pattern or decision deficiency,ℬ\\mathcal\{B\}specifies the corresponding corrected reasoning principle or behavioral guidance, and𝒱\\mathcal\{V\}defines explicit verification instructions for checking whether the current decision process satisfies the intended task requirements\. During execution, reasoning skills are injected into the model context as structured guidance, helping the agent avoid recurring reasoning errors and make more reliable decisions in related scenarios\.

##### Scenario Descriptor

Each skill entry\(k,dk\)∈𝒦all\(k,d\_\{k\}\)\\in\\mathcal\{K\}^\{\\mathrm\{all\}\}contains a skillkkand its associated scenario descriptor

dk=⟨𝒰k,𝒲k⟩,d\_\{k\}=\\langle\\mathcal\{U\}\_\{k\},\\mathcal\{W\}\_\{k\}\\rangle,where𝒰k\\mathcal\{U\}\_\{k\}denotes the applicable web context of skillkk, such as URL patterns or page\-level environment cues, and𝒲k\\mathcal\{W\}\_\{k\}denotes the semantic task scenario captured by skillkk\. Given the current instructionqqand observationoto\_\{t\}, the agent usesdkd\_\{k\}to retrieve candidate skills and determine whether a skill should be invoked under the current context\.

### 3\.3Skill Induction from Successful Trajectories

A successful episodeτ∈Γsucc\\tau\\in\\Gamma\_\{\\mathrm\{succ\}\}contains reusable web operation patterns for accomplishing a task, but it remains tied to a specific task instance and page context\. Instead of replaying such instance\-specific experience directly, DRIVE abstracts it into a parameterized interaction skill that can be reused across related tasks and similar web scenarios\.

Given a successful episodeτ\\tau, we prompt the language model to synthesize an executable interaction skillkτik\_\{\\tau\}^\{i\}\. This process abstracts task\-specific entities in the instruction and observation history into input arguments in𝒳\\mathcal\{X\}, converting the original episode into a reusable procedure conditioned on context\. The induced interaction skill encapsulates the necessary closed\-loop execution logic and is represented as:

kτi​\(C,X\)→\(α,e\),k\_\{\\tau\}^\{i\}\(C,X\)\\rightarrow\(\\alpha,e\),where𝒞\\mathcal\{C\}denotes the space of web contexts required for execution,𝒳\\mathcal\{X\}denotes the space of parameterized task\-specific inputs,α∈𝒜∗\\alpha\\in\\mathcal\{A\}^\{\*\}is the dynamic action trace generated during interaction, ande∈𝒴e\\in\\mathcal\{Y\}is the final execution outcome\.

Along with the executable procedure, the language model also generates a scenario descriptordkτid\_\{k\_\{\\tau\}^\{i\}\}that summarizes the task semantics and page\-context conditions under which the skill applies\. This descriptor serves as the retrieval key for later online invocation\.

The induced interaction skill and its associated scenario descriptor are then accumulated into a new interaction skill pool for the current round, rather than updating the main library immediately:

𝒦newi←𝒦newi∪\{\(kτi,dkτi\)\}\.\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{i\}\\leftarrow\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{i\}\\cup\\\{\(k\_\{\\tau\}^\{i\},d\_\{k\_\{\\tau\}^\{i\}\}\)\\\}\.

### 3\.4Counterfactual Failure Attribution from Failed Trajectories

Failed episodes provide useful signals for self\-evolution, but their underlying causes are often entangled\. To identify the source of failure in a failed episodeτ∈Γfail\\tau\\in\\Gamma\_\{\\mathrm\{fail\}\}, we introduce an LM\-based attribution module

c=fattr​\(τ\),c=f\_\{\\mathrm\{attr\}\}\(\\tau\),wherec∈ℰfail=\{erri,errr\}c\\in\\mathcal\{E\}\_\{\\mathrm\{fail\}\}=\\\{\\mathrm\{err\}^\{i\},\\mathrm\{err\}^\{r\}\\\}denotes the inferred failure type\. Instead of treating each failed episode as a uniform negative example, we prompt the language model to perform a counterfactual analysis of the error trace by asking whether the task could have been completed if the agent had preserved its high\-level intent but followed a different interaction procedure\. If the answer is yes, the failure is attributed to an interaction\-level error \(c=erric=\\mathrm\{err\}^\{i\}\); otherwise, it is attributed to a reasoning\-level error \(c=errrc=\\mathrm\{err\}^\{r\}\)\.

Based on the attributed failure type, the failed episode is converted into a corresponding skill entry:

Gfail​\(τ,c\)=\{\(kτi,dkτi\),c=erri,\(kτr,dkτr\),c=errr,G\_\{\\mathrm\{fail\}\}\(\\tau,c\)=\\begin\{cases\}\(k\_\{\\tau\}^\{i\},d\_\{k\_\{\\tau\}^\{i\}\}\),&c=\\mathrm\{err\}^\{i\},\\\\ \(k\_\{\\tau\}^\{r\},d\_\{k\_\{\\tau\}^\{r\}\}\),&c=\\mathrm\{err\}^\{r\},\\end\{cases\}whereGfail​\(⋅\)G\_\{\\mathrm\{fail\}\}\(\\cdot\)maps a failed episode to either an interaction\-skill entry or a reasoning\-skill entry\.

A reasoning\-level error indicates that the failure comes from incorrect task understanding, flawed decomposition, or invalid assumptions about the environment, rather than from the execution procedure itself\. In this case, the failed episode is converted into a reasoning skill

kτr=⟨ℳ,ℬ,𝒱⟩,k\_\{\\tau\}^\{r\}=\\langle\\mathcal\{M\},\\mathcal\{B\},\\mathcal\{V\}\\rangle,whereℳ\\mathcal\{M\}summarizes the observed reasoning error pattern,ℬ\\mathcal\{B\}specifies the corrected reasoning principle or behavioral guidance, and𝒱\\mathcal\{V\}defines explicit verification instructions\. These skills serve as structured corrective guidance and are injected into the model context in future related tasks to reduce similar reasoning failures\. The induced reasoning skill and its associated scenario descriptor are then accumulated into a new reasoning skill pool via

𝒦newr←𝒦newr∪\{\(kτr,dkτr\)\}\.\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{r\}\\leftarrow\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{r\}\\cup\\\{\(k\_\{\\tau\}^\{r\},d\_\{k\_\{\\tau\}^\{r\}\}\)\\\}\.
By contrast, an interaction\-level error indicates that the agent follows an appropriate high\-level strategy but fails during web interaction because of page\-specific interruptions, missing elements, selector mismatch, or brittle interaction paths\. In this case, we synthesize an interaction skillkτik\_\{\\tau\}^\{i\}that encodes a corrected executable procedure for the failed task pattern\. Compared with interaction skills induced from successful episodes, these skills may also include recovery logic, such as condition checking, exception handling, or conditional retries, to improve robustness under similar web conditions\. The induced interaction skill and its associated scenario descriptor are then added to the new interaction skill pool via

𝒦newi←𝒦newi∪\{\(kτi,dkτi\)\}\.\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{i\}\\leftarrow\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{i\}\\cup\\\{\(k\_\{\\tau\}^\{i\},d\_\{k\_\{\\tau\}^\{i\}\}\)\\\}\.
For both failure types, the newly induced skill is further associated with a scenario descriptor, denoted bydkτrd\_\{k\_\{\\tau\}^\{r\}\}ordkτid\_\{k\_\{\\tau\}^\{i\}\}, which supports future retrieval and conditional invocation under similar task semantics and page\-context conditions\. In this way, failed episodes are not discarded as unsuccessful experience, but are explicitly transformed into reusable knowledge that supports continual self\-evolution\.

### 3\.5Scenario\-Aware Skill Retrieval and Invocation

To reuse accumulated experience in a new task, the agent performs scenario\-aware retrieval and coordinated invocation over the dual\-level skill library

𝒦=\(𝒦r,𝒦i\)\.\\mathcal\{K\}=\(\\mathcal\{K\}^\{r\},\\mathcal\{K\}^\{i\}\)\.At environment steptt, the current observationoto\_\{t\}provides the full environmental signals\. From this, we extract the local page context, including the active URLutu\_\{t\}and the current interface state, denoted by

Ct=ExtractContext​\(ot\)\.C\_\{t\}=\\mathrm\{ExtractContext\}\(o\_\{t\}\)\.Each skill entry\(k,dk\)∈𝒦all\(k,d\_\{k\}\)\\in\\mathcal\{K\}^\{\\mathrm\{all\}\}contains a scenario descriptor

dk=⟨𝒰k,𝒲k⟩,d\_\{k\}=\\langle\\mathcal\{U\}\_\{k\},\\mathcal\{W\}\_\{k\}\\rangle,where𝒰k\\mathcal\{U\}\_\{k\}characterizes the page\-context conditions under which skillkkapplies, and𝒲k\\mathcal\{W\}\_\{k\}summarizes the task\-semantic conditions under which skillkkis relevant\.

Given the current task instructionqqand observationoto\_\{t\}, we use a two\-stage mechanism for both reasoning and interaction skills\. We first perform structural filtering by retaining only the skills whose applicable page context is compatible with the current webpage state, yielding candidate sets

𝒦cand,tr=\{\(k,dk\)∈𝒦r∣Match​\(Ct,𝒰k\)=1\},\\mathcal\{K\}^\{r\}\_\{\\mathrm\{cand\},t\}=\\\{\(k,d\_\{k\}\)\\in\\mathcal\{K\}^\{r\}\\mid\\mathrm\{Match\}\(C\_\{t\},\\mathcal\{U\}\_\{k\}\)=1\\\},𝒦cand,ti=\{\(k,dk\)∈𝒦i∣Match​\(Ct,𝒰k\)=1\}\.\\mathcal\{K\}^\{i\}\_\{\\mathrm\{cand\},t\}=\\\{\(k,d\_\{k\}\)\\in\\mathcal\{K\}^\{i\}\\mid\\mathrm\{Match\}\(C\_\{t\},\\mathcal\{U\}\_\{k\}\)=1\\\}\.We then derive a compact semantic representation of the current task context fromqqandoto\_\{t\}and compare it with the task\-semantic descriptors𝒲k\\mathcal\{W\}\_\{k\}associated with the candidate skills\. This descriptor\-level matching step ranks candidate skills without requiring the full contents of all skills to be injected into the model context\.

To avoid conflicts among multiple applicable skills, DRIVE treats the retrieved skills as candidates rather than actions to be executed jointly\. At each step, the agent applies a singleton activation policy within each skill level: it selects at most one reasoning skill and at most one interaction skill according to their compatibility with the current task\-page scene\. Formally,

\(ktr,dktr\)=Selectr​\(q,ot,𝒦cand,tr\),\(kti,dkti\)=Selecti​\(q,ot,𝒦cand,ti\),\(k\_\{t\}^\{r\},d\_\{k\_\{t\}^\{r\}\}\)=\\mathrm\{Select\}^\{r\}\(q,o\_\{t\},\\mathcal\{K\}\_\{\\mathrm\{cand\},t\}^\{r\}\),\\qquad\(k\_\{t\}^\{i\},d\_\{k\_\{t\}^\{i\}\}\)=\\mathrm\{Select\}^\{i\}\(q,o\_\{t\},\\mathcal\{K\}\_\{\\mathrm\{cand\},t\}^\{i\}\),where

Selectℓ​\(q,ot,𝒦cand,tℓ\)∈𝒦cand,tℓ∪\{∅\},ℓ∈\{r,i\}\.\\mathrm\{Select\}^\{\\ell\}\(q,o\_\{t\},\\mathcal\{K\}\_\{\\mathrm\{cand\},t\}^\{\\ell\}\)\\in\\mathcal\{K\}\_\{\\mathrm\{cand\},t\}^\{\\ell\}\\cup\\\{\\varnothing\\\},\\qquad\\ell\\in\\\{r,i\\\}\.The selection is based on the match between the current task semantics, page context, and the scenario descriptordkd\_\{k\}\. If no candidate is sufficiently compatible with the current scene, the selector returns∅\\varnothing, and the agent falls back to primitive action generation\. Thus, multiple skills may be retrieved as candidates, but they are not invoked simultaneously\.

When a reasoning skill is selected, it has the form

ktr=⟨ℳt,ℬt,𝒱t⟩k\_\{t\}^\{r\}=\\langle\\mathcal\{M\}\_\{t\},\\mathcal\{B\}\_\{t\},\\mathcal\{V\}\_\{t\}\\rangleand is injected into the model context as structured guidance for decision making\. Specifically, it reminds the agent of the relevant reasoning error patternℳt\\mathcal\{M\}\_\{t\}, provides the corresponding corrected strategyℬt\\mathcal\{B\}\_\{t\}, and highlights the verification instructions𝒱t\\mathcal\{V\}\_\{t\}to guide subsequent reasoning and action planning\.

When an interaction skillktik\_\{t\}^\{i\}is selected, the agent first checks whether its scenario descriptor is consistent with the current reasoning intent and page context\. If the selected interaction skill is compatible, the agent instantiates its required input argumentsXtX\_\{t\}from the current task and observation, and then executes the skill under the current web context:

\(αt,et\)=kti​\(Ct,Xt\),αt∈𝒜∗,et∈𝒴,\(\\alpha\_\{t\},e\_\{t\}\)=k\_\{t\}^\{i\}\(C\_\{t\},X\_\{t\}\),\\qquad\\alpha\_\{t\}\\in\\mathcal\{A\}^\{\*\},\\ e\_\{t\}\\in\\mathcal\{Y\},whereCtC\_\{t\}denotes the current page context,XtX\_\{t\}denotes the instantiated task\-specific arguments,αt\\alpha\_\{t\}denotes the realized finite sequence of primitive actions executed during the invocation, andete\_\{t\}denotes the corresponding execution outcome\. If the selected interaction skill is incompatible or fails its applicability check, the agent does not invoke it and instead falls back to primitive action generation\.

Compared with token\-level generation of primitive actions, this invocation mode directly reuses a context\-conditioned executable interaction procedure for the current task pattern and returns both the action trace and the local execution result\. In this way, scenario\-aware retrieval enables the agent to use the complementary roles of the two skill types together: reasoning skills guide what to do, while interaction skills support how to do it under the current page context\.

### 3\.6Skill Management and Continual Refinement

As experience accumulates, the dual\-level skill library

𝒦=\(𝒦r,𝒦i\)\\mathcal\{K\}=\(\\mathcal\{K\}^\{r\},\\mathcal\{K\}^\{i\}\)continues to grow\. DRIVE manages the library at each evolution round to remove invalid interaction skills and redundant reasoning skills\. Since the two skill types have different feedback signals, they are updated with different operators\.

##### Skill\-level feedback

When a skill is invoked, DRIVE records its execution feedback as

ϕ=⟨kϕ,dkϕ,qϕ,Ct,ϕ,Xt,ϕ,et,ϕ,Δt,ϕ,yϕ⟩\\phi=\\left\\langle k\_\{\\phi\},d\_\{k\_\{\\phi\}\},q\_\{\\phi\},C\_\{t,\\phi\},X\_\{t,\\phi\},e\_\{t,\\phi\},\\Delta\_\{t,\\phi\},y\_\{\\phi\}\\right\\ranglewherekϕk\_\{\\phi\}is the invoked skill,dkϕd\_\{k\_\{\\phi\}\}is its descriptor,qϕq\_\{\\phi\}is the task instruction,Ct,ϕC\_\{t,\\phi\}is the current page context,Xt,ϕX\_\{t,\\phi\}is the instantiated input,et,ϕe\_\{t,\\phi\}is the skill execution outcome,yϕy\_\{\\phi\}is the final task label, andΔt,ϕ\\Delta\_\{t,\\phi\}is the execution log\. The feedback collected before roundnnis denoted by

Φn=\{ϕ1,ϕ2,…,ϕm\}\.\\Phi\_\{n\}=\\\{\\phi\_\{1\},\\phi\_\{2\},\\ldots,\\phi\_\{m\}\\\}\.For interaction skills,Δt,ϕ\\Delta\_\{t,\\phi\}records local signals, including selector matching, URL or page\-state changes, returned results, and execution exceptions\. For reasoning skills, no clear local error is usually available, so DRIVE mainly uses the final task outcome as a weak utility signal\.

##### Interaction skill repair

Interaction skills in𝒦i\\mathcal\{K\}^\{i\}are executable procedures \(ki:𝒞×𝒳→𝒜∗×𝒴k^\{i\}:\\mathcal\{C\}\\times\\mathcal\{X\}\\rightarrow\\mathcal\{A\}^\{\*\}\\times\\mathcal\{Y\}\)\. They may fail when the UI changes, selectors become invalid, or a modal blocks the intended action\. To reduce this brittleness, DRIVE does not bind each operation to a single selector\. Instead, the internal executable procedure of each interaction skill is parameterized as a sequence of operation templates:

Θ​\(ki\)=\[\(r1,Σ1\),\(r2,Σ2\),…,\(rP,ΣP\)\],\\Theta\(k^\{i\}\)=\\left\[\(r\_\{1\},\\Sigma\_\{1\}\),\(r\_\{2\},\\Sigma\_\{2\}\),\\ldots,\(r\_\{P\},\\Sigma\_\{P\}\)\\right\],whererjr\_\{j\}is the intended operation and

Σj=\(σj,1,σj,2,…,σj,Lj\)\\Sigma\_\{j\}=\(\\sigma\_\{j,1\},\\sigma\_\{j,2\},\\ldots,\\sigma\_\{j,L\_\{j\}\}\)is an ordered sequence of selectors for grounding its target element\. These selectors may come from successful executions on similar pages and may include CSS selectors, XPath, DOM identifiers, text\-based locators, accessibility attributes, or nearby\-label locators\. During execution, DRIVE tries the selectors inΣj\\Sigma\_\{j\}in order and uses the first valid one\.

DRIVE treats a case as interaction\-skill failure when the skill is retrieved and invoked, but its local effect is missing\. This includes four common cases: no selector inΣj\\Sigma\_\{j\}matches the target element; the action is executed but the URL or page state does not change as expected; no expected result is returned; or the final page state violates the skill check\. In contrast, if the skill is not retrieved due to descriptor mismatch, the case is not treated as interaction\-skill failure\.

Given interaction feedbackΦni⊆Φn\\Phi\_\{n\}^\{i\}\\subseteq\\Phi\_\{n\}, DRIVE repairs interaction skills by local patching:

Θ​\(ki\)\+=Patchi​\(Θ​\(ki\),Φni\)\.\{\\Theta\(k^\{i\}\)\}^\{\+\}=\\mathrm\{Patch\}^\{i\}\(\\Theta\(k^\{i\}\),\\Phi\_\{n\}^\{i\}\)\.The patch operator mainly rewrites selector sets rather than regenerating the whole skill\. For an operationrjr\_\{j\}, the updated selector set is

Σj\+=UpdateSelector​\(Σj,Ct,ϕ,rj,Δt,ϕ\)\.\\Sigma\_\{j\}^\{\+\}=\\mathrm\{UpdateSelector\}\\left\(\\Sigma\_\{j\},C\_\{t,\\phi\},r\_\{j\},\\Delta\_\{t,\\phi\}\\right\)\.A failed selector is demoted or removed\. A working selector found in later traces is added to the set\. If needed, DRIVE also adds a small fallback branch, such as retrying with another locator or closing an unexpected modal\. When a skill still fails after repeated selector\-level patches, DRIVE removes it from𝒦i\\mathcal\{K\}^\{i\}\.

##### Reasoning skill consolidation

Reasoning skills in𝒦r\\mathcal\{K\}^\{r\}are natural\-language rules for task decisions\. They do not directly execute actions, so their errors are harder to localize\. DRIVE therefore does not rewrite a reasoning skill after a single failure\. Instead, it manages reasoning skills through addition, merging, and deletion\.

For each reasoning skillkrk^\{r\}, DRIVE records its usage count

N​\(kr\)=∑ϕ∈Φn𝟏​\[kϕ=kr\],N\(k^\{r\}\)=\\sum\_\{\\phi\\in\\Phi\_\{n\}\}\\mathbf\{1\}\[k\_\{\\phi\}=k^\{r\}\],and its successful usage count

S​\(kr\)=∑ϕ∈Φn𝟏​\[kϕ=kr\]⋅yϕ\.S\(k^\{r\}\)=\\sum\_\{\\phi\\in\\Phi\_\{n\}\}\\mathbf\{1\}\[k\_\{\\phi\}=k^\{r\}\]\\cdot y\_\{\\phi\}\.The smoothed utility score is

ρ​\(kr\)=S​\(kr\)\+λN​\(kr\)\+2​λ,\\rho\(k^\{r\}\)=\\frac\{S\(k^\{r\}\)\+\\lambda\}\{N\(k^\{r\}\)\+2\\lambda\},whereλ\\lambdais a smoothing constant\. Since task success may also depend on interaction skills,ρ​\(kr\)\\rho\(k^\{r\}\)is only a weak signal\. DRIVE uses it to identify reasoning skills that are rarely useful\.

When several reasoning skills describe similar scenarios and give similar guidance, DRIVE merges them:

\(k¯r,dk¯r\)=Merger​\(𝒢r\),\(\\bar\{k\}^\{r\},d\_\{\\bar\{k\}^\{r\}\}\)=\\mathrm\{Merge\}^\{r\}\(\\mathcal\{G\}^\{r\}\),where

𝒢r=\{\(k1r,dk1r\),…,\(klr,dklr\)\}\\mathcal\{G\}^\{r\}=\\\{\(k\_\{1\}^\{r\},d\_\{k\_\{1\}^\{r\}\}\),\\ldots,\(k\_\{l\}^\{r\},d\_\{k\_\{l\}^\{r\}\}\)\\\}is a group of similar reasoning skills\. The merged skill keeps the shared mistake pattern, correction rule, and verification condition, while removing repeated or overly specific content\. Reasoning skills with low utility, low usage, or strong overlap with a better skill are deleted\. New reasoning\-level failures identified byfattrf\_\{\\mathrm\{attr\}\}are added as new skills\.

##### Library update

Let𝒦newi\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{i\}and𝒦newr\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{r\}denote the pools of new skills induced from successful and failed trajectories during the current round\. At the end of roundnn, DRIVE updates the interaction library by applying a batch patching operator across applicable skills and pruning the invalid ones:

𝒦i,\(n\+1\)=Prunei​\(BatchPatchi​\(𝒦i,\(n\)∪𝒦newi,Φni\),Φni\)\.\\mathcal\{K\}^\{i,\(n\+1\)\}=\\mathrm\{Prune\}^\{i\}\\left\(\\mathrm\{BatchPatch\}^\{i\}\\left\(\\mathcal\{K\}^\{i,\(n\)\}\\cup\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{i\},\\Phi\_\{n\}^\{i\}\\right\),\\Phi\_\{n\}^\{i\}\\right\)\.It updates the reasoning library by applying a batch merging operator to groups of similar skills and pruning those with low utility scores based on accumulated feedback:

𝒦r,\(n\+1\)=Pruner​\(BatchMerger​\(𝒦r,\(n\)∪𝒦newr\),Φnr\)\.\\mathcal\{K\}^\{r,\(n\+1\)\}=\\mathrm\{Prune\}^\{r\}\\left\(\\mathrm\{BatchMerge\}^\{r\}\\left\(\\mathcal\{K\}^\{r,\(n\)\}\\cup\\mathcal\{K\}\_\{\\mathrm\{new\}\}^\{r\}\\right\),\\Phi\_\{n\}^\{r\}\\right\)\.Thus, interaction\-skill management focuses on selector repair and small execution patches, while reasoning\-skill management focuses on adding, merging, and deleting high\-level rules\. This keeps the library compact while allowing it to adapt over time\.

## 4Experiments

### 4\.1Experimental Setup

#### 4\.1\.1Environment and Metrics

We evaluate DRIVE on WebArena[zhou2023webarena](https://arxiv.org/html/2605.23939#bib.bib21), a high\-fidelity benchmark with 812 multi\-step web tasks spanning five domains:Shopping, CMS, Forum, Gitlab, Map\. WebArena is well suited to our setting because its tasks require both long\-horizon reasoning and accurate page\-level grounding\. Following the standard benchmark protocol, we use the automated functional\-correctness evaluator and report success rate as the primary metric\.

#### 4\.1\.2Baselines

We compare DRIVE with a standard prompting agent and several recent web\-agent frameworks\. The standard prompting baseline is the vanilla WebArena agent[zhou2023webarena](https://arxiv.org/html/2605.23939#bib.bib21), which maps observations directly to actions without retaining past experience\. For observation/action alignment, we include AgentOccam[yang2024agentoccam](https://arxiv.org/html/2605.23939#bib.bib30), a strong baseline that improves zero\-shot performance through LLM\-friendly textual representations and standardized action spaces\. For data\-centric exploration, we compare with Go\-Browse[gandhi2025go](https://arxiv.org/html/2605.23939#bib.bib31), which uses offline structured exploration to collect interaction trajectories for policy refinement\. We also include two skill\-based or hierarchical methods: SteP[sodhi2023step](https://arxiv.org/html/2605.23939#bib.bib32), which uses a stack of LLM policies for dynamic control flow, and SkillWeaver[zheng2504skillweaver](https://arxiv.org/html/2605.23939#bib.bib16), which discovers and distills experience into executable APIs\.

Implementation details\. To ensure a fair comparison, all experiments, including DRIVE and the reproduced baselines, use GPT\-4\.1 as the backbone and are evaluated under the same environment configuration\.

#### 4\.1\.3Experiment details

We limit the web agent to 20 environment steps per episode\. For dataset construction and evaluation, we first manually group tasks on each website by task type\. We then use a within\-website stratified split, assigning 30% of tasks to training and 70% to testing, while keeping the task\-type coverage of both partitions as broad as possible\. This task\-type\-aware stratification is designed to test generalization over a diverse and representative task distribution and to reduce bias from skewed task compositions\.

Unless otherwise specified, our main agent uses GPT\-4\.1 as the backbone language model with a decoding temperature of 0\.1\. For fair comparison, all baseline methods use the same model and decoding configuration whenever applicable, and all methods are evaluated on the same test tasks, with identical task IDs and evaluation protocol\.

### 4\.2Main Results

We evaluate DRIVE on WebArena and report success rates across five website domains in Table[1](https://arxiv.org/html/2605.23939#S4.T1)\. DRIVE achieves the best overall performance, with an average success rate of 52\.8%, and consistently outperforms all baselines on Shopping, CMS, Forum, Gitlab, and Map\.

Compared with AgentOccam, the strongest baseline in our evaluation, DRIVE improves the average success rate from 45\.5% to 52\.8%, a gain of 7\.3 percentage points and a relative improvement of 16\.0%\. DRIVE also outperforms Go\-Browse and SkillWeaver by 10\.6 and 23\.1 percentage points, respectively\. These results show the value of dual\-level skill modeling and continual skill refinement for web agents in dynamic environments\.

Table 1:Performance comparison on the WebArena benchmark\. We report task success rates \(%\) across five diverse website domains\. The best performance is marked inboldand the second\-best isunderlined\. Our proposed DRIVE framework consistently outperforms all baseline models, achieving the highest average success rate\.AgentShoppingCMSForumGitlabMapAverageWebArena21\.415\.48\.114\.112\.714\.3SteP37\.024\.059\.032\.030\.036\.4Go\-Browse41\.238\.451\.641\.238\.442\.2SkillWeaver25\.123\.841\.332\.026\.229\.7AgentOccam37\.241\.560\.441\.446\.845\.5Ours45\.946\.166\.750\.055\.452\.8We also include SteP[sodhi2023step](https://arxiv.org/html/2605.23939#bib.bib32), which achieves an average success rate of 36\.4%\. However, SteP relies on manually designed, website\-specific workflows, which require additional effort to transfer across websites and maintain as the environment changes\. In contrast, DRIVE automatically induces dual\-level skills from historical trajectories and continually refines the skill library using skill\-level feedback, without relying on handcrafted domain\-specific policies\.

### 4\.3Self\-Evolution with Increasing Training Experience

To evaluate whether DRIVE continues to benefit from additional experience, we progressively increase the number of training trajectories used for skill induction and refinement while keeping the test set fixed\. This experiment examines how the dual\-level skill library evolves as more interaction experience becomes available\.

Figure[2](https://arxiv.org/html/2605.23939#S4.F2)shows that DRIVE benefits consistently from additional training experience\. As the number of training trajectories increases from 42 to 251, both task success rate and interaction\-skill invocation success improve across most WebArena domains, indicating that self\-evolution can steadily transform accumulated trajectories into reusable capabilities\.

![Refer to caption](https://arxiv.org/html/2605.23939v1/x2.png)Figure 2:Continuous capability accumulation in DRIVE\. As the number of training trajectories increases from 42 to 251, the framework demonstrates steady self\-evolution across five WebArena domains\. The concurrent improvement in overall task success rate \(left\) and interaction skill usage success rate \(right\) validates that DRIVE effectively translates accumulated experience into robust, reusable capabilities\.The left panel shows that task success improves most clearly on Shopping, GitLab, and Map as more training trajectories are incorporated\. These domains involve relatively complex workflows and substantial interface variation, and therefore benefit more from the joint improvement of interaction skills and reasoning skills\. Interaction skills encode reusable executable procedures, while reasoning skills support task understanding, planning, and verification\. By contrast, Forum starts from a comparatively strong level and exhibits smaller marginal gains at higher trajectory budgets, suggesting that its dominant workflows are captured relatively early and that later updates mainly improve long\-tail cases\.

The right panel further shows that the invocation success rate of interaction skills generally increases with additional trajectories, suggesting that self\-evolution improves not only skill coverage but also execution reliability\. The upward trend is especially clear on Shopping, CMS, Map, and GitLab, where successful execution is more sensitive to page layout variation and page\-specific interaction details\. Forum maintains the highest invocation success rate throughout, indicating that its recurrent workflows are easier to consolidate into stable reusable interaction skills\. Although moderate fluctuations remain at intermediate trajectory budgets, the overall trend suggests that the interaction skill library becomes progressively more robust through repeated failure\-driven refinement\.

### 4\.4Ablation Study

Table 2:Ablation study of DRIVE on the WebArena benchmark\. “Succ\.\-Int\.” denotes interaction skills distilled from successful trajectories\. “Fail\.\-Int\.” and “Fail\.\-Reas\.” represent interaction and reasoning skills refined from failed trajectories, specifically designed to address operation errors and decision errors, respectively\. The full dual\-level framework provides the strongest performance\.MethodShop\.\(%\)CMS\(%\)Forum\(%\)GitLab\(%\)Map\(%\)Avg\.\(%\)Δ\\DeltaAvg\.\(pp\)Baseline37\.241\.560\.441\.446\.845\.5–\+ Succ\.\-Int\.41\.344\.965\.048\.249\.549\.8\+4\.3\+ Fail\.\-Int\.39\.742\.363\.445\.551\.848\.5\+3\.0\+ Fail\.\-Reas\.40\.142\.863\.142\.650\.147\.8\+2\.3DRIVE45\.946\.166\.750\.055\.452\.8\+7\.3Table[2](https://arxiv.org/html/2605.23939#S4.T2)details the performance contributions of different skill induction pathways in DRIVE\. The baseline agent, devoid of any skill library, achieves an average success rate of 45\.5%\. Introducing interaction skills distilled solely from successful trajectories \(Succ\.\-Int\.\) provides a solid initial gain, raising the success rate to 49\.8% \(\+4\.3 pp\)\. Crucially, the remaining ablations validate our failure\-driven reflection mechanism: incorporating interaction skills and reasoning skills refined from failed trajectories yields 48\.5% \(\+3\.0 pp\) and 47\.8% \(\+2\.3 pp\), respectively\. The complete DRIVE framework integrates all induction pathways, achieving the highest performance of 52\.8% \(\+7\.3 pp\)\. These macro\-level gains indicate that extracting knowledge from both successful executions and historical failures is indispensable for robust performance\.

### 4\.5Validation of Failure Attribution

DRIVE abstracts interaction procedures from successful executions, but its skill evolution also relies on learning corrective information from failed trajectories\. This requires deciding whether a failure stems primarily from reasoning or from interaction\. Because incorrect attribution can introduce poorly aligned knowledge into the skill library, we examine both the reliability of the LM\-based attribution module and the effect of attribution quality on downstream performance\.

We randomly sample 30 tasks from each of the five WebArena domains, yielding a 150\-task evaluation subset\. On this subset, the baseline agent achieves a 46\.0% success rate and produces 81 failed trajectories\. Two human annotators independently reviewed the execution traces and labeled each failure as either a reasoning error or an interaction error\. Disagreements were resolved through discussion\. We then compared the LM\-based predictions with these human\-adjudicated labels\.

![Refer to caption](https://arxiv.org/html/2605.23939v1/x3.png)

SettingSuccessRate \(%\)Δ\\Delta\(pp\)No skill46\.0–Human attr\.54\.0\+8\.0LM attr\.49\.3\+3\.3Random attr\.46\.7\+0\.7Reversed attr\.44\.7\-1\.3
\(a\) Attribution confusion matrix

\(b\) Effect of attribution quality

Figure 3:Validation of failure attribution\. \(a\) Confusion matrix between human\-adjudicated and LM\-based failure labels, where the LM\-based attribution module achieves 69\.1% overall accuracy\. \(b\) Success rates under different attribution settings on the 150\-task subset\. Human attribution performs best, while reversed attribution performs worse than the no\-skill baseline\.As shown in Figure[3](https://arxiv.org/html/2605.23939#S4.F3)\(a\), the LM\-based module achieves 69\.1% overall accuracy against human annotations\. It correctly identifies 36 of 47 reasoning errors, corresponding to 76\.6% recall, and 20 of 34 interaction errors, corresponding to 58\.8% recall\. These results suggest that the module captures useful failure\-type signals, especially for reasoning errors\. At the same time, the confusion matrix shows that failure disentanglement remains difficult\. In particular, 14 interaction errors are misclassified as reasoning failures\. This often happens in cascading failures, where a UI constraint, such as an unhandled pop\-up, causes the agent to try implausible alternative actions that the LM interprets as a reasoning error\. Conversely, the 11 reasoning errors misclassified as interaction errors often involve premature task termination: the agent incorrectly assumes that the task has been completed and stops, while the LM treats the failure as a localized execution problem\.

To assess how attribution quality affects downstream skill evolution, we compare five attribution settings on the 150\-task subset\. The No skill setting does not use skills derived from failures\. Human attribution routes failures according to the human\-adjudicated labels, whereas LM attribution uses the default automated module\. Random attribution assigns each failed trajectory to a category at random\. Reversed attribution deliberately flips the LM\-predicted labels, sending predicted reasoning errors to the interaction library and predicted interaction errors to the reasoning library\.

Figure[3](https://arxiv.org/html/2605.23939#S4.F3)\(b\) shows the effect of attribution quality on task performance\. Human attribution achieves the highest success rate, 54\.0%, which is 8\.0 percentage points above the baseline\. This result indicates that failed trajectories are most useful when they are routed to the appropriate skill type\. LM attribution reaches 49\.3%, a gain of 3\.3 percentage points, and outperforms both the no\-skill baseline and random attribution, which improves by only 0\.7 percentage points\. In contrast, reversed attribution reduces the success rate to 44\.7%, or 1\.3 percentage points below the baseline\. This drop supports the separation of skill representations: placing abstract reasoning failures into programmatic interaction skills, or turning UI\-specific failures into general natural\-language reasoning guidance, contaminates the corresponding skill libraries and weakens later task execution\. Although LM\-based attribution is still an imperfect approximation of human judgment, it offers a practical mechanism for supporting skill evolution at both levels\.

### 4\.6Error\-Type Analysis of Dual\-Level Skills

Beyond the overall ablation results, we conduct a fine\-grained error analysis to examine how the two types of skills affect different failure modes\. We use the same 100 randomly sampled test tasks for all skill configurations\. Each configuration is evaluated three times under the same evaluation protocol, and we report the mean and standard deviation across runs\.

We define a fixed failure taxonomy, shown in Table[3](https://arxiv.org/html/2605.23939#S4.T3), that distinguishesreasoning errorsfrominteraction errors\. Reasoning errors reflect failures in task interpretation, exploration, or decision making, whereas interaction errors reflect failures in executing the intended web operation\. Successful trajectories are counted separately\. Each failed trajectory is assigned to exactly one error category, so the three proportions in Table[4](https://arxiv.org/html/2605.23939#S4.T4)are normalized over the same task set and sum to 100% up to rounding\. GPT\-4\.1 is used only as a post\-hoc classifier: given the task instruction, execution trace, and final outcome, it maps each failed trajectory to the fixed human\-designed taxonomy\.

Table 3:Failure taxonomy used for error\-type analysis\. Each failed trajectory is assigned to one of the two mutually exclusive categories\.Error CategorySpecific Failure ModesReasoning ErrorsAnswer omission; insufficient exploration; task misunderstanding; false no\-data judgment; empty answer\.Interaction ErrorsRepeated execution; multi\-field form\-filling failure; oscillatory interaction loop; repeated backtracking\.Table 4:Error\-type analysis on 100 randomly sampled test tasks\. Results are reported as mean±\\pmstandard deviation over three runs\. Correct tasks and the two error categories are normalized over the same task set\.MethodCorrecttasks \(%\)Reasoningerrors \(%\)Interactionerrors \(%\)No skill42\.3±3\.0642\.3\\pm 3\.0635\.7±2\.5235\.7\\pm 2\.5222\.0±3\.5722\.0\\pm 3\.57Reasoning\-only skills48\.0±1\.8148\.0\\pm 1\.8130\.3±2\.3630\.3\\pm 2\.3621\.7±2\.0821\.7\\pm 2\.08Interaction\-only skills47\.7±2\.5247\.7\\pm 2\.5233\.3±0\.5833\.3\\pm 0\.5819\.0±2\.1719\.0\\pm 2\.17Reasoning \+ Interaction skills52\.3±2\.52\\mathbf\{52\.3\}\\pm 2\.5230\.0±2\.03\\mathbf\{30\.0\}\\pm 2\.0317\.7±1\.53\\mathbf\{17\.7\}\\pm 1\.53Table[4](https://arxiv.org/html/2605.23939#S4.T4)shows that the two skill types affect different parts of the failure distribution\. Reasoning\-only skills mainly reduce reasoning errors, while leaving the interaction\-error rate close to that of the no\-skill setting\. In contrast, interaction\-only skills primarily reduce interaction errors, but do not provide the same reduction in reasoning failures\. This pattern is consistent with the intended separation between task\-level decision support and page\-level execution support\.

The full DRIVE configuration combines these two effects\. On the sampled subset, it achieves the highest correctness rate while maintaining a reasoning\-error rate comparable to the reasoning\-only setting and further reducing interaction errors\. These results provide evidence that the two skill levels are complementary: reasoning skills help the agent avoid failures in task understanding, exploration, and stopping\-condition judgment, whereas interaction skills improve the reliability of executable operations under concrete page constraints\.

### 4\.7Case Analysis

We present two representative cases to illustrate why web agents require separate modeling of reasoning skills and interaction skills\. In web tasks, failures do not all call for the same kind of correction\. Some errors stem from high\-level task understanding and stopping\-condition judgment, and therefore require reasoning\-level guidance\. Others arise from brittle execution on dynamic interfaces and therefore require reliable executable procedures\. A single skill form does not handle both cases well\. Natural\-language guidance is often not enough for complex page interactions, while programmatic skills are not suitable for high\-level strategy and intent judgment\. For this reason, reasoning skills and interaction skills should be modeled separately and invoked according to the demands of the current task\.

![Refer to caption](https://arxiv.org/html/2605.23939v1/x4.png)Figure 4:Case study resolving aninteraction error\. The baseline agent fails due to brittle execution on dynamic interface elements \(e\.g\., state and zip\-code field dependencies\)\. In contrast, DRIVE invokes a programmatic interaction skill that encodes a robust procedure for structured form completion, successfully bypassing the page\-specific constraints\.![Refer to caption](https://arxiv.org/html/2605.23939v1/x5.png)Figure 5:Case study resolving areasoning error\. The baseline agent exhibits flawed stopping\-condition judgment, unnecessarily clicking into a product page when the search results already satisfy the query\. DRIVE leverages a natural\-language reasoning skill to correctly frame the task intent and halt execution at the appropriate state\.As shown in Fig\.[4](https://arxiv.org/html/2605.23939#S4.F4), the first case reflects a typical interaction\-level problem\. The agent reaches the target form and fills in several fields correctly, which suggests that it has understood the task intent\. However, it repeatedly fails on the state and zip\-code fields because success depends on handling field dependencies and dropdown interactions reliably\. In this case, additional natural\-language guidance is not enough\. What is needed is an interaction skill that encodes a robust executable procedure for structured form completion under page constraints\.

Fig\.[5](https://arxiv.org/html/2605.23939#S4.F5)shows the opposite case\. After submitting the query, the agent incorrectly continues to a product page even though the search results page already satisfies the user request\. Here, the problem lies not in execution but in task framing and stopping\-condition judgment\. A more robust interaction procedure would not fix this error, because the agent is pursuing the wrong objective\. What is needed instead is a reasoning skill that helps the agent interpret the task correctly and stop at the appropriate point\.

Taken together, these two cases show that neither natural\-language guidance nor programmatic skills alone are sufficient for web tasks\. Effective adaptation requires separate modeling and coordinated use of both reasoning skills and interaction skills\.

### 4\.8Efficiency and Overhead

We further analyze the efficiency overhead introduced by DRIVE\. Since skill retrieval and invocation mainly affect the input context, we focus on input\-token consumption and interaction cost\. As shown in Table[5](https://arxiv.org/html/2605.23939#S4.T5), DRIVE improves the Shopping success rate from 37\.23% to 45\.99%, corresponding to an absolute gain of 8\.76 percentage points\. Although DRIVE increases the average input\-token consumption per task by 13\.9%, the input\-token cost per successful completion decreases from 164\.7k to 151\.9k, yielding a 7\.8% reduction\. This suggests that the additional context overhead introduced by dual\-level skill retrieval is offset by the improved task\-completion rate\.

MethodSR \(%\)Avg\. inputper taskInputper successAvg\. stepsper taskAvg\. callsper taskBaseline37\.2361,316164\.7k9\.310\.8DRIVE45\.9969,858151\.9k9\.113\.3Relative change\+8\.76 pp\+13\.9%\-7\.8%\-2\.2%\+23\.2%Table 5:Overall efficiency and overhead comparison on the Shopping test set \(N=137N=137\)\. SR denotes task success rate\. Input per success is computed by normalizing total input\-token consumption by the number of successful completions\.MethodTotal inputtokensTotalstepsAvg\. inputper taskAvg\. stepsper taskInputper stepBaseline3\.35M46455\.8k7\.77\.22kDRIVE2\.98M40449\.7k6\.77\.38kRelative change\-10\.9%\-12\.9%\-10\.9%\-13\.0%\+2\.3%Table 6:Matched\-subset efficiency comparison on Shopping\. The comparison uses the same matched task subset extracted from the analysis logs, so both methods are evaluated on identical task instances\. Token counts are shown in compact form, where M and k denote millions and thousands, respectively\. All relative changes are computed with respect to the Baseline\.DRIVE also introduces additional orchestration overhead, as reflected by the increase in average calls per task\. However, the average number of environment steps slightly decreases from 9\.3 to 9\.1 on the full Shopping test set\. To further examine execution cost under comparable task conditions, Table[6](https://arxiv.org/html/2605.23939#S4.T6)reports statistics on the same matched task subset extracted from the analysis logs\. On this subset, DRIVE reduces total interaction steps by 12\.9%, average steps per task by 13\.0%, and total input\-token consumption by 10\.9%\. Although the input tokens per step increase slightly, the reduction in trajectory length dominates, leading to lower overall token consumption on the matched subset\.

### 4\.9Generalization to an Open\-Source Backbone

We also evaluate DRIVE with the smaller open\-source backbone Qwen2\.5\-7B\-Instruct[yang2024qwen25](https://arxiv.org/html/2605.23939#bib.bib33)to test whether its gains extend beyond strong proprietary models\. As shown in Table[7](https://arxiv.org/html/2605.23939#S4.T7), DRIVE improves success rates on all five WebArena domains, with an average relative gain of 52%\. The largest improvement appears on Forum, and the other domains also show clear and consistent gains\.

Table 7:Task success rates on WebArena using Qwen2\.5\-7B\-Instruct\. TheΔ\\Deltarow reports relative improvement over the baseline\.MethodShoppingCMSForumGitlabMapAverageBaseline8\.46\.19\.98\.77\.88\.2Ours12\.38\.916\.912\.711\.712\.5Δ\\Delta↑46%\\uparrow 46\\%↑46%\\uparrow 46\\%↑71%\\uparrow 71\\%↑46%\\uparrow 46\\%↑50%\\uparrow 50\\%↑52%\\uparrow 52\\%Although the absolute success rates remain lower than those achieved with stronger backbones, this is not surprising for long\-horizon web tasks\. Such tasks require models to handle large, noisy contexts and make reliable multi\-step decisions under partial observability\. Smaller models are more likely to lose track of relevant state information and make suboptimal decisions during extended interaction, which also reduces the reliability of induced multi\-step skills\. Even so, the consistent gains in Table[7](https://arxiv.org/html/2605.23939#S4.T7)show that the effectiveness of DRIVE is not tied to a particular high\-capacity backbone and can transfer to open\-source models as well\.

## 5Discussion

### 5\.1Why Dual\-Level Skill Modeling Matters

A key implication of this work is that web\-agent experience should not be treated as a single homogeneous memory[dong2022lifelong](https://arxiv.org/html/2605.23939#bib.bib34)\. Web tasks contain two kinds of reusable knowledge: reasoning knowledge, which supports task\-level decisions, and interaction knowledge, which grounds those decisions on concrete pages\. The former must generalize across tasks, whereas the latter must remain executable under local page constraints\. A unified representation is therefore unlikely to preserve both transferability and executability\.

This distinction is especially important in dynamic web environments\. Interface changes may break an execution procedure while leaving the underlying task strategy valid\. Conversely, similar page actions may serve different task goals\. Dual\-level skill modeling addresses this tension by separating abstract decision knowledge from grounded execution knowledge, allowing the agent to reuse both without forcing them into the same memory form\.

### 5\.2Continual Improvement Through Skill Maintenance

Continual improvement requires not only accumulating experience, but also keeping the skill library useful as it grows\. More trajectories can expose broader task patterns and page operations, helping the agent adapt to recurring web scenarios\. However, this benefit can diminish if outdated or redundant skills impair retrieval\.

DRIVE treats skill learning as an iterative maintenance process\. New trajectories expand the library, while skill\-level feedback revises interaction skills that no longer execute reliably and merges reasoning skills with overlapping guidance\. This maintenance reflects the different roles of the two skill types: interaction skills must stay aligned with changing page conditions, whereas reasoning skills must remain sufficiently general for future decisions\. In this way, historical experience becomes an evolving capability base rather than a static archive\.

### 5\.3Limitations and Broader Impacts

Several limitations remain\. First, DRIVE depends on reliable failure attribution; incorrect attribution may update the wrong part of the skill library\. Second, retrieval may become less precise as the library grows, which calls for more scalable selection mechanisms\. Third, our experiments use a DOM\-based benchmark and do not fully capture the visual complexity of real websites[yin2025context](https://arxiv.org/html/2605.23939#bib.bib35)\. Extending DRIVE to multimodal web agents is therefore an important direction for future work[koh2024visualwebarena](https://arxiv.org/html/2605.23939#bib.bib36),[zheng2024gpt](https://arxiv.org/html/2605.23939#bib.bib1),[he2024webvoyager](https://arxiv.org/html/2605.23939#bib.bib2)\.

Continual web\-agent learning also raises safety concerns\. Agents that improve through experience may become more capable, but they may also reinforce unsafe behaviors if experience is reused without constraints[qian2026zero](https://arxiv.org/html/2605.23939#bib.bib37)\. Future work should combine continual learning with safety checks, sandboxed evaluation, and human oversight for high\-risk operations\.

## 6Conclusion

We introduced DRIVE, a dual\-level skill modeling framework for continual learning in web agents\. DRIVE separates historical experience into transferable reasoning skills and executable interaction skills, and introduces specific mechanisms for scenario\-aware retrieval and continuous skill evolution\. This design improves both decision making and execution reliability in dynamic web environments\. Experiments on five WebArena domains show that DRIVE achieves an average success rate of 52\.8%, outperforming existing baselines\. In addition, our ablation studies verify that reasoning skills and interaction skills play distinct but complementary roles in the overall framework\. Overall, our results suggest that separating task\-level reasoning from page\-level execution is an effective way to build web agents that adapt and improve over time\.

## Declaration of generative AI and AI\-assisted technologies in the manuscript preparation process

During the preparation of this work, the authors used ChatGPT to improve language expression and refine the cover letter and parts of the manuscript text\. After using this tool, the authors carefully reviewed and edited the content as needed and take full responsibility for the content of the published article\.

## Acknowledgments

This work was supported by the Huxiang Young Talents in Science and Technology Innovation Project \(No\. 2024RC3148\)\.

## References

- \[1\]B\. Zheng, B\. Gou, J\. Kil, H\. Sun, Y\. Su, Gpt\-4v \(ision\) is a generalist web agent, if grounded, arXiv preprint arXiv:2401\.01614 \(2024\)\.
- \[2\]H\. He, W\. Yao, K\. Ma, W\. Yu, Y\. Dai, H\. Zhang, Z\. Lan, D\. Yu, Webvoyager: Building an end\-to\-end web agent with large multimodal models, in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), 2024, pp\. 6864–6890\.
- \[3\]H\. Lai, X\. Liu, I\. L\. Iong, S\. Yao, Y\. Chen, P\. Shen, H\. Yu, H\. Zhang, X\. Zhang, Y\. Dong, et al\., Autowebglm: A large language model\-based web navigating agent, in: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024, pp\. 5295–5306\.
- \[4\]C\. Liu, Y\. Wang, D\. Li, X\. Wang, Domain\-incremental learning without forgetting based on random vector functional link networks, Pattern recognition 151 \(2024\) 110430\.
- \[5\]B\. Zheng, B\. Gou, S\. Salisbury, Z\. Du, H\. Sun, Y\. Su, Webolympus: An open platform for web agents on live websites, in: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2024, pp\. 187–197\.
- \[6\]Y\. Pan, D\. Kong, S\. Zhou, C\. Cui, Y\. Leng, B\. Jiang, H\. Liu, Y\. Shang, S\. Zhou, T\. Wu, et al\., Webcanvas: Benchmarking web agents in online environments, 2024, URL https://arxiv\.org/abs/2406\.12373\.
- \[7\]S\. Ye, H\. Shi, D\. Shih, H\. Yun, T\. Roosta, T\. Shu, Realwebassist: A benchmark for long\-horizon web assistance with real\-world users, arXiv preprint arXiv:2504\.10445 \(2025\)\.
- \[8\]T\. Xue, W\. Qi, T\. Shi, C\. H\. Song, B\. Gou, D\. Song, H\. Sun, Y\. Su, An illusion of progress? assessing the current state of web agents, arXiv preprint arXiv:2504\.01382 \(2025\)\.
- \[9\]D\. Lee, J\. Lee, K\. Kim, J\. Tack, J\. Shin, Y\. W\. Teh, K\. Lee, Learning to contextualize web pages for enhanced decision making by llm agents, arXiv preprint arXiv:2503\.10689 \(2025\)\.
- \[10\]B\. Gou, R\. Wang, B\. Zheng, Y\. Xie, C\. Chang, Y\. Shu, H\. Sun, Y\. Su, Navigating the digital world as humans do: Universal visual grounding for gui agents, arXiv preprint arXiv:2410\.05243 \(2024\)\.
- \[11\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, Y\. Cao, React: Synergizing reasoning and acting in language models, in: The eleventh international conference on learning representations, 2022\.
- \[12\]Y\.\-n\. Han, J\.\-w\. Liu, Adaptive instance similarity embedding for online continual learning, Pattern Recognition 149 \(2024\) 110238\.
- \[13\]A\. Zhao, D\. Huang, Q\. Xu, M\. Lin, Y\.\-J\. Liu, G\. Huang, Expel: Llm agents are experiential learners, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol\. 38, 2024, pp\. 19632–19642\.
- \[14\]Z\. Z\. Wang, J\. Mao, D\. Fried, G\. Neubig, Agent workflow memory, arXiv preprint arXiv:2409\.07429 \(2024\)\.
- \[15\]Y\. Liu, C\. Si, K\. R\. Narasimhan, S\. Yao, Contextual experience replay for self\-improvement of language agents, in: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), 2025, pp\. 14179–14198\.
- \[16\]B\. Zheng, M\. Y\. Fatemi, X\. Jin, Z\. Z\. Wang, A\. Gandhi, Y\. Song, Y\. Gu, J\. Srinivasa, G\. Liu, G\. Neubig, et al\., Skillweaver: Web agents can self\-improve by discovering and honing skills, 2025, URL https://arxiv\.org/abs/2504\.07079\.
- \[17\]Y\. Zhou, Q\. Yang, K\. Lin, M\. Bai, X\. Zhou, Y\.\-X\. Wang, S\. Levine, L\. E\. Li, Proposer\-agent\-evaluator \(pae\): Autonomous skill discovery for foundation model internet agents, in: Forty\-second International Conference on Machine Learning, 2025\.
- \[18\]V\. Prabhu, Y\. Dai, M\. Fernandez, J\. Gu, K\. Ramakrishnan, Y\. Luo, S\. Savarese, C\. Xiong, J\. Li, Z\. Chen, et al\., Walt: Web agents that learn tools, arXiv preprint arXiv:2510\.01524 \(2025\)\.
- \[19\]H\. Zhong, F\. Faisal, L\. França, T\. Leesatapornwongsa, A\. Szekeres, K\. Rong, S\. Nath, Actionengine: From reactive to programmatic gui agents via state machine memory, arXiv preprint arXiv:2602\.20502 \(2026\)\.
- \[20\]W\. Liu, X\.\-J\. Wu, F\. Zhu, M\.\-M\. Yu, C\. Wang, C\.\-L\. Liu, Class incremental learning with self\-supervised pre\-training and prototype learning, Pattern Recognition 157 \(2025\) 110943\.
- \[21\]S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, et al\., Webarena: A realistic web environment for building autonomous agents, arXiv preprint arXiv:2307\.13854 \(2023\)\.
- \[22\]Z\. Wang, Y\. Sun, X\. Zhang, B\. Xu, Z\. Yang, H\. Lin, Continual learning with high\-order experience replay for dynamic network embedding, Pattern Recognition 159 \(2025\) 111093\.
- \[23\]G\. Liu, S\. Geng, S\. Li, H\. Cui, S\. Zhang, X\. Liu, T\. Liu, Webcoach: Self\-evolving web agents with cross\-session memory guidance, arXiv preprint arXiv:2511\.12997 \(2025\)\.
- \[24\]W\. Zhang, X\. Zhang, H\. Yu, S\. Nie, B\. Wu, J\. Yue, T\. Liu, Y\. Li, Expseek: Self\-triggered experience seeking for web agents, arXiv preprint arXiv:2601\.08605 \(2026\)\.
- \[25\]X\. Huang, J\. Chen, Y\. Fei, Z\. Li, P\. Schwaller, G\. Ceder, Cascade: Cumulative agentic skill creation through autonomous development and evolution, arXiv preprint arXiv:2512\.23880 \(2025\)\.
- \[26\]J\. Qiu, X\. Qi, T\. Zhang, X\. Juan, J\. Guo, Y\. Lu, Y\. Wang, Z\. Yao, Q\. Ren, X\. Jiang, et al\., Alita: Generalist agent enabling scalable agentic reasoning with minimal predefinition and maximal self\-evolution, arXiv preprint arXiv:2505\.20286 \(2025\)\.
- \[27\]Z\. Qi, X\. Liu, I\. L\. Iong, H\. Lai, X\. Sun, W\. Zhao, Y\. Yang, X\. Yang, J\. Sun, S\. Yao, T\. Zhang, W\. Xu, J\. Tang, Y\. Dong, Webrl: Training llm web agents via self\-evolving online curriculum reinforcement learning \(2025\)\.[arXiv:2411\.02337](http://arxiv.org/abs/2411.02337)\.
- \[28\]Z\. Wei, W\. Yao, Y\. Liu, W\. Zhang, Q\. Lu, L\. Qiu, C\. Yu, P\. Xu, C\. Zhang, B\. Yin, et al\., Webagent\-r1: Training web agents via end\-to\-end multi\-turn reinforcement learning, in: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp\. 7920–7939\.
- \[29\]P\. Putta, E\. Mills, N\. Garg, S\. Motwani, C\. Finn, D\. Garg, R\. Rafailov, Agent q: Advanced reasoning and learning for autonomous ai agents, arXiv preprint arXiv:2408\.07199 \(2024\)\.
- \[30\]K\. Yang, Y\. Liu, S\. Chaudhary, R\. Fakoor, P\. Chaudhari, G\. Karypis, H\. Rangwala, Agentoccam: A simple yet strong baseline for llm\-based web agents, arXiv preprint arXiv:2410\.13825 \(2024\)\.
- \[31\]A\. Gandhi, G\. Neubig, Go\-browse: Training web agents with structured exploration, arXiv preprint arXiv:2506\.03533 \(2025\)\.
- \[32\]P\. Sodhi, S\. Branavan, Y\. Artzi, R\. McDonald, Step: Stacked llm policies for web actions, arXiv preprint arXiv:2310\.03720 \(2023\)\.
- \[33\]A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, et al\., Qwen2\.5 technical report, arXiv preprint arXiv:2412\.15115 \(2024\)\.
- \[34\]J\. Dong, Y\. Cong, G\. Sun, T\. Zhang, Lifelong robotic visual\-tactile perception learning, Pattern Recognition 121 \(2022\) 108176\.
- \[35\]J\. Yin, X\. Zhang, L\. Wu, X\. Wang, Context\-aware prompt learning for test\-time vision recognition with frozen vision\-language model, Pattern Recognition 162 \(2025\) 111359\.
- \[36\]J\. Y\. Koh, R\. Lo, L\. Jang, V\. Duvvur, M\. Lim, P\.\-Y\. Huang, G\. Neubig, S\. Zhou, R\. Salakhutdinov, D\. Fried, Visualwebarena: Evaluating multimodal agents on realistic visual web tasks, in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), 2024, pp\. 881–905\.
- \[37\]Y\. Qian, K\. Qian, X\. He, L\. Chen, J\. Zhang, T\. Zhang, H\. Wei, L\. Wang, H\. Wu, B\. Mao, Zero\-permission manipulation: Can we trust large multimodal model powered gui agents?, arXiv preprint arXiv:2601\.12349 \(2026\)\.

Similar Articles

Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval

arXiv cs.AI

This paper proposes SGDR (State-Grounded Dynamic Retrieval), an online skill learning method for web agents that enables stepwise, state-aware skill reuse rather than static task-level retrieval. Experiments on WebArena show SGDR achieves 37.5% success rate with GPT-4.1, a ~10.6% relative gain over strong baselines.