web-agents

Tag

Cards List
#web-agents

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

arXiv cs.AI · 2026-05-29 Cached

This paper introduces GTA, a scalable framework for automatically generating long-horizon, multi-hop web agent tasks with executable trajectories, addressing the lack of process-level supervision in web agent benchmarks. The framework integrates crawling, retrieval-based seeding, and automated quality control to produce realistic tasks across multiple websites.

0 favorites 0 likes
#web-agents

@googledevs: Modern Web Guidance + Chrome DevTools for agents = A powerful new workflow. Matthias Rohmer takes you inside the #Googl…

X AI KOLs Following · 2026-05-26 Cached

Google demonstrated at Google I/O a new workflow of Chrome DevTools with AI agents, including APIs such as WebMCP and HTML-in-Canvas, aiming to make it easy for developers to expose web page functionality to AI agents while maintaining semantics, accessibility, and security boundaries.

0 favorites 0 likes
#web-agents

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

arXiv cs.AI · 2026-05-26 Cached

DRIVE proposes a dual-level skill modeling framework that separates reasoning knowledge from interaction knowledge for web agents under continual learning, achieving a 52.8% task success rate on WebArena, outperforming the skill-free baseline by 7.3 percentage points.

0 favorites 0 likes
#web-agents

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

arXiv cs.LG · 2026-05-21 Cached

Weasel is a trajectory selection method for offline training of web agents that improves out-of-domain generalization by balancing importance and diversity. It achieves up to 12.5x training speedups and improved performance across several benchmarks.

0 favorites 0 likes
#web-agents

Skim: Speculative Execution for Fast and Efficient Web Agents

arXiv cs.AI · 2026-05-19 Cached

Accio is a speculative execution framework that reduces cost and latency for web agents by leveraging offline site-structure profiling and online selection of fast paths, achieving a 1.9x reduction in per-task cost and 33.4% latency reduction while maintaining accuracy.

0 favorites 0 likes
#web-agents

ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

arXiv cs.AI · 2026-05-18 Cached

ShopGym is a framework that converts live e-commerce storefronts into self-contained sandbox shops for realistic, controllable, and reproducible benchmarking of web agents, with synthetic tasks across seven skill categories.

0 favorites 0 likes
#web-agents

SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents

arXiv cs.AI · 2026-05-15 Cached

SimPersona learns discrete buyer personas from raw clickstreams using a VQ-VAE and maps them to persona tokens for LLM-based web agents, achieving high conversion-rate alignment across many live storefronts.

0 favorites 0 likes
#web-agents

WebHarbor - We "dock" the real websites into local for web agents! [R]

Reddit r/MachineLearning · 2026-05-14

WebHarbor packages 15 real websites (Amazon, GitHub, BBC, etc.) as self-contained Flask+SQLite apps in a single Docker image with sub-second reset, designed for reproducible web agent evaluation and training. The project invites community contributions to expand to 100+ sites, with co-authorship opportunities.

0 favorites 0 likes
#web-agents

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

Hugging Face Daily Papers · 2026-05-12 Cached

This paper introduces LongMemEval-V2, a benchmark for evaluating long-term memory systems in web agents, along with two memory methods: AgentRunbook-R and AgentRunbook-C.

0 favorites 0 likes
#web-agents

@AdinaYakup: Qwen released WebWorld an open world model series for web agents 8B/14B/32B+Dataset Apache2.0 +9.9% MiniWob++, +10.9% W…

X AI KOLs Following · 2026-05-11 Cached

Qwen released WebWorld, an open-source model series for web agents (8B/14B/32B) under Apache 2.0, which improves performance on MiniWob++ and WebArena benchmarks.

0 favorites 0 likes
#web-agents

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

arXiv cs.AI · 2026-05-11 Cached

Apple Research introduces Weblica, a framework for creating scalable and reproducible training environments for visual web agents using HTTP caching and LLM-based synthesis.

0 favorites 0 likes
#web-agents

Region4Web: Rethinking Observation Space Granularity for Web Agents

arXiv cs.CL · 2026-05-11 Cached

This paper introduces Region4Web, a framework that improves web agent performance by organizing observation spaces into functional regions rather than individual elements. It demonstrates that this approach reduces observation length and increases task success rates on the WebArena benchmark.

0 favorites 0 likes
#web-agents

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Hugging Face Daily Papers · 2026-04-08 Cached

This paper introduces WebStep, a benchmark and framework for process-level evaluation of web agents using semantic state tracking. It reveals detailed performance differences and error localization beyond terminal success metrics.

0 favorites 0 likes
#web-agents

WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Papers with Code Trending · 2025-07-20 Cached

WebShaper is a formalization-driven framework for synthesizing information-seeking datasets using set theory and Knowledge Projections, achieving state-of-the-art performance on GAIA and WebWalkerQA benchmarks among open-source agents.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback