Tag
This paper proposes action-conditioned bisimulation over an empirical predictive state graph to decide when GUI agent memories of two web pages should be merged, improving success on MiniWoB++ over memoryless baselines without any training.
The tweet announces Browser Use Ultrafast, a cloud-based web agent tool that is 10x faster and cheaper, enabling tasks like flight comparison for $0.004 in under 20 seconds.
The paper introduces KNOWS, a benchmark for evaluating web agents on complex, long-horizon tasks that involve synthesizing and organizing knowledge into artifacts, revealing that current agents struggle with visual steps and long-horizon reasoning.
X-Tree recovers reusable skill hierarchies directly from agent trajectories (no LLM calls) and integrates them into offline RL, online RLVR, and on-policy self-distillation, improving success rates on WebArena, ScienceWorld, and WebShop by up to 5.8% over standard training recipes at matched data and budget.
AutoTailor is a meta-agentic framework that converts web trajectories into compact, user-aligned browser automation APIs, improving correctness and reducing token cost and latency in web tasks.
EconSkills introduces a skill library and evaluation framework for web agents to transfer and retrieve procedural knowledge for live economic data retrieval, showing improved efficiency in controlled transfer and competitive performance at library scale.
The author critiques web agent demos for failing on real-world URLs due to fetch-layer issues, emphasizing that data engineering is the harder problem than agent reasoning.
Dhruv Batra explains how web agents can navigate changing websites by using visual information from the screen, learning from interactions, and adapting to layout changes without manual updates.
AutoTailor is a meta-agentic framework that automatically selects and adapts compact sets of browser-automation APIs for web agents, improving accuracy and efficiency through offline filtering and dynamic reselection.
OdoBot is a novel web-agent architecture that uses application behavior modeling to reduce token consumption and improve task success rates, outperforming agents like Agent-E and WebVoyager on the Canvas LMS.
This paper proposes a budget-aware online teaching framework for web agents that reduces teacher calls and compute costs while maintaining performance.
Scaffold is a self-improving framework for visual web agents that induces parametric skills, maintains a recursive hierarchy, and distills skills into model weights, achieving significant performance improvements on benchmarks like WebArena.
Astra achieved 77.3% on the Browser Use Benchmark v2, far surpassing Opus 5 (50.5%) and GPT-5.6 Sol xhigh (49.1%), with 22 of 60 tasks earning full marks compared to zero for Opus 5.
Browser Use partners with Link to enable AI web agents to make credit card purchases using single-use cards, enhancing security and functionality for automated transactions.
The paper proposes a method for monitoring web agents without access to internal model signals, using observable trajectories and key-step supervision to predict failures early. It demonstrates competitive performance with internal-signal baselines across benchmarks.
The author discusses an edge case in their agent-readable-manifest API where webpage elements are not currently interactable due to visibility, highlighting the distinction between element existence and current interactability for web agents.
This paper introduces a framework for constructing verified synthetic web environments to improve the training of web agents, demonstrating enhanced performance and transferability across benchmarks.
The author discusses the problem of context window waste in web agents when scraping bad URLs and asks about methods to filter scrapes using metadata to improve efficiency.
The article explains that web agents are blocked due to inconsistent browser fingerprinting rather than the model or headless setup, and introduces pydoll, a Python library using Chrome DevTools Protocol for undetected automation.
The article discusses challenges in training browser-agent models for sequential decision-making, such as error recovery and memory, and seeks input on optimizing training objectives. The author mentions working on the 'mako' model at tinyfish and invites community feedback.