Tag
The paper introduces GUI-SD-v2, a two-stage on-policy self-distillation framework for multi-turn GUI agents that enhances privilege following and guidance, achieving superior performance on AndroidWorld and MobileWorld benchmarks.
This paper empirically investigates the susceptibility of LLM-based GUI agents to digital nudges, finding that reasoning configuration redirects rather than reduces nudge effects, positioning interface design as a governance concern for autonomous AI.
This paper introduces EvoSkill-GUI, a training-free framework that allows GUI agents to improve skills through in-execution reflection, revision, and reuse, demonstrating performance gains on multiple benchmarks without retraining.
EchoPath introduces a model-agnostic memory system for GUI agents that replays validated execution trajectories, significantly reducing token cost and execution time for enterprise recurrent tasks.
LLaDA-UI is a 16.7B-parameter mixture-of-experts diffusion vision-language agent that achieves strong multimodal GUI performance with block-parallel decoding efficiency, outperforming existing models on benchmarks.
TRACE is a training-free framework that optimizes GUI agent efficiency by ranking visual evidence based on utility and diversity, reducing latency and memory usage through adaptive token management and KV contraction.
The paper introduces ConflictGUI, a benchmark for conflict-aware termination in GUI agents, and proposes ConflictGuard, an inference-time framework to reduce over-compliance and improve performance on conflicting instructions.
This survey examines efficient GUI agents through a systems lens, focusing on observation, memory, action, and runtime optimization, and identifies key recurring ideas like selective reading and hybrid runtimes.
This paper proposes Evidence-First Reflection (EFR) to improve reflection in desktop GUI agents by decoupling action-induced visual difference extraction from outcome verification, yielding accuracy gains of 7.11% on benchmarks.
The paper proposes Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive reinforcement learning framework that incorporates trajectory-level quality signals to improve agent performance by addressing reward-gradient misalignment.
This paper introduces AnTrap, a benchmark for evaluating the robustness of Android GUI agents against runtime anomalies, revealing universal vulnerabilities and differentiating between learnable traps and intrinsic reasoning limitations.
UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstrations to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks.
CoAdapt-GUI is a test-time adaptation framework for mobile GUI agents that jointly adapts workflow context and policy, improving performance on unseen-app benchmarks like AndroidWorld-Generalization and AndroidWorld Plus.
Introduces AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents, with a taxonomy of real-world interface perturbations and a training-free Test-Time Introspective Recovery (TTIR) mechanism.
This paper investigates when hybrid computer-use agents actually choose to use MCP tools versus screenshots, finding that tool availability alone does not guarantee adoption: a reasoning model improves while a non-reasoning model degrades. It also explores training and context compression strategies to close the adoption gap and reduce token costs.
FocusMem introduces a latent memory interface for GUI agents that separates content retention, state-conditioned readout, and a trust gate to improve memory reliability. It consistently outperforms fixed-memory baselines across five GUI-agent benchmarks.
This paper introduces Maga, a method for consolidating domain-specific GUI agents into a single cross-platform policy via structured action distillation, reallocating training signals to focus on erroneous actions. It achieves strong success rates across mobile, web, and desktop benchmarks.
Qwen-UI-Agent is a new foundation GUI agent from Alibaba's Qwen team that handles mobile, computer, web, and DeepSearch tasks with state-of-the-art performance on mobile-use benchmarks and competitive results on computer/browser tasks, combining GUI and CLI actions in a unified action space.
This paper introduces VG-GUIBench, a benchmark to evaluate MLLM-based GUI agents' ability to follow video tutorials, and proposes TASKER, a keyframe extraction method that improves performance on VideoQA and video-guided agentic tasks.
This paper proposes a reinforcement learning framework for computer-use agents that uses autonomous vision-language evaluation as a scalable reward signal, modeling evaluator noise to improve task success rates across desktop environments.