gui-agents

Tag

Cards List
#gui-agents

Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

arXiv cs.AI ↗ · 5d ago Cached

The paper introduces GUI-SD-v2, a two-stage on-policy self-distillation framework for multi-turn GUI agents that enhances privilege following and guidance, achieving superior performance on AndroidWorld and MobileWorld benchmarks.

0 favorites 0 likes
#gui-agents

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

arXiv cs.AI ↗ · 2026-09-18 Cached

This paper empirically investigates the susceptibility of LLM-based GUI agents to digital nudges, finding that reasoning configuration redirects rather than reduces nudge effects, positioning interface design as a governance concern for autonomous AI.

0 favorites 0 likes
#gui-agents

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

arXiv cs.LG ↗ · 2026-09-17 Cached

This paper introduces EvoSkill-GUI, a training-free framework that allows GUI agents to improve skills through in-execution reflection, revision, and reuse, demonstrating performance gains on multiple benchmarks without retraining.

0 favorites 0 likes
#gui-agents

EchoPath: Execution-Level Replayable Memory for GUI Agents

arXiv cs.AI ↗ · 2026-09-16 Cached

EchoPath introduces a model-agnostic memory system for GUI agents that replays validated execution trajectories, significantly reducing token cost and execution time for enterprise recurrent tasks.

0 favorites 0 likes
#gui-agents

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Hugging Face Daily Papers ↗ · 2026-09-09 Cached

LLaDA-UI is a 16.7B-parameter mixture-of-experts diffusion vision-language agent that achieves strong multimodal GUI performance with block-parallel decoding efficiency, outperforming existing models on benchmarks.

0 favorites 0 likes
#gui-agents

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

Hugging Face Daily Papers ↗ · 2026-09-09 Cached

TRACE is a training-free framework that optimizes GUI agent efficiency by ranking visual evidence based on utility and diversity, reducing latency and memory usage through adaptive token management and KV contraction.

0 favorites 0 likes
#gui-agents

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

arXiv cs.AI ↗ · 2026-09-04 Cached

The paper introduces ConflictGUI, a benchmark for conflict-aware termination in GUI agents, and proposes ConflictGuard, an inference-time framework to reduce over-compliance and improve performance on conflicting instructions.

0 favorites 0 likes
#gui-agents

Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

arXiv cs.CL ↗ · 2026-09-03 Cached

This survey examines efficient GUI agents through a systems lens, focusing on observation, memory, action, and runtime optimization, and identifies key recurring ideas like selective reading and hybrid runtimes.

0 favorites 0 likes
#gui-agents

Reflection with Action-Induced Visual Differences for Desktop GUI Agents

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper proposes Evidence-First Reflection (EFR) to improve reflection in desktop GUI agents by decoupling action-induced visual difference extraction from outcome verification, yielding accuracy gains of 7.11% on benchmarks.

0 favorites 0 likes
#gui-agents

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

arXiv cs.AI ↗ · 2026-08-25 Cached

The paper proposes Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive reinforcement learning framework that incorporates trajectory-level quality signals to improve agent performance by addressing reward-gradient misalignment.

0 favorites 0 likes
#gui-agents

Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

This paper introduces AnTrap, a benchmark for evaluating the robustness of Android GUI agents against runtime anomalies, revealing universal vulnerabilities and differentiating between learnable traps and intrinsic reasoning limitations.

0 favorites 0 likes
#gui-agents

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

Hugging Face Daily Papers ↗ · 2026-08-16 Cached

UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstrations to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks.

0 favorites 0 likes
#gui-agents

CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

arXiv cs.AI ↗ · 2026-08-13 Cached

CoAdapt-GUI is a test-time adaptation framework for mobile GUI agents that jointly adapts workflow context and policy, improving performance on unseen-app benchmarks like AndroidWorld-Generalization and AndroidWorld Plus.

0 favorites 0 likes
#gui-agents

AndroidReality: How Far Are Mobile Agents from the Real World?

arXiv cs.AI ↗ · 2026-08-11 Cached

Introduces AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents, with a taxonomy of real-world interface perturbations and a training-free Test-Time Introspective Recovery (TTIR) mechanism.

0 favorites 0 likes
#gui-agents

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

arXiv cs.AI ↗ · 2026-08-05 Cached

This paper investigates when hybrid computer-use agents actually choose to use MCP tools versus screenshots, finding that tool availability alone does not guarantee adoption: a reasoning model improves while a non-reasoning model degrades. It also explores training and context compression strategies to close the adoption gap and reduce token costs.

0 favorites 0 likes
#gui-agents

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

Hugging Face Daily Papers ↗ · 2026-08-05 Cached

FocusMem introduces a latent memory interface for GUI agents that separates content retention, state-conditioned readout, and a trust gate to improve memory reliability. It consistently outperforms fixed-memory baselines across five GUI-agent benchmarks.

0 favorites 0 likes
#gui-agents

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper introduces Maga, a method for consolidating domain-specific GUI agents into a single cross-platform policy via structured action distillation, reallocating training signals to focus on erroneous actions. It achieves strong success rates across mobile, web, and desktop benchmarks.

0 favorites 0 likes
#gui-agents

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

Qwen-UI-Agent is a new foundation GUI agent from Alibaba's Qwen team that handles mobile, computer, web, and DeepSearch tasks with state-of-the-art performance on mobile-use benchmarks and competitive results on computer/browser tasks, combining GUI and CLI actions in a unified action space.

0 favorites 0 likes
#gui-agents

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Hugging Face Daily Papers ↗ · 2026-06-28 Cached

This paper introduces VG-GUIBench, a benchmark to evaluate MLLM-based GUI agents' ability to follow video tutorials, and proposes TASKER, a keyframe extraction method that improves performance on VideoQA and video-guided agentic tasks.

0 favorites 0 likes
#gui-agents

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

arXiv cs.AI ↗ · 2026-06-24 Cached

This paper proposes a reinforcement learning framework for computer-use agents that uses autonomous vision-language evaluation as a scalable reward signal, modeling evaluator noise to improve task success rates across desktop environments.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback