OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Summary
OmniGUI introduces a step-level benchmark for GUI agents that integrates static images, synchronous audio, and video clips to simulate real smartphone interactions. Evaluation shows current models struggle with temporal and auditory inputs, highlighting the need for omni-modal capabilities.
View Cached Full Text
Cached at: 05/20/26, 02:36 AM
Paper page - OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Source: https://huggingface.co/papers/2605.18758
Abstract
OmniGUI presents a novel multimodal benchmark for GUI agents that incorporates simultaneous audio, video, and image inputs to better simulate real smartphone interactions.
Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to process transient audio cues and temporal video dynamics that are tightly coupled with the moment of action. To bridge this gap, we introduce OmniGUI, the first step-level benchmark designed to evaluateGUI agentsin omni-modalsmartphone environments. OmniGUI provides continuous, interleavedmultimodal inputscomprising static images, synchronous audio, and video clips at every action step. The dataset encompasses 709 expert-demonstrated episodes (2,579 action steps) across 29 applications, systematically annotated with objective multimodal dependency levels. Because dedicated omni-modal GUI agent frameworks are currently in their nascent stage, we select foundational omni-modal models capable of natively processing interleaved inputs to serve as agent proxies for our initial baselines. Our empirical evaluation reveals that while current models exhibit competency on visually static tasks, theiraction predictionperformance degrades significantly in environments requiring synchronous temporal and auditory signals. Furthermore, ablation studies isolate specific operational bottlenecks, notablycross-modal interferencewhen processing task-irrelevant environmental noise. The complete dataset, evaluation pipeline, and baseline prompts are provided in the supplementary material. Project page: https://omni-gui.github.io.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2605\.18758
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.18758 in a model README.md to link it from this page.
Datasets citing this paper1
#### OmniGUI/OmniGUI Viewer• Updated31 minutes ago • 2.61k • 6.21k • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.18758 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
SimuWoB is a synthetic benchmark with 120 challenging tasks for mobile GUI agents, using high-fidelity virtual environments and automatic reward generation. Experiments reveal that current agents achieve only 27.92% average success rate, dropping to 17.82% on long-horizon tasks, indicating substantial weaknesses in complex scenarios.
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
OmniInteract introduces a streaming benchmark for real-time omnimodal LLMs, evaluating online audio-visual processing with temporal grounding and interactive response requirements. Experiments show that current models perform poorly, with the best overall IA-QTF1 score reaching only 0.368.
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
Introduces GUI-CC, a benchmark for evaluating the contextual consistency of GUI world models when used as environments for GUI agents, through offline and online interaction tracks.
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research
MobileGym is a browser-based simulation platform for mobile GUI agent research, featuring deterministic state evaluation and scalable parallel execution. It includes a benchmark of 416 tasks and demonstrates gains using GRPO on Qwen3-VL-4B.
Omni-IO Skills: Harnessing Your Agent Omni-Native
Omni-IO Skills presents a plug-and-play agent harness that enables omni-native capabilities across multiple modalities, significantly improving performance on multimodal benchmarks like UniM-90 without changing the agent's core reasoning.