@HuggingPapers: Microsoft just released Phi-Ground-Any on Hugging Face A 4B parameter vision model for GUI grounding that achieves SOTA…
Summary
Microsoft has released Phi-Ground-Any, a 4B parameter vision model for GUI grounding on Hugging Face that achieves state-of-the-art results, enabling AI agents to precisely interact with screen elements.
View Cached Full Text
Cached at: 05/09/26, 02:10 PM
Microsoft just released Phi-Ground-Any on Hugging Face
A 4B parameter vision model for GUI grounding that achieves SOTA results on
ScreenSpot-pro and UI-Vision, enabling AI agents to precisely click screen elements. https://t.co/VAgTlRPUbB
Similar Articles
@HuggingPapers: Microsoft just released Lens on Hugging Face A 3.8B parameter text-to-image model delivering efficient training and hig…
Microsoft released Lens, a 3.8B parameter text-to-image model on Hugging Face, capable of efficient training and high-resolution generation up to 1440×1440.
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
DRS-GUI proposes a training-free dynamic region search framework for GUI grounding, using a lightweight UI Perceptor with human-like perceptual actions and Monte Carlo Tree Search to progressively locate instruction-relevant elements. Experiments show a 14% improvement on ScreenSpot-Pro for both general and GUI-specific MLLMs.
VISTA: View-Consistent Self-Verified Training for GUI Grounding
VISTA introduces a view-consistent self-verified training method for GUI grounding that improves GRPO-based coordinate prediction by using multiple target-preserving views, achieving consistent accuracy gains across benchmarks.
MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
The MAI-UI technical report presents a family of foundation GUI agents in multiple sizes, addressing real-world deployment challenges with a self-evolving data pipeline, device-cloud collaboration, and online RL, achieving state-of-the-art results on GUI grounding and mobile navigation benchmarks.
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding
InnerZoom proposes a single-forward framework for cross-layer evidence bridging in GUI grounding, achieving state-of-the-art performance on multiple benchmarks while reducing latency by up to 31.8%.