Tag
This paper introduces VG-GUIBench, a benchmark to evaluate MLLM-based GUI agents' ability to follow video tutorials, and proposes TASKER, a keyframe extraction method that improves performance on VideoQA and video-guided agentic tasks.
Introduces Teach VLM, a model that extracts step-by-step operational knowledge from mobile screen demonstrations, and the Teach-and-Repeat paradigm that uses this knowledge to guide GUI agents, achieving state-of-the-art performance on a new benchmark.