Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
Summary
This paper compares gaze behavior in the MapTask and MUNDEX corpora to understand common grounding in collaborative tasks, finding that task-directed gaze is associated with aligned references and understood judgments, but the effects are modest.
View Cached Full Text
Cached at: 09/17/26, 02:51 AM
Paper page - Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
Source: https://huggingface.co/papers/2609.18011 In collaborative tasks with asymmetric information — maps with different landmarks, or a board game only one side knows — mutual understanding has to be built through the interaction, and gaze is one of the few observable traces of that process. We compare two such tasks: HCRC MapTask, where perspectivist labels mark a reference as aligned only when speaker and addressee interpretations match, and MUNDEX, German game explanations with retrospective understanding judgments from both sides. Both annotate gaze as discrete behavioral categories rather than eye-tracking coordinates, but with different category sets (up/down/off; partner/table/away), so we map them into a shared partner/task/away vocabulary and compute gaze features around each grounding-labeled unit.
The associations point the same way in both: aligned references and “understood” judgments come with more task-directed gaze, less partner-directed gaze, lower gaze entropy, and fewer gaze transitions — clearest for whoever leads the task (givers, explainers). In same-speaker MapTask reference chains, speaker gaze entropy drops at the mention where a referent becomes aligned. Effects are small and prediction gains over role/condition controls are modest, so we read gaze as one contributing cue to grounding rather than a standalone signal. Because the representation only needs discrete gaze labels, it should port to other corpora — happy to discuss, especially whether shared category names pick out the same interactional function across quite different tasks.
Similar Articles
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
This paper investigates whether vision-language models can distinguish potential from established common ground in asymmetric dialogue. Experiments on MapTask data show that providing task-relevant map content (visual or textual) biases models toward over-predicting alignment, as they rely on static referential cues rather than tracking grounding through dialogue history.
ReGround: Grounding Reviewer Comments in Multimodal Evidence
ReGround is a large-scale dataset for grounding reviewer comments in multimodal evidence from scientific papers, revealing the challenges and importance of integrating text, tables, and figures in retrieval tasks.
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.
Towards One-to-Many Temporal Grounding
This paper introduces One-to-Many Temporal Grounding (OMTG), a new task for localizing multiple disjoint video segments from a single text query, along with a benchmark, evaluation metrics, a 56k-sample dataset, and novel reward functions that achieve state-of-the-art results, outperforming Gemini 2.5 Pro and Seed-1.8.
GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
The paper introduces GUI-Primitives, a benchmark of 994 contrastive instruction pairs to diagnose spatial reasoning failures in vision-language models for GUI grounding, revealing that most failures stem from candidate localization rather than relation understanding.