Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Hugging Face Daily Papers Papers

Summary

This paper compares gaze behavior in the MapTask and MUNDEX corpora to understand common grounding in collaborative tasks, finding that task-directed gaze is associated with aligned references and understood judgments, but the effects are modest.

In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee's gaze. In same-speaker MapTask reference chains, the speaker's gaze entropy is lower at the mention where a previously non-aligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX. Because effects are small and several weaken when recurring participants rather than dialogues are the unit of inference, we treat gaze as one contributing cue to grounding, to be interpreted alongside task and dialogue context.
Original Article
View Cached Full Text

Cached at: 09/17/26, 02:51 AM

Paper page - Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Source: https://huggingface.co/papers/2609.18011 In collaborative tasks with asymmetric information — maps with different landmarks, or a board game only one side knows — mutual understanding has to be built through the interaction, and gaze is one of the few observable traces of that process. We compare two such tasks: HCRC MapTask, where perspectivist labels mark a reference as aligned only when speaker and addressee interpretations match, and MUNDEX, German game explanations with retrospective understanding judgments from both sides. Both annotate gaze as discrete behavioral categories rather than eye-tracking coordinates, but with different category sets (up/down/off; partner/table/away), so we map them into a shared partner/task/away vocabulary and compute gaze features around each grounding-labeled unit.

The associations point the same way in both: aligned references and “understood” judgments come with more task-directed gaze, less partner-directed gaze, lower gaze entropy, and fewer gaze transitions — clearest for whoever leads the task (givers, explainers). In same-speaker MapTask reference chains, speaker gaze entropy drops at the mention where a referent becomes aligned. Effects are small and prediction gains over role/condition controls are modest, so we read gaze as one contributing cue to grounding rather than a standalone signal. Because the representation only needs discrete gaze labels, it should port to other corpora — happy to discuss, especially whether shared category names pick out the same interactional function across quite different tasks.

Similar Articles

Towards One-to-Many Temporal Grounding

Hugging Face Daily Papers

This paper introduces One-to-Many Temporal Grounding (OMTG), a new task for localizing multiple disjoint video segments from a single text query, along with a benchmark, evaluation metrics, a 56k-sample dataset, and novel reward functions that achieve state-of-the-art results, outperforming Gemini 2.5 Pro and Seed-1.8.