Tag
WildCity introduces a large-scale multimodal dataset for city-scale urban navigation and spatial representation, collected by autonomous fleets. It provides 18 long trajectories and establishes baselines for reconstruction and closed-loop simulation to advance AI systems that can perceive and reason about city-scale environments.
Introduces DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions, and evaluates state-of-the-art models' ability to teach software skills. Models tend to give direct instructions without explanations or visual grounding, while human coaches provide more adaptive guidance.
Introduces MatMMExtract pipeline to decompose compound scientific figures into panels and annotate them using LLMs, creating the MatSciFig dataset of over 390,000 image-text pairs for vision-language learning in materials science.
KITScenes Multimodal is a high-fidelity European autonomous driving dataset with synchronized sensors, complete 3D HD maps, and four benchmarks for spatial learning and embodied AI research.
This paper introduces GroupAffect-4, a multimodal dataset of 40 participants in 10 four-person groups performing collaborative tasks. It includes aligned physiology, eye-tracking, audio, self-report, and personality data, along with benchmark targets for within-person, between-person, and group-level analysis.
BEACON is a large-scale multimodal dataset capturing behavioral signals from Valorant gameplay for continuous authentication and behavioral biometrics research.