Tag
Introduces a new problem domain for fine-grained analysis of children's gait behaviors from standard RGB video, along with a new dataset of over 1,100 high-frame-rate sequences and a unified framework, demonstrating that current SOTA methods and MLLMs fail on this clinical task.
This paper investigates object-driven shortcuts that hinder compositional generalization in zero-shot compositional action recognition, proposing RCORE to mitigate verb-collapse and improve unseen composition generalization.