Tag
Proposes FATE, a frame-level audio-visual temporal embedding method that aligns frame sequences on a physical timeline, enabling joint semantic and temporal understanding. It outperforms baselines on temporal retrieval, event localization, and generation evaluation metrics.
Proposes a hierarchical semantic-constrained heterogeneous graph model for open-vocabulary audio-visual event localization, addressing cross-modal consistency at multiple temporal scales and hierarchical semantic constraints between segment and video levels. Achieves state-of-the-art results on OV-AVEL benchmark.