Tag
This paper investigates semantic leakage in audio-video diffusion models through the 'attention triangle' of cross-attention mechanisms, and presents methods to enhance semantic grounding during generation.