Tag
V-RAE proposes a video representation autoencoder that builds semantically organized latents from frozen vision representations to enhance video generation quality, convergence speed, and predictive modeling.