Tag
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.