Tag
The paper introduces RiT, a vanilla Diffusion Transformer trained on frozen DINOv2 features using flow matching with x-prediction, achieving competitive FID scores on ImageNet 256×256 with fewer parameters and fast sampling without distillation.
This paper proposes Sphere Latent Encoder, an efficient few-step image generation framework that performs denoising entirely in a spherical latent space, achieving high-quality 256×256 images with significantly reduced computational cost and improved FID scores on ImageNet-1K.