Tag
This paper proposes a latent-to-pixel training strategy for pixel-space text-to-image diffusion models, accelerating convergence and improving inference speed while matching or surpassing latent-space counterparts.
Proposes Self Gradient Forcing (SGF), a two-pass training strategy for autoregressive video diffusion models that provides missing supervision for writing useful context memory, enabling strong long-video extrapolation even from short training windows.
ILLUME-X is a unified multimodal model for free-form interleaved text-image generation, featuring improved data efficiency, stable training, and a comprehensive evaluation metric called ILScore. It outperforms previous models on tasks like style transfer, image decomposition, and storytelling.