Scaling and Distilling Text Embeddings for Better Diffusibility
Summary
The paper shows that scaling text embeddings (e.g., swapping T5-small for T5Gemma-2-270M) greatly improves continuous diffusion language models, and that distilling the scaled embeddings into a more connected latent space makes them easier to generate — achieving Gen. PPL 17.8 that outperforms GPT-2-M.
View Cached Full Text
Cached at: 10/02/26, 04:28 AM
Paper page - Scaling and Distilling Text Embeddings for Better Diffusibility
Source: https://huggingface.co/papers/2610.01016 Glad to share our recent paper on the latent space for continuous diffusion language models (DLMs)!
In this paper, we find thatscalingtext embeddings can greatly boost the performance of continuous DLMs; for instance, by replacing theT5-smallembeddings used in the recentELFmodels with the advancedT5Gemma-2-270M, we can reduce Gen. PPL by about 40%.
But scaling alone is not enough, as the scaled embeddings can be hard to generate. They are so distinctive and informative that embeddings of similar and interchangeable words are far apart. Thus, continuous diffusion struggles to generate such separated and discrete targets. To mitigate this, wedistillthem into a more connected and robust latent space, making them easier for diffusion to generate.
As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL. Our results shed light on language representation learning, and also on a reliable way to scale DLMs.
Code and project page are released too:code|project page. Welcome to check them out!
Similar Articles
TextLDM: Language Modeling with Continuous Latent Diffusion
This paper introduces TextLDM, a method that adapts visual latent diffusion transformers for language modeling by mapping discrete tokens to continuous latents. It demonstrates that this approach, enhanced by representation alignment, matches GPT-2 performance and unifies visual and text generation architectures.
Scaling Properties of Text Conditioning in Visual Generation
This paper studies empirical scaling properties for text conditioning in visual generation, showing that converged diffusion loss scales with structured language in prompts, and introduces methods to improve diffusability and promptability.
Abra: Scaling Diffusion Image Training
This paper presents a systematic scaling law study for text-to-image diffusion models, showing they scale predictably but require significantly more data per parameter than language models for optimal training.
On the Diffusibility of High-Dimensional Latents
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
This paper proposes a latent-to-pixel training strategy for pixel-space text-to-image diffusion models, accelerating convergence and improving inference speed while matching or surpassing latent-space counterparts.