Abra: Scaling Diffusion Image Training
Summary
This paper presents a systematic scaling law study for text-to-image diffusion models, showing they scale predictably but require significantly more data per parameter than language models for optimal training.
View Cached Full Text
Cached at: 08/19/26, 03:58 AM
Paper page - Abra: Scaling Diffusion Image Training
Source: https://huggingface.co/papers/2608.17286
Abstract
Scaling laws for text-to-image diffusion models reveal predictable compute-optimal training requiring far more data per parameter than language models, with robust overtraining behavior and universal curve shapes.
Compute-optimalscaling lawsguide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study fortext-to-image diffusion modelsusing Abra, a controlled family offlow-matching transformerstrained across three orders of magnitude worth of compute (10^{19} to 10^{22} FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately 200 image tokens per parameter, ten times theChinchillacompute-optimalprescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss togenerative quality metrics, optimalCFGsettings,representation quality, and even the shape of the training curves, which collapse onto a universal form.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.17286 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.17286 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.17286 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
This paper proposes a latent-to-pixel training strategy for pixel-space text-to-image diffusion models, accelerating convergence and improving inference speed while matching or surpassing latent-space counterparts.
Scaling Properties of Text Conditioning in Visual Generation
This paper studies empirical scaling properties for text conditioning in visual generation, showing that converged diffusion loss scales with structured language in prompts, and introduces methods to improve diffusability and promptability.
Scaling and Distilling Text Embeddings for Better Diffusibility
The paper shows that scaling text embeddings (e.g., swapping T5-small for T5Gemma-2-270M) greatly improves continuous diffusion language models, and that distilling the scaled embeddings into a more connected latent space makes them easier to generate — achieving Gen. PPL 17.8 that outperforms GPT-2-M.
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis
This paper proposes Nemotron-Labs-Diffusion-Image, a masked discrete diffusion model for high-resolution text-to-image synthesis, introducing a token-editing mechanism and grouped cross-entropy objective to improve token refinement and training efficiency.
Scaling laws for neural language models
Foundational empirical study demonstrating power-law scaling relationships between language model performance and model size, dataset size, and compute budget, with implications for optimal training allocation and sample efficiency.