Scaling Properties of Text Conditioning in Visual Generation
Summary
This paper studies empirical scaling properties for text conditioning in visual generation, showing that converged diffusion loss scales with structured language in prompts, and introduces methods to improve diffusability and promptability.
View Cached Full Text
Cached at: 08/03/26, 05:30 AM
Paper page - Scaling Properties of Text Conditioning in Visual Generation
Source: https://huggingface.co/papers/2607.29679
Abstract
Westudyempiricalscalingpropertiesfortextconditioninginvisualgeneration.Suchpropertieshaverarelybeenmeasuredbecausediffusionlossdoesnotscalewiththenumberoftokensinnatural-languageprompts.Surprisingly,wefindthattheconvergeddiffusionlossscaleswiththeamountofstructuredlanguageintheprompt.Toquantifystructuredlanguage,weadapttwocomplementarymeasures:awhite-boxlikelihoodmetric(GPG)andablack-boxattributemetric(ED).Acrosscontrolledtrainingruns,theconvergeddiffusionlossdecreasesapproximatelylinearlywithGPGandfollowsapowerlawwithED.Guidedbythesescalingproperties,weimprovediffusabilitybyconstructingstructuredpromptswithsemanticandgeometricannotationsderivedfromimages,andimprovepromptabilitybytrainingaprompterthroughsupervisedfine-tuning,cold-start,andverifier-gatedon-policydistillation.Theresultingsystemoutperformsallevaluatedopen-weightmodelsonnearlyeverycompositional,reasoning,andworld-knowledgebenchmark,whilematchingorsurpassingthestrongestclosed-weightmodelsonmostevaluations.
View arXiv pageView PDFProject pageGitHub2Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.29679 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.29679 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.29679 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Scaling and Distilling Text Embeddings for Better Diffusibility
The paper shows that scaling text embeddings (e.g., swapping T5-small for T5Gemma-2-270M) greatly improves continuous diffusion language models, and that distilling the scaled embeddings into a more connected latent space makes them easier to generate — achieving Gen. PPL 17.8 that outperforms GPT-2-M.
Abra: Scaling Diffusion Image Training
This paper presents a systematic scaling law study for text-to-image diffusion models, showing they scale predictably but require significantly more data per parameter than language models for optimal training.
Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference
Visual Concept Fusion (VCF) enables dual conditioning on both an image and text prompt in diffusion models at inference time without retraining, using a lightweight aligner and fusion strategy.
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
This paper proposes a latent-to-pixel training strategy for pixel-space text-to-image diffusion models, accelerating convergence and improving inference speed while matching or surpassing latent-space counterparts.
Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
This paper proposes a novel approach that conditions diffusion models on Multimodal Large Language Models (MLLMs) for subject-driven image generation, using VAE-based identity conditioning and a Dual Layer Aggregation module to improve both semantic understanding and identity preservation while mitigating copy-paste artifacts.