AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Hugging Face Daily Papers Papers

Summary

AURORA-LM introduces a continuous-latent diffusion language model that separates decodable text representation from distribution modeling, achieving strong performance on OpenWebText and XSum while scaling to 1B parameters.

Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces not designed for joint generation and decoding, or compress autoencoded latents to ease diffusion, sacrificing token-level fidelity. Instead of simplifying the representation to suit the generative model, we preserve a high-capacity, decodable text latent and design the diffusion model to learn its distribution directly. We introduce AURORA-LM, a continuous-latent diffusion language model that separates the construction of a decodable text representation from the modeling of its distribution. A Query-based Encoder-Decoder organizes text into a high-capacity, prefix-aligned latent sequence, and a Block-causal Diffusion Transformer learns its distribution through flow matching, generating blocks left to right while denoising positions within each block in parallel. Because such a latent is harder for diffusion to model, AURORA-LM restricts only the noisy-input pathway while retaining the full clean-latent prediction target, accommodating full-width latents without reducing decoder-facing capacity. We further calibrate the noise-level distribution to the latent width, and introduce self-trajectory consistency to bridge independently sampled training noise and iterative denoising at inference. AURORA-LM achieves the strongest performance among evaluated continuous and diffusion-based language models on OpenWebText free generation and XSum summarization. Scaling to 1B parameters with about 1500 EFLOPs of total compute yields further gains, surpassing a larger publicly released latent-diffusion language model under a matched evaluation protocol. All experiments are conducted on Ascend NPUs.
Original Article
View Cached Full Text

Cached at: 08/05/26, 05:43 AM

Paper page - AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Source: https://huggingface.co/papers/2608.02602 Published on Aug 3

#2 Paper of the day Authors:

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Languageremainsanoutlieringenerativemodeling:whileimages,video,andaudioareincreasinglymodeledincontinuouslatentspaces,textgenerationstillreliespredominantlyondiscretetokens.Existingcontinuouslanguagemodelseitherinheritembeddingspacesnotdesignedforjointgenerationanddecoding,orcompressautoencodedlatentstoeasediffusion,sacrificingtoken-levelfidelity.Insteadofsimplifyingtherepresentationtosuitthegenerativemodel,wepreserveahigh-capacity,decodabletextlatentanddesignthediffusionmodeltolearnitsdistributiondirectly.WeintroduceAURORA-LM,acontinuous-latentdiffusionlanguagemodelthatseparatestheconstructionofadecodabletextrepresentationfromthemodelingofitsdistribution.AQuery-basedEncoder-Decoderorganizestextintoahigh-capacity,prefix-alignedlatentsequence,andaBlock-causalDiffusionTransformerlearnsitsdistributionthroughflowmatching,generatingblockslefttorightwhiledenoisingpositionswithineachblockinparallel.Becausesuchalatentisharderfordiffusiontomodel,AURORA-LMrestrictsonlythenoisy-inputpathwaywhileretainingthefullclean-latentpredictiontarget,accommodatingfull-widthlatentswithoutreducingdecoder-facingcapacity.Wefurthercalibratethenoise-leveldistributiontothelatentwidth,andintroduceself-trajectoryconsistencytobridgeindependentlysampledtrainingnoiseanditerativedenoisingatinference.AURORA-LMachievesthestrongestperformanceamongevaluatedcontinuousanddiffusion-basedlanguagemodelsonOpenWebTextfreegenerationandXSumsummarization.Scalingto1Bparameterswithabout1500EFLOPsoftotalcomputeyieldsfurthergains,surpassingalargerpubliclyreleasedlatent-diffusionlanguagemodelunderamatchedevaluationprotocol.AllexperimentsareconductedonAscendNPUs.

View arXiv pageView PDFProject pageGitHub7Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.02602 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.02602 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.02602 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

TextLDM: Language Modeling with Continuous Latent Diffusion

Hugging Face Daily Papers

This paper introduces TextLDM, a method that adapts visual latent diffusion transformers for language modeling by mapping discrete tokens to continuous latents. It demonstrates that this approach, enhanced by representation alignment, matches GPT-2 performance and unifies visual and text generation architectures.

Continuous Latent Diffusion Language Model

Hugging Face Daily Papers

Cola DLM is a hierarchical latent diffusion language model that uses text-to-latent mapping and conditional decoding to achieve efficient, non-autoregressive text generation.