AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
Summary
AURORA-LM introduces a continuous-latent diffusion language model that separates decodable text representation from distribution modeling, achieving strong performance on OpenWebText and XSum while scaling to 1B parameters.
View Cached Full Text
Cached at: 08/05/26, 05:43 AM
Paper page - AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
Source: https://huggingface.co/papers/2608.02602 Published on Aug 3
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Languageremainsanoutlieringenerativemodeling:whileimages,video,andaudioareincreasinglymodeledincontinuouslatentspaces,textgenerationstillreliespredominantlyondiscretetokens.Existingcontinuouslanguagemodelseitherinheritembeddingspacesnotdesignedforjointgenerationanddecoding,orcompressautoencodedlatentstoeasediffusion,sacrificingtoken-levelfidelity.Insteadofsimplifyingtherepresentationtosuitthegenerativemodel,wepreserveahigh-capacity,decodabletextlatentanddesignthediffusionmodeltolearnitsdistributiondirectly.WeintroduceAURORA-LM,acontinuous-latentdiffusionlanguagemodelthatseparatestheconstructionofadecodabletextrepresentationfromthemodelingofitsdistribution.AQuery-basedEncoder-Decoderorganizestextintoahigh-capacity,prefix-alignedlatentsequence,andaBlock-causalDiffusionTransformerlearnsitsdistributionthroughflowmatching,generatingblockslefttorightwhiledenoisingpositionswithineachblockinparallel.Becausesuchalatentisharderfordiffusiontomodel,AURORA-LMrestrictsonlythenoisy-inputpathwaywhileretainingthefullclean-latentpredictiontarget,accommodatingfull-widthlatentswithoutreducingdecoder-facingcapacity.Wefurthercalibratethenoise-leveldistributiontothelatentwidth,andintroduceself-trajectoryconsistencytobridgeindependentlysampledtrainingnoiseanditerativedenoisingatinference.AURORA-LMachievesthestrongestperformanceamongevaluatedcontinuousanddiffusion-basedlanguagemodelsonOpenWebTextfreegenerationandXSumsummarization.Scalingto1Bparameterswithabout1500EFLOPsoftotalcomputeyieldsfurthergains,surpassingalargerpubliclyreleasedlatent-diffusionlanguagemodelunderamatchedevaluationprotocol.AllexperimentsareconductedonAscendNPUs.
View arXiv pageView PDFProject pageGitHub7Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.02602 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.02602 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.02602 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TextLDM: Language Modeling with Continuous Latent Diffusion
This paper introduces TextLDM, a method that adapts visual latent diffusion transformers for language modeling by mapping discrete tokens to continuous latents. It demonstrates that this approach, enhanced by representation alignment, matches GPT-2 performance and unifies visual and text generation architectures.
Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion
This paper introduces a diffusion language model that treats text as a continuous process over binary bitstreams, using entropy-gated stochastic sampling to close the performance gap with autoregressive models. It achieves state-of-the-art results on LM1B and OWT benchmarks while reducing memory footprint.
Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
This paper introduces Diffusion Language Models (DLMs) as a new inference paradigm for lossless text compression, aiming to overcome the throughput bottlenecks of autoregressive LLM-based compressors while achieving state-of-the-art compression ratios.
Continuous Latent Diffusion Language Model
Cola DLM is a hierarchical latent diffusion language model that uses text-to-latent mapping and conditional decoding to achieve efficient, non-autoregressive text generation.
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
R²-dLLM introduces spatio-temporal redundancy reduction techniques that cut diffusion LLM decoding steps by up to 75% while preserving generation quality, addressing a key deployment bottleneck.