Efficient One-to-Many Translation with Joint Multi-Stream Diffusion
Summary
This paper proposes a discrete diffusion framework for efficient one-to-many machine translation, achieving sublinear latency scaling with the number of target languages and supporting zero-shot transfer to unseen source languages.
View Cached Full Text
Cached at: 09/16/26, 08:45 AM
# Efficient One-to-Many Translation with Joint Multi-Stream Diffusion
Source: [https://arxiv.org/html/2609.16312](https://arxiv.org/html/2609.16312)
Jacob WhitehillAffiliation:yguan2@wpi\.edu, jrwhitehill@wpi\.edu
###### Abstract
One\-to\-many machine translation \(MT\) is computationally expensive for autoregressive \(AR\) systems, which suffer from linear latency scaling with both sequence length and the number of target languages\. We explore how diffusion can enable multilingual translation with a discrete diffusion framework that refines all target languages in parallel, achieving sublinear latency scaling with the number of targets, and supports deployment as a single unified model to replace multiple independent systems\. Conditioned on a continuous semantic anchor rather than source tokens, our framework supports zero\-shot transfer to unseen source languages without retraining, maintaining approximately75%75\\%of its supervised translation quality on zero\-shot sources\. We investigate the quality\-latency frontier and find that with accelerated sampling, it achieves comparable supervised quality to AR baselines with a2×2\\timesspeedup and11\.9%11\.9\\%better zero\-shot BLEU\. These results highlight the potential of joint multi\-stream diffusion as a practical and flexible alternative for efficient one\-to\-many translation\.
## 1Introduction
Real\-world multilingual translation, such as in live interpreting for international meetings \(e\.g\., the United Nations, the European Parliament\), frequently requires different target languages simultaneously from a single source stream\. In these scenarios, the set of requested languages can change dynamically as different client nodes join or leave the session, demanding a flexible and low\-latency framework\. However, existing approaches predominantly rely on Transformer\-based\([Vaswani et al\., 2017](https://arxiv.org/html/2609.16312#bib.bib1)\)autoregressive \(AR\) architectures, typically either deploying separate one\-to\-one models for each language pair, or exploiting a unified model managing multilingualism via language\-specific control tokens\([Johnson et al\., 2017](https://arxiv.org/html/2609.16312#bib.bib2);[Fan et al\., 2021](https://arxiv.org/html/2609.16312#bib.bib3)\)\. While they attain high accuracy, AR models generate tokens sequentially and require separate decoding passes for each target language\. This computational bottleneck makes AR models expensive for parallel one\-to\-many translation, especially in resource\-constrained applications, as inference latency scales linearly with both the sequence length and the number of targets\([Gu et al\., 2018](https://arxiv.org/html/2609.16312#bib.bib4);[Kasai et al\., 2020](https://arxiv.org/html/2609.16312#bib.bib5)\)\.
To mitigate this issue, non\-autoregressive \(NAR\) translation has been recently investigated\([Gu et al\., 2018](https://arxiv.org/html/2609.16312#bib.bib4);[Ghazvininejad et al\., 2019](https://arxiv.org/html/2609.16312#bib.bib32);[Xiao et al\., 2023](https://arxiv.org/html/2609.16312#bib.bib6)\)\. By generating tokens in parallel, NAR methods offer compelling speed\-quality trade\-offs[Qian et al\. \(2021\)](https://arxiv.org/html/2609.16312#bib.bib7)\. However, most existing methods assume conditional independence among target tokens, resulting in the “multimodality problem”\([Gu et al\., 2018](https://arxiv.org/html/2609.16312#bib.bib4)\)\. Furthermore, naively adapting these frameworks to one\-to\-many translation \(e\.g\., running multiple models in parallel\) is both memory intensive and computationally inefficient\.
As an alternative, diffusion models \(DMs\) have emerged as a powerful NAR paradigm for sequence generation through iterative parallel refinement\([Ho et al\., 2020](https://arxiv.org/html/2609.16312#bib.bib11)\)\. Continuous DMs map tokens into a continuous embedding space\([Li et al\., 2022](https://arxiv.org/html/2609.16312#bib.bib12);[Gong et al\., 2023](https://arxiv.org/html/2609.16312#bib.bib8)\), while discrete DMs operate directly on the categorical vocabulary\([Austin et al\., 2021](https://arxiv.org/html/2609.16312#bib.bib14);[Nie et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib15)\)\. Despite their potential, applying DMs to machine translation \(MT\) remains underexplored\. XDLM[Chen et al\. \(2023\)](https://arxiv.org/html/2609.16312#bib.bib31)uses discrete diffusion with cross\-lingual pretraining but is limited to one\-to\-one tasks\. Scaling such diffusion frameworks raises important design questions of the optimal input representation, attention mechanism, and diffusion inference schedule\.
In this work, we explore the viability and design space of parallel multilingual MT\. As an instantiation, we proposePrismDiff, a discrete diffusion framework for parallel one\-to\-many generation\. Metaphorically, PrismDiff acts like a prism: it “refracts” a shared source representation into multiple target outputs through parallel diffusion refinement\. In multilingual MT tasks, PrismDiff harnesses a language\-agnostic semantic anchor to guide the diffusion process, rather than conditioning directly on source tokens\. This shared anchor not only enables zero\-shot transfer capability to unseen source languages but also improves translation accuracy\.
Our contributions are summarized as follows:
1. 1\.We propose PrismDiff, a multi\-stream discrete diffusion framework for one\-to\-many translation, which refines multiple target languages in parallel from a shared semantic anchor\. Joint optimization over all target streams provides implicit cross\-lingual regularization, achieving better translation quality than running independent one\-to\-one MT systems\.
2. 2\.We explore the quality\-latency frontier of diffusion systems against AR baselines\. We show that under supervised settings, PrismDiff achieves a competitive frontier and outperforms AR systems at matched low\-latency settings\. Furthermore, under source\-side zero\-shot settings, the framework forms a stronger performance frontier than AR baselines\.
3. 3\.Through systematic ablations, we show that effective zero\-shot transfer requires a high\-quality cross\-lingual semantic space rather than an arbitrary multilingual encoder, and the anchor mechanism is not replaceable by increasing the number of diffusion steps alone\.
## 2Related Work
### 2\.1Non\-autoregressive Translation
Non\-autoregressive translation \(NAT\) aims to reduce the sequential bottleneck of autoregressive decoding by predicting target tokens in parallel\([Gu et al\., 2018](https://arxiv.org/html/2609.16312#bib.bib4);[Guo et al\., 2019](https://arxiv.org/html/2609.16312#bib.bib36);[Ghazvininejad et al\., 2019](https://arxiv.org/html/2609.16312#bib.bib32)\)\. This set of methods has produced strong speed\-quality trade\-offs, especially through conditional masked language modeling and iterative refinement\. Several strong NAT systems have been proposed for one\-to\-one translation, for example, conditional masked language models repeatedly update masked positions conditioned on the source sentence\([Ghazvininejad et al\., 2019](https://arxiv.org/html/2609.16312#bib.bib32)\), while glancing training exposes the model to a curriculum of partially observed target tokens\([Qian et al\., 2021](https://arxiv.org/html/2609.16312#bib.bib7)\)\. Although these methods establish compelling quality\-latency trade\-offs in bilingual settings, they are not directly designed for parallel one\-to\-many generation, and do not provide a natural mechanism for sharing a single semantic representation across multiple target streams\. Similarly, XDLM[Chen et al\. \(2023\)](https://arxiv.org/html/2609.16312#bib.bib31)explores discrete diffusion for cross\-lingual generation focusing on one\-to\-one translation\. Our work targets different settings: a unified model decoding into multiple target languages in parallel, which motivates different design choices and baselines\. Previous systems for multilingual translation rely on task prompts or specialized architectures to coordinate target languages\([Johnson et al\., 2017](https://arxiv.org/html/2609.16312#bib.bib2);[Fan et al\., 2021](https://arxiv.org/html/2609.16312#bib.bib3);[Azpiazu and Pera, 2020](https://arxiv.org/html/2609.16312#bib.bib9);[Guan and Whitehill, 2025](https://arxiv.org/html/2609.16312#bib.bib10)\), which are limited in deployment flexibility\. In contrast, our work uses a single diffusion canvas in which multiple target streams are optimized together, allowing for joint one\-to\-many decoding\.
### 2\.2Diffusion Language Model
In the context of language modeling, diffusion language models \(DLMs\) generally fall into two categories: continuous and discrete\. Continuous DLMs operate in embedding space and learn to denoise latent token representations\([Li et al\., 2022](https://arxiv.org/html/2609.16312#bib.bib12);[Gong et al\., 2023](https://arxiv.org/html/2609.16312#bib.bib8)\)\. Discrete DLMs instead define the corruption and reverse processes directly over categorical tokens, often using masking or transition matrices\([Austin et al\., 2021](https://arxiv.org/html/2609.16312#bib.bib14);[He et al\., 2023](https://arxiv.org/html/2609.16312#bib.bib35);[Nie et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib15);[Bie et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib16);[Ye et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib17)\)\. Our method follows the discrete masked diffusion paradigm, which naturally supports parallel token reconstruction and variable sampling budgets\.
Diffusion\-based translation remains less explored than AR and other NAT methods\. XDLM\([Chen et al\., 2023](https://arxiv.org/html/2609.16312#bib.bib31)\)studies cross\-lingual discrete diffusion with pretraining, but focuses primarily on one\-to\-one generation\. In contrast, we target parallel one\-to\-many translation, where the model must generate multiple target languages for the same source sentence concurrently; we also emphasize the quality\-latency frontier exposed by changing the number of sampling steps on the same model\.
### 2\.3Inference\-time Acceleration for Diffusion Models
A practical limitation of DMs is that generation quality depends on the number of denoising steps, creating a trade\-off between latency and output quality\. Prior work has explored techniques to accelerate discrete diffusion inference speed while maintaining quality\([Chen et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib18);[Liu et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib19);[Wu et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib20);[Bartosh et al\., 2026](https://arxiv.org/html/2609.16312#bib.bib37)\), including training\-free acceleration strategies, such as non\-uniform schedules and jumpy sampling\([Yeh et al\., 2024](https://arxiv.org/html/2609.16312#bib.bib22);[Chen et al\., 2024](https://arxiv.org/html/2609.16312#bib.bib21);[Irwin et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib23)\)\.
In our work, we adopt the logarithm\-uniform schedule applied in[Irwin et al\. 2025](https://arxiv.org/html/2609.16312#bib.bib23)as a control knob at inference time, allowing a single trained model to operate at multiple quality\-latency points by varying the number of sampling steps, without modifying model capacity or retraining the model\. This is the key mechanism we use to compare with AR baselines, where latency reduction requires shrinking model depth\.
Figure 1:An illustration of PrismDiff\.“Reset PE” is short for Reset Positional Encoding\. PrismDiff takes one source language sentenceXXas input, generating multiple target languages\{Y\(1\),\(2\),…,Y\(K\)\}\\\{Y^\{\(1\)\},^\{\(2\)\},\\dots,Y^\{\(K\)\}\\\}concurrently\. The source language can be replaced with other zero\-shot languages, such as Spanish\. For simplicity, we use En\-\{De, Fr\} translation as an example\. In practice, the total diffusion refinement steps can be reduced to accelerate\.
## 3Methodology
### 3\.1Preliminaries
We adopt the framework of existing discrete diffusion models, which generate outputs through iterative denoising of masked sequences\([Austin et al\., 2021](https://arxiv.org/html/2609.16312#bib.bib14);[Nie et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib15);[Bie et al\., 2025](https://arxiv.org/html/2609.16312#bib.bib16)\)\. However, we focus on one\-to\-many generation, conditioned on a sequence of latent representations instead of source tokens\.
Given a source sequenceXXwith lengthLsrcL\_\{\\mathrm\{src\}\}, our goal is to generateKKtarget translations𝒴=\{Y\(1\),Y\(2\),…,Y\(K\)\}\\mathcal\{Y\}=\\\{Y^\{\(1\)\},Y^\{\(2\)\},\\dots,Y^\{\(K\)\}\\\}in parallel, where thekk\-th target language sequenceY\(k\)=\(y1k,…,yLkk\)Y^\{\(k\)\}=\(y\_\{1\}^\{k\},\\dots,y\_\{L\_\{k\}\}^\{k\}\)consists of discrete tokens from a vocabulary𝒱\\mathcal\{V\}\. Following existing discrete DLMs, we represent each example as a concatenate a promptpp\(conditioned information\) and a responserr\(target to be generated\)\. In our setting,ppremains unmasked, whilerris progressively corrupted and reconstructed\.
The forward process is a discrete masked corruption process defined overTTdiscrete timestepst∈\{1,…,T\}t\\in\\\{1,\\ldots,T\\\}, whereTTcontrols the total amount of noise applied during training\. Letr0r\_\{0\}be the clean response andrtr\_\{t\}be the corrupted response at timesteptt, wherert,ir\_\{t,i\}denotes the token at the position indexiiin the generated sequence at diffusion timesteptt\. The forward process independently applies\[MASK\]s to the clean responser0r\_\{0\}to produce the corrupted statertr\_\{t\}, based on a scheduleγ\(t\)\\gamma\(t\)\(i\.e\., the probability of each token being masked at timesteptt\), whereγ\(T\)=1\\gamma\(T\)=1results in fully masked sequence\.
The reverse process employs a neural networkpθ\(⋅\)p\_\{\\theta\}\(\\cdot\)to denoise the sequence, which takes the concatenation\[p;rt\]\[p;\\ r\_\{t\}\]of the promptppand the corrupted sequencertr\_\{t\}, and predicts the probability of the original response tokenpθ\(r0,i\|p,rt\)p\_\{\\theta\}\(r\_\{0,i\}\\ \|\\ p,\\ r\_\{t\}\)for each masked positioni∈ℳti\\in\\mathcal\{M\}\_\{t\}, whereℳt\\mathcal\{M\}\_\{t\}denotes all masked positions at timett\. The training objective is to maximize the variational lower bound of data log\-likelihood, optimized via a cross\-entropy loss restricted to masked tokens:
ℒ\(θ\)≜−𝔼t,p,r0,rt\[∑i∈ℳtlogpθ\(r0,i\|p,rt\)\]\\mathcal\{L\}\(\\theta\)\\triangleq\-\\mathbb\{E\}\_\{t,p,r\_\{0\},r\_\{t\}\}\\left\[\\sum\_\{i\\in\\mathcal\{M\}\_\{t\}\}\\,\\mathrm\{log\}\\,p\_\{\\theta\}\\big\(r\_\{0,i\}\\ \|\\ p,\\ r\_\{t\}\\big\)\\right\]\(1\)
The inference stage begins from a fully masked sequence, and iteratively denoises it over multiple refinement steps\. At each step, the model first predicts the clean sequence from the current noisy state, then performs a remasking process that selects a fraction of tokens to be masked again, ensuring the sequence becomes progressively less noisy\.
### 3\.2Proposed Method
An illustration of our framework is in Figure[1](https://arxiv.org/html/2609.16312#S2.F1)\.
#### Language\-Agnostic Semantic Anchor\.
Conditioning on source tokens would bind the model to a specific source language, preventing zero\-shot transfer\. To decouple the generation process from source language surface forms and enable zero\-shot transfer, we condition the diffusion process on a continuous representation rather than on source tokens directly\. Specifically, we extract the last hidden statesH∈ℝLsrc×dmodelH\\in\\mathbb\{R\}^\{L\_\{\\mathrm\{src\}\}\\times d\_\{\\mathrm\{model\}\}\}from the source inputXXusing LaBSE \(Language\-agnostic BERT Sentence Embedding\)\([Feng et al\., 2022](https://arxiv.org/html/2609.16312#bib.bib13)\), a pretrained multilingual encoder optimized to align parallel sentences across languages into a shared embedding space\. ThenHHis projected into the diffusion model’s dimensiondmodeld\_\{\\mathrm\{model\}\}via a learnable linear projection to obtain the anchorZ∈ℝLsrc×dmodelZ\\in\\mathbb\{R\}^\{L\_\{\\mathrm\{src\}\}\\times d\_\{\\mathrm\{model\}\}\}\.
#### Multi\-Stream Generation\.
To enable parallel generation, we construct a unified “canvas”CtC\_\{t\}at timestept∈\[0,T\]t\\in\[0,T\]\. This canvas concatenates the anchorZZwith allKKnoisy target streams:Ct=\[p0;rt\]=\[Z;Yt\(1\);…;Yt\(K\)\]C\_\{t\}=\[p\_\{0\};\\ r\_\{t\}\]=\[Z;\\ Y\_\{t\}^\{\(1\)\};\\ \\dots;\\ Y\_\{t\}^\{\(K\)\}\]whereYt\(k\)Y\_\{t\}^\{\(k\)\}represents thekk\-th target language stream at diffusion steptt\. The neural network processes the entireCtC\_\{t\}simultaneously and outputs predictions for all masked positions across all streams\. Masked tokens in thekk\-th stream are replaced by a language\-specific mask token\[MASK\]k\\texttt\{\[MASK\]\}\_\{k\}, added as special tokens to the shared BPE vocabulary, rather than a generic mask\.
Full cross\-stream attention introduces interference between target languages during generation\. We therefore replace full cross\-stream attention with exclusive attention within each stream\. While the attention is isolated, all streams share the same model parameters and are jointly optimized, providing implicit cross\-lingual regularization that benefits each individual stream\.
Since the canvas concatenates multiple sequences, standard positional encoding would assign monotonically increasing indices\. Instead, we employ Reset Positional Encoding \(EposE\_\{\\mathrm\{pos\}\}\), as illustrated in the upper left corner of Figure[1](https://arxiv.org/html/2609.16312#S2.F1)\. For the anchorZZ, position indices range from 0 toLsrc−1L\_\{src\}\-1\. For each target streamY\(k\)Y^\{\(k\)\}, we restart the position indices from 0 toLtgt−1L\_\{tgt\}\-1\. This design allows for the dynamic reordering, removal or addition of target languages without affecting performance, and encourages the model to correctly identify token positions within individual streams\.
#### Training\.
During training, we sample a timestept∼𝒰\(1,T\)t\\sim\\mathcal\{U\}\(1,T\)and corrupt the target streams, while the anchorZZremains uncorrupted\. The model predicts the original tokens from all masked positions across all streams in parallel\. The loss function is adapted from Eq\.[1](https://arxiv.org/html/2609.16312#S3.E1):
ℒ\(θ\)=−𝔼t,Z,r0,rt\[∑i∈ℳtlogpθ\(r0,i\|Z,rt\)\]\\mathcal\{L\}\(\\theta\)=\-\\mathbb\{E\}\_\{t,Z,r\_\{0\},r\_\{t\}\}\\left\[\\sum\_\{i\\in\\mathcal\{M\}\_\{t\}\}\\,\\mathrm\{log\}\\,p\_\{\\theta\}\\big\(r\_\{0,i\}\|Z,r\_\{t\}\\big\)\\right\]\(2\)wherert=\[Yt\(1\);…;Yt\(K\)\]r\_\{t\}=\[Y\_\{t\}^\{\(1\)\};\\ \\dots;\\ Y\_\{t\}^\{\(K\)\}\]is the concatenation of all target streams\.
#### Inference\.
In the inference phase, we start with a fully masked canvasCTC\_\{T\}with lengthLLcontaining only the anchorZZand language\-specific masks\. At each refinement step, the model retains the top\(1−m\)×L\(1\-m\)\\times Lmost confident tokens and remasks the remainder, where the masking ratiommis determined by the schedule\. We experiment on both the vanilla schedule and a logarithm\-uniform schedule following[Irwin et al\. 2025](https://arxiv.org/html/2609.16312#bib.bib23)to progressively unmask tokens\. Specifically, the logarithm\-uniform mask distribution over the masking ratiommis defined as:
logm∼𝒰\(log1L,0\)\\log m\\sim\\mathcal\{U\}\\left\(\\log\\frac\{1\}\{L\},0\\right\)\(3\), the masking ratiommis bounded within\[1/L,1\]\[1/L,1\], and1/L1/Lcorresponds to the single\-token mask limit\. See more details in[A\.1](https://arxiv.org/html/2609.16312#A1.SS1)\.
Since discrete diffusion models lack explicit autoregressive constraints, they may generate consecutive repetitive tokens\. To mitigate this, we apply Connectionist Temporal Classification \(CTC\) decoding[Graves et al\. \(2006\)](https://arxiv.org/html/2609.16312#bib.bib24)as a post\-processing step without training on CTC loss\. Specifically, we use the CTC collapse rule \(removing consecutive duplicates and blanks\) on the final generated sequence, which effectively acts as a structural regularization for NAR output\.
## 4Experimental Setup
### 4\.1Baselines
We compare PrismDiff with the following AR and diffusion baselines:
- •AR: Decoder\-only AR Transformers conditioned on the same source\-side representations \(anchorZZ\), and trained on all En\-X pairs with task prompts\. We vary the number of decoder layers to trace the AR quality\-latency frontier\. Our preliminary experiments \(detailed in[A\.3](https://arxiv.org/html/2609.16312#A1.SS3)\) show that this architecture outperforms other AR settings such as encoder\-decoder\.
- •Diff\-Indep: Independent diffusion decoders conditioned on the same anchorZZ, but trained separately for each En\-X pair\. This isolates the contribution of joint multi\-stream optimization from the anchor mechanism\.
- •Diff\-SrcTokens: A multi\-stream diffusion model identical to PrismDiff, but conditioned on source tokens instead of anchorZZ\. This tests whether the gains of PrismDiff come from anchor\-based semantic conditioning\.
- •Transformer\-Encoder Trees \(TET\): A strong concurrent one\-to\-many NAT model[Guan and Whitehill \(2025\)](https://arxiv.org/html/2609.16312#bib.bib10)that generates multiple targets with a unified encoder\-tree architecture with CTC\. TET is trained on En\-X tasks but lacks zero\-shot transfer capability\.
We do not directly compare against one\-to\-one NAT systems such as GLAT and XDLM in a one\-to\-many setting, as our focus is on isolating specific design decisions within a unified diffusion framework\. The Diff\-Indep is designed for controllable comparisons: it uses the same architecture and anchor as PrismDiff but generates each languages independently\.
### 4\.2Datasets and Evaluation Metrics
#### Datasets\.
We utilize Multi30K[Elliott et al\. \(2016\)](https://arxiv.org/html/2609.16312#bib.bib25)as our primary benchmark\. Multi30K is a multilingual dataset of image descriptions with short, visually grounded sentences\. We use the English\-to\-German and English\-to\-French \(En\-\{De, Fr\}\) pairs \(29k\) for training, and 1k for testing\. To evaluate zero\-shot transfer performance, we create a synthetic Spanish\-to\-X \(Es\-X\) test set by translating English into Spanish using NLLB\-200\([Costa\-Jussà et al\., 2022](https://arxiv.org/html/2609.16312#bib.bib27)\)\.
To further verify the effectiveness of anchor mechanism, we conduct supplementary experiments on Europarl\([Koehn, 2005](https://arxiv.org/html/2609.16312#bib.bib26)\), a large\-scale, aligned parallel corpus from European Parliament proceedings\. We construct a training set of 260k English\-to\-French, Dutch, Romanian, and Danish examples \(En\-\{Fr, Nl, Ro, Da\}\), and 33k for testing, with Spanish serving as the unseen source language\. Europarl’s Spanish evaluation set comes from original human\-authored corpus, allowing us to rigorously verify our zero\-shot transfer claims\.
#### Metrics\.
Translation quality and cross\-lingual semantic transfer are measured with SacreBLEU111The SacreBLEU signature isnrefs:1\|case:mixed\|eff:no\|tok:13a\|smooth:exp\|version:2\.5\.1\.\([Post, 2018](https://arxiv.org/html/2609.16312#bib.bib28)\)and COMET222COMET is with[Unbabel/wmt22\-comet\-da](https://unbabel/wmt22-comet-da)\(v2\.2\.7\)\([Rei et al\., 2020](https://arxiv.org/html/2609.16312#bib.bib29)\)\. Inference efficiency is reported as the average end\-to\-end wall\-clock latency \(ms\) for generating all requested target sentences for a source sentence on a single model instance\. Note that AR systems can reduce wall\-clock latency by running multiple model replicas in parallel, which increases memory and deployment complexity\. Our comparison focuses on one\-to\-many latency in a single model\.
### 4\.3Implementation Details
All diffusion systems use an 8\-layer Transformer encoder backbone \(dmodel=768d\_\{model\}=768,nhead=8n\_\{head\}=8\) to serve as the diffusion decoder; each contains about 230M parameters\. Unless otherwise noted, the diffusion systems use a total number of diffusion stepsT=100T=100on Multi30K andT=40T=40on Europarl with accelerated sampling at different numbers of sampling steps\. This allows a single trained model to exhibit multiple operating points by adjusting the number of sampling stepstt\. We also reportT=10T=10diffusion results as supporting comparisons\. We set the max length of diffusion blocksLsrc=40L\_\{\\mathrm\{src\}\}=40for Multi30K andLsrc=80L\_\{\\mathrm\{src\}\}=80for Europarl\.
The decoder\-only AR models use similar Transformer settings, contain about 247M, 216M, and 208M parameters in the 8\-layer, 4\-layer, 3\-layer versions, respectively\. AR models use the same max length as diffusion systems\.
The Transformer\-Encoder Tree \(TET\) model[Guan and Whitehill \(2025\)](https://arxiv.org/html/2609.16312#bib.bib10)for Europarl has 4 leaf nodes \(corresponding to the targets\), each with a depth of 6\. It contains 16 distinct Transformer encoder layers in total, distributed across its tree topology, resulting in a model size of 556M\.
LaBSE is kept frozen during training; only the linear projection layer is trained as an adapter\. The models are optimized using AdamW\([Loshchilov and Hutter, 2019](https://arxiv.org/html/2609.16312#bib.bib30)\)with a learning rate of 5e\-5\. Models are trained for 60 epochs, and models with the best validation loss are used for evaluation\. All models are trained and evaluated on a single NVIDIA L40S GPU, using batch size 8 during training and batch size 1 for evaluation\.
#### Decoding Protocol\.
All models use a BPE vocabulary across all target languages\. In the inference phase of all diffusion models, CTC collapse \(as mentioned in Sec\.[3\.2](https://arxiv.org/html/2609.16312#S3.SS2)\) is applied after tokenization decoding\. We find this technique consistently outperforms \(by an average of 2\.5%\) vanilla decoding methods by removing consecutive duplicate tokens at the subword level\. See[A\.3](https://arxiv.org/html/2609.16312#A1.SS3)for detailed results\. For both diffusion and AR systems, we truncate each decoded sequence at the first EOS token\.
Table 1:Main results on Multi30K\.Metrics are averaged on De and Fr\. Es is a zero\-shot source language\.ttis the accelerated sampling steps using logarithm\-uniform scheduling\. All diffusion systems are trained withT=100T=100\. Latency is reported in ms\. Bold text denotes the best results under comparable latency budgets \(or equivalent sampling stepstt\)\. See[A\.3](https://arxiv.org/html/2609.16312#A1.SS3)for details\.Src\.ModelttBLEUCOMETLatencyEnTET26\.640\.63012\.1Diff\-SrcTokens4026\.230\.728224\.82024\.480\.73593\.7PrismDiff4029\.540\.778239\.32028\.840\.785112\.7EsTET0\.140\.26511\.4Diff\-SrcTokens400\.030\.320225\.6200\.330\.32095\.8PrismDiff4023\.470\.745237\.02022\.100\.745112\.1
Table 2:Main results on Europarl\.Metrics are averaged on Fr, Nl, Ro, and Da\. Es is a zero\-shot source language\. All diffusion systems are trained withT=40T=40\. TET generates all targets in a single CTC pass without sequential or iterative steps, thus achieving low latency\.
## 5Results and Analysis
Table[1](https://arxiv.org/html/2609.16312#S4.T1)and Table[2](https://arxiv.org/html/2609.16312#S4.T2)summarize the performance on Multi30K and Europarl\. Rather than comparing a single decoding configuration, we analyze the quality–latency frontier induced by diffusion sampling steps and AR decoder depth\. Compared with strong AR baselines, PrismDiff achieves a competitive quality–latency trade\-off: it reaches comparable SacreBLEU and COMET scores with substantially lower latency, while showing advantages in source\-side zero\-shot transfer and parallel one\-to\-many generation\.
### 5\.1Supervised Quality\-Latency Frontier
Figure[2](https://arxiv.org/html/2609.16312#S5.F2)\(left\) compares PrismDiff with strong decoder\-only AR baselines\. On supervised En\-X tasks, PrismDiff reaches most of its BLEU gains within 25–30 sampling steps \(tt\), corresponding to roughly 70–100 ms latency\. Beyond this regime, additional sampling steps only bring marginal BLEU gains, while PrismDiff remains competitive with the best AR operating points over a wide latency range\.
The key distinction between the two frontiers is the underlying trade\-off mechanisms\. To achieve faster inference, AR models must reduce decoder depth \(model capacity\), leading to a substantial drop in BLEU\. In contrast, diffusion\-based systems obtain faster operating points by reducing the number of refinement steps at inference time, without sacrificing model capacity\. The left portion of the curve highlights this advantage: under strict low\-latency budgets, PrismDiff gracefully maintains near\-peak translation quality \(∼\\sim37 BLEU\), whereas the AR frontier degrades sharply\. Hence, this comparison does not claim that diffusion is superior to AR in every supervised setting; rather, it shows that diffusion systems provide a competitive performance frontier and a flexible control knob\.
Figure 2:Quality\-latency frontiers on Multi30K\.Solid curves denote supervised En\-X results, and dashed curves denote source\-side zero\-shot Es\-X results\.Left: Comparison between AR models and PrismDiff\.Right: Comparison among diffusion\-based systems\. We omit the Es\-X results for Diff\-SrcTokens as its BLEU is<1<1\.TTindicates the total diffusion steps, and the latency operating points are obtained by varying the number of sampling steps\. For AR models, operating points are obtained by varying the decoder depth from 3 to 8\.
### 5\.2Source\-Side Zero\-Shot Transfer
The architectural advantages become even more pronounced under source\-side zero\-shot transfer\. As shown in Figure[2](https://arxiv.org/html/2609.16312#S5.F2), when evaluating the English\-trained systems on unseen Spanish sources \(Es\-X\), the PrismDiff frontier consistently outperforms the AR baselines across all points\. Because both models utilize the identical source\-side semantic representation, this performance gap cannot be attributed to differences in input features\. Instead, it suggests that iterative refinement over a shared semantic anchor is more robust to source\-side domain shift than traditional left\-to\-right decoding\.
We speculate that this is due to the cumulative effect of generation errors\. In AR models, out\-of\-distribution inputs may trigger early decoding errors, which can cascade due to autoregression\. In contrast, PrismDiff’s bidirectional context and iterative refinement mitigate such error accumulation, resulting in a more stable generation process\.
This robustness is important for real\-world deployment, where models trained on high\-resource source languages encounter input from unseen sources\. While PrismDiff does not completely bridge the supervised\-to\-zero\-shot gap, it shows a stronger zero\-shot frontier, suggesting that iterative refinement can more effectively exploit the semantic meaning regardless of the source language\.
### 5\.3Ablation Studies
#### Effect of Semantic Anchor\.
Figure[2](https://arxiv.org/html/2609.16312#S5.F2)and Table[1](https://arxiv.org/html/2609.16312#S4.T1)isolate the effect of the anchor from increasingTT\. A Diff\-SrcTokens model trained withT=10T=10peaks at 33\.84 BLEU; increasing toT=100T=100improves this to 35\.07\. Under the sameT=100T=100setting, PrismDiff consistently outperforms Diff\-SrcTokens across matched latency budgets in supervised scenarios, confirming that a largerTTimproves the sampling frontier but cannot fully explain PrismDiff’s gains\. The most important divergence emerges in zero\-shot scenarios: without the anchor, Diff\-SrcTokens degrades severely \(BLEU<1<1on Es\-X\), whereas PrismDiff maintains strong performance\. Thus, a largerTTand the semantic anchor play complementary roles: the former improves latency\-quality trade\-offs, whereas the latter provides the necessary cross\-lingual mapping for zero\-shot transfer\.
Interestingly, PrismDiff performs better with exclusive attention than full attention, whereas Diff\-SrcTokens shows the opposite trend\. This suggests that without the anchor, cross\-stream attention serves as a necessary channel for implicit cross\-lingual alignment; once the anchor fulfills this alignment, cross\-stream attention becomes redundant and may introduce interference between language\-specific surface forms\. \(See Table[5](https://arxiv.org/html/2609.16312#A1.T5)\)
#### Effect of Joint Multi\-Stream Optimization\.
While Diff\-Indep uses the same diffusion mechanism and semantic anchor under matchedT=100T=100setting, it generates each target language independently\. This ablation therefore clarifies whether PrismDiff’s gains come solely from the anchored diffusion process\. As shown in Figure[2](https://arxiv.org/html/2609.16312#S5.F2)and Table[1](https://arxiv.org/html/2609.16312#S4.T1), PrismDiff consistently outperforms Diff\-Indep in both supervised and zero\-shot settings: att=30t=30, it maintains a∼\\sim1\.5 and∼\\sim1\.4 BLEU advantage over Diff\-Indep on supervised and zero\-shot setting, respectively\. This suggests that jointly optimizing all streams provides not only deployment convenience for one\-to\-many translation, but also implicit cross\-lingual regularization beyond what the anchor alone achieves\.
#### Effect of Reset Positional Encoding\.
We also ablate the Reset Positional Encoding by replacing it with standard absolute positional encoding\. As shown in Figure[2](https://arxiv.org/html/2609.16312#S5.F2), this leads to a substantial performance drop \(23\.8 vs\. 33\.5 BLEU att=10t=10\), indicating the necessity for the model to correctly identify token positions within each stream\.
Figure 3:Encoder ablationfor PrismDiff \(T=100T=100\) on Multi30K\. Sold curves denote supervised En\-X frontiers, and dashed curves denote zero\-shot Es\-X frontiers\.
### 5\.4Encoder Choice and Anchor Quality
We further investigate whether the zero\-shot transfer capability depends on the specific choice of multilingual encoder by comparing LaBSE against mBERT\([Devlin et al\., 2019](https://arxiv.org/html/2609.16312#bib.bib33)\)and XLM\-R\([Conneau et al\., 2020](https://arxiv.org/html/2609.16312#bib.bib34)\)\.
Figure[3](https://arxiv.org/html/2609.16312#S5.F3)demonstrates that the effectiveness of the semantic anchor relies on the specific properties of the multilingual encoder\. When integrating LaBSE, mBERT, and XLM\-R into the same PrismDiff architecture, all three encoders yield comparable performance on supervised En\-X tasks\. However, their zero\-shot Es\-X frontiers diverge drastically\. While LaBSE maintains robust zero\-shot performance of near 28 BLEU, both XLM\-R and mBERT collapse to significantly lower performance levels \(near 13 and 10 BLEU, respectively\)\.
Importantly, we do not claim LaBSE is universally the best encoder choice across all decoding architectures\. Experiments on AR models reveal that mBERT can achieve higher supervised quality \(∼\\sim39 BLEU\) under traditional AR decoding, despite still failing catastrophically in Es\-X zero\-shot transfer \(∼\\sim8\.5 BLEU\)\. This suggests that the optimal encoder choice is dependent on both the decoding mechanism and the specific task\. For PrismDiff, specifically, LaBSE is more effective because its language\-agnostic sentence\-level representations perfectly meet the need of a shared anchor to guide the iterative diffusion process\.
### 5\.5Deployment Flexibility and Scaling
Our experiments further validate PrismDiff’s architectural flexibility for real\-world deployment\. First, the framework is inherently invariant to the generation order of target languages\. For example, generating En\-\{De, Fr\} versus En\-\{Fr, De\} yields identical performance\. Second, dynamically omitting target streams at inference time impacts translation quality by less than 1%, proving that the system allows for flexible configuration without retraining\. A core motivation of our framework is efficient one\-to\-many generation\. We analyze the latency and memory scaling behavior as the number of target languages \(KK\) increases \(see[A\.2](https://arxiv.org/html/2609.16312#A1.SS2)\)\. Traditional AR systems suffer from strict linear latency scaling, as generatingKKlanguages sequentially multiplies the temporal bottleneck\. In contrast, although the canvas expands with more target streams, PrismDiff exhibits a sublinear latency growth thanks to the parallel diffusion refinement\.
## 6Conclusion
In this work, we explore the design space of parallel multilingual machine translation\. To illustrate this, we propose PrismDiff, a discrete diffusion framework for parallel one\-to\-many translation\. By guiding the generative process with a language\-agnostic semantic anchor, PrismDiff enables parallel one\-to\-many translation with high semantic consistency, controllable quality\-latency trade\-offs, and robust source\-side zero\-shot transfer\. These results suggest that parallel multi\-stream diffusion with a semantic anchor is a viable and flexible approach that offers a different quality\-latency trade\-off compared to AR systems\. In supervised settings, PrismDiff with accelerated sampling achieves comparable translation quality to strong AR baselines while delivering a 2×\\timesspeedup\. More notably, PrismDiff establishes a stronger zero\-shot frontier: the semantic anchor design enables the framework to maintain approximately 75% of its supervised performance on unseen source languages, outperforming its AR counterparts in zero\-shot BLEU\.
Looking forward, our framework holds the potential to generalize to other one\-to\-many or many\-to\-many generation tasks beyond translation\. Exploring more robust anchor alignment strategies could further improve performance\.
## Limitations
Despite its advantages in latency and flexibility, PrismDiff has several limitations\.
First, as a diffusion\-based NAR model, it does not uniformly outperform strong AR baselines on supervised translation quality\. Its advantage lies in a controllable frontier: fewer refinement steps reduce latency, while additional steps enhance quality\. Diffusion decoding also remains slower than single\-step NAR models, necessitating further acceleration strategies\.
Second, our current approach relies on a pretrained language\-agnostic sentence encoder \(LaBSE\) to extract the semantic anchor\. Although this anchor enables strong zero\-shot transfer in PrismDiff, encoder ablations show that not every multilingual encoder provides a sufficiently aligned semantic space\. A poorly represented language, domain, or a truly distant low\-resource language in the pretrained encoder may greatly degrade transfer performance\. Moreover, the best source representation can be decoder\-dependent; different multilingual encoders may benefit different decoder architectures\. Thus, our conclusion about LaBSE should be interpreted within PrismDiff rather than as a universal claim for all translation decoders\.
Third, for experiment simplicity, we employ several vanilla implementation strategies of MDMs, such as simple noise schedules and fixed language block lengths\. Although our settings \(lmulti30k=40l\_\{\\mathrm\{multi30k\}\}=40,leuroparl=80l\_\{\\mathrm\{europarl\}\}=80\) can cover most examples, there still exist longer sentences that will overflow, which considerably contributes to the poor performance on Europarl\. Better performance may be achieved if these strategies and hyperparameters are more carefully tuned\.
While exclusive attention already produces strong results through joint parameter optimization, exploring lightweight cross\-stream interaction mechanisms, such as gated cross\-attention between target streams, could be a promising direction but may also introduce more computation and inter\-stream interference\.
Additionally, our efficiency evaluation focuses on a standard single\-GPU, single\-instance, batch\-size=1 latency, which closely reflects serving settings with extremely high latency requirements where synchronous batch is impractical\. Optimizing batch throughput remains a direction to be explored in the future\.
## References
- Austinet al\.\(2021\)J\. Austin, D\. D\. Johnson, J\. Ho, D\. Tarlow, and R\. Van Den BergStructured denoising diffusion models in discrete state\-spaces\.Advances in neural information processing systems34,pp\. 17981–17993\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1),[§3\.1](https://arxiv.org/html/2609.16312#S3.SS1.p1.1)\.
- Azpiazu and Pera \(2020\)I\. M\. Azpiazu and M\. S\. PeraA framework for hierarchical multilingual machine translation\.arXiv preprint arXiv:2005\.05507\.Cited by:[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Bartoshet al\.\(2026\)G\. Bartosh, T\. Pandeva, S\. Karmalkar, and J\. ZazoForward\-learned discrete diffusion: learning how to noise to denoise faster\.arXiv preprint arXiv:2605\.18204\.Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1)\.
- Bieet al\.\(2025\)T\. Bie, M\. Cao, K\. Chen, L\. Du, M\. Gong, Z\. Gong, Y\. Gu, J\. Hu, Z\. Huang, Z\. Lan,et al\.Llada2\. 0: scaling up diffusion language models to 100b\.arXiv preprint arXiv:2512\.15745\.Cited by:[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1),[§3\.1](https://arxiv.org/html/2609.16312#S3.SS1.p1.1)\.
- Chenet al\.\(2023\)L\. Chen, A\. Feng, B\. Yang, and Z\. LiXdlm: cross\-lingual diffusion language model for machine translation\.arXiv preprint arXiv:2307\.13560\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p2.1)\.
- Chenet al\.\(2025\)X\. Chen, S\. Huang, C\. Guo, C\. Wei, Y\. He, J\. Zhang, H\. Li, Y\. Chen,et al\.Dpad: efficient diffusion language models with suffix dropout\.arXiv preprint arXiv:2508\.14148\.Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1)\.
- Chenet al\.\(2024\)Z\. Chen, H\. Yuan, Y\. Li, Y\. Kou, J\. Zhang, and Q\. GuFast sampling via discrete non\-markov diffusion models with predetermined transition time\.Advances in Neural Information Processing Systems37,pp\. 106870–106905\.Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1)\.
- Conneauet al\.\(2020\)A\. Conneau, K\. Khandelwal, N\. Goyal, V\. Chaudhary, G\. Wenzek, F\. Guzmán, E\. Grave, M\. Ott, L\. Zettlemoyer, and V\. StoyanovUnsupervised cross\-lingual representation learning at scale\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 8440–8451\.Cited by:[§5\.4](https://arxiv.org/html/2609.16312#S5.SS4.p1.1)\.
- Costa\-Jussàet al\.\(2022\)M\. R\. Costa\-Jussà, J\. Cross, O\. Çelebi, M\. Elbayad, K\. Heafield, K\. Heffernan, E\. Kalbassi, J\. Lam, D\. Licht, J\. Maillard,et al\.No language left behind: scaling human\-centered machine translation\.arXiv preprint arXiv:2207\.04672\.Cited by:[§4\.2](https://arxiv.org/html/2609.16312#S4.SS2.SSS0.Px1.p1.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota\.Cited by:[§5\.4](https://arxiv.org/html/2609.16312#S5.SS4.p1.1)\.
- Elliottet al\.\(2016\)D\. Elliott, S\. Frank, K\. Sima’an, and L\. SpeciaMulti30k: multilingual english\-german image descriptions\.InProceedings of the 5th Workshop on Vision and Language,pp\. 70–74\.Cited by:[§4\.2](https://arxiv.org/html/2609.16312#S4.SS2.SSS0.Px1.p1.1)\.
- Fanet al\.\(2021\)A\. Fan, S\. Bhosale, H\. Schwenk, Z\. Ma, A\. El\-Kishky, S\. Goyal, M\. Baines, O\. Celebi, G\. Wenzek, V\. Chaudhary,et al\.Beyond english\-centric multilingual machine translation\.Journal of Machine Learning Research22\(107\),pp\. 1–48\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Fenget al\.\(2022\)F\. Feng, Y\. Yang, D\. Cer, N\. Arivazhagan, and W\. WangLanguage\-agnostic bert sentence embedding\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 878–891\.Cited by:[§3\.2](https://arxiv.org/html/2609.16312#S3.SS2.SSS0.Px1.p1.1)\.
- Ghazvininejadet al\.\(2019\)M\. Ghazvininejad, O\. Levy, Y\. Liu, and L\. ZettlemoyerMask\-predict: parallel decoding of conditional masked language models\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 6112–6121\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Gonget al\.\(2023\)S\. Gong, M\. Li, J\. Feng, Z\. Wu, and L\. KongDiffuSeq: sequence to sequence text generation with diffusion models\.InInternational Conference on Learning Representations \(ICLR 2023\)\(01/05/2023\-05/05/2023, Kigali, Rwanda\),Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1)\.
- Graveset al\.\(2006\)A\. Graves, S\. Fernández, F\. Gomez, and J\. SchmidhuberConnectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks\.InProceedings of the 23rd international conference on Machine learning,pp\. 369–376\.Cited by:[§3\.2](https://arxiv.org/html/2609.16312#S3.SS2.SSS0.Px4.p2.1)\.
- Guet al\.\(2018\)J\. Gu, J\. Bradbury, C\. Xiong, V\. Li, and R\. SocherNon\-autoregressive neural machine translation\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p1.1),[§1](https://arxiv.org/html/2609.16312#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Guan and Whitehill \(2025\)Y\. Guan and J\. WhitehillTransformer\-encoder trees for efficient multilingual machine translation and speech translation\.arXiv preprint arXiv:2509\.17930\.Cited by:[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1),[4th item](https://arxiv.org/html/2609.16312#S4.I1.i4.p1.1),[§4\.3](https://arxiv.org/html/2609.16312#S4.SS3.p3.1)\.
- Guoet al\.\(2019\)J\. Guo, X\. Tan, D\. He, T\. Qin, L\. Xu, and T\. LiuNon\-autoregressive neural machine translation with enhanced decoder input\.InProceedings of the AAAI conference on artificial intelligence,Vol\.33,pp\. 3723–3730\.Cited by:[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Heet al\.\(2023\)Z\. He, T\. Sun, Q\. Tang, K\. Wang, X\. Huang, and X\. QiuDiffusionbert: improving generative masked language models with diffusion models\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 4521–4534\.Cited by:[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1)\.
- Hoet al\.\(2020\)J\. Ho, A\. Jain, and P\. AbbeelDenoising diffusion probabilistic models\.Advances in neural information processing systems33,pp\. 6840–6851\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p3.1)\.
- Irwinet al\.\(2025\)R\. Irwin, A\. Tibo, J\. P\. Janet, and S\. OlssonSemlaFlow–efficient 3d molecular generation with latent attention and equivariant flow matching\.InInternational Conference on Artificial Intelligence and Statistics,pp\. 3772–3780\.Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p2.1),[§3\.2](https://arxiv.org/html/2609.16312#S3.SS2.SSS0.Px4.p1.1)\.
- Johnsonet al\.\(2017\)M\. Johnson, M\. Schuster, Q\. Le, M\. Krikun, Y\. Wu, Z\. Chen, N\. Thorat, F\. Viégas, M\. Wattenberg, G\. Corrado,et al\.Google’s multilingual neural machine translation system: enabling zero\-shot translation\.Transactions of the Association for Computational Linguistics5,pp\. 339–351\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Kasaiet al\.\(2020\)J\. Kasai, N\. Pappas, H\. Peng, J\. Cross, and N\. A\. SmithDeep encoder, shallow decoder: reevaluating non\-autoregressive machine translation\.arXiv preprint arXiv:2006\.10369\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p1.1)\.
- Koehn \(2005\)P\. KoehnEuroparl: a parallel corpus for statistical machine translation\.InThe Tenth Machine Translation Summit Proceedings of Conference,pp\. 79–86\.Cited by:[§4\.2](https://arxiv.org/html/2609.16312#S4.SS2.SSS0.Px1.p2.1)\.
- Liet al\.\(2022\)X\. Li, J\. Thickstun, I\. Gulrajani, P\. S\. Liang, and T\. B\. HashimotoDiffusion\-lm improves controllable text generation\.Advances in neural information processing systems35,pp\. 4328–4343\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1)\.
- Liuet al\.\(2025\)Z\. Liu, Y\. Yang, Y\. Zhang, J\. Chen, C\. Zou, Q\. Wei, S\. Wang, and L\. ZhangDllm\-cache: accelerating diffusion large language models with adaptive caching\.arXiv preprint arXiv:2506\.06295\.Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1)\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.InInternational Conference on Learning Representations,Cited by:[§4\.3](https://arxiv.org/html/2609.16312#S4.SS3.p4.1)\.
- Nieet al\.\(2025\)S\. Nie, F\. Zhu, Z\. You, X\. Zhang, J\. Ou, J\. Hu, J\. ZHOU, Y\. Lin, J\. Wen, and C\. LiLarge language diffusion models\.InICLR 2025 Workshop on Deep Generative Model in Machine Learning: Theory, Principle and Efficacy,Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1),[§3\.1](https://arxiv.org/html/2609.16312#S3.SS1.p1.1)\.
- Post \(2018\)M\. PostA call for clarity in reporting bleu scores\.InProceedings of the Third Conference on Machine Translation: Research Papers,pp\. 186\.Cited by:[§4\.2](https://arxiv.org/html/2609.16312#S4.SS2.SSS0.Px2.p1.1)\.
- Qianet al\.\(2021\)L\. Qian, H\. Zhou, Y\. Bao, M\. Wang, L\. Qiu, W\. Zhang, Y\. Yu, and L\. LiGlancing transformer for non\-autoregressive neural machine translation\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 1993–2003\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.16312#S2.SS1.p1.1)\.
- Reiet al\.\(2020\)R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. LavieCOMET: a neural framework for mt evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 2685–2702\.Cited by:[§4\.2](https://arxiv.org/html/2609.16312#S4.SS2.SSS0.Px2.p1.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.Advances in neural information processing systems30\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p1.1)\.
- Wuet al\.\(2025\)C\. Wu, H\. Zhang, S\. Xue, Z\. Liu, S\. Diao, L\. Zhu, P\. Luo, S\. Han, and E\. XieFast\-dllm: training\-free acceleration of diffusion llm by enabling kv cache and parallel decoding\.arXiv preprint arXiv:2505\.22618\.Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1)\.
- Xiaoet al\.\(2023\)Y\. Xiao, L\. Wu, J\. Guo, J\. Li, M\. Zhang, T\. Qin, and T\. LiuA survey on non\-autoregressive generation for neural machine translation and beyond\.IEEE Transactions on Pattern Analysis and Machine Intelligence45\(10\),pp\. 11407–11427\.Cited by:[§1](https://arxiv.org/html/2609.16312#S1.p2.1)\.
- Yeet al\.\(2025\)J\. Ye, Z\. Xie, L\. Zheng, J\. Gao, Z\. Wu, X\. Jiang, Z\. Li, and L\. KongDream 7b: diffusion large language models\.arXiv preprint arXiv:2508\.15487\.Cited by:[§2\.2](https://arxiv.org/html/2609.16312#S2.SS2.p1.1)\.
- Yehet al\.\(2024\)C\. Yeh, C\. Chen, C\. Hsu, and J\. ChienCross\-modality diffusion modeling and sampling for speech recognition\.\.InINTERSPEECH,Cited by:[§2\.3](https://arxiv.org/html/2609.16312#S2.SS3.p1.1)\.
## Appendix AAppendix
### A\.1Logarithm\-Uniform Masking Schedule
LetLLdenote the maximum sequence length\. The masking ratiommis bounded within\[1/L,1\]\[1/L,1\], where1/L1/Lcorresponds to the single\-token mask limit\. At each refinement step, tokens are remasked based on predicted confidence scores: the model retains the top\(1−m\)×L\(1\-m\)\\times Ltokens with highest predicted probability and remasks the remainder, wheremmis determined by the schedule\. We define a log\-uniform distribution overmmby sampling:
logm∼𝒰\(log1L,0\)\\log m\\sim\\mathcal\{U\}\\left\(\\log\\frac\{1\}\{L\},0\\right\)This formulation assigns equal probability mass to equal intervals on the logarithmic scale\. To adapt this continuous prior for fixed\-step training, we discretize the schedule over the total training stepsNN\. In practice, the log\-uniform schedule is implemented deterministically by mapping training steps to masking ratios\. For example, if usingL=40L=40withK=4K=4discrete levels, the corresponding time step sequencetit\_\{i\}is given by:
ti=exp\(log\(1L\)×i\)×Nt\_\{i\}=\\exp\\left\(\\log\(\\frac\{1\}\{L\}\)\\times i\\right\)\\times N, which are:
t0=N,t1=0\.4N,t2=0\.16N,t3=0\.063Nt\_\{0\}=N,\\ t\_\{1\}=0\.4N,\\ t\_\{2\}=0\.16N,\\ t\_\{3\}=0\.063NThis ensures that the training process effectively covers the full range of masking ratios with the desired log\-uniform distribution\. We found this masking schedule consistently outperforms other simple strategies in our task\.
### A\.2Scaling Analysis: Latency and Memory
Figure 4:Latency–language scaling graph\.The PrismDiff models are trained withT=100T=100, and accelerated sampling is performed during the inference stage usingt=50t=50ort=25t=25\.A detailed latency–language scaling analysis is provided in Figure[4](https://arxiv.org/html/2609.16312#A1.F4), comparing decoder\-only AR \(8\-layer and 4\-layer\) against PrismDiff \(T=100T=100,t=50t=50, andt=25t=25\)\. PrismDiff exhibits sublinear latency scaling with respect to the number of target streamsKK\. With exclusive attention, each refinement step requires one forward pass over the full canvas of sizeLL, where each stream attends only to itself and the anchor with costO\(L2\)O\(L^\{2\}\)\. The per\-step cost is thereforeO\(K⋅L2\)O\(K\\cdot L^\{2\}\), and the total cost overttdiffusion refinement steps isO\(t⋅K⋅L2\)O\(t\\cdot K\\cdot L^\{2\}\)\. GeneratingKKlanguages in a single forward pass therefore isK×K\\timesthe cost of a single\-language diffusion pass, butK×K\\timescheaper than runningKKseparate passes\.
In contrast, AR systems decoding forKKlanguages requiresKKsequential passes, each withO\(L2\)O\(L^\{2\}\)attention cost overLLautoregressive steps, resulting inO\(K⋅L3\)O\(K\\cdot L^\{3\}\)in total\. PrismDiff’s total cost ofO\(t⋅K⋅L2\)O\(t\\cdot K\\cdot L^\{2\}\)is therefore favorable whent≪Lt\\ll L\. We note that AR systems could reduce wall\-clock latency forKKtargets by runningKKmodel replicas simultaneously\. However, this would multiply the memory consumption byKKwhereas PrismDiff generates allKKtargets with a single model instance\.
Table 3:Peak GPU memory footprint\(in GB\) of PrismDiff during parallel multi\-stream inference\. The memory scales gracefully even under extreme stress tests \(e\.g\.,K=8K=8targets withl=300l=300tokens\)\.Table[3](https://arxiv.org/html/2609.16312#A1.T3)reports peak GPU memory footprint \(GB\) of PrismDiff under different block lengths \(ll\) and number of target streams \(KK\)\. Withl=40l=40block length, memory increases by less than 10% asKKscales from 1 to 8 \(2\.66 GB to 2\.91 GB\), confirming that the canvas\-based multi\-stream design does not incur significant memory overhead in practice\. The memory growth becomes more pronounced at larger block lengths \(e\.g\.,l=300l=300\), which is expected as the total canvas size scales asLsrc\+K×LtgtL\_\{\\mathrm\{src\}\}\+K\\times L\_\{\\mathrm\{tgt\}\}\. The memory scales gracefully even under extreme stress tests \(e\.g\.,K=8K=8withL=300L=300\), confirming that while the hyper\-sequence design inherently costs more memory for long texts, it easily fits in the capacity of standard GPUs without requiring complex multi\-GPU model parallelism\. Note that AR memory does not scale withKKsince each language is decoded sequentially; for reference, an 8\-layer decoder\-only AR model on Multi30K \(max length<40<40\) has a peak memory of approximately 2\.69 GB, comparable to PrismDiff atK=1K=1\(2\.66 GB\)\. The key difference lies not in memory but in latency: the cost of AR models is reflected in a linear latency growth, as shown in Figure[4](https://arxiv.org/html/2609.16312#A1.F4)\.
### A\.3Full Results
Full experimental results on Multi30K are reported in Tables[5](https://arxiv.org/html/2609.16312#A1.T5)–[8](https://arxiv.org/html/2609.16312#A1.T8)\. The results include comparing different implementation setups and architectures \(Table[5](https://arxiv.org/html/2609.16312#A1.T5)\), using different sampling budgets \(Table[6](https://arxiv.org/html/2609.16312#A1.T6)\), decoder\-only AR baselines with various numbers of decoder layers \(Table[7](https://arxiv.org/html/2609.16312#A1.T7)\), and applying different multilingual encoders for extracting source\-side representation \(anchor\) as the input of PrismDiff \(Table[8](https://arxiv.org/html/2609.16312#A1.T8)\)\.
### A\.4The Role of Cross\-Lingual Alignment in Attention
The performance gap of Diff\-Indep using full attention versus using exclusive attention suggests that full cross\-stream attention plays different roles depending on the conditioning signal\. When conditioned directly on source tokens \(Diff\-SrcTokens\), the model lacks a centralized, language\-agnostic representation\. Without a semantic anchor, full attention can partially compensate for weak source conditioning: each target stream essentially treats the noisy states of other streams as auxiliary context, forming a pseudo multi\-source prediction problem to extract semantic clues\. However, when PrismDiff is conditioned on a cross\-lingually aligned semantic anchor, direct target\-stream attention can introduce surface\-form interference across languages\. Therefore, exclusive stream attention encourages each language to perform its own denoising process, while semantic coordination is pre\-solved by the shared anchor\.
### A\.5Inference Examples
One example of the accelerated inference process of PrismDiff trained on Multi30K with a total diffusion steps ofT=100T=100is shown in Table[9](https://arxiv.org/html/2609.16312#A1.T9), and Table[10](https://arxiv.org/html/2609.16312#A1.T10)contains another example trained on Europarl withT=40T=40\. All target languages are refined simultaneously and generated together at each step\. The canvas size and the refinement steps of each target language is truncated in the examples for better readability\. A larger stepttmeans an earlier stage in the inference time, so more mask tokens \(\[M\]\) will be presented in the example\. The translation displayed in the last step \(t=0t=0\) is CTC\-collapsedt=0t=0prediction, which removes repetitive tokens \(only if there exist\)\.
### A\.6Zero\-Shot Transfer Examples
We present representative examples illustrating typical failure modes of AR systems under zero\-shot transfer in Table[11](https://arxiv.org/html/2609.16312#A1.T11)\. We provide a qualitative analysis using zero\-shot Spanish\-to\-French \(Es→\\rightarrowFr\) translation examples from Multi30K\. The original task on which the model is trained on is En\-\{De, Fr\}\. In Example 1, the AR model suffers from error cascading and hallucination\. In Example 2 and Example 3, the AR baselines exhibits structural collapse, generating non\-existing words\.
### A\.7Implementation of TET
Our implemented TET model on Europarl contains 4 leaf nodes \(corresponding to the target languages\), each has a depth of 6\. We list the detailed paths of all languages in TET in Table[4](https://arxiv.org/html/2609.16312#A1.T4)\. In all, the TET model contains 16 Transformer encoder layers, resulting in a total parameter count of 556M\.
TET’s extremely low latency is attributed to its single\-pass CTC decoding\. Unlike sequential autoregressive decoding, or iterative diffusion refinements, TET generates all target languages in a single forward pass\. The higher parameter count does not proportionally increase latency due to the parallel tree structure\.
Table 4:Paths of all languages in TET on Europarl\. All paths begin with "common1–common2–⋯\\cdots”, which is omitted in this table for simplicity\.
### A\.8Licenses and Usage of Artifacts
The Multi30K dataset is publicly available and licensed under the CC BY\-SA 4\.0\. Europarl dataset is publicly available as open data and widely permitted for research purposes\. We utilize the pre\-trained LaBSE \(Language\-Agnostic BERT Sentence Embeddings\) and mBERT \(multilingual BERT for model design and evaluation, both of which are officially released by Google and distributed under the Apache License 2\.0\. We also employ XLM\-R \(XLM\-RoBERTa\), developed by Meta, which is released under the MIT License\. We also use the NLLB \(No Language Left Behind, e\.g\., NLLB\-200\) model, which is distributed by Meta under the CC\-BY\-NC 4\.0, strictly limiting its use to non\-commercial research purposes\. We confirm that our use of the datasets is consistent with their intended use for machine translation research\. Similarly, all models are employed for feature extraction and data synthesis, respectively, aligning with their original research purposes\.
BLEUCOMETAvg\.LatencySrc\.ModelArchitectureAttn\.TTDeFrDeFrBLEUCOMET\(ms\)EnDiff\-Indep1028\.4739\.510\.620\.7333\.990\.6826\.910031\.0740\.740\.660\.7635\.910\.71254\.5Diff\-SrcTokensfull1030\.3537\.320\.610\.7033\.840\.6625\.6\- CTCfull1029\.0536\.310\.580\.6732\.680\.6324\.5full10032\.1240\.760\.650\.7436\.440\.70234\.4excl\.10030\.9538\.640\.650\.7434\.800\.695247\.4ARenc\-dec29\.9035\.830\.670\.7332\.870\.7088\.9enc\-dec \+ Z20\.8522\.650\.610\.6621\.750\.6490\.1dec\-only \+ Z32\.8241\.670\.680\.7637\.250\.72148\.3PrismDifffull1027\.9137\.400\.610\.7132\.660\.6627\.8\- CTCfull1027\.1636\.620\.580\.6831\.890\.6332\.5excl\.1028\.4838\.470\.620\.7233\.480\.6726\.5\- CTCexcl\.1027\.2937\.090\.590\.7032\.190\.6526\.7Abs\. PEexcl\.1019\.5628\.040\.530\.6223\.800\.5826\.2full10031\.9240\.240\.670\.7536\.080\.71259\.0excl\.10033\.1641\.100\.680\.7637\.130\.72278\.3EsDiff\-SrcTokensfull100\.280\.160\.280\.290\.220\.2921\.9full1000\.140\.310\.290\.300\.230\.30234\.0excl\.1000\.300\.310\.260\.320\.310\.29238\.6ARdec\-only \+ Z20\.4829\.950\.600\.7125\.220\.66154\.5PrismDifffull1019\.4330\.430\.560\.6824\.930\.6229\.5excl\.1020\.3331\.510\.570\.6925\.920\.6327\.2Abs\. PEexcl\.1014\.7521\.530\.500\.5918\.140\.5525\.8full10022\.3733\.130\.620\.7327\.750\.6725\.6excl\.10023\.2533\.170\.620\.7328\.210\.68284\.0
Table 5:Results on Multi30K\. In architecture ablations, “enc\-dec” is encoder\-decoder and “dec\-only” denotes decoder\-only; “\+Z” means using semantic anchor Z as input, otherwise the AR model directly uses source tokens as input; “\-CTC” means without CTC decoding; “Abs\. PE” is using absolute positional encoding instead of reset positional encoding\. “Attn\.” represents the attention mode used by PrismDiff, in which “full” means using full attention across target language, and “excl\.” means using exclusive attention for each target language\.TTis the total diffusion steps the model is trained on, and all diffusion systems report the results usingTTfull sampling steps with uniform noise scheduler\. “Src\.” lists the source language input to the model, in which Spanish \(Es\) is a zero\-shot source language that the model is never trained on\.Table 6:Sampling steps vs\. performance on Multi30K\. The PrismDiff model is trained with a total diffusion steps ofT=100T=100, and is evaluated under different accelerated sampling steps \(tt\) using logarithm\-uniform noise scheduler\.Table 7:Number of AR decoder layers vs\. performance on Multi30K\.Table 8:Performance comparison of using different multilingual encoders to extract the source\-side representation \(anchor\) for PrismDiff on Multi30K\. All PrismDiff models are trained with total diffusion steps ofT=100T=100\.StepGermanFrencht=54\.1Ein \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]Un homme \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]t=18\.4Ein Mann schläft in einem grünen \[M\] auf \[M\] \[M\]fa \[M\]Un homme dormant dans une \[M\] \[M\] un can \[M\]é \[M\] \[M\]t=5\.4Ein Mann schläft in einem grünen \[M\] auf einem Sofa\.Un homme dormant dans une pièce sur un canapé \[M\]\.t=0Ein Mann schläft in einem grünen Zimmer auf einem Sofa\.Un homme dormant dans une pièce sur un canapé vert\.\[\]GroundtruthEin Mann schläft in einem grünen Raum auf einer Couch\.Un homme dormant dans une chambre verte sur un canapé\.Table 9:An example of accelerated \(with 25 refinement steps\) translation process of PrismDiff \(T=100T=100\) trained on Multi30K\. The canvas size and the number of the steps are reduced in this example for better readability\. The “t=0” is the result after CTC collapse\. The source English sentence is: “A man sleeping in a green room on a couch\.”StepFrenchRomanianDanisht=30Le Parlement européen \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]Parlamentul European a votat \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]\[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]t=20Le Parlement européen a voté \[M\] \[M\] résolution \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]Parlamentul European a votat o rezolu \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]Parlamentet har stemt for \[M\] beslutning om \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\] \[M\]t=10Le Parlement européen a voté pour \[M\] résolution \[M\] l’avenir du Fonds social \[M\] \[M\]Parlamentul European a votat \[M\] rezolu \[M\] referitoare la viitorul Fondului \[M\] \[M\] \[M\]Parlamentet har stemt for en beslutning om Den Europæiske Social \[M\] \[M\] \[M\] \[M\] \[M\]t=0Le Parlement européen a voté pour une résolution sur l’avenir du Fonds social européen\.Parlamentul European a votat o rezolu ie referitoare la viitorul Fondului social european\.Parlamentet har stemt for en beslutning om Den Europæiske Socialfonds fremtid\.\[\]GroundtruthLe Parlement européen a voté une résolution sur l’avenir du Fonds social européen\.Parlamentul European a votat pentru o rezolu ie referitoare la viitorul Fondului social european\.Parlamentet har stemt for beslutningen om Den Europæiske Socialfonds fremtid\.Table 10:An example of translation process of PrismDiff \(T=40T=40\) trained on Europarl\. Here, only 3 target languages are displayed for simplicity\. The canvas size and the number of the steps are reduced in this example for better readability\. The “t=0” is the result after CTC collapse\. The source English sentence is: “The European Parliament has voted for a resolution on the future of the European Social Fund\.”Table 11:Zero\-shot examples on Multi30K\. The input source is Spanish \(Es\), and target languages are German \(De\) and French \(Fr\)\. We only show French sentences for illustration\. The English source sentences in this table are only for reference \(not used as input\)\.Similar Articles
Multi-Block Diffusion Language Models
This paper proposes Multi-Block Diffusion Language Models (MBD-LMs), extending single-block diffusion to concurrent multi-block decoding with improved training strategies like Multi-block Teacher Forcing and an optimized Block Buffer decoding algorithm. Experiments show increased tokens per forward pass and improved accuracy on benchmarks.
Less Uniform Discrete Diffusion is More Powerful and Scalable
This paper proposes LUDI, a less uniform diffusion language modeling framework that fixes over-uniform training objectives and condition-target confusion in uniform diffusion LMs, enabling a 7B-scale UDLM with 3x-token-per-step speedup over autoregressive decoding and competitive complex reasoning performance.
Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
This paper proposes Decoupled Residual Denoising Diffusion Models (DRDD) for unified and data-efficient image-to-image translation, decoupling noise diffusion for domain harmonization from residual diffusion for semantic mapping.
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
This paper proposes JUMP, a single-pass membership inference attack for fine-tuned discrete diffusion language models that exploits their any-order and parallel decodability to improve detection accuracy with fewer queries.
@alec_helbling: Diffusion LMs generate multiple tokens in parallel. However, iterative unmasking repeatedly updates token states, limit…
The post describes how diffusion language models face KV-cache reuse limitations due to iterative unmasking, and introduces Block Diffusion as a solution that decodes blocks left-to-right for efficient caching while generating in parallel.