All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

arXiv cs.CL Papers

Summary

This paper proposes a causality-aware framework for LLM-based simultaneous speech-to-speech translation, introducing a novel data pipeline and adaptive policy to improve quality-latency trade-off, achieving state-of-the-art results with reduced latency.

arXiv:2609.30416v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.
Original Article
View Cached Full Text

Cached at: 09/28/26, 09:36 AM

# All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation
Source: [https://arxiv.org/html/2609.30416](https://arxiv.org/html/2609.30416)
Amir Hussein12, Enas Albasiri2, Travis M\. Bartley2, Nourchene Ferchichi2, Ke Hu2, Harishchandra Dubey2, Myungjong Kim2, Zhehuai Chen2, Oluwatobi Olabiyi2, Sanjeev Khudanpur1

###### Abstract

Large Language Models \(LLMs\) have shown strong performance in low\-resource offline translation; however, extending them to simultaneous speech\-to\-speech translation \(Simul\-S2ST\) remains challenging due to the scarcity of causally aligned training data with high cross\-lingual speaker fidelity\. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency\. We propose a causality\-aware Simul\-S2ST framework with a novel data pipeline that generates high\-fidelity, causally aligned segments with improved voice transfer\. The framework introduces \(i\) a factorized S2ST architecture \(FAST\), \(ii\) a causality\-aware adaptive policy \(CAP\), and \(iii\) causality\-aware latency metric\. Experiments on CVSS Spanish, German, and French show that FAST\-CAP consistently improves the quality–latency trade\-off, achieving up to \+1\.2 BLEU and a 26% relative latency reduction over a fixed policy\. Despite using substantially less training data than existing systems, FAST\-CAP achieves state\-of\-the\-art results in speech translation quality and speaker fidelity while yielding up to a 38\.8% relative reduction in latency\.

###### Index Terms:

simultaneous speech\-to\-speech translation, large language models, causal alignment, adaptive translation\.

©2026 IEEE\. Personal use of this material is permitted\. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works\.

## IIntroduction

Simultaneous speech\-to\-speech translation \(Simul\-S2ST\) aims to translate speech from one language into another in real time, allowing natural cross\-lingual interactions\. Beyond low latency, human\-like dialogue requires cross\-lingual voice transfer that preserves speaker identity while appropriately transferring paralinguistic cues across languages\[[1](https://arxiv.org/html/2609.30416#bib.bibx1)\]\. It must also handle conversational phenomena such as overlapping speech and interruptions, making Simul\-S2ST particularly challenging\[[2](https://arxiv.org/html/2609.30416#bib.bibx2)\]\. Another major challenge is the scarcity of paired S2ST training data\[[3](https://arxiv.org/html/2609.30416#bib.bibx3),[4](https://arxiv.org/html/2609.30416#bib.bibx4)\], especially compared to the abundance of transcribed speech and translated text\. This disparity motivates a key question:How can Simul\-S2ST systems leverage large pretrained models trained on abundant speech and text data?

Recently, there has been growing interest in leveraging pretrained large language models \(LLMs\) as translation backbones, motivated by their strong contextual reasoning, long\-context modeling capabilities, and improved translation quality in low\-resource and zero\-shot settings\[[5](https://arxiv.org/html/2609.30416#bib.bibx5),[6](https://arxiv.org/html/2609.30416#bib.bibx6),[7](https://arxiv.org/html/2609.30416#bib.bibx7),[8](https://arxiv.org/html/2609.30416#bib.bibx8)\]\. These advances have led to increased interest in LLM\-based simultaneous translation to exploit sequence modeling capacity of LLMs for low\-latency translation\. However, existing LLM\-based simultaneous translation approaches largely rely on fixed or heuristic read/write policies, such as wait\-kk, which may generate output before sufficient source context is available and degrade translation quality\[[9](https://arxiv.org/html/2609.30416#bib.bibx9),[10](https://arxiv.org/html/2609.30416#bib.bibx10),[11](https://arxiv.org/html/2609.30416#bib.bibx11),[12](https://arxiv.org/html/2609.30416#bib.bibx12)\]\. Moreover, most speech translation systems focus on lexical content and overlook rich paralinguistic cues such as prosody, emotion, attitude, and speaker intent\[[13](https://arxiv.org/html/2609.30416#bib.bibx13),[14](https://arxiv.org/html/2609.30416#bib.bibx14)\]\.

To jointly optimize speech perception and generation, recent work has explored direct Simul\-S2ST models that translate source speech into target speech while preserving expressive speech cues\[[14](https://arxiv.org/html/2609.30416#bib.bibx14)\]\. Hibiki\[[15](https://arxiv.org/html/2609.30416#bib.bibx15)\]extends this direction with an LLM\-based multi\-stream architecture and a dynamic translation policy driven by offline MT perplexity\. However, offline MT perplexity tends to favor longer source contexts, leading to high latency\. Moreover, using a single neural codec representation for both speech perception and generation introduces competing objectives and can result in suboptimal performance\[[16](https://arxiv.org/html/2609.30416#bib.bibx16)\]\.

To address the aforementioned limitations, we introduce a causality\-aware framework111Code:[https://github\.com/AmirHussein96/FAST\-CAP/tree/main\.](https://github.com/AmirHussein96/FAST-CAP/tree/main.)for LLM\-based Simul\-S2ST\. LLM\-based Simul\-S2ST requires aligned speech\-to\-speech data to learn when to wait and when to speak, but such data is unavailable\. To bridge this gap, we develop a unified data pipeline that constructs causally aligned training examples and synthesizes natural target speech with high\-fidelity cross\-lingual voice transfer\. Inspired by cognitive studies on chunking strategies used by professional interpreters\[[17](https://arxiv.org/html/2609.30416#bib.bibx17),[18](https://arxiv.org/html/2609.30416#bib.bibx18)\], we introduce a causality\-aware adaptive policy in which read/write decisions are guided by the availability of sufficient source information to generate the corresponding translation, resulting in optimal waiting strategy\. To mitigate the limitations of a shared codec representation, we propose a factorized architecture \(FAST\) that decouples lexical and acoustic modeling: lexical information is extracted by an ASR encoder, while acoustic information is modeled using neural audio codec\. FAST adopts a multi\-stream full\-duplex design\[[19](https://arxiv.org/html/2609.30416#bib.bibx19)\]that enables effective handling of multimodal inputs, overlapping speech, and conversational interruptions\. Finally, we propose CAAL, an alignment\-aware latency metric that measures only avoidable delay beyond the ideal causal policy\. Our key contributions are:

- •A data generation pipeline with causality\-aware adaptive policy \(CAP\)\.
- •A factorized architecture \(FAST\) that decouples lexical and acoustic representations while preserving speaker identity and vocal characteristics\.
- •A causality\-aware latency metric \(CAAL\) that distinguishes necessary linguistic delays from avoidable system delays\.

## IIProposed Approach

Our proposed factorized simultaneous speech\-to\-speech translation \(FAST\) model uses an LLM as the backbone for translation and adopts a multi\-stream full\-duplex design\[[19](https://arxiv.org/html/2609.30416#bib.bibx19)\]\. The multi\-stream design enables real\-time processing different modalities and overlapping speech as separate but temporally aligned streams\.

TABLE I:Comparison of speaker similarity and audio quality on the CVSS\-T Dataset, with audio quality measured by UTMOS\-V2\.### II\-AData Generation Pipeline

To enable simultaneous processing, the generated translations must be causal emitting outputs as soon as sufficient source information becomes available with minimal latency\. Figure[1](https://arxiv.org/html/2609.30416#S2.F1)illustrates the proposed data generation pipeline with causality\-aware adaptive chunking\. Improving cross\-lingual speaker fidelity:Speech\-to\-speech translation models are commonly trained on paired data with target speech synthesized using cross\-lingual TTS, as in CVSS\-T\[[20](https://arxiv.org/html/2609.30416#bib.bibx20)\]\. However, CVSS\-T exhibits low source–target speaker similarity, with an average ECAPA\-TDNN cosine similarity of 0\.24 as shown in Table[I](https://arxiv.org/html/2609.30416#S2.T1)\. We therefore resynthesize CVSS\-T using the zero\-shot TTS model A2Flow222[https://catalog\.ngc\.nvidia\.com/orgs/nvidia/teams/nvigisdk/models/riva\-tts\-a2flow](https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nvigisdk/models/riva-tts-a2flow), which increases the average speaker similarity to 0\.61\. Alignment generation:Training an LLM\-based Simul\-S2ST system requires causally aligned source–target pairs so the model can learn when to wait and when to generate translation\. We first align the source speech𝐬s​r​c1:T\\mathbf\{s\}^\{src\}\_\{1:T\}with its transcript𝐟=\(f1,…,fJ\)\\mathbf\{f\}=\(f\_\{1\},\\dots,f\_\{J\}\)using Montreal Forced Aligner \(MFA\)\[[21](https://arxiv.org/html/2609.30416#bib.bibx21)\]\. We then use Awesome\-align\[[22](https://arxiv.org/html/2609.30416#bib.bibx22)\]with multilingual BERT to obtain word alignments between𝐟\\mathbf\{f\}and the target translation𝐞=\(e1,…,eI\)\\mathbf\{e\}=\(e\_\{1\},\\dots,e\_\{I\}\)\. Following the IBM alignment formulation, we define the raw alignment set as𝒜=\{\(j,i\)∣j∈\{1,…,J\},i∈\{1,…,I\}\}\\mathcal\{A\}=\\\{\(j,i\)\\mid j\\in\\\{1,\\dots,J\\\},\\ i\\in\\\{1,\\dots,I\\\}\\\}, where\(j,i\)\(j,i\)indicates that source wordfjf\_\{j\}is aligned to target wordeie\_\{i\}; see Figure[1](https://arxiv.org/html/2609.30416#S2.F1)\.

Fig\. 1:Overview of the data generation pipeline\. The process includes \(1\) alignments generation, \(2\) unique mapping and causal pivot extraction, and \(3\) causality\-aware adaptive chunks construction\.Causality\-Aware Adaptive Policy:We remove crossing dependencies by converting the raw text\-to\-text alignment into a monotonic alignmentΠ\\Pi\. This is non\-trivial due to word\-order differences, many\-to\-one alignments and omitted words in translation\. To address these challenges, Algorithm[1](https://arxiv.org/html/2609.30416#alg1)introducesMonotonicAlign\(⋅\)\(\\cdot\), which builds a monotonic target\-to\-source alignment by sorting alignments by target index, retaining the highest source index for each target word, and enforcing non\-decreasing source indices with a running maximum\. Missing target alignments are then filled with a backward pass\. The resulting pivot alignments𝒫\\mathcal\{P\}serve as anchors for causal segmentation, as shown in Figure[1](https://arxiv.org/html/2609.30416#S2.F1)\. We optionally applyMergeChunks\(⋅\)\(\\cdot\)to merge adjacent short chunks into longer causal segments\. TheMonotonicAlign\(⋅\)\(\\cdot\)function returns a set of pivot alignments with unique indices, denoted by𝒫\\mathcal\{P\}, which serve as anchors for causal segmentation\. The behavior ofMonotonicAlign​\(⋅\)\\textsc\{MonotonicAlign\}\(\\cdot\)is illustrated in the*Unique Mapping & Causal Pivots*step of Figure[1](https://arxiv.org/html/2609.30416#S2.F1)\. We then optionally applyCombine​\(⋅\)\\textsc\{Combine\}\(\\cdot\)to merge adjacent causal chunks into longer causal segments, analogous to phrase\-based translation from statistical machine translation\. Finally, the causal emission policy is implemented throughCausalPairedChunks​\(⋅\)\\textsc\{CausalPairedChunks\}\(\\cdot\), which constructs causal source\-target chunks around these pivots\. On the source side, each chunk spans from the end of the previous pivot to the end of the current pivot \(left aggregation\)\. On the target side, each chunk spans from the current pivot’s target start to the next pivot’s target start \(right aggregation\)\. This construction ensures each target token is emitted only after the necessary source evidence has been observed, forming a Causality\-Aware Adaptive Policy \(CAP\) for Simul\-S2ST training\. Formally, for each target tokeneie\_\{i\}, letR⁡\(i\):=\{j∣\(j,i\)∈Π\}R\(i\):=\\\{j\\mid\(j,i\)\\in\\Pi\\\}denote the source indices required to generateeie\_\{i\}and and letts​r​ce​n​\[j\]t^\{en\}\_\{src\}\[j\]denote the end time of source wordfjf\_\{j\}\. We define the source\-side pivot end timeτs​r​ce​n​\[i\]\\tau^\{en\}\_\{src\}\[i\]recursively as:

τs​r​ce​n​\[i\]=max⁡\(τs​r​ce​n​\[i−1\],maxj∈R⁡\(i\)⁡ts​r​ce​n​\[j\]\)\.\\tau^\{en\}\_\{src\}\[i\]=\\max\\left\(\\tau^\{en\}\_\{src\}\[i\-1\],\\ \\max\_\{j\\in R\(i\)\}t^\{en\}\_\{src\}\[j\]\\right\)\.\(1\)Generation is therefore restricted to the observed speech prefix, such thatP\(ei∣𝐬1:T\)=P\(ei∣𝐬1:τs​r​ce​n​\[i\]\)P\(e\_\{i\}\\mid\\mathbf\{s\}\_\{1:T\}\)=P\(e\_\{i\}\\mid\\mathbf\{s\}\_\{1:\\tau^\{en\}\_\{src\}\[i\]\}\)\. This ensures causal generation with minimal latency\. When a target speech chunk \(e\.g\., “The”\) is shorter than the next aligned source segment \(e\.g\., “misterio”\), we append silence to preserve causal timing\. To reduce boundary artifacts, silence gaps are smoothed using mirrored edges from neighboring chunks with a Hamming window\. Alignment processing and causal chunk construction are implemented with Lhotse\[[23](https://arxiv.org/html/2609.30416#bib.bibx23)\]\.

Algorithm 1Causality\-Aware Adaptive Policy1:Source words

𝐟=\(f1,…,fJ\)\\mathbf\{f\}=\(f\_\{1\},\\dots,f\_\{J\}\),

𝐭s​r​c=\{\(tjst,tjen\)\}j=1J\\mathbf\{t\}\_\{src\}=\\\{\(t\_\{j\}^\{\\mathrm\{st\}\},t\_\{j\}^\{\\mathrm\{en\}\}\)\\\}\_\{j=1\}^\{J\}
2:Target words

𝐞=\(e1,…,eI\)\\mathbf\{e\}=\(e\_\{1\},\\dots,e\_\{I\}\),

𝐭t​g​t=\{\(tist,tien\)\}i=1I\\mathbf\{t\}\_\{tgt\}=\\\{\(t\_\{i\}^\{\\mathrm\{st\}\},t\_\{i\}^\{\\mathrm\{en\}\}\)\\\}\_\{i=1\}^\{I\}
3:Text alignments

𝒜⊆\{1\.\.J\}×\{1\.\.I\}\\mathcal\{A\}\\subseteq\\\{1\.\.J\\\}\\times\\\{1\.\.I\\\}
4:Minimum source chunk duration

dmind\_\{\\min\}
5:functionMonotonicAlign\(

𝒜,𝐟,𝐞\\mathcal\{A\},\\mathbf\{f\},\\mathbf\{e\}\)

6:

Π←Sort​\(𝒜,by target index​i,then source index j\)\\Pi\\leftarrow\\textsc\{Sort\}\(\\mathcal\{A\},\\text\{by target index \}i,\\text\{then source index j\}\)
7:Retain highest source index for each target index

8:for

k=2k=2to

\|Π\|\|\\Pi\|do

9:

πk​\[s​r​c\]←max⁡\(πk​\[s​r​c\],πk−1​\[s​r​c\]\)\\pi\_\{k\}\[src\]\\leftarrow\\max\(\\pi\_\{k\}\[src\],\\pi\_\{k\-1\}\[src\]\)⊳\\trianglerightenforce monotonicity

10:for

i=I−1i=I\-1to

11do⊳\\trianglerightloop over target indices

11:if

∄π∈Π:π\[tgt\]=i\\nexists\\ \\pi\\in\\Pi:\\pi\[tgt\]=ithen

12:

Π←Π∪\{\(πi\+1​\[s​r​c\],i\)\}\\Pi\\leftarrow\\Pi\\cup\\\{\(\\pi\_\{i\+1\}\[src\],i\)\\\}⊳\\trianglerightadd missing targets

13:

𝒫←GetPivots⁡\(Π\)\\mathcal\{P\}\\leftarrow\\mathrm\{GetPivots\(\\Pi\)\}⊳\\trianglerightextract pivots for causal chunks

14:return

𝒫\\mathcal\{P\}
15:functionCausalPairedChunks\(

𝒫∗\\mathcal\{P\}^\{\*\},

𝐭s​r​c,𝐭t​g​t,𝐟,𝐞\\mathbf\{t\}\_\{src\},\\mathbf\{t\}\_\{tgt\},\\mathbf\{f\},\\mathbf\{e\}\)

16:

Csrc,Ctgt=\{∅\}k=1\|𝒫∗\|C^\{\\text\{src\}\},C^\{\\text\{tgt\}\}=\\\{\\emptyset\\\}\_\{k=1\}^\{\|\\mathcal\{P\}^\{\*\}\|\};⊳\\trianglerightinitialize source and target chunks dict

17:

τs​r​ce​n←\\tau^\{en\}\_\{src\}\\leftarrowPivotSrcEn\(𝒫∗​\[s​r​c\]\\mathcal\{P\}^\{\*\}\[src\],𝐭s​r​c\\mathbf\{t\}\_\{src\}\)⊳\\trianglerightget pivot src end time

18:

τt​g​ts​t←\\tau^\{st\}\_\{tgt\}\\leftarrowPivotTgtSt\(𝒫∗​\[t​g​t\]\\mathcal\{P\}^\{\*\}\[tgt\],𝐭t​g​t\\mathbf\{t\}\_\{tgt\}\)⊳\\trianglerightget pivot tgt start time

19:for

r=1r=1to

\|𝒫∗\|\|\\mathcal\{P\}^\{\*\}\|do⊳\\trianglerightiterate over pivot indices

20:

Crs​r​c←C\_\{r\}^\{src\}\\leftarrow\[

𝐟r−1:r\\mathbf\{f\}\_\{r\-1:r\},

\(τs​r​ce​n​\[r−1\],τs​r​ce​n​\[r\]\)\(\\tau^\{en\}\_\{src\}\[r\-1\],\\tau^\{en\}\_\{src\}\[r\]\)\]⊳\\trianglerightsrc left aggregation

21:

Crt​g​t←C\_\{r\}^\{tgt\}\\leftarrow\[

𝐞r:r\+1\\mathbf\{e\}\_\{r:r\+1\},

\(τt​g​ts​t​\[r\],τt​g​ts​t​\[r\+1\]\)\(\\tau^\{st\}\_\{tgt\}\[r\],\\tau^\{st\}\_\{tgt\}\[r\+1\]\)\]⊳\\trianglerighttgt right aggregation

22:return

Cs​r​cC^\{src\},

Ct​g​tC^\{tgt\}
23:

𝒫←\\mathcal\{P\}\\leftarrowMonotonicAlign\(𝒜,𝐟,𝐞\)\(\\mathcal\{A\},\\mathbf\{f\},\\mathbf\{e\}\)

24:

𝒫∗←\\mathcal\{P\}^\{\*\}\\leftarrowMergeChunks\(𝒫,dm​i​n\)\(\\mathcal\{P\},d\_\{min\}\)

25:

Cs​r​c,Ct​g​t←C^\{src\},C^\{tgt\}\\leftarrowCausalPairedChunks\(𝒫∗,𝐭s​r​c,𝐭t​g​t,𝐟,𝐞\)\(\\mathcal\{P\}^\{\*\},\\mathbf\{t\}\_\{src\},\\mathbf\{t\}\_\{tgt\},\\mathbf\{f\},\\mathbf\{e\}\)

26:return

Cs​r​c,Ct​g​tC^\{src\},C^\{tgt\}

### II\-BModel Architecture

The FAST architecture adopts a multi\-stream design\[[19](https://arxiv.org/html/2609.30416#bib.bibx19)\], as illustrated in Figure[2](https://arxiv.org/html/2609.30416#S2.F2)\. Unlike prior work\[[15](https://arxiv.org/html/2609.30416#bib.bibx15)\], which uses a single neural codec representation for both perception and generation, FAST decouples lexical and acoustic representations inspired by\[[24](https://arxiv.org/html/2609.30416#bib.bibx24)\]\. Lexical information is provided to the LLM through an ASR encoder, while acoustic information is modeled by an autoregressive TTS module using codec tokens\. This design improves translation accuracy while preserving high\-quality speech synthesis\. To enable cross\-lingual voice transfer, we extract source\-speaker embeddings using TitaNet333[https://huggingface\.co/nvidia/speakerverification\_en\_titanet\_large](https://huggingface.co/nvidia/speakerverification_en_titanet_large)and add them to the codec\-token representations\. Speech is tokenized with streaming NanoCodec\[[25](https://arxiv.org/html/2609.30416#bib.bibx25)\], which uses Finite Scalar Quantization \(FSQ\)\. Unlike RVQ\-based codecs\[[15](https://arxiv.org/html/2609.30416#bib.bibx15)\], FSQ uses independent codebooks, enabling parallel codebook prediction at each timestep\. Given a causally aligned segment from Section[II\-A](https://arxiv.org/html/2609.30416#S2.SS1), let𝐐=\(𝐪1,…,𝐪K\)\\mathbf\{Q\}=\(\\mathbf\{q\}\_\{1\},\\dots,\\mathbf\{q\}\_\{K\}\)denote the target speech codec sequence, where each frame containsNqN\_\{q\}discrete code indices,𝐪k∈\{1,…,\|𝒱\(q\)\|\}Nq\\mathbf\{q\}\_\{k\}\\in\\\{1,\\dots,\|\\mathcal\{V\}^\{\(q\)\}\|\\\}^\{N\_\{q\}\}\. The target text sequence𝐞=\(e1,…,ei,⟨pad⟩\)\\mathbf\{e\}=\(e\_\{1\},\\dots,e\_\{i\},\\langle\\text\{pad\}\\rangle\)is padded to match the codec length, such that\|𝐐\|=\|𝐞\|\|\\mathbf\{Q\}\|=\|\\mathbf\{e\}\|\. FAST processes source speech prefix𝐬i:=𝐬1:τs​r​ce​n​\[i\]\\mathbf\{s\}\_\{i\}:=\\mathbf\{s\}\_\{1:\\tau^\{en\}\_\{src\}\[i\]\}incrementally and generates both text tokens𝐞\\mathbf\{e\}and codec tokens𝐐\\mathbf\{Q\}in a streaming fashion\. We factorize the joint conditional distributionP⁡\(𝐐,𝐞∣𝐬\)P\(\\mathbf\{Q\},\\mathbf\{e\}\\mid\\mathbf\{s\}\)as

P⁡\(𝐐,𝐞∣𝐬\)\\displaystyle P\(\\mathbf\{Q\},\\mathbf\{e\}\\mid\\mathbf\{s\}\)=P⁡\(𝐞∣𝐬\)​P​\(𝐐∣𝐞,𝐬\)\\displaystyle=P\(\\mathbf\{e\}\\mid\\mathbf\{s\}\)\\;P\(\\mathbf\{Q\}\\mid\\mathbf\{e\},\\mathbf\{s\}\)\(2\)whereP⁡\(𝐞∣𝐬\)P\(\\mathbf\{e\}\\mid\\mathbf\{s\}\)is the speech to text translation module \(Simul\-S2T\), andP⁡\(𝐐∣𝐞,𝐬\)P\(\\mathbf\{Q\}\\mid\\mathbf\{e\},\\mathbf\{s\}\)is the text\-to\-speech synthesis module \(Simul\-T2S\)\. FAST parameterizesP⁡\(𝐞∣𝐬\)P\(\\mathbf\{e\}\\mid\\mathbf\{s\}\)with an ASR\-pretrained streaming speech encoder, followed by a pretrained LLM\. The speech encoder maps𝐬\\mathbf\{s\}to subsampled speech representations𝐇∈ℝK×d\\mathbf\{H\}\\in\\mathbb\{R\}^\{K\\times d\}, matching the codec frame rate\. The causally aligned target text embeddings𝐄∈ℝK×d\\mathbf\{E\}\\in\\mathbb\{R\}^\{K\\times d\}derived from𝐞\\mathbf\{e\}are fused with the speech embeddings𝐇\\mathbf\{H\}by element\-wise addition,𝐆=𝐄\+𝐇\\mathbf\{G\}=\\mathbf\{E\}\+\\mathbf\{H\}\. The fused representation is then passed to the LLM, which produces logits𝐳ktxt\\mathbf\{z\}^\{\\text\{txt\}\}\_\{k\}for the next token, yieldingP\(ek∣e<k,𝐆1:κ⁡\(k\)\)=Softmax\(𝐳ktxt\)\.P\(e\_\{k\}\\mid e\_\{<k\},\\mathbf\{G\}\_\{1:\\kappa\(k\)\}\)=\\mathrm\{Softmax\}\(\\mathbf\{z\}^\{\\text\{txt\}\}\_\{k\}\)\.The Simul\-S2T objective is the autoregressive negative log\-likelihood:

ℒs2t=−∑k=1\|𝐐\|logP\(ek∣e<k,𝐆1:κ⁡\(k\)\),\\mathcal\{L\}\_\{\\text\{s2t\}\}=\-\\sum\_\{k=1\}^\{\|\\mathbf\{Q\}\|\}\\log P\(e\_\{k\}\\mid e\_\{<k\},\\mathbf\{G\}\_\{1:\\kappa\(k\)\}\),\(3\)
![Refer to caption](https://arxiv.org/html/2609.30416v1/fast.png)Fig\. 2:Overview of the proposed FAST architecture with multistream joint sequence modeling\.whereκ⁡\(k\)\\kappa\(k\)denotes the frame index up to which the source speech has been consumed when predicting thekk\-th target token\. For the Simul\-T2S module, the translation speech waveform is first converted into discrete tokens using streaming NanoCodec\. The speech decoder is a decoder\-only architecture that autoregressively predicts next codec token𝐪k\\mathbf\{q\}\_\{k\}conditioned on previous codec tokens𝐪<k\\mathbf\{q\}\_\{<k\}, previous text translatione≤ke\_\{\\leq k\}, and fixed\-dimensional speaker embedding𝐮∈ℝd\\mathbf\{u\}\\in\\mathbb\{R\}^\{d\}extracted fromsis\_\{i\}\. The Simul\-T2S training objective is the autoregressive negative log\-likelihood:

ℒt2s=−∑k=1\|𝐐\|logP\(𝐪k∣𝐪<k,e≤k,𝐮\)\.\\mathcal\{L\}\_\{\\text\{t2s\}\}=\-\\sum\_\{k=1\}^\{\|\\mathbf\{Q\}\|\}\\log P\(\\mathbf\{q\}\_\{k\}\\mid\\mathbf\{q\}\_\{<k\},e\_\{\\leq k\},\\mathbf\{u\}\)\.\(4\)The overall multitask objectiveℒt​o​t\\mathcal\{L\}\_\{tot\}is the weighted sum of the two objectives:

ℒt​o​t=αs​2​t​ℒs​2​t\+αt​2​s​ℒt​2​s,\\mathcal\{L\}\_\{tot\}=\\alpha\_\{s2t\}\\mathcal\{L\}\_\{s2t\}\+\\alpha\_\{t2s\}\\mathcal\{L\}\_\{t2s\},\(5\)whereαs​2​t\\alpha\_\{s2t\}andαt​2​s\\alpha\_\{t2s\}are hyperparameters balance the contributions of the two tasks\. During training, we employ teacher forcing\. At inference time, the decoder conditions on the autoregressive LLM predictions obtained via greedy decoding\. Finally, the streaming codec decoder reconstructs the target waveform from the predicted discrete speech tokens\.

### II\-CCausality\-Aware Average Lagging \(CAAL\)

Fig\. 3:Illustration of LAAL and CAAL computation\. Oracle delays \(green\) and system delays \(blue\) are paired according to the alignment mapping \(arrows\), and lagging values \(red\) are computed as the difference between each system delay and its aligned oracle delay\.In this paper, we highlight key limitations of existing translation latency metrics, which do not incorporate translation alignments\[[26](https://arxiv.org/html/2609.30416#bib.bibx26),[27](https://arxiv.org/html/2609.30416#bib.bibx27),[28](https://arxiv.org/html/2609.30416#bib.bibx28),[29](https://arxiv.org/html/2609.30416#bib.bibx29)\]\. These metrics assume uniform word timing and penalize all delays equally, regardless of whether they are linguistically necessary\. This leads to inconsistent and misleading latency estimates\. We primarily compare against LAAL, a recent and widely used latency metric\. Figure[3](https://arxiv.org/html/2609.30416#S2.F3)illustrates this issue using an example from Spanish–English CVSS\-T\. System 1 follows a fixed wait\-kkpolicy withk=120k=120ms, emitting translations word by word without modeling reordering\. System 2, adopts a causality\-aware dynamic chunking while consuming speech frames of 40 ms, emitting each target token only after sufficient source evidence has been observed\. LAAL computes the ideal delay under a uniform timing assumption asdi∗=\(i−1\)\|ss​r​c1:T\|max⁡\(\|𝐞r​e​f\|,\|𝐞h​y​p\|\)=9947≈142d\_\{i\}^\{\*\}=\(i\-1\)\\frac\{\|s^\{src\}\_\{1:T\}\|\}\{\\max\(\|\\mathbf\{e\}\_\{ref\}\|,\|\\mathbf\{e\}\_\{hyp\}\|\)\}=\\frac\{994\}\{7\}\\approx 142ms\. Because LAAL ignores alignments, it matches oracle and system emission times incorrectly\. As a result, insertions may produce negative lagging values \(e\.g\.,“little”assigned−208\-208and“the”assigned−230\-230\), which artificially lower the overall latency score\. Consequently, LAAL incorrectly reports System 1 as nearly four times faster than System 2\. Consequently, LAAL incorrectly reports System 1 as nearly four times faster than System 2\.

In contrast, System 2 closely follows the true source end times: it incurs only 20 ms delay for“The”, 0 ms for“misterio”and“Duró muy poco”, and 6 ms for“ayer”\. Its initial 480 ms wait only slightly exceeds the 460 ms required to reorder“Duró muy poco el”before emitting“the”\. The first 460 ms waiting is linguistically justified and should not be counted as system\-induced delay\. CAAL addresses this by using causal alignments to account for reordering: it neither rewards premature emissions nor penalizes necessary waiting\. Instead, it isolates excess delay beyond the ideal causal policy and correctly identifies System 2 as substantially faster than System 1\.

Formally, CAAL measures the difference between the system emission timeti′t^\{\\prime\}\_\{i\}for target wordeie\_\{i\}and its ideal causal delayτs​r​ce​n​\[i\]\\tau^\{en\}\_\{src\}\[i\], derived from the alignment procedure in Section[II\-A](https://arxiv.org/html/2609.30416#S2.SS1)\. It uses two alignment types:

- •Causal alignment𝒫\\mathcal\{P\}from Section[II\-A](https://arxiv.org/html/2609.30416#S2.SS1)
- •Hypothesis alignmentℬ⊆\{1,…,Ihyp\}×\{1,…,I\}\\mathcal\{B\}\\subseteq\\\{1,\\dots,I\_\{\\text\{hyp\}\}\\\}\\times\\\{1,\\dots,I\\\}, from hypothesis translation to reference translation\.

LetRRbe the set of reference word indices, and letM⊆RM\\subseteq Rdenote the subset aligned to hypothesis words throughℬ\\mathcal\{B\}\. For aligned words, the effective delay is

di=max⁡\(ti′−τs​r​ce​n​\[i\],0\)\.d\_\{i\}=\\max\\left\(t^\{\\prime\}\_\{i\}\-\\tau^\{en\}\_\{src\}\[i\],0\\right\)\.\(6\)The maximum prevents systems from artificially lowering latency by emitting reordered words before the ideal causal policy permits them\. For deletions,ei∈R∖Me\_\{i\}\\in R\\setminus M, we approximate the delay using theρ\\rho\-quantile of the sentence\-level aligned empirical delay distribution:

did​e​l=Qρ​\(\{dj∣j∈M\}\),d^\{del\}\_\{i\}=Q\_\{\\rho\}\\left\(\\\{d\_\{j\}\\mid j\\in M\\\}\\right\),\(7\)which penalizes omissions while remaining robust to outliers\. If deletions should be ignored, one may instead setdid​e​l=0d^\{del\}\_\{i\}=0\. The CAAL metric is finally computed as

CAAL=1\|R\|​\(∑i∈Mdi\+∑i∈R∖Mdid​e​l\)\.\\mathrm\{CAAL\}=\\frac\{1\}\{\|R\|\}\\left\(\\sum\_\{i\\in M\}d\_\{i\}\+\\sum\_\{i\\in R\\setminus M\}d^\{del\}\_\{i\}\\right\)\.\(8\)Intuitively, CAAL measures the average additional waiting introduced by the system beyond an ideal causal policy that emits each target word as soon as all aligned source evidence has been observed\.

## IIIExperimental Setup

Data & pre\-processing:We evaluate Spanish, German, and French speech\-to\-speech translation into English\. Using the pipeline in Section[II\-A](https://arxiv.org/html/2609.30416#S2.SS1), we construct causally aligned S2ST training data from CVSS\-T and additional internal datasets, totaling approximately 2\.7K hours per language pair\. We extract 80\-dimensional mel\-spectrograms with a 25 ms window and 10 ms frame shift\. During training, we apply SpecAugment with two frequency masks of width up to 27 bins and ten time masks capped at 5% of the sequence length\.

TABLE II:Ablation study on the CVSS\-T development set comparing alignment methods \(NFA vs\. MFA\), multimodal representations \(interleaved, IL, vs\. multistream, MS\), and TTS conditioning \(text vs\. LLM latent representations, Lat\)\. Translation quality is evaluated with BLEU, chrF\+\+, and COMET, with speech metrics computed on ASR transcriptions\.Architecture details:The speech encoder is initialized from a multilingual streaming FastConformer ASR model\[[30](https://arxiv.org/html/2609.30416#bib.bibx30)\], with 17 layers, 8 attention heads, 1024 attention dimension, and 2048 feed\-forward dimension\. The LLM backbone is initialized from Qwen2\.5\-1\.5B\-Instruct444[https://huggingface\.co/Qwen/Qwen2\.5\-1\.5B\-Instruct](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)\. The TTS component is initialized from MagpieTTS model\[[31](https://arxiv.org/html/2609.30416#bib.bibx31)\], which uses streaming NanoCodec\[[32](https://arxiv.org/html/2609.30416#bib.bibx32)\]withNq=13N\_\{q\}=13codebooks\. The TTS decoder has 12 attention layers with dimension 768, 12 attention heads per layer, and 3072 feed\-forward dimension\. Both the speech encoder and the codec operate at a frame rate of 12\.5 frames per second\. FAST model has a total of 2B parameters\. Training:Experiments are implemented in PyTorch with NeMo framework\[[33](https://arxiv.org/html/2609.30416#bib.bibx33)\]and trained on 32 NVIDIA A100 \(80GB\) GPUs\. During FAST training, all parameters are optimized except the frozen codec and speaker encoder\. We use AdamW with learning rate10−410^\{\-4\}, inverse square\-root annealing with 4K warmup steps, and train for 25 epochs\. The multitask weights in Eq\. \([5](https://arxiv.org/html/2609.30416#S2.E5)\) areαs​2​t=3\\alpha\_\{s2t\}=3andαt​2​s=2\\alpha\_\{t2s\}=2\. Evaluation:We evaluate FAST on the CVSS\-T development and test sets\. To ensure a comprehensive translation evaluation, we report BLEU\[[34](https://arxiv.org/html/2609.30416#bib.bibx34)\], chrF\+\+, and COMET\[[35](https://arxiv.org/html/2609.30416#bib.bibx35)\]555We use theUnbabel/wmt22\-comet\-damodel\.\. For speech translation evaluation, we transcribe the generated speech using a pretrained ASR666[https://huggingface\.co/nvidia/stt\_en\_fastconformer\_transducer\_large](https://huggingface.co/nvidia/stt_en_fastconformer_transducer_large)and compute the same translation metrics on the resulting transcripts\. All translation evaluations are case\-insensitive and punctuation\-free\. Latency is measured with LAAL and our proposed CAAL, both implemented using the SimulEval framework\[[36](https://arxiv.org/html/2609.30416#bib.bibx36)\]\. For CAAL deletion penalties, we setρ=0\.9\\rho=0\.9in Eq\. \([7](https://arxiv.org/html/2609.30416#S2.E7)\), a robust high\-quantile that mitigates the influence of outliers\. Perceptual speech quality is evaluated with UTMOS\-V2\[[37](https://arxiv.org/html/2609.30416#bib.bibx37)\], while speaker similarity is computed as the cosine similarity between ECAPA\-TDNN embeddings\[[38](https://arxiv.org/html/2609.30416#bib.bibx38)\]extracted from reference and generated speech\. Inference:During inference, the LLM decodes incrementally from speech encoder outputs arriving every 80 ms\. Text predictions are generated with greedy decoding, converted to text embeddings, and passed to the TTS decoder\.

## IVResults

### IV\-AAblation analysis

In this section, we assess the impact of key components of the proposed FAST approach, including alignment quality, the multimodal representation modeling, and TTS conditioning, as shown in Table[II](https://arxiv.org/html/2609.30416#S3.T2)\. All experiments are conducted on the Spanish–English language pair using a fixed 2 s chunking policy\. The baseline \(Row 1\) represents a common class of prior LLM\-based Simul\-S2T approaches that rely on interleaved speech–text multimodal representations\[[39](https://arxiv.org/html/2609.30416#bib.bibx39),[11](https://arxiv.org/html/2609.30416#bib.bibx11),[12](https://arxiv.org/html/2609.30416#bib.bibx12)\]\. In this baseline, TTS is conditioned on the text predicted by the LLM\. We first assess speech\-to\-text alignment quality by comparing CTC\-based alignments from NeMo Forced Aligner \(NFA\) with HMM\-based alignments from Montreal Forced Aligner \(MFA\)\. Replacing NFA with MFA improves performance by \+4\.4 BLEU on text and \+1\.0 BLEU on speech, which we attribute to MFA’s more precise word\-boundary estimates\. Next, we compare aligned interleaving \(Row 2\) and aligned multistream joint modeling \(Row 4\)\. Multistream modeling consistently outperforms interleaving, yielding gains of \+3\.6 BLEU, \+4\.9 chrF\+\+, and \+5\.5 COMET on text translation, and \+7\.3 BLEU, \+7\.5 chrF\+\+, and \+5\.7 COMET on speech translation\. We then evaluate two TTS conditioning strategies: LLM latent representations \(FAST\-Lat, Row 3\) and argmax\-decoded text predictions \(Row 4\)\. Under our limited\-data setting, latent conditioning substantially degrades speech translation quality, resulting in drops of 16\.3 BLEU, 25\.3 chrF\+\+, and 12\.2 COMET\. Finally, to assess modularity, we replace the Qwen2\-1\.5B multilingual multimodal LLM with TinyLLaMA\-1\.2B, a text\-only multilingual model\. Despite the lack of multimodal pretraining, TinyLLaMA achieves comparable performance, indicating that the proposed approach is largely backbone\-agnostic and does not require large\-scale speech–text translation pretraining\.

TABLE III:Comparison of FAST on CVSS\-T development set using the causality\-aware adaptive policy \(CAP\) and a fixed policy\. Translation quality is reported using BLEU, chrF\+\+, and COMET while latency is measured using LAAL and CAAL\.
### IV\-BCausality\-Aware Adaptive Policy

TABLE IV:Comparison of translation performance on the CVSS\-T test set\. Translation quality is evaluated with BLEU, chrF\+\+, and COMET, with speech metrics computed on ASR transcriptions\. Audio quality is measured using UTMOS\-V2, and speaker similarity is measured using cosine similarity\.In this section, we evaluate whether the proposed causality\-aware adaptive policy \(CAP\) improves the quality–latency trade\-off over the commonly used fixed chunking policy adopted by several prior LLM\-based Simul\-S2T approaches\[[9](https://arxiv.org/html/2609.30416#bib.bibx9),[10](https://arxiv.org/html/2609.30416#bib.bibx10),[11](https://arxiv.org/html/2609.30416#bib.bibx11),[12](https://arxiv.org/html/2609.30416#bib.bibx12)\]\. In all experiments, we use the best\-performing FAST configuration from Table[II](https://arxiv.org/html/2609.30416#S3.T2)\. Without chunk merging \(0\.4 s average input\), performance degrades substantially despite sufficient source information being theoretically available to the LLM\. Increasing the average input chunk duration to 0\.9 s through chunk merging improves text translation quality to near\-baseline levels, suggesting that the LLM benefits from additional contextual grounding\. At an average input duration of 1\.5 s, FAST\-CAP surpasses the fixed 2 s policy, improving text translation by \+1\.1 BLEU, \+3\.0 chrF\+\+, and \+2\.5 COMET\.

Next, we analyze latency using both LAAL and the proposed CAAL metric\. As expected, latency increases with larger input chunks; however, LAAL becomes less discriminative at longer chunk durations\. For example, on Spanish–English, FAST\-CAP \(1\.5 s\) and FAST\-Fixed \(2 s\) differ by only 1\.7% in LAAL, whereas CAAL captures a more substantial 8% relative reduction\. This suggests that CAAL provides a more discriminative measure of latency\. To evaluate the generalization of FAST\-CAP beyond Spanish, we extend the comparison to German and French\. On German, FAST\-CAP improves text translation by \+1\.2 BLEU, \+2\.8 chrF\+\+, and \+0\.6 COMET, while reducing CAAL by 16% relative, compared to only 4\.7% under LAAL\. On French, translation performance remains comparable to the fixed baseline, while FAST\-CAP achieves a 26% relative reduction in CAAL, compared to 16\.6% under LAAL\. These results demonstrate that FAST\-CAP consistently improves the quality–latency trade\-off across language pairs\. Moreover, CAAL provides a more discriminative assessment of latency than LAAL\.

### IV\-CComparison with state\-of\-the\-art models

In this section, we compare FAST\-CAP with Hibiki\-Zero\[[40](https://arxiv.org/html/2609.30416#bib.bibx40)\]and streaming SeamlessM4T\[[14](https://arxiv.org/html/2609.30416#bib.bibx14)\], as shown in Table[IV](https://arxiv.org/html/2609.30416#S4.T4)\. To better match their scale, we replace the 0\.1B multilingual ASR encoder with a 0\.6B encoder777[https://huggingface\.co/nvidia/nemotron\-3\.5\-asr\-streaming\-0\.6b](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b), yielding a 2\.5B\-parameter model, and increase the average input chunk size from 1\.5 s to 2 s\. FAST\-CAP is trained on total of 8K hours of multilingual data across French, Spanish, and German, compared with 160K hours for Hibiki\-Zero and 145K/245K hours of speech\-to\-speech/speech\-to\-text data for SeamlessM4T\. Despite this data gap, FAST\-CAP outperforms Hibiki\-Zero in text translation quality, improving by \+2\.3 BLEU, \+1\.3 chrF\+\+, and \+3\.2 COMET, while achieving comparable speaker similarity and reducing latency by 38\.8% in CAAL and 40\.8% in LAAL\. Compared with SeamlessM4T, FAST\-CAP achieves comparable text translation and audio quality, higher speaker similarity \(0\.44 vs\. 0\.31\), and relative latency reductions of 30\.9% in CAAL and 33\.9% in LAAL\. In both comparisons, the main remaining gap for FAST\-CAP lies in generated speech translation quality\. To improve generated speech quality, FAST\-CAP\-L adopts the TTS model from Audio Flamingo 3\-Chat\[[41](https://arxiv.org/html/2609.30416#bib.bibx41)\], which clones the speaker’s voice using an audio prompt\. FAST\-CAP\-L also uses a larger translation\-focused LLM888[https://huggingface\.co/nvidia/Riva\-Translate\-4B\-Instruct\-v1\.1](https://huggingface.co/nvidia/Riva-Translate-4B-Instruct-v1.1)to broaden language coverage, bringing the total model size to 5B parameters\. FAST\-CAP\-L retains FAST\-CAP’s latency while achieving the best speech translation scores across BLEU, chrF\+\+, and COMET, with the largest gain in COMET: \+2\.1 points above Hibiki\-Zero\. It also achieves the highest speaker similarity \(0\.53 versus 0\.47 for Hibiki\-Zero\) while maintaining competitive audio quality\.

## VConclusion

In this work, we introduced a causality\-aware framework for LLM\-based simultaneous speech\-to\-speech translation that integrates causal constraints across data generation, source–target alignment, translation policy, and latency evaluation\. Our novel data pipeline derives a target\-driven adaptive policy that generates each target segment once sufficient source context is available\. The factorized S2ST architecture \(FAST\) decouples perception and generation representations, improving translation accuracy without compromising synthesis fidelity\. We also proposed an alignment\-aware latency metric \(CAAL\) that separates linguistically necessary reordering from avoidable system\-induced delay\. Experiments on multilingual benchmarks show that FAST combined with the causality\-aware adaptive policy \(CAP\) provides better quality–latency trade\-offs than fixed\-policy baselines\. Despite using limited training data, FAST\-CAP achieves state\-of\-the\-art speech translation quality while providing lower latency and higher speaker fidelity\. CAAL also offers a more accurate and discriminative latency assessment than existing latency metrics\. Future work will focus on further improving cross\-lingual voice transfer, extending the framework to longer conversational contexts, and supporting multi\-party and multimodal interactions\.

## References

## References

- \[1\]Haopeng Xie, Ismail Ulgen, Sofia Son, Berrak Sisman and Philipp Koehn“Is Prosody Lost in Translation? Fine\-Grained Cross\-Lingual Prosody Similarity Across Languages”In*arXiv preprint arXiv:2608\.27848*, 2026
- \[2\]Ibrahim Ahmad et al\.“Findings of the IWSLT 2024 evaluation campaign”In*Proceedings of the 21st International Conference on Spoken Language Translation \(IWSLT 2024\)*, 2024, pp\. 1–11DOI:[10\.18653/v1/2024\.iwslt\-1\.1](https://dx.doi.org/10.18653/v1/2024.iwslt-1.1)
- \[3\]Milind Agarwal et al\.“FINDINGS OF THE IWSLT 2023 EVALUATION CAMPAIGN”In*Proceedings of the 20th International Conference on Spoken Language Translation \(IWSLT 2023\)*, 2023, pp\. 1–61DOI:[10\.18653/v1/2023\.iwslt\-1\.1](https://dx.doi.org/10.18653/v1/2023.iwslt-1.1)
- \[4\]Amir Hussein et al\.“JHU IWSLT 2023 Dialect Speech Translation System Description”In*Proceedings of the 20th International Conference on Spoken Language Translation \(IWSLT 2023\)*, 2023, pp\. 283–290DOI:[10\.18653/v1/2023\.iwslt\-1\.26](https://dx.doi.org/10.18653/v1/2023.iwslt-1.26)
- \[5\]Zhichao Huang et al\.“Speech Translation with Large Language Models: An Industrial Practice”In*Proc\. EMNLP*, 2023
- \[6\]Baban Gain, Dibyanayan Bandyopadhyay, Asif Ekbal and Trilok Singh“Bridging the linguistic divide: a survey on leveraging large language models for machine translation”In*arXiv preprint arXiv:2504\.01919*, 2025
- \[7\]Chenyang Lyu et al\.“A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models”In*Proc\. COLING*, 2024, pp\. 1339–1352
- \[8\]Roman Koshkin, Katsuhito Sudoh and Satoshi Nakamura“TransLLaMa: LLM\-based Simultaneous Translation System”In*Proc\. ACL*, 2024, pp\. 461–476DOI:[10\.18653/v1/2024\.findings\-emnlp\.27](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.27)
- \[9\]Victor Agostinelli, Max Wild, Matthew Raffel, Kazi Fuad and Lizhong Chen“Simul\-LLM: A Framework for Exploring High\-Quality Simultaneous Translation with Large Language Models”In*Proc\. ACL*, 2024, pp\. 10530–10541DOI:[10\.18653/v1/2024\.acl\-long\.567](https://dx.doi.org/10.18653/v1/2024.acl-long.567)
- \[10\]Roman Koshkin, Katsuhito Sudoh and Satoshi Nakamura“LLMs Are Zero\-Shot Context\-Aware Simultaneous Translators”In*Proc\. EMNLP*, 2024, pp\. 1192–1207
- \[11\]Hayato Futami et al\.“Scheduled Interleaved Speech\-Text Training for Speech\-to\-Speech Translation with LLMs”In*Proc\. Interspeech*, 2025, pp\. 36–40DOI:[10\.21437/Interspeech\.2025\-1595](https://dx.doi.org/10.21437/Interspeech.2025-1595)
- \[12\]Siqi Ouyang, Xi Xu and Lei Li“InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model”In*Proc\. ACL*, 2025, pp\. 3032–3046DOI:[10\.18653/v1/2025\.findings\-acl\.157](https://dx.doi.org/10.18653/v1/2025.findings-acl.157)
- \[13\]Ye Jia, Michelle Ramanovich, Tal Remez and Roi Pomerantz“Translatotron 2: High\-quality direct speech\-to\-speech translation with voice preservation”In*International conference on machine learning*, 2022, pp\. 10120–10134
- \[14\]Loı̈c Barrault et al\.“Seamless: Multilingual Expressive and Streaming Speech Translation”In*arXiv preprint arXiv:2312\.05187*, 2023
- \[15\]Tom Labiausse, Laurent Mazaré, Edouard Grave, Alexandre Défossez and Neil Zeghidour“High\-Fidelity Simultaneous Speech\-To\-Speech Translation”In*Proc\. ICML*267, Proceedings of Machine Learning Research, 2025, pp\. 32116–32129
- \[16\]Amir Hussein, Sameer Khurana, Gordon Wichern, François\. Germain and Jonathan Le“HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement”In*Proc\. Interspeech*, 2025, pp\. 5393–5397DOI:[10\.21437/Interspeech\.2025\-2063](https://dx.doi.org/10.21437/Interspeech.2025-2063)
- \[17\]Dan Huang, Fang Li and Hang Guo“Chunking in simultaneous interpreting: The impact of task complexity and translation directionality on lexical bundles”In*Frontiers in Psychology*14, 2023, pp\. 01–14
- \[18\]S\. Song and D\. Li“Aptitude for interpreting: The predictive value of cognitive fluency”In*The Interpreter and Translator Trainer*17\.1, 2023, pp\. 155–172
- \[19\]Ke Hu et al\.“Efficient and Direct Duplex Modeling for Speech\-to\-Speech Language Model”In*Proc\. Interspeech*, 2025, pp\. 2715–2719
- \[20\]Ye Jia, Mikhail\. Ramanovich, Qian Wang and Heiga Zen“CVSS Corpus and Massively Multilingual Speech\-to\-Speech Translation”In*Proceedings of the Thirteenth Language Resources and Evaluation Conference \(LREC 2022\)*, 2022, pp\. 6691–6703
- \[21\]Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner and Morgan Sonderegger“Montreal Forced Aligner: Trainable Text\-Speech Alignment Using Kaldi”In*Proc\. Interspeech*, 2017, pp\. 498–502DOI:[10\.21437/Interspeech\.2017\-1386](https://dx.doi.org/10.21437/Interspeech.2017-1386)
- \[22\]Zi\-Yi Dou and Graham Neubig“Word Alignment by Fine\-tuning Embeddings on Parallel Corpora”In*Proc\. ACL*, 2021
- \[23\]Piotr Żelasko, Daniel Povey, Jan Trmal and Sanjeev Khudanpur“Lhotse: a speech data representation library for the modern deep learning ecosystem”In*NeurIPS Data\-Centric AI Workshop*, 2021
- \[24\]Amir Hussein, Sameer Khurana, Gordon Wichern, François\. Germain and Jonathan Le Roux“HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement”In*Proc\. Interspeech*, 2025, pp\. 5393–5397DOI:[10\.21437/Interspeech\.2025\-2063](https://dx.doi.org/10.21437/Interspeech.2025-2063)
- \[25\]Edresson Casanova et al\.“NanoCodec: Towards High\-Quality Ultra Fast Speech LLM Inference”In*Proc\. Interspeech*, 2025, pp\. 5028–5032DOI:[10\.21437/Interspeech\.2025\-827](https://dx.doi.org/10.21437/Interspeech.2025-827)
- \[26\]Kyunghyun Cho and Masha Esipova“Can neural machine translation do simultaneous translation?”In*arXiv preprint arXiv:1606\.02012*, 2016
- \[27\]Mingbo Ma et al\.“STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix\-to\-Prefix Framework”In*Proc\. ACL*, 2019, pp\. 3025–3036DOI:[10\.18653/v1/P19\-1289](https://dx.doi.org/10.18653/v1/P19-1289)
- \[28\]Sara Papi, Marco Gaido, Matteo Negri and Marco Turchi“Over\-generation cannot be rewarded: Length\-adaptive average lagging for simultaneous speech translation”In*Proceedings of the Third Workshop on Automatic Simultaneous Translation*, 2022, pp\. 12–17
- \[29\]Colin Cherry and George Foster“Thinking slow about latency evaluation for simultaneous machine translation”In*arXiv preprint arXiv:1906\.00048*, 2019
- \[30\]Vahid Noroozi, Somshubra Majumdar, Ankur Kumar, Jagadeesh Balam and Boris Ginsburg“Stateful conformer with cache\-based inference for streaming automatic speech recognition”In*Proc\. ICASSP*, 2024, pp\. 12041–12045
- \[31\]Shehzeen Hussain et al\.“Koel\-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance”In*Proc\. EMNLP*, 2025, pp\. 21219–21234DOI:[10\.18653/v1/2025\.emnlp\-main\.1076](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1076)
- \[32\]Edresson Casanova et al\.“NanoCodec: Towards High\-Quality Ultra Fast Speech LLM Inference”In*Proc\. Interspeech*, 2025, pp\. 5028–5032
- \[33\]Oleksii Kuchaiev et al\.“Nemo: a toolkit for building ai applications using neural modules”In*arXiv preprint arXiv:1909\.09577*, 2019
- \[34\]Kishore Papineni, Salim Roukos, Todd Ward and Wei\-Jing Zhu“Bleu: a method for automatic evaluation of machine translation”In*Proc\. ACL*, 2002, pp\. 311–318
- \[35\]Ricardo Rei, Craig Stewart, Ana Farinha and Alon Lavie“COMET: A Neural Framework for MT Evaluation”In*Proc\. EMNLP*, 2020, pp\. 2685–2702
- \[36\]Xutai Ma, Mohammad Dousti, Changhan Wang, Jiatao Gu and Juan Pino“SIMULEVAL: An evaluation toolkit for simultaneous translation”In*Proc\. EMNLP*, 2020, pp\. 144–150
- \[37\]Kaito Baba, Wataru Nakata, Yuki Saito and Hiroshi Saruwatari“The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high\-quality synthetic speech”In*Proc\. SLT*, 2024, pp\. 818–824
- \[38\]Brecht Desplanques, Jenthe Thienpondt and Kris Demuynck“ECAPA\-TDNN: Emphasized Channel Attention, propagation and aggregation in TDNN based speaker verification”, 2020, pp\. 3830–3834
- \[39\]Biao Fu et al\.“Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture”In*arXiv preprint arXiv:2504\.11809*, 2025
- \[40\]Tom Labiausse, Romain Fabre, Yannick Estève, Alexandre Défossez and Neil Zeghidour“Simultaneous speech\-to\-speech translation without aligned data”In*arXiv preprint arXiv:2602\.11072*, 2026
- \[41\]Sreyan Ghosh et al\.“Audio flamingo 3: Advancing audio intelligence with fully open large audio language models”In*Advances in Neural Information Processing Systems*38, 2026, pp\. 41819–41886

Similar Articles

Streaming Speech-to-Text Translation with a SpeechLLM

arXiv cs.CL

Presents a SpeechLLM architecture for streaming speech-to-text translation that adaptively decides when to output tokens based on audio, achieving 1-2 second latency with quality close to non-streaming baselines.