M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

arXiv cs.CL Papers

Summary

M3-DuplexBench is a new multi-turn, multilingual, multidomain benchmark for evaluating full-duplex spoken dialogue systems, supporting English and Japanese across casual conversation and question answering domains.

arXiv:2607.29125v1 Announce Type: new Abstract: Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. However, fair comparisons in multi-turn conversations remain a challenge. In addition, existing benchmarks provide limited coverage of languages and dialogue domains. We propose M3-DuplexBench, a multi-turn, multilingual, multidomain benchmark for FDSDSs. M3-DuplexBench supports English and Japanese and covers both casual conversation and multi-turn question answering. In addition, we evaluate models under multiple dialogue context settings, including single-turn, user-only, and teacher-forced full-context settings, to analyze how dialogue history affects model behavior. Experiments with recent FDSDSs reveal model-specific turn-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:35 AM

# M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
Source: [https://arxiv.org/html/2607.29125](https://arxiv.org/html/2607.29125)
###### Abstract

Full\-duplex spoken dialogue systems \(FDSDSs\) can listen while speaking, enabling natural behaviors such as smooth turn\-taking, backchannel handling, and user barge\-in handling\. However, fair comparisons in multi\-turn conversations remain a challenge\. In addition, existing benchmarks provide limited coverage of languages and dialogue domains\. We propose M3\-DuplexBench, a multi\-turn, multilingual, multidomain benchmark for FDSDSs\. M3\-DuplexBench supports English and Japanese and covers both casual conversation and multi\-turn question answering\. In addition, we evaluate models under multiple dialogue context settings, including single\-turn, user\-only, and teacher\-forced full\-context settings, to analyze how dialogue history affects model behavior\. Experiments with recent FDSDSs reveal model\-specific turn\-taking characteristics, clear performance gaps across languages and domains, and mixed effects of dialogue context\.

## IIntroduction

Spoken dialogue systems \(SDSs\) aim to communicate with users through speech\. Conventional SDSs often follow a half\-duplex communication style: they wait until the user finishes speaking before generating a response\[[5](https://arxiv.org/html/2607.29125#bib.bib5),[23](https://arxiv.org/html/2607.29125#bib.bib6)\]\. In contrast, full\-duplex spoken dialogue systems \(FDSDSs\), which can listen while speaking, have recently attracted attention\[[16](https://arxiv.org/html/2607.29125#bib.bib7),[10](https://arxiv.org/html/2607.29125#bib.bib8)\]\. FDSDSs can support more natural and interactive conversations by handling user interruptions, continuing to speak during user backchannels, and producing natural overlapping speech\.

Along with the development of FDSDSs, many benchmarks have been proposed to automatically evaluate their behavior\. Existing benchmarks evaluate various aspects of full\-duplex dialogue, such as turn\-taking\[[3](https://arxiv.org/html/2607.29125#bib.bib9),[14](https://arxiv.org/html/2607.29125#bib.bib10),[13](https://arxiv.org/html/2607.29125#bib.bib11)\], response naturalness\[[12](https://arxiv.org/html/2607.29125#bib.bib12)\], and instruction following\[[12](https://arxiv.org/html/2607.29125#bib.bib12),[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. Several studies have also evaluated multi\-turn conversations\. For example, recent work evaluates multi\-turn interaction using an automated examiner\[[12](https://arxiv.org/html/2607.29125#bib.bib12)\]or pre\-collected spoken dialogue data\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. These studies provide important tools for evaluating full\-duplex behavior\.

For a useful benchmark, one important requirement is comparability: multiple models should be evaluated under fair and consistent conditions\. Many FDSDS benchmarks perform single\-turn evaluation, where the system is evaluated on its response to one user input\. This setting enables fair comparison because every model receives the same input\. However, comparable evaluation in multi\-turn conversations is not straightforward\. Dynamic evaluation with an automated examiner can provide adaptive interaction, but different systems may receive different dialogue histories\. This makes fine\-grained comparison difficult\. Static evaluation with fixed user inputs improves comparability, but it can create a mismatch between the fixed user utterances and the system’s previous responses\. This context mismatch can make the dialogue history unnatural and affect the validity of the evaluation\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\.

Depending on the evaluation objective, domain and language coverage may also be important\. Previous works used several domains, such as casual conversations\[[14](https://arxiv.org/html/2607.29125#bib.bib10),[9](https://arxiv.org/html/2607.29125#bib.bib13)\]and task\-oriented dialogues \(e\.g\., question answering, mental health, and reservations\)\[[12](https://arxiv.org/html/2607.29125#bib.bib12),[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. However, cross\-domain comparisons remain limited\. This is important because appropriate full\-duplex behavior can vary greatly depending on the type of conversation\. For example, casual conversations and question answering may have different turn\-taking patterns and user expectations\. Language coverage is also limited\. Most existing FDSDS benchmarks are designed mainly for English dialogues, although models for non\-English languages are also being studied\[[17](https://arxiv.org/html/2607.29125#bib.bib14),[1](https://arxiv.org/html/2607.29125#bib.bib15),[29](https://arxiv.org/html/2607.29125#bib.bib16)\]\. Benchmarks for evaluating non\-English models, including Japanese models, remain limited\.

In this work, we propose M3\-DuplexBench, a multi\-turn, multilingual, multidomain benchmark for FDSDSs\.It supports English and Japanese, covers casual conversation and multi\-turn question answering, and evaluates full\-duplex behavior across timing and content metrics\. A key feature of M3\-DuplexBench is controlled multi\-turn evaluation using teacher\-forced inference, which enables fine\-grained model comparison under the same coherent dialogue context\. The evaluation revealed several key findings:

- •First, dialogue history helped FDSDSs understand the current question and generate more accurate answers in multi\-turn QA\.
- •Second, the results suggested that FDSDSs adjusted their response timing to match the patterns in the context\.
- •Third, the cross\-domain comparison showed that smooth turn\-taking was easier in the task\-oriented domain than in the chat domain\.
- •Finally, the cross\-lingual comparison showed that Japanese models underperformed English models, and that their main limitation lay in language understanding and generation rather than timing behavior\.

TABLE I:Multi\-turn evaluation protocols used in existing FDSDS benchmarks\. FDB and MTR denote Full\-Duplex\-Bench and MTR\-DuplexBench, respectively\.†\\daggerFull\-context conditioning is applied only to dialogue quality evaluation\.Multi\-turnFDBTalkingFD\-BenchMTROursevaluation\[[14](https://arxiv.org/html/2607.29125#bib.bib10),[13](https://arxiv.org/html/2607.29125#bib.bib11),[12](https://arxiv.org/html/2607.29125#bib.bib12)\]Turns\[[3](https://arxiv.org/html/2607.29125#bib.bib9)\]\[[19](https://arxiv.org/html/2607.29125#bib.bib17)\]\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]None\(Single\-turn\)✓v1, 1\.5✓Dynamic evaluation✓v2✓Static evaluationUser context✓✓✓Full context✓†✓
## IIRelated Work

Benchmarks for FDSDSs have recently been proposed to evaluate different aspects of full\-duplex interaction\. Full\-Duplex\-Bench evaluates smooth turn\-taking, pause handling, backchanneling, and barge\-in handling\[[14](https://arxiv.org/html/2607.29125#bib.bib10)\]\. Full\-Duplex\-Bench v1\.5 extends the evaluation to overlapping speech scenarios, including user backchannels, barge\-ins, talking to others, and background speech\[[13](https://arxiv.org/html/2607.29125#bib.bib11)\]\. Several studies have also evaluated multi\-turn interactions through conversations with humans\[[3](https://arxiv.org/html/2607.29125#bib.bib9)\]or automated examiners\[[12](https://arxiv.org/html/2607.29125#bib.bib12)\], or using pre\-collected spoken dialogue data\[[19](https://arxiv.org/html/2607.29125#bib.bib17),[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\.

### II\-AMulti\-turn Evaluation

Existing FDSDS benchmarks can be grouped by multi\-turn evaluation protocol, as shown in[TableI](https://arxiv.org/html/2607.29125#S1.T1)\. Single\-turn benchmarks evaluate local responses to a user utterance or an overlap event\. This setting allows fair comparison because all systems receive the same input, but it does not test how dialogue history affects system behavior\. Dynamic evaluation addresses this limitation by letting an automated examiner or a human user interact with the target system in real time\[[12](https://arxiv.org/html/2607.29125#bib.bib12),[3](https://arxiv.org/html/2607.29125#bib.bib9)\]\. This protocol can provide adaptive interaction, but the dialogue history can differ across systems\. As a result, fine\-grained comparison between models becomes difficult\.

Static multi\-turn evaluation uses dialogue histories from pre\-collected conversations\. This makes the input condition more comparable across systems\. In some studies, only the user\-side context was fixed, and the model responded freely\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. In this setting, the user’s later utterances may not match the system’s previous responses\. This context mismatch can make the dialogue history unnatural and can affect the validity of the evaluation\. MTR\-DuplexBench reduces this issue by using full\-context conditioning, where previous system turns are fixed to reference speech by teacher\-forced inference, and the model is evaluated at the current turn\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. However, this full\-context setting is applied only to dialogue quality evaluation\. Timing behaviors were evaluated in user\-context settings using discontinuous multi\-turn contexts created by concatenating single\-turn utterance pairs, assuming that timing behavior is independent of conversational context\. However, context mismatch may also affect timing behavior\.

In this work, we apply three context conditions, including single\-turn evaluation and two static evaluation conditions: user\- and full\-context conditioning\. We analyzed how dialogue history affects model behavior and discussed the validity and limitations of these evaluation protocols\.

### II\-BLanguage and Domain Coverage

Most existing benchmarks focus on English dialogues\. Recently, FDSDSs have also been developed for non\-English languages, such as Chinese\[[28](https://arxiv.org/html/2607.29125#bib.bib18)\]and Japanese\[[17](https://arxiv.org/html/2607.29125#bib.bib14)\]\. However, benchmarks for non\-English FDSDSs remain limited\. Yan et al\.\[[28](https://arxiv.org/html/2607.29125#bib.bib18)\]extended Full\-Duplex\-Bench to Chinese\. The ICASSP 2026 HumDial Challenge also includes Chinese and English data for evaluating human\-like spoken dialogue systems\[[29](https://arxiv.org/html/2607.29125#bib.bib16)\]\. However, to our knowledge, no existing benchmark supports automatic evaluation of Japanese FDSDSs\. Domain coverage is another limitation\. Although existing benchmarks evaluate several full\-duplex behaviors and sometimes include multiple task settings\[[12](https://arxiv.org/html/2607.29125#bib.bib12),[14](https://arxiv.org/html/2607.29125#bib.bib10),[9](https://arxiv.org/html/2607.29125#bib.bib13)\], cross\-domain comparison of timing\-related behavior has not been sufficiently studied\.

We address these limitations by covering both language and domain variation\. M3\-DuplexBench includes English and Japanese dialogues, covers casual conversation and multi\-turn question answering, and evaluates multiple aspects of full\-duplex behavior\.

![Refer to caption](https://arxiv.org/html/2607.29125v1/x1.png)

\(a\)Events included in spoken dialogues: Turn shift \(SHIFT\), long pause \(PAUSE\), backchanneling \(BC\), and barge\-in \(BARGE\_IN\)\.
![Refer to caption](https://arxiv.org/html/2607.29125v1/x2.png)

\(b\)Temporal regions used in evaluation\.

Figure 1:Overview of the proposed M3\-DuplexBench evaluation framework\. \(a\) Four interaction event types extracted from continuous spoken dialogues\. \(b\) Evaluation protocol under three context conditions \(None,User, andFull\)\.

## IIIM3\-DuplexBench: Data

M3\-DuplexBench extracts evaluation events from natural and synthetic spoken dialogue datasets\. This section describes the event definitions \([SectionIII\-A](https://arxiv.org/html/2607.29125#S3.SS1)\) and the datasets used in the benchmark \([SectionIII\-B](https://arxiv.org/html/2607.29125#S3.SS2)\)\.

### III\-AEvent Definition

We focus on four events: turn shift \(SHIFT\), long pause \(PAUSE\), backchannel \(BC\), and barge\-in \(BARGE\_IN\), as shown in[Fig\.1a](https://arxiv.org/html/2607.29125#S2.F1.sf1)\. TURN and BC are inter\-pausal units \(IPUs\) separated by silences longer than 0\.5 seconds\. IPUs with a duration of 0\.8 seconds or less are regarded as BCs, while all others are considered TURNs\. PAUSE denotes the interval between TURNs produced by the same speaker\. SHIFT denotes the interval between TURNs produced by different speakers\. For SHIFTs, overlaps of up to 0\.4 seconds are allowed\. Speaker changes with an overlap longer than 0\.4 seconds are regarded as BARGE\_INs\.

Let a two\-party dialogue beD=\(a\(0\),a\(1\),ℰ\(0\),ℰ\(1\)\)D=\(a^\{\(0\)\},a^\{\(1\)\},\\mathcal\{E\}^\{\(0\)\},\\mathcal\{E\}^\{\(1\)\}\), wherea\(c\)a^\{\(c\)\}is the waveform of channelc∈\{0,1\}c\\in\\\{0,1\\\}andℰ\(c\)=\{e0\(c\),…,eN\(c\)\(c\)\}\\mathcal\{E\}^\{\(c\)\}=\\\{e\_\{0\}^\{\(c\)\},\.\.\.,e\_\{N^\{\(c\)\}\}^\{\(c\)\}\\\}is its event sequence of lengthN\(c\)N^\{\(c\)\}\. An evente∈ℰ\(0\)∪ℰ\(1\)e\\in\\mathcal\{E\}^\{\(0\)\}\\cup\\mathcal\{E\}^\{\(1\)\}is written ase=\(ℓ,s,t\)e=\(\\ell,s,t\), whereℓ∈\{SHIFT,PAUSE,BC,BARGE\_IN\}\\ell\\in\\\{\\text\{SHIFT\},\\text\{PAUSE\},\\text\{BC\},\\text\{BARGE\\\_IN\}\\\}is the event type ands,ts,tare the event’s start and end times\. SHIFT events may havet<st<sbecause overlaps are allowed\. We describe the procedure for evaluating models usingDDin[SectionIV](https://arxiv.org/html/2607.29125#S4)\.

TABLE II:Data statistics of M3\-DuplexBench\. BARGE and dur denote BARGE\_IN and average SHIFT duration, respectively\.
### III\-BDataset

M3\-DuplexBench covers two domains, chat and task\-oriented dialogue, and two languages, English and Japanese\. We use natural spoken dialogue data for the chat domain and synthetic spoken dialogue data for the task\-oriented domain\.[TableII](https://arxiv.org/html/2607.29125#S3.T2)summarizes the data statistics\.

For the English chat domain, we use Candor\[[21](https://arxiv.org/html/2607.29125#bib.bib19)\], following MTR\-DuplexBench\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. For the Japanese chat domain, we use MagicData, a natural speech conversation dataset of approximately 10 hours\[[4](https://arxiv.org/html/2607.29125#bib.bib3)\]\. Each dialogue is split into segments of at most 120 seconds, following\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\.

For the task\-oriented domain, we use TopiOCQA, an English conversational QA dataset\[[2](https://arxiv.org/html/2607.29125#bib.bib20)\]\. Each conversation consists of an average of 13 turns\. Unlike prior work that constructs dialogues by concatenating single\-turn QA pairs\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\], we use multi\-turn QA dialogues, enabling us to evaluate whether models can answer questions based on dialogue context\. We convert these text\-based task dialogues into synthetic spoken dialogues using the data generation pipeline described below\. In this work, we created English and Japanese data based on TopiOCQA\. The Japanese data are created by translating TopiOCQA, allowing us to test the same questions and knowledge across English and Japanese models\.

#### III\-B1Synthetic dialogue generation pipeline

Synthetic dialogue generation has recently been explored for training FDSDSs, including text\-based dialogue synthesis\[[24](https://arxiv.org/html/2607.29125#bib.bib21)\]and synthetic spoken dialogue generation\[[26](https://arxiv.org/html/2607.29125#bib.bib22),[11](https://arxiv.org/html/2607.29125#bib.bib23),[17](https://arxiv.org/html/2607.29125#bib.bib14)\]\. Following these works, we construct synthetic spoken dialogues from text dialogues and reference spoken dialogues through four steps\. We use TopiOCQA as the source text dialogues, and Candor and MagicData as English and Japanese reference spoken dialogues, respectively\.

1. 1\.We analyze the reference spoken dialogues to get statistics of pauses, turn shifts, backchannels, and overlaps\.
2. 2\.We edit the text dialogues\. Written dialogues are converted into spoken\-style dialogues using a large language model \(LLM\) and translated into Japanese when needed\. We use Gemma 4 31B\[[8](https://arxiv.org/html/2607.29125#bib.bib1)\]\.
3. 3\.We synthesize speech for each utterance using a CosyVoice2\-based TTS model\[[7](https://arxiv.org/html/2607.29125#bib.bib24)\]\. For speaker conditioning, we randomly select reference speech from the reference dialogues\.
4. 4\.Finally, we place the synthesized utterances on a timeline\. PAUSE, SHIFT, and BC events are sampled from the statistics obtained from the reference dialogues\. This produces synthetic spoken dialogues with timing characteristics similar to those observed in natural dialogues\.

BARGE\_IN events are created separately because they rarely appear in ordinary spoken dialogue data\. Following Hu et al\.\[[10](https://arxiv.org/html/2607.29125#bib.bib8)\], we create a BARGE\_IN sample by cutting off the system utterance at a sampled interruption point and shifting the user utterance to start from that point\. This creates a controlled overlap where the user starts speaking while the system is still speaking\.

## IVM3\-DuplexBench: Evaluation Framework

### IV\-AInference

For each eventee, we define three regions: contextCC, pre\-event user’s turnPP, and evaluation windowWW, as shown in[Fig\.1b](https://arxiv.org/html/2607.29125#S2.F1.sf2)\. The evaluation windowWWis the region where we measure model behavior\. It starts atssand has a fixed lengthΔ\\Delta, i\.e\.,W=\[s,s\+Δ\]W=\[s,s\+\\Delta\]\. The Pre\-event user’s turnPPis the short segment before the event, i\.e\.,P=\[b,s\]P=\[b,s\]\. For SHIFT and PAUSE,bbis the start time of a turn immediately preceding or containing the event, respectively\. For BC and BARGE\_IN, the user is speaking during a system turn, and thus there are no pre\-event user’s turn, soP=∅P=\\emptysetand we setb=sb=s\. The contextCCis the dialogue history beforePP\. Given a context lengthLL,C=\[τ​\(L\),b\]C=\[\\tau\(L\),b\], whereτ​\(L\)\\tau\(L\)is the start time of the earliest turn overlapping the time window\[b−L,b\]\[b\-L,b\]\. In the experiment, we setΔ=10\\Delta=10seconds andL=120L=120seconds\. We consider three context conditions:

- •Noneprovides no dialogue context and uses only the current event input\. The model receives user\-side audio in\(P,W\)\(P,W\)\. This corresponds to a single\-turn evaluation setting in previous works\[[14](https://arxiv.org/html/2607.29125#bib.bib10),[13](https://arxiv.org/html/2607.29125#bib.bib11)\]\. BC and BARGE\_IN are not evaluated under this condition becausePPdoes not exist\.
- •Userprovides user\-side speech history as a context\. The model receives user\-side audio in\(C,P,W\)\(C,P,W\), while system\-side audio is not provided\. The model state before the target event is induced by the model’s own generated responses\.
- •Fullprovides both user\- and system\-side speech history through teacher\-forced inference\. The model state is forced using both speaker channels in\(C,P\)\(C,P\)before receiving user\-side audio inWW\.

For all conditions, the user\-side audio inWWis retained only during the target event and muted outside the event interval, so that model behavior is evaluated only with respect to the target event\.

After inference, we evaluate the speech generated by the model using word\-level timestamps\. We first use Whisper ASR\[[20](https://arxiv.org/html/2607.29125#bib.bib25),[25](https://arxiv.org/html/2607.29125#bib.bib2)\]to transcribe the generated speech, then apply Montreal Forced Aligner\[[15](https://arxiv.org/html/2607.29125#bib.bib26)\]to align the ASR transcript with the generated audio and obtain timestamps\.

TABLE III:Evaluation dimensions and metrics in M3\-DuplexBench\.
### IV\-BEvaluation metrics

Table[III](https://arxiv.org/html/2607.29125#S4.T3)summarizes the evaluation metrics used in M3\-DuplexBench\. We evaluate models based on two aspects: timing and content\.

#### IV\-B1Timing Metrics

For timing evaluation, each dimension corresponds to one event type:Smooth Turn Takingto SHIFT,Pause Handlingto PAUSE,User Backchannelto BC, andUser Barge\-into BARGE\_IN\. We follow Full\-Duplex\-Bench\[[14](https://arxiv.org/html/2607.29125#bib.bib10)\]and use the Takeover Rate \(TOR\) and latency forSmooth Turn TakingandPause Handling\. TOR is the fraction of samples in which the model takes the turn within the evaluation window\. Accounting for overlap, a takeover is triggered if system utterance begins after 0\.4 seconds prior to the evaluation window\. Latency is the time from the event onset to the onset of the takeover utterance111Our computation is not exactly the same as\[[14](https://arxiv.org/html/2607.29125#bib.bib10)\]\. For example, their definition can count speech in the pre\-event region as a takeover, which tends to increase TOR and reduce latency, sometimes yielding negative latency\.\. ForSmooth Turn Taking, a higher TOR and a lower latency are better, since the model should take the floor after the user turn\. ForPause Handling, a lower TOR and a higher latency are better, since the model should wait while the user keeps the floor\. ForUser BackchannelandUser Barge\-in, we follow Full\-Duplex\-Bench v1\.5\[[13](https://arxiv.org/html/2607.29125#bib.bib11)\]and use stop latency\. Stop latency is the time from the event onset to the point where the model stops speaking\. For backchannels, a higher stop latency is better because the model should continue speaking\. For barge\-ins, a lower stop latency is better because the model should stop quickly after the user interruption\.

#### IV\-B2Content metrics

For content evaluation, we evaluate model responses to SHIFT events\. Several studies have used LLMs to evaluate FDSDS responses\[[14](https://arxiv.org/html/2607.29125#bib.bib10),[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. We measure response relevance, context consistency, and QA accuracy on a 0–2 scale, using LLM\-as\-a\-Judge with OpenAI GPT\-5 nano\[[18](https://arxiv.org/html/2607.29125#bib.bib4)\]\. These scores are normalized to the range \[0, 1\] for reporting\. Response relevance judges whether the model response addresses the current user utterance\. Context consistency judges whether the response is consistent with the previous dialogue context\. QA accuracy judges whether the response matches the reference answer in task\-oriented samples\.

## VExperiments

### V\-AModels

We evaluate open\-source FDSDSs in English and Japanese\. We focus on Moshi\-based models\[[6](https://arxiv.org/html/2607.29125#bib.bib27)\]becauseFullsetting requires direct conditioning on parallel user and system speech streams\. Extending the benchmark to other model architectures, including cascaded systems such as Freeze\-Omni\[[27](https://arxiv.org/html/2607.29125#bib.bib29)\], is left for future work\. For English, we evaluate Moshi and PersonaPlex\[[22](https://arxiv.org/html/2607.29125#bib.bib28)\]\. For Japanese, we evaluate J\-Moshi\[[17](https://arxiv.org/html/2607.29125#bib.bib14)\]and LLM\-jp\-Moshi\[[1](https://arxiv.org/html/2607.29125#bib.bib15)\]\. Moshi is an end\-to\-end full\-duplex model based on a 7B text\-based LLM Helium, and a neural audio codec Mimi\. It models parallel streams of user speech tokens, system speech tokens, and system text tokens\.222Since Moshi is trained to start with greetings, it often overlaps with the user’s first utterance\. Therefore, inNoneandUsersettings, 5 seconds of silence is prepended to the user’s speech\.PersonaPlex is based on the Moshi architecture and supports speaker and role control through voice and text prompts\. J\-Moshi and LLM\-jp\-Moshi are Moshi\-based models fine\-tuned on different Japanese spoken dialogue datasets\.

TABLE IV:Results on the chat domain\. Values are reported as mean±\\pm95% confidence interval\.None= no dialogue context;User= user\-side context only;Full= user\- and system\-side context\.None&Fulldenotes the representative value calculated from theNoneandFullsettings\. The best representative values among models within the same language are highlighted in bold\.IDModelContextTimingContentSmoothPauseUser Back\-UserResponseContextTurn TakingHandlingchannelingBarge\-inRelevanceConsistencyTOR↑\\uparrowLatency \(s\)↓\\downarrowTOR↓\\downarrowLatency \(s\)↑\\uparrowStop Latency \(s\)↑\\uparrowStop Latency \(s\)↓\\downarrowScore↑\\uparrowScore↑\\uparrowEnglish chatC1MoshiNone0\.675±\\pm0\.0253\.117±\\pm0\.1250\.201±\\pm0\.0154\.717±\\pm0\.144––0\.781±\\pm0\.0140\.818±\\pm0\.014C2User0\.631±\\pm0\.0262\.533±\\pm0\.1280\.258±\\pm0\.0164\.742±\\pm0\.1601\.102±\\pm0\.1591\.349±\\pm0\.2140\.795±\\pm0\.0140\.868±\\pm0\.012C3Full0\.582±\\pm0\.0261\.315±\\pm0\.1510\.103±\\pm0\.0117\.199±\\pm0\.1512\.755±\\pm0\.2471\.959±\\pm0\.2300\.703±\\pm0\.0140\.887±\\pm0\.012C4None&Full0\.6292\.2160\.1525\.9582\.7551\.9590\.7420\.853C5PersonaPlexNone0\.950±\\pm0\.0120\.759±\\pm0\.0590\.330±\\pm0\.0183\.432±\\pm0\.139––0\.747±\\pm0\.0140\.731±\\pm0\.015C6User0\.898±\\pm0\.0161\.264±\\pm0\.1060\.335±\\pm0\.0183\.109±\\pm0\.1404\.177±\\pm0\.3512\.192±\\pm0\.3430\.738±\\pm0\.0140\.794±\\pm0\.014C7Full0\.944±\\pm0\.0120\.771±\\pm0\.0560\.282±\\pm0\.0173\.659±\\pm0\.1444\.574±\\pm0\.2732\.097±\\pm0\.2570\.681±\\pm0\.0140\.882±\\pm0\.012C8None&Full0\.9470\.7650\.3063\.5464\.5742\.0970\.7140\.807Japanese chatC9J\-MoshiNone0\.774±\\pm0\.0191\.637±\\pm0\.1170\.312±\\pm0\.0224\.788±\\pm0\.199––0\.805±\\pm0\.0130\.893±\\pm0\.010C10User0\.323±\\pm0\.0211\.622±\\pm0\.1830\.231±\\pm0\.0206\.865±\\pm0\.2222\.919±\\pm0\.5252\.550±\\pm0\.3720\.751±\\pm0\.0160\.912±\\pm0\.009C11Full0\.516±\\pm0\.0231\.322±\\pm0\.1410\.428±\\pm0\.0234\.447±\\pm0\.2332\.575±\\pm0\.2341\.656±\\pm0\.2070\.743±\\pm0\.0140\.888±\\pm0\.010C12None&Full0\.6451\.4800\.3704\.6182\.5751\.6560\.7740\.891C13LLM\-jp\-MoshiNone0\.945±\\pm0\.0102\.566±\\pm0\.1140\.299±\\pm0\.0223\.274±\\pm0\.140––0\.717±\\pm0\.0120\.808±\\pm0\.012C14User0\.868±\\pm0\.0151\.974±\\pm0\.1150\.643±\\pm0\.0231\.101±\\pm0\.1663\.496±\\pm0\.1922\.599±\\pm0\.2130\.678±\\pm0\.0130\.796±\\pm0\.012C15Full0\.804±\\pm0\.0181\.568±\\pm0\.1070\.407±\\pm0\.0233\.292±\\pm0\.1842\.503±\\pm0\.1721\.779±\\pm0\.1520\.709±\\pm0\.0130\.866±\\pm0\.011C16None&Full0\.8752\.0670\.3533\.2832\.5031\.7790\.7130\.837![Refer to caption](https://arxiv.org/html/2607.29125v1/x3.png)Figure 2:Case study showing the effect of dialogue context in task\-oriented QA\. Responses are obtained by PersonaPlex under the same context\-dependent question using three context conditions\. The gray region denotes the context region\. Without dialogue history \(None\), the model failed to answer the context\-dependent question\.Usercontext led to an inconsistent dialogue history and an incorrect answer, whereasfullcontext preserved the dialogue consistency and produced a correct answer\.![Refer to caption](https://arxiv.org/html/2607.29125v1/x4.png)

\(a\)English examples\.
![Refer to caption](https://arxiv.org/html/2607.29125v1/x5.png)

\(b\)Japanese examples\.

Figure 3:Case studies of model responses in multi\-turn QA\. In these examples, English models generated responses relevant to the current question, whereas Japanese models produced less relevant responses, showing the cross\-lingual gap in context\-dependent answer generation\.
### V\-BResults

Tables[IV](https://arxiv.org/html/2607.29125#S5.T4)and[V](https://arxiv.org/html/2607.29125#S5.T5)show the results on the chat and task\-oriented domains\.

#### V\-B1Effect of Context

Conditioning on theFullcontext tended to improve timing performance, especially smooth turn\-taking latency\. In both domains, most models showed lower smooth turn\-taking latency inFullthan inNone\. For example, in the task\-oriented domain, Moshi reduced latency from 2\.453 seconds to 0\.430 seconds \(T1 vs\. T3\), and J\-Moshi reduced it from 0\.933 to 0\.326 seconds \(T9 vs\. T11\)\. As shown in[TableII](https://arxiv.org/html/2607.29125#S3.T2), the average SHIFT duration \(i\.e\., smooth turn\-taking latency\) in the evaluation data was 0\.88 seconds, 0\.72 seconds, 0\.45 seconds, and 0\.35 seconds for English chat, Japanese chat, English task\-oriented dialogue, and Japanese task\-oriented dialogue, respectively\. We found that the latency inFullwas often closer to these values than inNone\. This suggests that theFDSDSs can adjust their response timing to match the patterns in the dialogue context\. The effect ofFullcontext on smooth turn\-taking TOR was mixed, and it often degraded performance in the chat domain \(e\.g\., C1 vs\. C3\)\. Chat conversations contain many ambiguous turn transitions, optional responses, and diverse interaction patterns\. This may have caused a gap with the model’s natural generation trajectories\. TheFullcontext also improved the content metrics of English models\. For example, in the task\-oriented domain, PersonaPlex improved QA accuracy from 0\.180 inNoneto 0\.306 inFull\(T5 vs\. T7\)\. This suggests thatdialogue history helped models understand the current question and generate more accurate answers in multi\-turn QA\.[Fig\.2](https://arxiv.org/html/2607.29125#S5.F2)shows an example of comparing the three contextual conditions in the task\-oriented domain, highlighting that thefullcontext preserved dialogue consistency and improved response quality\. In contrast, Japanese models showed only limited content improvements underFull\(e\.g\., C9 vs\. C11\), indicating that providing context did not necessarily lead to accurate answer generation\.

TABLE V:Results on the task\-oriented domain\. Metrics, notations, and reporting format are identical to those in[TableIV](https://arxiv.org/html/2607.29125#S5.T4)\.Conditioning only on user context \(User\) was unstable in many cases\. Across domains and models, smooth turn\-taking TOR generally decreased, and latency showed inconsistent trends\. Prior work evaluated multi\-turn timing underUsercontext conditioning and reported that timing metrics degraded as the number of turns increased\[[9](https://arxiv.org/html/2607.29125#bib.bib13)\]\. However, our results showed that theFullsetting often outperformedNonein timing metrics\. This suggests thatthe degradation observed underUsercontext conditioning may not only reflect the inherent difficulty of multi\-turn interaction, but may also be caused by context mismatch between the fixed user\-side history and the model’s own generated history\. TheFullsetting also has limitations: it requires direct conditioning on parallel user and system speech streams, which limits the range of applicable models, and may introduce a gap between the teacher\-forced trajectory and the model’s natural inference trajectory\. Nevertheless, given the instability of theUsersetting, we considerFullto be a more reliable condition for static multi\-turn evaluation\. In the following model, domain, and language comparisons, we therefore used the average of theNoneandFullsettings as the main representative value, excluding theUsersetting\. For metrics unavailable in theNonesetting, we used the values from theFullsetting as representative values\.

#### V\-B2Model Comparison

PersonaPlex was the most responsive model\. It achieved a smooth turn\-taking TOR of 0\.947 with a latency of 0\.765 seconds in the chat domain \(C8\)\. However, in the chat domain, its pause\-handling TOR was 0\.306, which was higher than Moshi’s 0\.152 \(C8 vs\. C4\)\. This indicates that PersonaPlex responded quickly, but was also more likely to take the turn during user\-held pauses\.

The stop latency of barge\-in was generally higher than that of back\-channeling, suggesting that the models’ ability to distinguish user backchannels and user barge\-ins to some extent\. However, barge\-in stop latency still remained around 1\.7–2\.6 seconds for several models, indicating that real\-time interruption handling remains challenging\.

Among the Japanese models, LLM\-jp\-Moshi showed better timing performance\. In the chat domain, its smooth turn\-taking TOR was 0\.875, outperforming J\-Moshi’s 0\.645\. In contrast, J\-Moshi achieved better content scores in both domains\.

#### V\-B3Cross\-domain Difference

Across domains,smooth turn\-taking was easier in the task\-oriented domain than in the chat domain\.Most models achieved higher TOR and lower latency in task\-oriented dialogue\. For example, Moshi’s smooth turn\-taking TOR increased from 0\.629 in the chat domain to 0\.916 in the task\-oriented domain \(C4 vs\. T4\)\. This likely reflects the structure of task\-oriented QA, where user utterances are often explicit questions and the expected timing of system responses is clearer\. In contrast, chat contains more ambiguous turn transitions, optional responses, and diverse response candidates, making it more difficult for models to decide when to speak\.

#### V\-B4Cross\-lingual Difference

The cross\-lingual gap was clear in the content metrics for the task\-oriented domain\. English models achieved higher response relevance and QA accuracy than Japanese models\. For example, PersonaPlex obtained 0\.848 in response relevance and 0\.243 in QA accuracy, while J\-Moshi obtained 0\.557 and 0\.136 \(T8 vs\. T12\)\. Since the English and Japanese task\-oriented evaluation sets were generated from the same QA data, these results suggest that Japanese models still lag behind English models in language understanding and context\-dependent answer generation\.[Fig\.3](https://arxiv.org/html/2607.29125#S5.F3)shows examples of English and Japanese model responses to the same question\.

On the other hand, Japanese models showed competitive timing performance\. For example, J\-Moshi showed smaller latencies of smooth turn\-taking in both domains compared to Moshi \(e\.g\., T12 vs\.T4\)\. These results suggest thatthe main limitation of current Japanese full\-duplex models lies in content understanding and generation, rather than in timing behavior\.

## VIConclusion

In this work, we present M3\-DuplexBench, a novel benchmark for FDSDSs\. The benchmark covers English and Japanese, chat and task\-oriented dialogue, and evaluates both timing and content through event\-level samples extracted from continuous spoken conversations\. By comparing multiple context conditions, M3\-DuplexBench enables controlled analysis of how dialogue history affects FDSDSs\. Our experiments show that dialogue context greatly affects both timing and content evaluation\. In particular, user\-only conditioning can introduce context mismatch, whereas full\-context conditioning provides a more reliable protocol for controlled multi\-turn evaluation\. Our future work will extend this benchmark to include more languages, domains and models\.

## VIIGenerative AI Use Disclosure

This manuscript was edited and polished with the assistance of generative AI\. All experimental design, implementation, and analysis were conducted by the authors who take full responsibility for the content\.

## References

- \[1\]Y\. Abe, M\. Saeki, A\. Ohashi, S\. Takamichi, S\. Fujie, T\. Kobayashi, T\. Ogawa, and R\. Higashinaka\(2026\)Effects of dialogue corpora properties on fine\-tuning a moshi\-based spoken dialogue model\.In16th International Workshop on Spoken Dialogue System Technology \(IWSDS 2026\),pp\. 104–108\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p4.1),[§V\-A](https://arxiv.org/html/2607.29125#S5.SS1.p1.1)\.
- \[2\]V\. Adlakha, S\. Dhuliawala, K\. Suleman, H\. de Vries, and S\. Reddy\(2022\)TopiOCQA: open\-domain conversational question answering with topic switching\.Transactions of the Association for Computational Linguistics10,pp\. 468–483\.Cited by:[§III\-B](https://arxiv.org/html/2607.29125#S3.SS2.p3.1)\.
- \[3\]S\. Arora, Z\. Lu, C\. Chiu, R\. Pang, and S\. Watanabe\(2025\)Talking turns: benchmarking audio foundation models on turn\-taking dynamics\.In13th International Conference on Learning Representations \(ICLR 2025\),Cited by:[TABLE I](https://arxiv.org/html/2607.29125#S1.T1.3.3.2.3.1),[§I](https://arxiv.org/html/2607.29125#S1.p2.1),[§II\-A](https://arxiv.org/html/2607.29125#S2.SS1.p1.1),[§II](https://arxiv.org/html/2607.29125#S2.p1.1)\.
- \[4\]Beijing Magic Data Technology Co\., Ltd\.\(2025\)Japanese duplex conversation training dataset\.Note:https://magichub\.com/datasets/japanese\-duplex\-conversation\-training\-dataset/Cited by:[§III\-B](https://arxiv.org/html/2607.29125#S3.SS2.p2.1)\.
- \[5\]Y\. Chen and H\. Yu\(2025\)From turn\-taking to synchronous dialogue: a survey of full\-duplex spoken language models\.arXiv preprint arXiv:arXiv:2509\.14515\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p1.1)\.
- \[6\]A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour\(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.arXiv preprint arXiv:2410\.00037\.Cited by:[§V\-A](https://arxiv.org/html/2607.29125#S5.SS1.p1.1)\.
- \[7\]Z\. Du, Y\. Wang, Q\. Chen, X\. Shi, X\. Lv, T\. Zhao, Z\. Gao, Y\. Yang, C\. Gao, H\. Wang,et al\.\(2024\)Cosyvoice 2: scalable streaming speech synthesis with large language models\.arXiv preprint arXiv:2412\.10117\.Cited by:[item 3](https://arxiv.org/html/2607.29125#S3.I1.i3.p1.1)\.
- \[8\]Google DeepMind\(2026\)Gemma 4 31b\.Note:https://huggingface\.co/google/gemma\-4\-31BCited by:[item 2](https://arxiv.org/html/2607.29125#S3.I1.i2.p1.1)\.
- \[9\]Z\. He, W\. Cui, H\. Xu, X\. Li, L\. Zhu, H\. Bai, M\. Shaohua, and I\. King\(2026\)MTR\-DuplexBench: towards a comprehensive evaluation of multi\-round conversations for full\-duplex speech language models\.InFindings of the Association for Computational Linguistics \(ACL 2026\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),pp\. 5334–5351\.Cited by:[TABLE I](https://arxiv.org/html/2607.29125#S1.T1.3.3.2.5),[§I](https://arxiv.org/html/2607.29125#S1.p2.1),[§I](https://arxiv.org/html/2607.29125#S1.p3.1),[§I](https://arxiv.org/html/2607.29125#S1.p4.1),[§II\-A](https://arxiv.org/html/2607.29125#S2.SS1.p2.1),[§II\-B](https://arxiv.org/html/2607.29125#S2.SS2.p1.1),[§II](https://arxiv.org/html/2607.29125#S2.p1.1),[§III\-B](https://arxiv.org/html/2607.29125#S3.SS2.p2.1),[§III\-B](https://arxiv.org/html/2607.29125#S3.SS2.p3.1),[§IV\-B2](https://arxiv.org/html/2607.29125#S4.SS2.SSS2.p1.1),[§V\-B1](https://arxiv.org/html/2607.29125#S5.SS2.SSS1.p2.1)\.
- \[10\]K\. Hu, E\. Hosseini\-Asl, C\. Chen, E\. Casanova, S\. Ghosh, P\. Żelasko, Z\. Chen, J\. Li, J\. Balam, and B\. Ginsburg\(2025\)Efficient and Direct Duplex Modeling for Speech\-to\-Speech Language Model\.In26th Annual Conference of the International Speech Communication Association \(INTERSPEECH 2025\),pp\. 2715–2719\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p1.1),[§III\-B1](https://arxiv.org/html/2607.29125#S3.SS2.SSS1.p2.1)\.
- \[11\]S\. Lee, K\. Kim, and G\. Kim\(2025\)Behavior\-SD: behaviorally aware spoken dialogue generation with large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL 2025\),pp\. 9574–9593\.Cited by:[§III\-B1](https://arxiv.org/html/2607.29125#S3.SS2.SSS1.p1.1)\.
- \[12\]G\. Lin, S\. S\. Kuan, J\. Shi, K\. Chang, S\. Arora, S\. Watanabe, and H\. Lee\(2026\)Full\-duplex\-bench\-v2: a multi\-turn evaluation framework for duplex dialogue systems with an automated examiner\.In64th Annual Meeting of the Association for Computational Linguistics \(ACL 2026\),pp\. 27–36\.Cited by:[TABLE I](https://arxiv.org/html/2607.29125#S1.T1.3.3.2.2),[§I](https://arxiv.org/html/2607.29125#S1.p2.1),[§I](https://arxiv.org/html/2607.29125#S1.p4.1),[§II\-A](https://arxiv.org/html/2607.29125#S2.SS1.p1.1),[§II\-B](https://arxiv.org/html/2607.29125#S2.SS2.p1.1),[§II](https://arxiv.org/html/2607.29125#S2.p1.1)\.
- \[13\]G\. Lin, S\. S\. Kuan, Q\. Wang, J\. Lian, T\. Li, S\. Watanabe, and H\. Lee\(2026\)Full\-duplex\-bench v1\.5: evaluating overlap handling for full\-duplex speech models\.In2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP 2026\),pp\. 19447–19451\.Cited by:[TABLE I](https://arxiv.org/html/2607.29125#S1.T1.3.3.2.2),[§I](https://arxiv.org/html/2607.29125#S1.p2.1),[§II](https://arxiv.org/html/2607.29125#S2.p1.1),[1st item](https://arxiv.org/html/2607.29125#S4.I1.i1.p1.2),[§IV\-B1](https://arxiv.org/html/2607.29125#S4.SS2.SSS1.p1.1)\.
- \[14\]G\. Lin, J\. Lian, T\. Li, Q\. Wang, G\. Anumanchipalli, A\. H\. Liu, and H\. Lee\(2025\)Full\-duplex\-bench: a benchmark to evaluate full\-duplex spoken dialogue models on turn\-taking capabilities\.In2025 IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\),pp\. 1–8\.Cited by:[TABLE I](https://arxiv.org/html/2607.29125#S1.T1.3.3.2.2),[§I](https://arxiv.org/html/2607.29125#S1.p2.1),[§I](https://arxiv.org/html/2607.29125#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29125#S2.SS2.p1.1),[§II](https://arxiv.org/html/2607.29125#S2.p1.1),[1st item](https://arxiv.org/html/2607.29125#S4.I1.i1.p1.2),[§IV\-B1](https://arxiv.org/html/2607.29125#S4.SS2.SSS1.p1.1),[§IV\-B2](https://arxiv.org/html/2607.29125#S4.SS2.SSS2.p1.1),[footnote 1](https://arxiv.org/html/2607.29125#footnote1)\.
- \[15\]M\. McAuliffe, M\. Socolof, S\. Mihuc, M\. Wagner, and M\. Sonderegger\(2017\)Montreal Forced Aligner: Trainable Text\-Speech Alignment Using Kaldi\.InProceedings of the 18th Annual Conference of the International Speech Communication Association \(INTERSPEECH 2017\),pp\. 498–502\.Cited by:[§IV\-A](https://arxiv.org/html/2607.29125#S4.SS1.p2.1)\.
- \[16\]T\. A\. Nguyen, E\. Kharitonov, J\. Copet, Y\. Adi, W\. Hsu, A\. Elkahky, P\. Tomasello, R\. Algayres, B\. Sagot, A\. Mohamed,et al\.\(2023\)Generative spoken dialogue language modeling\.Transactions of the Association for Computational Linguistics11,pp\. 250–266\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p1.1)\.
- \[17\]A\. Ohashi, S\. Iizuka, J\. Jiang, and R\. Higashinaka\(2025\)Towards a japanese full\-duplex spoken dialogue system\.In26th Annual Conference of the International Speech Communication Association \(INTERSPEECH 2025\),pp\. 1783–1787\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29125#S2.SS2.p1.1),[§III\-B1](https://arxiv.org/html/2607.29125#S3.SS2.SSS1.p1.1),[§V\-A](https://arxiv.org/html/2607.29125#S5.SS1.p1.1)\.
- \[18\]OpenAI\(2026\)GPT\-5 nano Model\.Note:https://developers\.openai\.com/api/docs/models/gpt\-5\-nanoAccessed: 2026\-06\-15Cited by:[§IV\-B2](https://arxiv.org/html/2607.29125#S4.SS2.SSS2.p1.1)\.
- \[19\]Y\. Peng, Y\. Chao, D\. Ng, Y\. Ma, C\. Ni, B\. Ma, and E\. S\. Chng\(2025\)FD\-Bench: A Full\-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems\.In26th Annual Conference of the International Speech Communication Association \(INTERSPEECH 2025\),pp\. 176–180\.Cited by:[TABLE I](https://arxiv.org/html/2607.29125#S1.T1.3.3.2.4),[§II](https://arxiv.org/html/2607.29125#S2.p1.1)\.
- \[20\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever\(2022\)Robust speech recognition via large\-scale weak supervision\.arXiv preprint arXiv:2212\.04356\.Cited by:[§IV\-A](https://arxiv.org/html/2607.29125#S4.SS1.p2.1)\.
- \[21\]A\. Reece, G\. Cooney, P\. Bull, C\. Chung, B\. Dawson, C\. Fitzpatrick, T\. Glazer, D\. Knox, A\. Liebscher, and S\. Marin\(2023\)The candor corpus: insights from a large multimodal dataset of naturalistic conversation\.Science advances9\.Cited by:[§III\-B](https://arxiv.org/html/2607.29125#S3.SS2.p2.1)\.
- \[22\]R\. Roy, J\. Raiman, S\. Lee, T\. Ene, R\. Kirby, S\. Kim, J\. Kim, and B\. Catanzaro\(2026\)PersonaPlex: voice and role control for full duplex conversational speech models\.arXiv preprint arXiv:2602\.06053\.Cited by:[§V\-A](https://arxiv.org/html/2607.29125#S5.SS1.p1.1)\.
- \[23\]G\. Skantze\(2021\)Turn\-taking in conversational systems and human\-robot interaction: a review\.Computer Speech & Language67\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p1.1)\.
- \[24\]S\. K\. Suresh, W\. Mengjun, T\. Pranav, and E\. S\. Chng\(2025\)DiaSynth: synthetic dialogue generation framework for low resource dialogue applications\.InFindings of the Association for Computational Linguistics \(NAACL 2025\),pp\. 673–690\.Cited by:[§III\-B1](https://arxiv.org/html/2607.29125#S3.SS2.SSS1.p1.1)\.
- \[25\]SYSTRAN\(2023\)Faster\-whisper: faster whisper transcription with ctranslate2\.Note:https://github\.com/SYSTRAN/faster\-whisperCited by:[§IV\-A](https://arxiv.org/html/2607.29125#S4.SS1.p2.1)\.
- \[26\]M\. Wang, Y\. Bai, Y\. Wang, T\. Vu, E\. Shareghi, and G\. Haffari\(2025\)SpeechDialogueFactory: A Framework for Natural Speech Dialogue Generation\.In26th Annual Conference of the International Speech Communication Association \(INTERSPEECH 2025\),pp\. 1758–1762\.Cited by:[§III\-B1](https://arxiv.org/html/2607.29125#S3.SS2.SSS1.p1.1)\.
- \[27\]X\. Wang, Y\. Li, C\. Fu, Y\. Shen, L\. Xie, K\. Li, X\. Sun, and L\. Ma\(2024\)Freeze\-omni: a smart and low latency speech\-to\-speech dialogue model with frozen LLM\.arXiv preprint arXiv:2411\.00774\.Cited by:[§V\-A](https://arxiv.org/html/2607.29125#S5.SS1.p1.1)\.
- \[28\]R\. Yan, W\. Chen, Z\. Liu, Z\. Ma, H\. Lin, H\. Wen, H\. Xie, J\. Wu, Y\. Liang, Y\. Zhao, P\. Feng, J\. Qian, H\. Meng, Y\. Dai, S\. Yin, M\. Tao, L\. Xie, K\. Yu, X\. Wang, and X\. Chen\(2026\)SoulX\-duplug: plug\-and\-play streaming state prediction module for realtime full\-duplex speech conversation\.arXiv preprint arXiv:arXiv:2603\.14877\.Cited by:[§II\-B](https://arxiv.org/html/2607.29125#S2.SS2.p1.1)\.
- \[29\]Z\. Zhao, S\. Wang, G\. Li, H\. Xue, C\. Wang, S\. Wang, L\. Xiao, Z\. Zhang, H\. Bu, X\. Xu,et al\.\(2026\)The ICASSP 2026 HumDial challenge: benchmarking human\-like spoken dialogue systems in the LLM era\.In2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP 2026\),pp\. 21850–21852\.Cited by:[§I](https://arxiv.org/html/2607.29125#S1.p4.1),[§II\-B](https://arxiv.org/html/2607.29125#S2.SS2.p1.1)\.

Similar Articles

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

arXiv cs.CL

Introduces Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, evaluating 22 frontier models across six controlled modes (constraint memory, precise execution, constraint synthesis, object localization, action suppression, reference resolution) with 209 tasks spanning 12-76 turns. Even the strongest model, GPT-5.5, satisfies only 41.1% of responses.

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hugging Face Daily Papers

SpeechEditBench is a bilingual multi-attribute benchmark for evaluating instruction-guided speech editing across seven atomic tasks and compositional tasks, using an anchor-based evaluation protocol with three metrics. Evaluation of mainstream Speech LLMs reveals no single model excels across all dimensions, and compositional editing remains highly challenging.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

Hugging Face Daily Papers

This paper introduces Omni-DuplexEval, a benchmark and automatic evaluation framework for real-time duplex interaction in multimodal large language models, assessing continuous response generation and proactive event detection in streaming scenarios.