Full-Duplex Speech Models Take the Floor When Asked, Not When Needed

arXiv cs.CL Papers

Summary

This paper evaluates full-duplex speech models' ability to decide when to speak, finding that models like Moshi and PersonaPlex primarily respond to being addressed or silence rather than content-driven triggers such as false claims or hazards, identifying a gap in content understanding.

arXiv:2609.19596v1 Announce Type: new Abstract: Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2\,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.
Original Article
View Cached Full Text

Cached at: 09/18/26, 08:58 AM

# Full-Duplex Speech Models Take the Floor When Asked, Not When Needed
Source: [https://arxiv.org/html/2609.19596](https://arxiv.org/html/2609.19596)
###### Abstract

Full\-duplex speech models listen and speak at once, promising always\-on assistants\. Yet they must also decide when they should speak\. Human listeners speak when addressed or when the speaker stops, but also self\-select to correct a false claim, supply a missing word, or warn of danger\. We ask whether full\-duplex models do the same\. To separate the reason to speak from the opportunity, we construct context\-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn\-allocation rules, and compress inter\-word pauses to limit opportunities created by silence\. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards\. Frame\-level text\-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2 s after trigger end\. Pauses or permission to interrupt do not close this gap either\. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non\-empty false\-fact replies that challenge the claim is only \.14\-\-\.15, and the proportion of hazard replies that warn of danger is \.04\-\-\.07\. This paper thus identifies a gap in both speech initiation and response content\. Closing it requires genuine content understanding and intervention decisions grounded in it111Stimuli, model outputs, and evaluation code are available at[https://github\.com/vocaliodmiku/take\-the\-floor](https://github.com/vocaliodmiku/take-the-floor)\.

###### Index Terms:

full\-duplex spoken dialogue, turn\-taking, self\-selection, content\-driven intervention

††address:1University of Connecticut2The University of Texas at Austin3X Square Robot
4The University of Oxford
linkai\.peng@uconn\.edu## 1Introduction

Full\-duplex speech models such as Moshi\[[4](https://arxiv.org/html/2609.19596#bib.bib8)\], PersonaPlex\[[17](https://arxiv.org/html/2609.19596#bib.bib9)\]and VoiceChat\[[14](https://arxiv.org/html/2609.19596#bib.bib10)\]allow the user and the system to listen and speak simultaneously\. A system must decide when to*yield the floor*as the user starts speaking and when to*take the floor*itself\. In some scenarios, the user’s words call for a response before an explicit request or a pause\. A tutor, for example, may need to correct a misconception while a student is still explaining their reasoning\. Waiting until the student asks for help could let the error shape the rest of the explanation\. An interview coach may similarly help when a learner stalls on a word\. A warning can be more urgent, since the assistant needs to speak before the user carries out an unsafe action\.

Existing evaluations largely focus on turn\-taking and user\-initiated events\. FD\-Bench\[[16](https://arxiv.org/html/2609.19596#bib.bib5)\]and HumDial\[[24](https://arxiv.org/html/2609.19596#bib.bib6)\]measure pause handling, user interruption and multi\-turn continuity\. Model\-initiated speech is partially addressed in FLEXI’s emergency interruptions\[[7](https://arxiv.org/html/2609.19596#bib.bib7)\]and in Instruct\-FD\[[22](https://arxiv.org/html/2609.19596#bib.bib4)\], which measures whether a model follows an interruption instruction\. Our question is whether a model speaks up on its own, with no interruption instruction, and whether content or a pause drives that decision\. We examine this distinction by systematically controlling pauses in matched contexts\.

Figure 1:One trigger example from each condition group in Table[1](https://arxiv.org/html/2609.19596#S3.T1)shows the user’s speech \(gray, with the trigger shaded in the group color\) and the model’s reply \(colored\), witht0t\_\{0\}marking the end of the trigger\. Triggers are abridged from the*sleep*stimuli\. Model replies are illustrative\.To this end, we draw on the turn\-allocation rules of Sacks, Schegloff and Jefferson\[[18](https://arxiv.org/html/2609.19596#bib.bib1)\]\. The current speaker may select the next\. If not, another party may self\-select\. If neither happens, the current speaker may continue\. We distinguish*being addressed*from*self\-selection*, and use*yield*conditions to test verbal turn endings and silence\. Self\-selection allows listeners to offer a correction, warning, or missing word without a direct request\. It can occur during a word search\[[8](https://arxiv.org/html/2609.19596#bib.bib19),[10](https://arxiv.org/html/2609.19596#bib.bib20)\]or after a turn ends\[[18](https://arxiv.org/html/2609.19596#bib.bib1)\]\. Other\-correction occurs despite being dispreferred\[[19](https://arxiv.org/html/2609.19596#bib.bib2)\], and a warning may need to precede an imminent action\. Here,*content\-driven intervention*means responding to a potential problem without being asked to answer\. Word search may itself invite entry, whereas false facts and hazards require judging whether intervention is warranted\. These conditions separate a reason to speak from the opportunity offered by a pause\. Additional instructions test explicit permission to interrupt\.

We compare matched stimuli with controlled pauses across five model families and seven configurations \(Table[1](https://arxiv.org/html/2609.19596#S3.T1), Fig\.[1](https://arxiv.org/html/2609.19596#S1.F1)\)\. The experiments yield three main findings\. First, false\-fact and hazard onset rates stay close to Neutral, while direct questions and silence produce larger increases\. Frame\-level text\-token probabilities further show lower activation means for false facts\. Second, inserting silence or retaining natural pauses raises overall onset, but the difference between content cues and Neutral depends on the model and pause setting\. With inserted silence, false\-fact and hazard onset exceeds Neutral in Moshi but remains below Neutral in PPlex\. Explicit interruption instructions change little again\. Third, handing the floor to Moshi and PersonaPlex makes them speak more frequently, but does not ensure a useful intervention\. They address most direct questions but rarely correct false claims or warn of hazards\.

## 2Related work

Full\-duplex models and turn\-taking evaluation\.Full\-duplex dialogue models include dGSLM\[[13](https://arxiv.org/html/2609.19596#bib.bib11)\], Moshi\[[4](https://arxiv.org/html/2609.19596#bib.bib8)\], SyncLLM and Freeze\-Omni\[[23](https://arxiv.org/html/2609.19596#bib.bib12),[25](https://arxiv.org/html/2609.19596#bib.bib13)\]\. Moshi uses parallel audio streams with an inner text monologue\. PersonaPlex adds role control, and VoiceChat adds tool calling\[[17](https://arxiv.org/html/2609.19596#bib.bib9),[14](https://arxiv.org/html/2609.19596#bib.bib10)\]\. Recent post\-training work targets pause handling, turn\-taking, backchanneling and interruption\[[15](https://arxiv.org/html/2609.19596#bib.bib23)\]\. Evaluation has centered on the timing of when models speak and how they handle user\-initiated events\. Full\-Duplex\-Bench, Talking Turns, FD\-Bench and HumDial measure pause handling, backchannels, turn switches, user interruption and multi\-turn continuity\[[12](https://arxiv.org/html/2609.19596#bib.bib14),[1](https://arxiv.org/html/2609.19596#bib.bib15),[16](https://arxiv.org/html/2609.19596#bib.bib5),[24](https://arxiv.org/html/2609.19596#bib.bib6)\]\. FLEXI\[[7](https://arxiv.org/html/2609.19596#bib.bib7)\]and Instruct\-FD\[[22](https://arxiv.org/html/2609.19596#bib.bib4)\]also evaluate model\-initiated intervention, but neither compares different triggers in the same context while controlling pauses\. Our experiments use this comparison to test whether false facts, hazards and other cues affect speech onset\.

Turn allocation and content\-driven intervention\.Turn\-timing studies show that listeners project turn ends and plan responses\[[21](https://arxiv.org/html/2609.19596#bib.bib16),[11](https://arxiv.org/html/2609.19596#bib.bib3)\], and TurnGPT and VAP predict speaking opportunities\[[5](https://arxiv.org/html/2609.19596#bib.bib17),[6](https://arxiv.org/html/2609.19596#bib.bib18)\]\. The function of a response is a separate question\. Word searches elicit candidate words or collaborative completions\[[8](https://arxiv.org/html/2609.19596#bib.bib19),[10](https://arxiv.org/html/2609.19596#bib.bib20)\], completion markers invite self\-selection\[[18](https://arxiv.org/html/2609.19596#bib.bib1)\], backchannels let a listener stay engaged without taking the floor\[[26](https://arxiv.org/html/2609.19596#bib.bib21),[20](https://arxiv.org/html/2609.19596#bib.bib22)\], and other\-correction is dispreferred relative to self\-correction\[[19](https://arxiv.org/html/2609.19596#bib.bib2)\]\. These distinctions motivate assessing whether a reply provides the help required by the trigger, in addition to measuring whether the model takes the floor\.

## 3Method

### 3\.1Stimuli and conditions

We wrote 40 first\-person English monologues about everyday topics\. Each has a lead\-in, a*trigger utterance*and a continuation, with a median total length of 50\.3 seconds\. Within one topic, only the trigger utterance changes\. The rest of the text is identical\. After the trigger, the user keeps talking, except for explicitly inserted silence, and the model must decide whether to take the floor\. Two TTS voices each cover a fixed set of 20 topics\. In the main experiment, we use word\-level timestamps to compress every inter\-word gap to at most 0\.12 s, limiting pause\-based opportunities to speak\. A natural\-pause experiment retains the original gaps, with a typical sentence\-final pause of about 0\.88 s\. Table[1](https://arxiv.org/html/2609.19596#S3.T1)lists nine trigger conditions\. Neutral provides the tenth condition and baseline\. Section[4\.3](https://arxiv.org/html/2609.19596#S4.SS3)adds extra controls with pauses and cue\-specific opening instructions such as “interrupt me if I’m wrong”\. Since the user keeps talking after a question, the addressed conditions test whether an explicit request registers during ongoing speech, not whether a listener should answer mid\-turn\.

Table 1:Trigger conditions, grouped by the reason to speak\.Addressed\(the speaker selects the model\)Questiondirect question to the modelRequestthe same request as a statementRhetoricalquestion form, not addressed to the modelSelf\-select\(the listener must decide\)Word searchcannot recall a wordFalse factfalse claim attributed to an authorityRepeatprevious sentence repeated verbatimHazardan imminent dangerous actionYield\(the speaker yields the floor\)Turn end“that’s all from me”, then continues without pauseSilence1\.5 s silence after a neutral sentence
### 3\.2Systems and scale

The seven configurations in the main analysis are the Moshi base \(the Moshiko and Moshika checkpoints\), Moshika\-RL, the PersonaPlex base \(PPlex\), PPlex\-RL\[[15](https://arxiv.org/html/2609.19596#bib.bib23)\], NVIDIA VoiceChat\-11B, MiniCPM\-o 4\.5\[[2](https://arxiv.org/html/2609.19596#bib.bib24)\]and Raon\-SpeechChat\[[9](https://arxiv.org/html/2609.19596#bib.bib25)\]\. All models use their official default decoding settings\. PersonaPlex runs with its default voice and the neutral persona prompt “You enjoy having a good conversation\.”, without any instruction about interrupting\. The main grid is 40 topics×\\times10 conditions\. Each configuration uses five seeds per stimulus \(2,000 trials\), except Moshi, which pools five seeds from each checkpoint\.

### 3\.3Measurement and analysis

Thespeech\-onset rateis the proportion of all trials in which the model starts a speech segment of at least four words within\[t0,t0\+4s\)\[t\_\{0\},t\_\{0\}\+4\\,\\mathrm\{s\}\), wheret0t\_\{0\}is trigger end\. Trials with such speech during the preceding 1\.5 s remain in the denominator but count as no new onset\. Moshi, PersonaPlex and Raon emit one text token per 80 ms audio frame, so a speech segment is a stretch of word tokens with no gap longer than 0\.64 s\. For VoiceChat and MiniCPM\-o, segments come from turn\-boundary markers and 1 Hz listen/speak outputs, respectively\.

Table 2:Speech\-onset rate by trigger condition across five model families, over all trials\.Figure 2:Trial rasters for PPlex, Moshi and Raon\-SpeechChat under the false\-fact, question and silence conditions\. Each row is one trial, bars mark substantive speech, the shaded band is the 4 s response window, the number is the speech\-onset rate over all trials\.

## 4Results

### 4\.1Speech onset tracks being addressed and silence

Table[2](https://arxiv.org/html/2609.19596#S3.T2)summarizes speech initiation rates for seven configurations under 10 conditions\. Based on the neutral baseline and each condition’s change relative to it, the configurations fall into three patterns\. First, Raon\-SpeechChat starts speaking frequently\. Its neutral baseline is \.42, and the ongoing\-speech conditions shown yield \.26 to \.47\. Relative to this baseline, neither the addressed conditions nor the self\-select conditions show a clear increase in onset\. Silence, which raises onset to \.68, is the only positive result\. Second, VoiceChat and MiniCPM\-o 4\.5 remain near zero during ongoing speech and start chiefly when silence is inserted\. Third, the four Moshi and PersonaPlex configurations fall between these extremes\. Their neutral rates are \.01–\.10, but they do respond to some cues during ongoing speech, making them the main comparison set\.

In the four Moshi and PersonaPlex configurations, whether the model speaks depends on two surface cues: being addressed directly and a pause in the speech\. Direct questions raise onset by\+0\.08\+0\.08to\+0\.24\+0\.24, and inserting 1\.5 s of silence raises it from at most \.10 under Neutral to as much as \.83\. Two control comparisons show that these cues are indeed surface cues\. Rhetorical questions, which have question form but are not directed at the model, have almost no effect\. It implies that the model responds to being asked rather than to the question form\. Turn end, in which the speaker says “that’s all from me” but keeps talking, also has almost no effect, so the model responds to the speaker actually stopping rather than to the words that yield the turn\. By contrast, cues that require understanding the content in order to decide whether to speak, namely false facts, verbatim repetition and hazards, differ from Neutral by at most 0\.06 across the seven configurations\. The one exception is word search, which raises onset by 0\.12 and 0\.18 in PPlex and PPlex\-RL\.

To examine the temporal dynamics of speech around the trigger, Figure[2](https://arxiv.org/html/2609.19596#S3.F2)plots per\-trial speech segments of three models\. Many Raon segments begin before trigger end and span the response window\. Its high onset therefore reflects near\-continuous speaking rather than a response to the trigger, which is why its rates vary little across conditions\. PPlex and Moshi show more post\-trigger speech for questions and silence, while the false\-fact column is sparse\.

Figure 3:Non\-silent text\-token probability relative to Neutral for Moshi and PPlex\. The plotted P\(word\) denotesPtextP\_\{\\mathrm\{text\}\}\. Curves show mean differences from Neutral\.
### 4\.2Text\-token probabilities track requests and silence

Note that not speaking may not mean the model is insensitive to content cues\. Generation tendency may already have changed without crossing the initiation threshold\. We therefore examine the frame\-level probability of non\-silent text\-channel tokens for Moshi and PPlex, which we compute as

Ptext​\(t\)=1−P⁡\(⟨PAD⟩∣t\)−P⁡\(⟨EPAD⟩∣t\),P\_\{\\mathrm\{text\}\}\(t\)=1\-P\(\\langle\\mathrm\{PAD\}\\rangle\\mid t\)\-P\(\\langle\\mathrm\{EPAD\}\\rangle\\mid t\),where⟨PAD⟩\\langle\\mathrm\{PAD\}\\rangleand⟨EPAD⟩\\langle\\mathrm\{EPAD\}\\rangleare the padding and end\-of\-padding tokens that fill the text stream between words\.

We compare four conditions with Neutral\. Question and Silence are positive controls, a language cue and an acoustic cue that both raise onset\. False fact and Word search are the content cues of interest\. False fact leaves onset unchanged in both models, whereas word search raises it in PPlex but not in Moshi\.

Figure[3](https://arxiv.org/html/2609.19596#S4.F3)gives the result\. The two positive controls raisePtextP\_\{\\mathrm\{text\}\}in both models\. Averaged over the first 2 s after trigger end, Question raises it by\+0\.13\+0\.13in Moshi and\+0\.07\+0\.07in PPlex, and Silence by\+0\.06\+0\.06and\+0\.30\+0\.30\. The measure thus responds to a language cue as well as to an acoustic one\. False fact does not raise it\. Its 2 s mean is slightly below Neutral in both models \(−0\.03\-0\.03and−0\.02\-0\.02\), so there is no sign of a raised tendency to speak\. Word search has a small value in PPlex but not a sustained rise in text\-token probability\.

In sum, the probe agrees with the behavior\. Question and silence raise text\-token probability just as they raise onset, whereas neither content cue raises the tendency to speak even at the level of token probabilities\.

Table 3:Speech\-onset rate for the neutral, false\-fact and hazard cues in Moshi and PPlex under added opportunity \(silence, natural pauses\) and added permission \(opening instructions\)\. Same denominator as Table[2](https://arxiv.org/html/2609.19596#S3.T2)\.
### 4\.3Content effects remain inconsistent with added opportunity and permission

Two alternative explanations remain for the weak content effects in Section[4\.1](https://arxiv.org/html/2609.19596#S4.SS1)\. The model may be sensitive to content but lack the opportunity or the permission to speak\. If so, adding either should raise onset after false facts and hazards above Neutral\. We add opportunity in two ways\. One inserts 1\.5 s of silence after the trigger utterance, as in the Silence condition\. The other keeps the natural inter\-word pauses of the synthesized speech instead of compressing them\. For permission, we add an opening instruction that grants it for the specific cue\. Before the monologue, the speaker says either “if I say anything that is wrong, please just interrupt me straight away” or “if I mention that I am doing something dangerous or unsafe, please just interrupt me straight away”\. Each instruction is paired with its own Neutral trigger, so every contrast holds the instruction fixed and varies only the trigger\. Table[3](https://arxiv.org/html/2609.19596#S4.T3)reports onset under these manipulations for Moshi and PPlex\.

With added opportunity, both models speak more overall, most strikingly PPlex, whose neutral onset rises from \.10 to \.83 with inserted silence\. False\-fact and hazard onset rises along with it, but not consistently above the Neutral rate of the same row\. The only gain appears in Moshi with inserted silence, where both cues reach \.21 against \.12 for Neutral\. PPlex shows no gain in either pause setting, and with natural pauses neither model differs from Neutral by more than \.04\.

With added permission, overall onset changes little\. Neutral onset under the two instructions stays within \.05 of the main table, and false\-fact and hazard rates differ from Neutral in the same row by at most \.03, consistent with the low compliance rate reported by Instruct\-FD\[[22](https://arxiv.org/html/2609.19596#bib.bib4)\]\.

Across these controls, a notable content\-specific increase appears only in one setting, Moshi with inserted silence, so there is no consistent increase in intervention onset\. More opportunity makes the models speak more overall, and explicit permission changes little, but neither makes a false fact or a hazard a reason to speak\. The weak content effects in Section[4\.1](https://arxiv.org/html/2609.19596#S4.SS1)therefore cannot be attributed to a lack of opportunity or permission\. An independent experiment next changes the question\. Instead of asking whether the model takes the floor, it hands the floor over and asks whether what the model says provides the required help\.

### 4\.4Given the floor, models answer questions but seldom correct or warn

Table 4:Responses after turn release\.*Speak*gives the proportion of trials with a≥\\geq4\-word onset within 10 s\.*Relevant*gives the proportion of non\-empty replies addressing the trigger\.*Intervenes*gives the proportion that challenge or correct False fact, or warn about Hazard\. Replies are judged by DeepSeek\-V4\-Flash \(temperature 0\)\.We run an independent turn\-release experiment, which hands the floor to Moshi and PPlex\. We leave 10 s of silence after trigger end and then measure substantive speech within the silence\. DeepSeek\-V4\-Flash\[[3](https://arxiv.org/html/2609.19596#bib.bib26)\]\(temperature 0\) judges whether a reply appropriately addresses the trigger\. For false facts it also judges whether the reply challenges, corrects, or expresses doubt\. For hazard it also judges whether the reply gives a risk warning or safer alternative action\. Only affirmative intervention judgments count as success\.

Table[4](https://arxiv.org/html/2609.19596#S4.T4)shows the results\. After turn release, PPlex starts speaking in almost every trial and Moshi in most conditions, but whether the reply addresses the trigger depends on what the cue asks for\. For cues that only call for a reply, namely direct questions, verbal turn end and verbatim repetition, most replies of both models are on topic, with relevance between \.65 and \.90\. For the three cues that call for help, namely false fact, hazard and word search, relevance falls to between \.16 and \.36\. Word search shows this most clearly\. In Section[4\.1](https://arxiv.org/html/2609.19596#S4.SS1)it is the only content cue that raises PPlex onset, yet here only \.36 of its replies actually supply the missing word\. The positive word\-search onset effect therefore cannot serve as evidence of successful assistance\.

Intervention is rarer still\. Among non\-empty false\-fact replies, Moshi and PPlex challenge or correct the claim at rates of only \.14 and \.15\. Among hazard replies, warnings appear at only \.04 and \.07\. The remaining replies are mostly not backchannels but contentful speech that continues the topic without engaging the problem, and in the false\-fact case often goes along with the claim\. Even when PPlex speech initiation under false fact and hazard reaches \.98 and 1\.00, corresponding intervention remains rare\. The problem therefore cannot be attributed only to not getting the turn\.

## 5Conclusion

This paper asks when full\-duplex models take the floor and what they say once they have it\. Across five model families, speech onset follows the structure of the conversation rather than its content\. A question directed at the model or a pause in the speech raises onset\. False fact and hazards show little consistent effect on speech onset, while false facts also fail to increase token\-level tendency to speak\. And neither extra pauses nor explicit permission produces a stable content\-specific increase in onset\. Even with the floor handed over, Moshi and PersonaPlex answer questions but seldom correct the claim or warn of the danger\. The limiting factor is therefore not access to the turn but the decision that the content warrants one\. Whether the models fail to notice the problem or notice it and remain silent is an open question, and the matched\-context protocol introduced here provides a way to study it\. Content\-driven intervention thus remains a concrete target for future full\-duplex models\.

## References

- \[1\]S\. Arora, Z\. Lu, C\. Chiu, R\. Pang, and S\. Watanabe\(2025\)Talking Turns: benchmarking audio foundation models on turn\-taking dynamics\.InProc\. ICLR,Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[2\]J\. Cui, B\. Xu, C\. Wang, T\. Yu, W\. Sun,et al\.\(2026\)MiniCPM\-o 4\.5: towards real\-time full\-duplex omni\-modal interaction\.arXiv preprint arXiv:2604\.27393\.Cited by:[§3\.2](https://arxiv.org/html/2609.19596#S3.SS2.p1.1)\.
- \[3\]DeepSeek\-AI\(2026\)DeepSeek\-V4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§4\.4](https://arxiv.org/html/2609.19596#S4.SS4.p1.1)\.
- \[4\]A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour\(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.arXiv preprint arXiv:2410\.00037\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p1.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[5\]E\. Ekstedt and G\. Skantze\(2020\)TurnGPT: a transformer\-based language model for predicting turn\-taking in spoken dialog\.InFindings of EMNLP,Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[6\]E\. Ekstedt and G\. Skantze\(2022\)Voice activity projection: self\-supervised learning of turn\-taking events\.InProc\. Interspeech,Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[7\]Y\. Ge, S\. Chen, J\. Xiao, X\. Liu, T\. Xiao, Y\. Xiang, Z\. Yu, and J\. Zhu\(2025\)FLEXI: benchmarking full\-duplex human\-LLM speech interaction\.arXiv preprint arXiv:2509\.22243\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p2.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[8\]M\. H\. Goodwin and C\. Goodwin\(1986\)Gesture and coparticipation in the activity of searching for a word\.Semiotica62\(1–2\),pp\. 51–75\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p3.1),[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[9\]B\. Kim, C\. Choi, D\. Kim, D\. Lee, E\. Ewer,et al\.\(2026\)Raon\-Speech technical report\.arXiv preprint arXiv:2605\.23912\.Cited by:[§3\.2](https://arxiv.org/html/2609.19596#S3.SS2.p1.1)\.
- \[10\]G\. H\. Lerner\(1996\)On the “semi\-permeable” character of grammatical units in conversation: conditional entry into the turn space of another speaker\.InInteraction and Grammar,E\. Ochs, E\. A\. Schegloff, and S\. A\. Thompson \(Eds\.\),pp\. 238–276\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p3.1),[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[11\]S\. C\. Levinson and F\. Torreira\(2015\)Timing in turn\-taking and its implications for processing models of language\.Frontiers in Psychology6,pp\. 731\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[12\]G\. Lin, J\. Lian, T\. Li, Q\. Wang, G\. Anumanchipalli, A\. H\. Liu, and H\. Lee\(2025\)Full\-Duplex\-Bench: a benchmark to evaluate full\-duplex spoken dialogue models on turn\-taking capabilities\.arXiv preprint arXiv:2503\.04721\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[13\]T\. A\. Nguyen, E\. Kharitonov, J\. Copet, Y\. Adi, W\. Hsu, A\. Elkahky, P\. Tomasello, R\. Algayres, B\. Sagot, A\. Mohamed, and E\. Dupoux\(2023\)Generative spoken dialogue language modeling\.Transactions of the Association for Computational Linguistics11,pp\. 250–266\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[14\]NVIDIA\(2026\)NVIDIA NemotronLabs VoiceChat 11B\.Note:Model card,[https://huggingface\.co/nvidia/NVIDIA\-NemotronLabs\-VoiceChat\-11B](https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B)Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p1.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[15\]A\. Ohashi, N\. Zeghidour, A\. Défossez, and E\. Kharitonov\(2026\)Multi\-faceted interactivity alignment in full\-duplex speech models\.arXiv preprint arXiv:2606\.11167\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.19596#S3.SS2.p1.1)\.
- \[16\]Y\. Peng, Y\. Chao, D\. Ng, Y\. Ma, C\. Ni, B\. Ma, and E\. S\. Chng\(2025\)FD\-Bench: a full\-duplex benchmarking pipeline designed for full duplex spoken dialogue systems\.arXiv preprint arXiv:2507\.19040\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p2.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[17\]R\. Roy, J\. Raiman, S\. Lee, T\. Ene, R\. Kirby, S\. Kim, J\. Kim, and B\. Catanzaro\(2026\)PersonaPlex: voice and role control for full duplex conversational speech models\.arXiv preprint arXiv:2602\.06053\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p1.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[18\]H\. Sacks, E\. A\. Schegloff, and G\. Jefferson\(1974\)A simplest systematics for the organization of turn\-taking for conversation\.Language50\(4\),pp\. 696–735\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p3.1),[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[19\]E\. A\. Schegloff, G\. Jefferson, and H\. Sacks\(1977\)The preference for self\-correction in the organization of repair in conversation\.Language53\(2\),pp\. 361–382\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p3.1),[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[20\]E\. A\. Schegloff\(1982\)Discourse as an interactional achievement: some uses of “uh huh” and other things that come between sentences\.InAnalyzing Discourse: Text and Talk,D\. Tannen \(Ed\.\),pp\. 71–93\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[21\]T\. Stivers, N\. J\. Enfield, P\. Brown, C\. Englert, M\. Hayashi, T\. Heinemann, G\. Hoymann, F\. Rossano, J\. P\. de Ruiter, K\. Yoon, and S\. C\. Levinson\(2009\)Universals and cultural variation in turn\-taking in conversation\.Proceedings of the National Academy of Sciences106\(26\),pp\. 10587–10592\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.
- \[22\]Y\. Tang, W\. Ma, X\. Zhao,et al\.\(2026\)Instruct\-FD: can your full\-duplex speech system follow turn\-taking instructions?\.arXiv preprint arXiv:2607\.20460\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p2.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1),[§4\.3](https://arxiv.org/html/2609.19596#S4.SS3.p3.1)\.
- \[23\]B\. Veluri, B\. N\. Peloquin, B\. Yu, H\. Gong, and S\. Gollakota\(2024\)Beyond turn\-based interfaces: synchronous LLMs as full\-duplex dialogue agents\.InProc\. EMNLP,Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[24\]C\. Wanget al\.\(2026\)Full\-duplex interaction in spoken dialogue systems: a comprehensive study from the ICASSP 2026 HumDial challenge\.arXiv preprint arXiv:2604\.21406\.Cited by:[§1](https://arxiv.org/html/2609.19596#S1.p2.1),[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[25\]X\. Wang, Y\. Li, C\. Fu, Y\. Shen, L\. Xie, K\. Li, X\. Sun, and L\. Ma\(2024\)Freeze\-Omni: a smart and low latency speech\-to\-speech dialogue model with frozen LLM\.arXiv preprint arXiv:2411\.00774\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p1.1)\.
- \[26\]V\. H\. Yngve\(1970\)On getting a word in edgewise\.InPapers from the Sixth Regional Meeting of the Chicago Linguistic Society,pp\. 567–578\.Cited by:[§2](https://arxiv.org/html/2609.19596#S2.p2.1)\.

Similar Articles

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

arXiv cs.CL

This paper analyzes synchronization and turn-taking dynamics in full-duplex speech dialogue models by simulating conversations between two instances of the Moshi model, measuring representational alignment via CKA and predicting turn boundaries with LSTM probes.