Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework
Summary
This paper proposes a generalized style-aware full-duplex framework with a lightweight turn controller LPS-TC, introduces a large-scale dataset WildTurn for real-world conversations, and presents a two-tier evaluation scheme to enhance proactive spoken interactions and response quality in dialogue systems.
View Cached Full Text
Cached at: 09/01/26, 11:56 AM
# Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework
Source: [https://arxiv.org/html/2608.28630](https://arxiv.org/html/2608.28630)
Tianrui Pan, Qinglin Zhang1, Chong Deng1, Luyao Cheng1, Qian Chen1, Wen Wang1, Jie Tang, Gangshan Wu, Jie Liu∗ 1Token Foundry, Alibaba Group
###### Abstract\.
Compared with half\-duplex dialogue systems where the system waits for user turn completion before it responds, natural full\-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels\. This creates a key challenge: improving turn timing without sacrificing response quality\. To address limitations in realistic proactive turn\-taking, we build a generalized style\-aware full\-duplex framework with three key components\. Firstly, we propose LPS\-TC, a Lightweight Proactive Speech Turn Controller for plug\-and\-play integration\. It features a fine\-grained action space covering both reactive and proactive turn behaviors, enabling half\-duplex models with full\-duplex capabilities and enhancing existing full\-duplex models with superior timing control\. Secondly, we construct WildTurn, a large\-scale, real\-world English dataset containing approximately 2,981 hours of filtered multi\-turn stereo conversations from face\-to\-face and telephone conversations, annotated with five turn\-taking and five backchanneling styles\. Trained on WildTurn, LPS\-TC exhibits rich spoken dynamics that are not captured by existing static full\-duplex benchmarks\. Thirdly, we introduce a two\-tier evaluation scheme that assesses both chunk\-level timing precision and turn\-level interaction quality under realistic streaming constraints\. Our experiments, integrating LPS\-TC with half\-duplex models like Qwen2\.5\-Omni and full\-duplex models like Freeze\-Omni, showcase its superior performance in timing appropriateness and response quality\. Our framework also demonstrates fine\-grained style controllability and strong generalizability, enabling more natural and human\-like spoken interactions\.
style\-aware full\-duplex dialogue, proactive spoken interactions
††ccs:Human\-centered computing Collaborative interaction## 1\.Introduction
Figure 1\.Proposed generalized full\-duplex dialogue framework\. Our lightweight, plug\-and\-play turn controller LPS\-TC \(i\) integrates seamlessly with off\-the\-shelf half/full\-duplex Speech LLMs, \(ii\) supports style\-conditioned proactive spoken behaviors such as backchannels and interruptions\.Spoken dialogue systems have progressed from text\-based dialogue frameworks that emphasize contextual understanding\(Chenet al\.,[2022](https://arxiv.org/html/2608.28630#bib.bib179); Liaoet al\.,[2021](https://arxiv.org/html/2608.28630#bib.bib180); Wuet al\.,[2020](https://arxiv.org/html/2608.28630#bib.bib182)\)and response generation\(Rolleret al\.,[2021](https://arxiv.org/html/2608.28630#bib.bib181); Yeet al\.,[2022](https://arxiv.org/html/2608.28630#bib.bib183)\)to timing\-sensitive, expression\-rich conversational agents\. While some works develop turn\-based half\-duplex voice assistants\(Nguyenet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib161); Wuet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib110)\)into full\-duplex models that can listen and speak concurrently\(Wanget al\.,[2024c](https://arxiv.org/html/2608.28630#bib.bib113),[https://arxiv.org/html/2608.28630#bib.bib151](https://arxiv.org/html/2608.28630#bib.bib151)\), other lines of work investigate diverse spoken behavioral styles, including distinct personalities and fine\-grained emotional expression\(Tuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib174); Genget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib173); Cuiet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib175)\)\. Different from reactively following the user\-oriented conversation, natural full\-duplex systems should manage real\-time proactive turn behaviors such as interrupting for clarifications, offering information without being asked, or backchannels to show engagement\. Such behaviors enable assistants to support humans not only mechanically but also socially and emotionally\. However, research on proactivity in spoken dialogues remains underexplored\. Some works\(Denget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib178),[https://arxiv.org/html/2608.28630#bib.bib186](https://arxiv.org/html/2608.28630#bib.bib186)\)consider the semantics of text\-based responses, while others\(Nguyenet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib161); Mitsuiet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib187); Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)struggle with heterogeneous user preferences, since the same proactive responses may be perceived as supportive by some users yet intrusive by others\. Enabling strategic and motivational turns in natural spoken dialogues requires fine\-grained temporal grounding that jointly accounts for semantics, prosody, and interaction style, yet existing spoken dialogue models still lack this capability\(Changet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib164)\)\. Yet existing efforts remain fragmented across turn\-control methods, datasets, and evaluation protocols, leaving no unified support for proactive, style\-controllable full\-duplex interaction\. To bridge this gap, we propose a generalized full\-duplex framework with three key components\.
First, we proposeLPS\-TC, balancing the trade\-off between high response quality and natural spoken turn interactions with its architectual design\.While advanced half\-duplex speech large language models \(LLMs\)\(Wuet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib110); Yanget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib125)\)excel in response quality, their strictly sequential turn\-taking limits real\-time processing and natural spoken\-turn dynamics\. Meanwhile, developing effective full\-duplex models remains challenging\. Some proprietary commercial agents\(OpenAI,[2024](https://arxiv.org/html/2608.28630#bib.bib135); Intelligence,[2025](https://arxiv.org/html/2608.28630#bib.bib136)\)incur prohibitive costs for precise timing, while open\-source solutions\(Wanget al\.,[2024c](https://arxiv.org/html/2608.28630#bib.bib113); Chenet al\.,[2025b](https://arxiv.org/html/2608.28630#bib.bib165)\)suffer from limited interactional modeling and degradation in response quality\. OurLPS\-TC, depicted in Figure[1](https://arxiv.org/html/2608.28630#S1.F1), balances this trade\-off\. It is a lightweight, plug\-and\-play spoken turn controller for turn\-timing prediction and style\-conditioned spoken interactions\. It autoregressively processes dual\-channel audio streams from both user and assistant\. This allows it to capture critical paralinguistic cues and conversational dynamics, which are typically lost in text\-centric or single\-stream models\.LPS\-TCcan either equip half\-duplex speech LLMs with full\-duplex capability or enhance existing full\-duplex models with superior timing control\.
Second, we introduce WildTurn, an English dataset with fine\-grained proactive spoken turn dynamics\.Most available spoken dialogue datasets\(Leeet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib171); Linet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib170); Tuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib174); Genget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib173); Cuiet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib175)\)mainly focus on utterance\-level behavioral or paralinguistic styles\. They rarely capture interactional dynamics such as overlapping speech\. Although Behavior\-SD\(Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)incorporates turn behaviors, it is synthetic and may not capture the diversity of real\-world conversations\. To address these limitations, we curate WildTurn fromreal\-worldinteractions and annotate it with a comprehensive label space encompassing five action categories: Normal Turn Taking \(NTT\), Interruptive Turn Taking \(ITT\), Backchanneling \(BC\), Barge In \(BI\), and No Action \(NA\)\. Among these, NTT and BC are further refined into style labels: we define five turn\-taking styles for NTT and five backchannel styles for BC based on statistical metrics such as silence or overlap duration and action frequency\. We then use GPT\-5\.2\(Singhet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib172)\)to generate style\-conditioned instructions\. Moreover, to capture timing in human\-human conversations, we employ span\-based labeling informed by human\-computer interaction studies\(Wanget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib139); Chenet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib140)\)on realistic intent\-to\-speech latency, with each label spanning from cue to response onset\.
Third, we address the limitations of current full\-duplex benchmarks and build a new evaluation protocol towards more holistic and real\-time interactions\.Prior full\-duplex benchmarks\(Linet al\.,[2025c](https://arxiv.org/html/2608.28630#bib.bib155); Penget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib154); Aroraet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib123)\)often overlook two aspects\. First, they exclude advanced half\-duplex speech LLMs such as Qwen3\-Omni\(Yanget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib125)\): when adapted with proactive prompts, these models can serve as powerful proactive controllers for a fair comparison\. Second, their reliance on static, pre\-recorded dialogues creates a fundamentalcontext mismatch: in multi\-turn settings, each subsequent user turn is influenced by the previous ground\-truth assistant response in both timing and content\. To address these issues, we propose a two\-tier framework that evaluateschunk\-level timing precisionandturn\-level interaction qualityacross turn\-taking, backchanneling, and turn\-yielding underrealistic streaming constraints\. We benchmark against a broad spectrum of half\-duplex baselines via standardized proactive prompting\. To resolve the context mismatch, we reconstruct the test set by splitting multi\-round dialogues into individual turns\. Each turn is then presented to the model along with its real preceding context\. Experimental results show thatLPS\-TCoutperforms other turn controllers inchunk\-leveltiming accuracy based on our labeling space\. It also improvesturn\-levelinteraction quality when integrated with both half\-duplex and full\-duplex speech LLMs\. In addition,LPS\-TCdemonstrates superior style controllability and robust instruction following\.
## 2\.Related Work
Full\-Duplex Spoken Turn Controller\.End\-to\-end full\-duplex models\(Défossezet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib147); Zhanget al\.,[2025b](https://arxiv.org/html/2608.28630#bib.bib148);[Wanget al\.,](https://arxiv.org/html/2608.28630#bib.bib151); Yuet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib150)\)can handle overlapping speech, but their tightly coupled design is costly to train and may degrade response quality\(Xie and Wu,[2024](https://arxiv.org/html/2608.28630#bib.bib137); Chenet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib188); Aroraet al\.,[2025b](https://arxiv.org/html/2608.28630#bib.bib189); Royet al\.,[2026](https://arxiv.org/html/2608.28630#bib.bib193)\)\. To address this issue, other works add a turn controller, either VAD\-based\(Wanget al\.,[2024a](https://arxiv.org/html/2608.28630#bib.bib149); Fuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib145); Mai and Carson\-Berndsen,[2025](https://arxiv.org/html/2608.28630#bib.bib117); Wanget al\.,[2024c](https://arxiv.org/html/2608.28630#bib.bib113)\)for binary state modeling or hidden\-state\-based\(Changet al\.,[2022](https://arxiv.org/html/2608.28630#bib.bib168); Maet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib163); Chenet al\.,[2025c](https://arxiv.org/html/2608.28630#bib.bib146); Liuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib166); Luet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib142)\)for richer turn prediction\. Compared with standard VAD methods\(Team,[2024](https://arxiv.org/html/2608.28630#bib.bib121); Wuet al\.,[2025b](https://arxiv.org/html/2608.28630#bib.bib153); Xuet al\.,[2026](https://arxiv.org/html/2608.28630#bib.bib192)\), which only model binary turn states, recent controllers\(Zhanget al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib116); Liet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib124); Liaoet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib138)\)add states such aswaitandidleto better handle backchannels and background noise\. In contrast,LPS\-TCdecouples timing control from response generation, enabling fine\-grained assistant behaviors such as interruptions and backchannels without sacrificing response quality\.
Proactive Spoken Interactions\.Although proactivity in conversational agents has attracted growing attention\(Zarghamet al\.,[2022](https://arxiv.org/html/2608.28630#bib.bib198); Denget al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib200); Liaoet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib197); Zhanget al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib199)\), most prior work focuses on text, while spoken floor\-taking remains underexplored\(Viswanath and Buschmeier,[2026](https://arxiv.org/html/2608.28630#bib.bib195)\)\. Existing efforts study either real\-time agent frameworks\(Qiuet al\.,[2026](https://arxiv.org/html/2608.28630#bib.bib201); LiveKit,[2024](https://arxiv.org/html/2608.28630#bib.bib202); Daily\.co,[2024](https://arxiv.org/html/2608.28630#bib.bib203); Hugging Face,[2024](https://arxiv.org/html/2608.28630#bib.bib204)\)or sentence\-level speaking styles\(Tuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib174); Genget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib173); Cuiet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib175)\), but conversational fluency also depends ondiverse spoken turn behaviors\. We therefore focus on turn\-level proactive spoken interactions, including timely backchannels, predictive turn\-taking, and smooth turn\-yielding\. However, existing spoken dialogue datasets with proactive behaviors\(Nguyenet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib161); Mitsuiet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib187); Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169); Zhouet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib205)\)remain limited\. They rely on synthetic data, which fail to capture the nuanced dynamics of real conversations\(Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)\. We construct and annotate WildTurn from large\-scale real\-world face\-to\-face and telephone conversations with diverse turn\-taking and backchannel styles\.
Full\-duplex benchmarks\.Existing full\-duplex benchmarks mainly focus on isolated or limited spoken interactions\. Early work such as Full\-Duplex\-Bench\(Linet al\.,[2025c](https://arxiv.org/html/2608.28630#bib.bib155),[b](https://arxiv.org/html/2608.28630#bib.bib156)\)evaluates pause handling, interruption response, and overlap management, but is largely restricted to single\-round settings\. Later benchmarks extend to multi\-round scenarios, yet still lack two aspects: evaluating generalized Speech LLMs as real\-time turn\-taking models under streaming input, and modeling multi\-turn dynamic interaction under realistic streaming conditions\. For example, Talking Turns\(Aroraet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib123)\)focuses on timing prediction, FD\-Bench\(Penget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib154)\)emphasizes interruption\-heavy cases, and Full\-Duplex\-Bench\-v2\(Linet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib157)\)relies on a separate Speech LLM as examiner\. In contrast, we propose a two\-tier real\-time evaluation framework that assesses generalized Speech LLMs at both chunk and turn levels in realistic multi\-turn streaming interactions\.
Figure 2\.Decoupled full\-duplex pipeline\. LPS\-TC predicts actionztz\_\{t\}based on style instructionSSand history contexts of both user inputUa,<tU\_\{a,<t\}and SLM\-generated audioAa,<tA\_\{a,<t\}\. SLM generatesata\_\{t\}fromUa,<tU\_\{a,<t\}andztz\_\{t\}\. LPS\-TC offers an expanded action space with both reactive and proactive behaviors, facilitating precisely\-timed backchannels and interruptions\.
## 3\.Method
We present a unified full\-duplex framework for natural proactive spoken interaction with style\-aware, fine\-grained turn control\. Section[3\.1](https://arxiv.org/html/2608.28630#S3.SS1)formalize the differences between half\-duplex and full\-duplex interaction, Section[3\.2](https://arxiv.org/html/2608.28630#S3.SS2)introducesLPS\-TCfor turn control, and Section[3\.3](https://arxiv.org/html/2608.28630#S3.SS3)presents WildTurn data construction pipeline\.
### 3\.1\.Problem Formalization
Reactive half\-duplex and proactive full\-duplex systems mainly differ in the history contextHHthey condition on\. We discretize the audio stream into small chunks indexed by time steptt\. Ahalf\-duplex modelconditions on the complete user utterance as a fixed context\. LetUa=\{u0,…,uT−1\}U\_\{a\}=\\\{u\_\{0\},\\dots,u\_\{T\-1\}\\\}denote the full user audio utterance\. The model conditions on a history contextHhalfH\_\{\\text\{half\}\}that is available only after the user finishes speaking, i\.e\., at timeTT\. The assistant response generation ofAa=\{a1,…,aM\}A\_\{a\}=\\\{a\_\{1\},\\dots,a\_\{M\}\\\}is thus formalized as:
\(1\)P\(Aa\|Hhalf\)=∏m=1MP\(am\|a<m,Hhalf\),Hhalf=Context\(Ua\)\.P\(A\_\{a\}\|H\_\{\\text\{half\}\}\)=\\prod\_\{m=1\}^\{M\}P\(a\_\{m\}\|a\_\{<m\},H\_\{\\text\{half\}\}\),H\_\{\\text\{half\}\}=\\text\{Context\}\(U\_\{a\}\)\.whereContext\(⋅\)\\text\{Context\}\(\\cdot\)is a function mapping the available interaction history to a contextual representation\. This formulation is inherently reactive: the fixed contextHhalfH\_\{\\text\{half\}\}cannot capture real\-time dynamics and thus cannot support interruption or overlapping speech\. Afull\-duplex modeloperates on a dynamic contextHfull,tH\_\{\\text\{full\},t\}that evolves at each time steptt, comprising all available information: the incoming user audioUa,<tU\_\{a,<t\}and the model’s own audio historyAa,<tA\_\{a,<t\}\.
\(2\)Hfull,t=Context\(Ua,<t,Aa,<t\)\.H\_\{\\text\{full\},t\}=\\text\{Context\}\(U\_\{a,<t\},A\_\{a,<t\}\)\.At each time step, the model implicitly decides whether to speak or wait\. This decision can be formalized as a joint probability over a conceptual timing variablezt∈\{SPEAK,SILENCE\}z\_\{t\}\\in\\\{\\text\{SPEAK\},\\text\{SILENCE\}\\\}and the output audio tokenata\_\{t\}:
\(3\)P\(zt,at\|Hfull,t\)=P\(zt\|Hfull,t\)⋅P\(at\|zt,Hfull,t\)\.P\(z\_\{t\},a\_\{t\}\|H\_\{\\text\{full\},t\}\)=P\(z\_\{t\}\|H\_\{\\text\{full\},t\}\)\\cdot P\(a\_\{t\}\|z\_\{t\},H\_\{\\text\{full\},t\}\)\.In a fully integrated model, the conceptual variableztz\_\{t\}is not an explicit architectural component; instead, it implicitly guides the output generation\. Eq\.[3](https://arxiv.org/html/2608.28630#S3.E3)thus reveals a core challenge: an integrated model must simultaneously handle two interconnected tasks:timing control\(determiningztz\_\{t\}\) andcontent generation\(predictingata\_\{t\}\)\. This joint optimization causes a trade\-off, where improving low\-latency prediction ofztz\_\{t\}will compromise the complex reasoning required for high\-quality generation ofata\_\{t\}, and vice versa\.
### 3\.2\.Style\-Aware Decoupled Full\-Duplex Pipeline
To handle the trade\-off between timingztz\_\{t\}and response qualityata\_\{t\}, we introduceLPS\-TC, a Lightweight plug\-and\-play Proactive Speech Turn Controller\. It factorizes the joint decision in Equation[3](https://arxiv.org/html/2608.28630#S3.E3)by explicitly modeling the timing variableztz\_\{t\}separately\. Prior turn controllers\(Yuet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib150); Liaoet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib138)\)either serialize overlap into an interleaved user\-assistant token sequence or feed only user speech to the model\. These designs increase latency and fail to preserve paralinguistic cues that are critical for bidirectional interaction\. Moreover, because user preferences are heterogeneous, the same proactive action may be perceived as welcome by some users but unwelcome by others\.LPS\-TCintroduces two key innovations:
- •Low\-latency dual\-channel architecture, which directly processes raw audio from both user \(UaU\_\{a\}\) and assistant \(AaA\_\{a\}\) for simultaneous interaction without additional preprocessing\.
- •Style\-aware turn control, which incorporates an explicit style instructionSSto adapt turn\-taking and backchannel behavior to user preferences\.
For streaming inference, we adapt the Whisper encoder with causal convolutions\(Dielemanet al\.,[2016](https://arxiv.org/html/2608.28630#bib.bib143)\)and block causal attention\(Zenget al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib112)\)\. In Figure[2](https://arxiv.org/html/2608.28630#S2.F2), at each time steptt,LPS\-TC\(denoted byff\) explicitly predictsztz\_\{t\}conditioned on the accumulated context and the given style:
\(4\)zt=f\(Ua,<t,Aa,<t,Z<t,S\),z\_\{t\}=f\(U\_\{a,<t\},A\_\{a,<t\},Z\_\{<t\},S\),wherezt∈\{NA, NTT, ITT, BC, BI\}z\_\{t\}\\in\\\{\\text\{NA, NTT, ITT, BC, BI\}\\\}represents the predicted action: No Action \(NA\), Normal Turn Taking \(NTT\), Interruptive Turn Taking \(ITT\), Backchannel \(BC\), or Barge In \(BI\)\. The style instructionsSSconsist of two components: the turn\-taking styleSTTS\_\{\\text\{TT\}\}controls the turn\-taking tendency across five levels, and the backchannel styleSBCS\_\{\\text\{BC\}\}specifies one of five feedback patterns based on frequency and timing\. Details are given in Section[3\.3](https://arxiv.org/html/2608.28630#S3.SS3)\. The comprehensive action space and style conditioning together allowLPS\-TCto provide nuanced and adaptive turn control\. Then the SpeechLLM, denoted byFF, predicts response speech tokensata\_\{t\}as follows:
\(5\)at=F\(zt,Hhalf,t\|Hfull,t\),a\_\{t\}=F\\bigl\(z\_\{t\},H\_\{\\text\{half\},t\}\|H\_\{\\text\{full\},t\}\),where the predicted actionztz\_\{t\}fromLPS\-TCdetermines the behavior of SpeechLLM\. Specifically, NTT, ITT, and BC trigger SpeechLLM to generate speech responses, NA instructs SpeechLLM to maintain its current state \(i\.e\., either continue speaking or remain silent\), and BI instructs SpeechLLM to stop generating speech tokens\. After SpeechLLM executes the action at timett, the current user inpututu\_\{t\}and assistant outputata\_\{t\}are appended to the history \(formingAa,tA\_\{a,t\}in Equation[4](https://arxiv.org/html/2608.28630#S3.E4)\), which then predicts the next actionzt\+1z\_\{t\+1\}\. In this way,LPS\-TCand SpeechLLM form a closed real\-time control loop\. Theoretically,LPS\-TCcan equip half\-duplex models with full\-duplex capability by expanding their contextHhalfH\_\{\\text\{half\}\}\(Eq\.[1](https://arxiv.org/html/2608.28630#S3.E1)\) with the assistant’s audio responses\. It can also enhance full\-duplex models by replacing the implicit timing variableztz\_\{t\}\(Eq\.[3](https://arxiv.org/html/2608.28630#S3.E3)\) with a broader action space that includes ITT and BC in addition to NTT, BI and NA\. In summary, the decoupled and style\-aware design ofLPS\-TCenables low\-latency, fine\-grained turn control while supporting adaptive, personalized spoken interactions\.
Table 1\.Style definitions for five turn\-taking stylesSTTS\_\{\\text\{TT\}\}and five backchannel stylesSBCS\_\{\\text\{BC\}\}, based on statistical thresholds for turn\-boundary delays and action frequencies, which are derived from the spoken turn action labels \(NA, NTT, ITT, BC, BI\)\.Turn\-takingBackchannelStyleRatioBoundary timingStyleFrequencyOnset TimingPatientOnly NTT actions\.N/AHigh\-EarlyHigh BCs/turn or high BCs/minuteShortly after user startsMixedlow\\text\{Mixed\}\_\{\\text\{low\}\}Both ITT and NTT\.Higher NTT latency, shorter ITT lead\.High\-LateHigh BCs/turn or high BCs/minuteNear pause or user endMixedmedium\\text\{Mixed\}\_\{\\text\{medium\}\}Both ITT and NTT\.neutral NTT latency and ITT lead time\.Low\-EarlyLow BCs/turn and low BCs/minuteShortly after user startsMixedhigh\\text\{Mixed\}\_\{\\text\{high\}\}Both ITT and NTT\.Lower NTT latency, longer ITT lead\.Low\-LateLow BCs/turn and low BCs/minuteNear pause or user endAssertiveOnly ITT actions\.N/ANo BackchannelNo BC actions\.
### 3\.3\.WildTurn Dataset Construction
Existing dialogue datasets are insufficient for learning natural proactive spoken turns\. Most focus on utterance\-level expressive styles, lack fine\-grained reactive and proactive turn labels, and often rely on synthetic data\(Leeet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib171); Linet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib170); Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)\. We constructWildTurn, a real\-world dataset with multi\-turn conversations, fine\-grained spoken turn action labelsZZ, and corresponding style instructionsSS\. Figure[3](https://arxiv.org/html/2608.28630#S3.F3)summarizes the five\-step construction pipeline\.
Step a: Label spoken turn actions\.We assign fine\-grained action labels at the chunk level\. Following\(Aroraet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib123)\), each 40 ms audio chunk is annotated with one action label, matching the temporal resolution of the tokenized audio input\. We first identify assistant state transitions betweenSpeakandSilence\. We then assign NTT, ITT, BC, or BI at each transition boundary, while all remaining chunks are labeled as NA \(Figure[3](https://arxiv.org/html/2608.28630#S3.F3)a\)\. NTT, ITT, and BI are detected directly from VAD\-based\(Team,[2024](https://arxiv.org/html/2608.28630#bib.bib121)\)state changes, whereas BC is identified with a lexicon\-based procedure using a 66\-entry English backchannel lexicon expanded from\(Ekstedt and Skantze,[2022](https://arxiv.org/html/2608.28630#bib.bib122)\)\. On average, each dual\-channel audio sample in WildTurn contains 2\.26 interruptions \(ITT\), 1\.70 backchannels \(BC\), and 3\.64 natural turn\-takes \(NTT\)\.
Step b: Compute distributional metrics\.Based on the action labels from Step a, we define a small set of conversation\-level behavioral metrics\. For turn\-taking, these include the ITT\-to\-NTT ratio, NTT latency, and ITT lead time\. For backchanneling, they include backchannel frequency and onset timing\. Figure[3](https://arxiv.org/html/2608.28630#S3.F3)b shows the corresponding thresholds obtained by quantile\-based partitioning over their empirical distributions\.
Step c: Define conversation styles\.Using the metrics from Step b, we map each conversation to discrete style categories for turn\-taking and backchanneling, as shown in Figure[3](https://arxiv.org/html/2608.28630#S3.F3)c\. Forturn\-taking styles, the ratio and boundary timing metrics define a spectrum from fullyPatientbehavior to fullyAssertivebehavior\. We further define three intermediate categories:Mixedlow\\text\{Mixed\}\_\{\\text\{low\}\},Mixedmedium\\text\{Mixed\}\_\{\\text\{medium\}\}, andMixedhigh\\text\{Mixed\}\_\{\\text\{high\}\}\. Forbackchanneling styles, backchannel frequency and onset timing form a two\-dimensional grid\. This yields four active styles, namelyHigh\-Early,High\-Late,Low\-Early, andLow\-Late, together with theNo Backchannelstyle\. Table[1](https://arxiv.org/html/2608.28630#S3.T1)summarizes the resulting style definitions\.
Figure 3\.Overview of the WildTurn construction pipeline: chunk\-level action labeling, metric extraction, style categorization, instruction generation, and region\-based expansion\.Step d: Build style instructions\.Based on the style categories defined in Step c, we use GPT\-5\.2\(Singhet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib172)\)to verbalize the metric\-based style definitions in Table[1](https://arxiv.org/html/2608.28630#S3.T1)into concise natural\-language instructions for model conditioning\. These instructions serve as the style input paired with the action labels in WildTurn\.
Step e: Expand to region\-based labels\.Finally, we convert the point\-wise action labels from Step a into region\-based labels for temporally tolerant supervision\. Since point\-wise labels are sparse and sensitive to onset errors, we extend each action label backward from its onset to form an active region, following prior human\-computer interaction studies on the delay between speaking intention and speech onset\(Wanget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib139); Chenet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib140)\)\(Figure[3](https://arxiv.org/html/2608.28630#S3.F3)e\)\.
## 4\.Experiments
### 4\.1\.Datasets and Implementation Details
WildTurn is built from 1\.2k hours of face\-to\-face conversations from Seamless Interaction\(Agrawalet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib118)\), 2k hours of telephonic conversations from Fisher\(Cieriet al\.,[2004](https://arxiv.org/html/2608.28630#bib.bib119)\), and a small amount of synthetic data from Behavior\-SD\(Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)for training stability\. The dataset totals 2,981 hours of filtered stereo audio and 86,430 samples\. Original recordings are segmented into clips of up to 120 seconds using GPT\-4o\(OpenAI,[2024](https://arxiv.org/html/2608.28630#bib.bib135)\)to identify natural breakpoints such as topic shifts or speaker restarts\. We split the data into training, validation, and test sets with 85323, 100, and 1007 samples, respectively\.LPS\-TCcomprises an audio encoder initialized from Whisper\-large\-v3\(Radfordet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib127)\), an audio adapter, and a Qwen3\-0\.6B\(Yanget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib125)\)backbone\. For efficiency, we use a 20\-second sliding window over the most recent audio context\.
### 4\.2\.Two\-tier Real\-time Full\-Duplex Evaluation
We propose atwo\-tierreal\-time full\-duplex evaluation scheme that measureschunk\-leveltiming precision andturn\-levelinteraction quality for turn\-taking, backchanneling, and turn\-yielding\. Compared with prior full\-duplex benchmarks, our evaluation scheme has two key differences:
- •We treat generalized Speech LLMs as real\-time turn\-taking models by evaluating incremental chunk\-level action prediction over streaming audio\.
- •We segment the test set into streaming\-aligned evaluation instances, each paired with its preceding dialogue history\. This enables both chunk\-level and turn\-level evaluation under realistic streaming conditions\.
Assistant Turn\-taking Evaluation\.We formulate turn timing prediction for generalized Speech LLMs as incremental action prediction over audio chunks\. For eachuserturnutu\_\{t\}, given conversation historyHt−1=\{\(u1,a1\),…,\(ut−1,at−1\)\}H\_\{t\-1\}=\\\{\(u\_\{1\},a\_\{1\}\),\\ldots,\(u\_\{t\-1\},a\_\{t\-1\}\)\\\}and streaming audio chunksCt=\{c1,…,cn\}C\_\{t\}=\\\{c\_\{1\},\\ldots,c\_\{n\}\\\}, where each chunkcic\_\{i\}is a 640 ms segment chosen to balance temporal resolution and model robustness, the model predicts an action at each stepii:
\(6\)z^i=F\(Ht−1,c1:i\),\\hat\{z\}\_\{i\}=F\(H\_\{t\-1\},c\_\{1:i\}\),wherez^i∈\{wait,backchannel,response\}\\hat\{z\}\_\{i\}\\in\\\{\\textit\{wait\},\\textit\{backchannel\},\\textit\{response\}\\\}\. These actions are mapped to the interaction labels in Equation[4](https://arxiv.org/html/2608.28630#S3.E4):waitto NA,backchannelto BC, andresponseto either ITT or NTT\. Aresponseprediction is categorized as ITT if its first onseti∗i^\{\*\}occurs before the end of the user turn \(i\.e\.,i∗<ni^\{\*\}<n\), and as NTT otherwise\. Here,i∗=min\{i∣z^i=response\}i^\{\*\}=\\min\\\{i\\mid\\hat\{z\}\_\{i\}=\\textit\{response\}\\\}denotes the index of the first response\. This onset\-based criterion ensures that each user turn yields exactly one turn\-taking decision, while BC may occur multiple times or not at all\.
Assistant Turn\-yielding Evaluation\.For eachassistantturn, the model processes labeled assistant text and user audio, starting attat\_\{a\}andtut\_\{u\}, respectively\. At each time stepii, the model predicts:
\(7\)z^i=F\(Ht−1,at\[ta:tu\+Ti\],ut\[tu:tu\+Ti\]\),\\hat\{z\}\_\{i\}=F\(H\_\{t\-1\},a\_\{t\}\[t\_\{a\}:t\_\{u\}\+T\_\{i\}\],u\_\{t\}\[t\_\{u\}:t\_\{u\}\+T\_\{i\}\]\),wherez^i∈\{NA,BI\}\\hat\{z\}\_\{i\}\\in\\\{\\text\{NA\},\\text\{BI\}\\\}andTi=i×640msT\_\{i\}=i\\times 640\\text\{ms\}denotes the current temporal offset\. Here,at\[ta:tu\+Ti\]a\_\{t\}\[t\_\{a\}:t\_\{u\}\+T\_\{i\}\]represents assistant speech from its onset, whileut\[tu:tu\+Ti\]u\_\{t\}\[t\_\{u\}:t\_\{u\}\+T\_\{i\}\]denotes user audio chunks\. These predictions determine whether the assistant should continue speaking \(NA\) or yield the turn \(BI\)\.
Table 2\.Chunk\-level prediction results for each action label on the Switchboard testset\.Upper: Generalized speech LLMs as spoken turn judges\.Middle: Our lightweight turn controllerLPS\-TCcompared with task\-specific baselines\.Lower: Style Controllability over specific\-style subsets of Switchboard test set\.0indicates no predictions for this class\.Turn ControllerChunk\-level F1↑\\uparrowNANTTITTBCBIUpper: generalized speech LLMsFreeze\-Omni\(Wanget al\.,[2024c](https://arxiv.org/html/2608.28630#bib.bib113)\)0\.880\.340\.22\-0\.31GLM\-4\-Voice\(Zenget al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib112)\)0\.850\.430\.240\.120\.19Qwen3\-Omni\(Yanget al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib125)\)0\.860\.490\.140\.130\.20Middle: specific turn controllerFireRedVAD\(Xuet al\.,[2026](https://arxiv.org/html/2608.28630#bib.bib192)\)0\.860\.51\-\-\-RTTL\-DG\(Mai and Carson\-Berndsen,[2025](https://arxiv.org/html/2608.28630#bib.bib117)\)0\.900\.520\.62Ours w/o instructs0\.930\.660\.540\.510\.69Ours w/ instructs0\.920\.640\.600\.630\.71Lower: ours on specific subsets*Patient*0\.940\.6600\.490\.69*Assertive*0\.9200\.600\.540\.68*No Backchannel*0\.930\.690\.5500\.72
Table 3\.Turn\-level full\-duplex performance on the WildTurn testset for timing \(when\) and response quality \(what\)\.F1 score for turn\-level turn\-takingNTT/ITTand for turn\-level turn\-yieldingBIquantify timing accuracy relative to ground truth actions\. Using prompts from\(Wanget al\.,[2024b](https://arxiv.org/html/2608.28630#bib.bib158)\), we use Gemini\-2\.5\-Pro as an LLM\-as\-a\-Judge as\(Changet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib164)\)to provide binary score \(0/1\) for Timing and Response appropriateness across ITT, BC, and BI actions\.MethodTurn\-takingBackchannelsTurn\-yieldingResponse moduleJudgeNTT↑\\uparrowITT↑\\uparrowTimingITT\\text\{Timing\}\_\{\\text\{ITT\}\}↑\\uparrowResponseITT\\text\{Response\}\_\{\\text\{ITT\}\}↑\\uparrowTimingBC\\text\{Timing\}\_\{\\text\{BC\}\}↑\\uparrowResponseBC\\text\{Response\}\_\{\\text\{BC\}\}↑\\uparrowBI↑\\uparrowTimingBI\\text\{Timing\}\_\{\\text\{BI\}\}↑\\uparrowFull\-duplex Speech LLMsGPT\-4o\(OpenAI,[2024](https://arxiv.org/html/2608.28630#bib.bib135)\)native0\.540\.5064\.274\.678\.892\.00\.7473\.2Freeze\-Omni\(Wanget al\.,[2024c](https://arxiv.org/html/2608.28630#bib.bib113)\)native0\.440\.3247\.442\.1––0\.7271\.1\+ Ours10\.560\.5257\.043\.8––0\.7572\.9MiniCPM4\.5\(Yaoet al\.,[2024](https://arxiv.org/html/2608.28630#bib.bib177); Yuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib185)\)native0\.520\.4850\.268\.464\.490\.20\.7068\.8Half\-duplex Speech LLMsStep\-Audio 2\(Wuet al\.,[2025a](https://arxiv.org/html/2608.28630#bib.bib110)\)native0\.340\.4936\.840\.040\.556\.10\.4340\.1VAD0\.520\.2443\.539\.162\.569\.70\.4340\.1\+ Ours0\.600\.5656\.857\.380\.791\.70\.7767\.7Qwen2\.5\-Omni\(Xuet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib111)\)native0\.400\.5250\.453\.055\.152\.10\.5357\.1VAD0\.560\.4056\.461\.161\.066\.00\.5357\.1\+ Ours\-w/o20\.620\.6060\.468\.472\.289\.40\.7071\.2\+ Ours\-w/0\.600\.6463\.672\.878\.893\.00\.7271\.4
- 1We replace the original prediction head with our proposed model for unified turn controller\.
- 2w/o and w/ denote without and with style instructions, respectively\.
### 4\.3\.Chunk\-level Turn Timing Results
As shown in Table[2](https://arxiv.org/html/2608.28630#S4.T2), we evaluate chunk\-level turn prediction on the Switchboard test set to enable direct comparison with prior work\. Switchboard is a widely used benchmark for turn\-taking prediction, and the chunk\-level labels are derived from its publicly available annotations\.
The upper sectionassesses how well existing large speech LLMs predict turn changes\. Following MThread\(Wanget al\.,[2024b](https://arxiv.org/html/2608.28630#bib.bib158)\), we modify the system prompts of the half\-duplex models GLM\-4\-Voice and Qwen3\-Omni for streaming speech, adapting them to a full\-duplex setting\. The results indicate that these models perform well on NTT, where the system responds after the end of user turn\. However, they exhibit poor performance or a complete lack of ability in predicting appropriately timed ITT and BC\. And they exhibit poor turn\-yielding capabilities and struggle to stop speaking when the user interrupts \(poor BI\)\.
The middle sectionbenchmarks our model against specialized turn\-prediction baselines, with all evaluations standardized to a 160ms label resolution to ensure comparability\. By contrast, the generalized Speech LLMs in the upper section are evaluated at 640 ms resolution\. Since 160 ms is stricter, the specialized models would be expected to perform even better under the coarser 640 ms setting\. For FireRedVAD\(Xuet al\.,[2026](https://arxiv.org/html/2608.28630#bib.bib192)\), which mainly distinguishes between complete and incomplete user audio, we map its outputs to NTT and NA labels\. We also evaluate RTTL\-DG\(Mai and Carson\-Berndsen,[2025](https://arxiv.org/html/2608.28630#bib.bib117)\), an audio\-LLM baseline with a similar architecture to ours\.OurLPS\-TCoutperforms all baselines among both generalized speech LLMs and specialized turn controllers on chunk\-level turn timing accuracy for all action labels\. Specifically, our model without style instructions achieves the highest NA \(0\.93\) and NTT \(0\.66\) scores, and incorporating style instructions further boosts ITT to 0\.60, BC to 0\.63, and BI to 0\.71, demonstrating superior precision and control granularity\. Furthermore, ablation results confirm that style\-specific prompts are indispensable for replicating natural conversational dynamics\.
The lower sectionevaluates style instruction\-following across subsets of the Switchboard test set with distinct conversational patterns, each containing approximately 100 samples\. For example, onPatientsubset, the model should not predict any ITT action label \(i\.e\.,0for ITT\)\. No prediction of ITT for Patient subset, of NTT for Assertive subset, and of BC for No\-Backchannel subset in Table[2](https://arxiv.org/html/2608.28630#S4.T2)further demonstratesLPS\-TC’s precise fine\-grained, proactive style controllability and its ability to disable actions based on style constraints\. Collectively, these findings underscore the model’s superior chunk\-level precision for fine\-grained actions and its high fidelity in personalized style control\.
### 4\.4\.Turn\-level Full\-duplex Evaluation Results
To demonstrate the versatility of our spoken turn controller, we integrate it with various speech LLMs and present a comprehensive turn\-level evaluation on the WildTurn test set in Table[3](https://arxiv.org/html/2608.28630#S4.T3)\. We assess performance across three aspects:
- •turn accuracy\(F1 of NTT, ITT, BI\),
- •timing appropriateness\(TimingITT\\text\{Timing\}\_\{\\text\{ITT\}\},TimingBC\\text\{Timing\}\_\{\\text\{BC\}\},TimingBI\\text\{Timing\}\_\{\\text\{BI\}\}\),
- •response quality\(ResponseITT\\text\{Response\}\_\{\\text\{ITT\}\},ResponseBC\\text\{Response\}\_\{\\text\{BC\}\}\)
We use Gemini\-2\.5\-Pro as an LLM\-as\-a\-Judge like\(Changet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib164)\)to provide binary score \(0/1\) for Timing and Response appropriateness across ITT, BC, and BI\. We omitResNTT\\text\{Res\}\_\{\\text\{NTT\}\}, as turn\-end response content is unchanged from the underlying SpeechLLMs\. As defined in Section[4\.2](https://arxiv.org/html/2608.28630#S4.SS2), each turn contains one action, which is either turn\-taking \(NTT, ITT, MISSED\) or turn\-yielding \(BI, NA\)\. Any number of backchannel \(BC\) actions can occur within the same turn\.
Forfull\-duplex models, proprietary systems such as GPT\-4o set a strong benchmark, particularly in response quality \(ResBC=92\.0\\text\{Res\}\_\{\\text\{BC\}\}=92\.0\)\. For open\-source models, our controller significantly enhances existing full\-duplex systems for a more natural proactive actions\. For instance, integratingLPS\-TCwith Freeze\-Omni boosts turn\-level interruption accuracy \(ITT F1\) from 0\.32 to 0\.52\. This demonstrates its effectiveness in refining proactive behaviors\. Freeze\-Omni lacks backchanneling capabilities\. We therefore omit its BC metrics\. MiniCPM shows moderate native performance but remains below our enhanced models\. Forhalf\-duplex models, we evaluate whetherLPS\-TCcan enable full\-duplex interaction\. Native half\-duplex models such as Step\-Audio 2 struggle with proactive turn\-taking, as reflected by a lowNTT F1of 0\.34\. A simple VAD\-integrated baseline improves some metrics\. However, it cannot reliably distinguish true turn endings from user backchannels or noise\. Its turn\-yielding decisions are therefore unreliable, so we exclude it from BI comparisons\. In contrast, our controller substantially improves performance across the board\. Notably, Qwen2\.5\-Omni with our controller achieves the best overall performance, reachingNTT F1=0\.60\\text\{NTT F1\}=0\.60andITT F1=0\.64\\text\{ITT F1\}=0\.64\. Furthermore, an ablation study on Qwen2\.5\-Omni shows that adding style instructions \(\+ Ours\-w/\) improves key metrics such as response appropriatenessResITT\\text\{Res\}\_\{\\text\{ITT\}\}from 68\.4 to 72\.8\. This confirms the value of explicit style guidance\. Semantic evaluation with Gemini\-2\.5\-Pro, following\(Wanget al\.,[2024a](https://arxiv.org/html/2608.28630#bib.bib149); Changet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib164)\), shows that our controller improves both timing appropriateness and response quality in ITT and BC scenarios\. Overall, these results show thatour plug\-and\-play turn controller refines native full\-duplex systems and enables natural, controllable full\-duplex interaction for half\-duplex LLMs\.
We also evaluate end\-to\-end system latency, sinceLPS\-TCand the SpeechLLM operate as an integrated real\-time control loop\. Following FireRedChat\(Chenet al\.,[2025b](https://arxiv.org/html/2608.28630#bib.bib165)\), we measure end\-to\-first response time: the wall\-clock time from the end of a user’s utterance to the system’s first audio output\. All measurements use a single server with an NVIDIA A100\-80GB GPU, an Intel Xeon CPU, and 400 GiB of RAM\. Integrating speechLLM withLPS\-TCachieves latencies of 1\.4s \(Freeze\-Omni\+Ours\) and 1\.8s \(Qwen2\.5\-Omni\+Ours\),both within the 2\-4s range reported for SOTA systems in FireRedChat\(Chenet al\.,[2025b](https://arxiv.org/html/2608.28630#bib.bib165)\)\. This low latency is enabled by proactive turn\-taking, which anticipates user turns instead of waiting for end\-of\-speech\.
### 4\.5\.Out\-of\-Distribution Generalizability
We evaluate OOD generalization on two real\-world datasets: cross\-dataset English CANDOR\(Reeceet al\.,[2023](https://arxiv.org/html/2608.28630#bib.bib190)\)and cross\-lingual Mandarin AliMeeting\(Yuet al\.,[2022](https://arxiv.org/html/2608.28630#bib.bib191)\)\. Specifically, when annotating the chunk\-level behavior labels for backchannels on AliMeeting dataset, we identify candidates via exact matching or n\-gram heuristics up to three Mandarin words\. We compare our model with BeDLM\(Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\), which also targets various spoken turn behaviors but is trained on constructed synthetic English data\. We argue that our diverse real\-world training data inherently contains more complex and authentic spoken interaction patterns than synthetic data, providing generalizability advantages toLPS\-TC\.
Table 4\.OOD generalization across datasets and languages\.We compare our model with the baseline on chunk\-level F1 on 100 real\-world English samples from CANDOR and 100 Mandarin two\-speaker samples from AliMeeting\.MethodChunk\-level Labels \(F1 score↑\\uparrow\)NANTTITTBCBICross\-Dataset OOD \(CANDOR\)BeDLM\(Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)0\.800\.510\.180\.220\.58Ours0\.860\.570\.510\.490\.67Cross\-Lingual OOD \(AliMeeting\)BeDLM\(Sehun Lee,[2025](https://arxiv.org/html/2608.28630#bib.bib169)\)0\.400\.370\.100\.130\.55Ours0\.510\.460\.250\.290\.60As shown in Table[4](https://arxiv.org/html/2608.28630#S4.T4), BeDLM degrades markedly on CANDOR, especially on ITT and BC\. This supports our claim that incorporating realistic spoken interaction behaviors is crucial\. ForLPS\-TCon AliMeeting, NTT and BI are affected the least since NTT mainly reflects the assistant’s decision to take the turn after a pause, while BI reflects the decision to stop when overlap is detected as the user begins to interrupt, where paralinguistic cues \(e\.g\., pauses, intonation, and overlap\) provide critical signals beyond contextual semantics\. In contrast, ITT and BC depend more on semantic context, hence cross\-lingual discrepancy leads to a larger degradation from English to Mandarin\.
## 5\.Ablation and Analysis
Ablation 1: Effect of Style Instructions on Behavioral Distribution Shifts\.Different from Table[2](https://arxiv.org/html/2608.28630#S4.T2)Lower section that measures timing accuracy on specific\-style subsets, we also explicitly assess instruction\-following controllability\. In Figure[4](https://arxiv.org/html/2608.28630#S5.F4), we systematically override the original style instruction for every sample in Switchboard testset to track the resulting shifts in behavioral distributions\.
The top paneldemonstrates precise control over theturn\-takingtrade\-off between patience and assertiveness\. As the instruction shifts fromPatienttoAssertive, we observe a clear inverse relationship: the turn\-wait time \(NTT latency time\) plummets from 1,520 ms to 490 ms\. Concurrently, metrics for proactiveness, ITT lead time and ITT ratio of ITT/\(ITT\+NTT\), rise significantly, with the interruption lead time peaking at 630 ms under theAssertivestyle\. These results showour model can quantitatively interpret qualitative style instructions and modulate its turn\-taking strategy accordingly\.
The lower panelreveals fine\-grained, two\-dimensional control overbackchanneling\. The model successfully decouples frequency and timing\. Instructions likeHigh\-EarlyandHigh\-Lateyield a much higher backchannel rate \(up to 5\.17 per minute\) than theirLowcounterparts, while theNo BC\.prompt correctly suppresses BC entirely\. Simultaneously, the model precisely controls the onset timing:Earlyprompts trigger backchannels around 200 ms, far sooner than the ~1000 ms onset forLateprompts\. This ability to independently manage “how often” and “when” to provide listener feedback is crucial for natural interaction\.
Figure 4\.Style instructions induce clear behavioral distribution shifts\. Top: A shift from patient waiting to assertive interruption\. Bottom: Decoupled shifts in backchannel frequency and timing\.Ablation 2: Sample\-level Style Consistency\.We introduceStyle Consistency Accuracy\(SCA\) to measure whether a generated sample’s style matches its ground\-truth style\. Each sample is labeled following the procedure in Section[3\.3](https://arxiv.org/html/2608.28630#S3.SS3)\. For each target style, SCA is the proportion of test samples mapped back to the same style, i\.e\., per\-style recall under our labeling scheme\. We use SCA instead of F1 because the goal is to measure adherence to each target style, rather than performance on a balanced multi\-class classification task\. As shown in Table[5](https://arxiv.org/html/2608.28630#S5.T5), extreme styles achieve high consistency: the Patient and Assertive turn\-taking styles both exceed 0\.95, and No Backchannel reaches 0\.91\. In contrast, for turn\-taking styles, the intermediate mixed styles degrade noticeably, ranging from 0\.62 to 0\.88, suggesting difficulty in maintaining a fine\-grained balance between NTT pause durations and ITT interrupt lead durations, consistent with\(Changet al\.,[2025](https://arxiv.org/html/2608.28630#bib.bib164)\)\. For backchannel styles, we observe a clear asymmetry: late onset achieves 0\.89, whereas early onset drops to 0\.67, reflecting the challenge of early\-stage prediction under limited context\. These findings underscore a critical yet often overlooked challenge: designing methodologies tailored to achieve fine\-grained control over the timing and frequency of turn behaviors\.
Table 5\.Sample\-level style consistency\.The model shows high consistency on extreme styles \(e\.g\., Patient, Assertive\) but struggles with more nuanced intermediate styles, particularly for early\-onset backchannels\.Turn\-takingBackchannelStyleConsistencyStyleConsistencyPatient0\.95No BC\.0\.91Mixedlow\\text\{Mixed\}\_\{\\text\{low\}\}0\.62Freq\. high0\.94Mixedmedium\\text\{Mixed\}\_\{\\text\{medium\}\}0\.88Freq\. low0\.82Mixedhigh\\text\{Mixed\}\_\{\\text\{high\}\}0\.74Onset early0\.67Assertive0\.96Onset late0\.89Ablation 3: Visualization\.Figure[5](https://arxiv.org/html/2608.28630#S5.F5)compares our framework with baselines and highlights two advantages\.
First, within the real\-time streaming paradigm, our method enables more sophisticated and natural interaction\. The “SLM with Native” baseline, lacking temporal awareness, prematurely completes the user’s utterance \(“He should pay attention…”\) based on incomplete context\. The “SLM with VAD” setting avoids this error by waiting for silence, but remains purely reactive and cannot produce proactive behaviors such as backchanneling\. In contrast, “SLM with Ours” leverages a proactive turn controller that integrates both semantic and paralinguistic cues\. This allows it to make nuanced, context\-aware decisions, such as providing a timely backchannel \(¡BC¿ Mhm\.\) and later executing a strategic interruption \(¡ITT¿\), thus facilitating a fluid, human\-like conversational flow\.
Second, compared with the non\-streaming mode, our framework better balances responsiveness and quality\. The “Non\-streaming” approach, by processing the user’s full utterance, generates a high\-quality, comprehensive response\. However, this quality comes at the cost of high latency, which disrupts the conversational flow and defeats the purpose of a real\-time agent\. Conversely, our streaming framework, guided by the turn controller, engages in meaningful, real\-time interaction through semantically rich turn\-taking behaviors \(e\.g\., backchanneling, interruption\)\. This maintains conversational flow without sacrificing final response quality\. Overall, these results show that our spoken turn controller improves both the timing and content of system turns, enabling more natural and efficient full\-duplex interaction than the baselines\.
Figure 5\.We visualize how different controllers affect the turn signals and responses of a single Speech LLM Qwen2\.5\-Omni: a standalone Speech LLM, the SpeechLLM with VAD, and the SpeechLLM with our proposed turn controller\.Ablation 4: Human Evaluation\.Table[6](https://arxiv.org/html/2608.28630#S5.T6)presents the correlation between the LLM\-as\-a\-Judge \(Gemini\-2\.5\-Pro\) scores and human evaluations regarding the timing appropriateness of ITT, BC, and BI\. Notably, BC exhibits the highest alignment with human judgment \(ρ=0\.719\\rho=0\.719\)\. This suggests that the decision to backchannel is judged more consistently because it relies on explicit and localized cues, such as brief pauses\. In contrast, the lower correlation for ITT \(ρ=0\.637\\rho=0\.637\) reflects the inherent difficulty of interruption, a task requiring a delicate trade\-off between waiting \(NTT\) and acting, which in turn depends on a complex interplay of semantic and paralinguistic cues\. This inherent ambiguity in judging ITT timing highlights the critical role of our style instructions, as they provide the model with a clear policy to navigate such uncertain scenarios\. The moderate correlation for BI reflects its nature as a more straightforward task than ITT, primarily relying on speech overlap detection rather than complex semantic reasoning\.
Table 6\.Correlation between human and LLM\-as\-a\-Judge \(Gemini\-2\.5\-Pro\) binary scores, calculated on 50 samples per spoken turn behavior from Qwen2\.5\-Omni\+Ours\.Spoken Turn ActionsITTBCBIPearson’srr0\.6230\.7110\.671Spearman’sρ\\rho0\.6370\.7190\.673
## 6\.Conclusion
We presentLPS\-TC, a lightweight, decoupled turn controller that resolves the trade\-off between reasoning quality and interactional timing in spoken dialogue systems by enabling precise, style\-aware control\. Evaluated on our new WildTurn dataset with a novel two\-tier framework,LPS\-TCsignificantly improves timing precision and interactional fluidity, advancing the development of natural, proactive full\-duplex agents\.
## References
- V\. Agrawal, A\. Akinyemi, K\. Alvero, M\. Behrooz, J\. Buffalini, F\. M\. Carlucci, J\. Chen, J\. Chen, Z\. Chen, S\. Cheng,et al\.\(2025\)Seamless interaction: dyadic audiovisual motion modeling and large\-scale dataset\.arXiv preprint arXiv:2506\.22554\.Cited by:[§4\.1](https://arxiv.org/html/2608.28630#S4.SS1.p1.1)\.
- Talking turns: benchmarking audio foundation models on turn\-taking dynamics\.arXiv preprint arXiv:2503\.01174\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p4.1),[§2](https://arxiv.org/html/2608.28630#S2.p3.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p2.1)\.
- S\. Arora, J\. Tian, H\. Futami, J\. Shi, Y\. Kashiwagi, E\. Tsunoo, and S\. Watanabe \(2025b\)Chain\-of\-thought reasoning in streaming full\-duplex end\-to\-end spoken dialogue systems\.arXiv preprint arXiv:2510\.02066\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- K\. Chang, E\. Hu, C\. Kuan, W\. Ren, W\. Chen, G\. Lin, Y\. Tsao, S\. Sun, H\. Lee, and J\. Glass \(2025\)Game\-time: evaluating temporal dynamics in spoken language models\.arXiv preprint arXiv:2509\.26388\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§4\.4](https://arxiv.org/html/2608.28630#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2608.28630#S4.SS4.p2.6),[Table 3](https://arxiv.org/html/2608.28630#S4.T3),[§5](https://arxiv.org/html/2608.28630#S5.p4.1)\.
- S\. Chang, B\. Li, T\. N\. Sainath, C\. Zhang, T\. Strohman, Q\. Liang, and Y\. He \(2022\)Turn\-taking prediction for natural conversational speech\.arXiv preprint arXiv:2208\.13321\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- C\. Chen, K\. Hu, C\. H\. Yang, A\. Pasad, E\. Casanova, W\. Wang, S\. Fu, J\. Li, Z\. Chen, J\. Balam,et al\.\(2025a\)Reinforcement learning enhanced full\-duplex spoken dialogue language models for conversational interactions\.InSecond Conference on Language Modeling,Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- J\. Chen, C\. Gu, J\. Zhang, Z\. Liu, and S\. Konomi \(2024\)Sensing the intentions to speak in vr group discussions\.Sensors24\(2\)\.External Links:[Link](https://www.mdpi.com/1424-8220/24/2/362),ISSN 1424\-8220,[Document](https://dx.doi.org/10.3390/s24020362)Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p6.1)\.
- J\. Chen, Y\. Hu, J\. Li, K\. Li, K\. Liu, W\. Li, X\. Li, Z\. Li, F\. Shen, X\. Tang,et al\.\(2025b\)Fireredchat: a pluggable, full\-duplex voice interaction system with cascaded and semi\-cascaded implementations\.arXiv preprint arXiv:2509\.06502\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p2.1),[§4\.4](https://arxiv.org/html/2608.28630#S4.SS4.p3.1),[§4\.4](https://arxiv.org/html/2608.28630#S4.SS4.p3.1.3)\.
- Q\. Chen, Y\. Chen, Y\. Chen, M\. Chen, Y\. Chen, C\. Deng, Z\. Du, R\. Gao, C\. Gao, Z\. Gao,et al\.\(2025c\)Minmo: a multimodal large language model for seamless voice interaction\.arXiv preprint arXiv:2501\.06282\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- Z\. Chen, L\. Chen, B\. Chen, L\. Qin, Y\. Liu, S\. Zhu, J\. Lou, and K\. Yu \(2022\)UniDU: towards a unified generative dialogue understanding framework\.InProceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue,pp\. 442–455\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- C\. Cieri, D\. Miller, and K\. Walker \(2004\)The fisher corpus: a resource for the next generations of speech\-to\-text\.\.InLREC,Vol\.4,pp\. 69–71\.Cited by:[§4\.1](https://arxiv.org/html/2608.28630#S4.SS1.p1.1)\.
- W\. Cui, D\. Yu, X\. Jiao, Z\. Meng, G\. Zhang, Q\. Wang, S\. Y\. Guo, and I\. King \(2025\)Recent advances in speech language models: a survey\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13943–13970\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- Daily\.co \(2024\)Pipecat: open source framework for voice and multimodal conversational ai\.External Links:[Link](https://github.com/pipecat-ai/pipecat)Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour \(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.arXiv preprint arXiv:2410\.00037\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- Y\. Deng, W\. Lei, W\. Lam, and T\. Chua \(2023\)A survey on proactive dialogue systems: problems, methods, and prospects\.arXiv preprint arXiv:2305\.02750\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- Y\. Deng, L\. Liao, W\. Lei, G\. H\. Yang, W\. Lam, and T\. Chua \(2025\)Proactive conversational ai: a comprehensive survey of advancements and opportunities\.ACM Transactions on Information Systems43\(3\),pp\. 1–45\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- \[17\]Y\. Deng, W\. Zhang, W\. Lam, S\. Ng, and T\. ChuaPlug\-and\-play policy planner for large language model powered dialogue agents\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- S\. Dieleman, H\. Zen, K\. Simonyan, O\. Vinyals, A\. Graves, N\. Kalchbrenner, A\. Senior, K\. Kavukcuoglu,et al\.\(2016\)Wavenet: a generative model for raw audio\.arXiv preprint arXiv:1609\.0349912,pp\. 1\.Cited by:[§3\.2](https://arxiv.org/html/2608.28630#S3.SS2.p1.6)\.
- E\. Ekstedt and G\. Skantze \(2022\)Voice Activity Projection: Self\-supervised Learning of Turn\-taking Events\.InProc\. Interspeech 2022,pp\. 5190–5194\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2022-10955)Cited by:[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p2.1)\.
- C\. Fu, H\. Lin, X\. Wang, Y\. Zhang, Y\. Shen, X\. Liu, H\. Cao, Z\. Long, H\. Gao, K\. Li,et al\.\(2025\)Vita\-1\.5: towards gpt\-4o level real\-time vision and speech interaction\.arXiv preprint arXiv:2501\.01957\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- X\. Geng, Q\. Shao, H\. Xue, S\. Wang, H\. Xie, Z\. Guo, Y\. Zhao, G\. Li, W\. Tian, C\. Wang,et al\.\(2025\)Osum\-echat: enhancing end\-to\-end empathetic spoken chatbot via understanding\-driven spoken dialogue\.arXiv preprint arXiv:2508\.09600\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- Hugging Face \(2024\)Speech\-to\-Speech: an open\-source pipeline for real\-time voice assistants\.External Links:[Link](https://github.com/huggingface/speech-to-speech)Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- A\. A\. G\. Intelligence \(2025\)Amazon nova sonic: technical report and model card\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p2.1)\.
- K\. Lee, K\. Park, and D\. Kim \(2023\)Dailytalk: spoken dialogue dataset for conversational text\-to\-speech\.InICASSP 2023\-2023 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 1–5\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p1.2)\.
- G\. Li, C\. Wang, H\. Xue, S\. Wang, D\. Gao, Z\. Zhang, Y\. Lin, W\. Li, L\. Xiao, Z\. Fu,et al\.\(2025\)Easy turn: integrating acoustic and linguistic modalities for robust turn\-taking in full\-duplex spoken dialogue systems\.arXiv preprint arXiv:2509\.23938\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- B\. Liao, Y\. Xu, J\. Ou, K\. Yang, W\. Jian, P\. Wan, and D\. Zhang \(2025\)FlexDuo: a pluggable system for enabling full\-duplex capabilities in speech dialogue systems\.arXiv preprint arXiv:2502\.13472\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[§3\.2](https://arxiv.org/html/2608.28630#S3.SS2.p1.3)\.
- L\. Liao, L\. H\. Long, Y\. Ma, W\. Lei, and T\. Chua \(2021\)Dialogue state tracking with incremental reasoning\.Transactions of the Association for Computational Linguistics9,pp\. 557–569\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- L\. Liao, G\. H\. Yang, and C\. Shah \(2023\)Proactive conversational agents in the post\-chatgpt world\.InProceedings of the 46th international ACM SIGIR conference on research and development in information retrieval,pp\. 3452–3455\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- G\. Lin, C\. Chiang, and H\. Lee \(2024\)Advancing large language models to capture varied speaking styles and respond properly in spoken conversations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6626–6642\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p1.2)\.
- G\. Lin, S\. S\. Kuan, J\. Shi, K\. Chang, S\. Arora, S\. Watanabe, and H\. Lee \(2025a\)Full\-duplex\-bench\-v2: a multi\-turn evaluation framework for duplex dialogue systems with an automated examiner\.arXiv preprint arXiv:2510\.07838\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p3.1)\.
- G\. Lin, S\. S\. Kuan, Q\. Wang, J\. Lian, T\. Li, S\. Watanabe, and H\. Lee \(2025b\)Full\-duplex\-bench v1\. 5: evaluating overlap handling for full\-duplex speech models\.arXiv preprint arXiv:2507\.23159\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p3.1)\.
- G\. Lin, J\. Lian, T\. Li, Q\. Wang, G\. Anumanchipalli, A\. H\. Liu, and H\. Lee \(2025c\)Full\-duplex\-bench: a benchmark to evaluate full\-duplex spoken dialogue models on turn\-taking capabilities\.arXiv preprint arXiv:2503\.04721\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p4.1),[§2](https://arxiv.org/html/2608.28630#S2.p3.1)\.
- Z\. Liu, Y\. Duan, M\. Wang, P\. Feng, H\. Zhang, X\. Xing, Y\. Shan, H\. Zhu, Y\. Dai, C\. Lu,et al\.\(2025\)X\-talk: on the underestimated potential of modular speech\-to\-speech dialogue system\.arXiv preprint arXiv:2512\.18706\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- LiveKit \(2024\)LiveKit Agents: build real\-time multimodal ai applications\.External Links:[Link](https://github.com/livekit/agents)Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- Y\. Lu, Y\. Niu, S\. Hu, and H\. Wang \(2025\)CleanS2S: single\-file framework for proactive speech\-to\-speech interaction\.arXiv preprint arXiv:2506\.01268\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- Z\. Ma, Y\. Song, C\. Du, J\. Cong, Z\. Chen, Y\. Wang, Y\. Wang, and X\. Chen \(2025\)Language model can listen while speaking\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 24831–24839\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- L\. Mai and J\. Carson\-Berndsen \(2025\)Real\-time textless dialogue generation\.arXiv preprint arXiv:2501\.04877\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[§4\.3](https://arxiv.org/html/2608.28630#S4.SS3.p3.1),[Table 2](https://arxiv.org/html/2608.28630#S4.T2.1.1.9.1)\.
- K\. Mitsui, Y\. Hono, and K\. Sawada \(2023\)Towards human\-like spoken dialogue generation between ai agents from written dialogue\.arXiv preprint arXiv:2310\.01088\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- T\. A\. Nguyen, E\. Kharitonov, J\. Copet, Y\. Adi, W\. Hsu, A\. Elkahky, P\. Tomasello, R\. Algayres, B\. Sagot, A\. Mohamed,et al\.\(2023\)Generative spoken dialogue language modeling\.Transactions of the Association for Computational Linguistics11,pp\. 250–266\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- OpenAI \(2024\)GPT\-4o system card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p2.1),[§4\.1](https://arxiv.org/html/2608.28630#S4.SS1.p1.1),[Table 3](https://arxiv.org/html/2608.28630#S4.T3.20.16.19.1.1)\.
- Y\. Peng, Y\. Chao, D\. Ng, Y\. Ma, C\. Ni, B\. Ma, and E\. S\. Chng \(2025\)FD\-bench: a full\-duplex benchmarking pipeline designed for full duplex spoken dialogue systems\.arXiv preprint arXiv:2507\.19040\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p4.1),[§2](https://arxiv.org/html/2608.28630#S2.p3.1)\.
- J\. Qiu, Z\. Chen, L\. Yang, M\. Zhu, Z\. Liu, J\. Tan, W\. Zhao, R\. Murthy, R\. Ram, A\. Prabhakar,et al\.\(2026\)Building enterprise realtime voice agents from scratch: a technical tutorial\.arXiv preprint arXiv:2603\.05413\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever \(2023\)Robust speech recognition via large\-scale weak supervision\.InInternational conference on machine learning,pp\. 28492–28518\.Cited by:[§4\.1](https://arxiv.org/html/2608.28630#S4.SS1.p1.1)\.
- A\. Reece, G\. Cooney, P\. Bull, C\. Chung, B\. Dawson, C\. Fitzpatrick, T\. Glazer, D\. Knox, A\. Liebscher, and S\. Marin \(2023\)The candor corpus: insights from a large multimodal dataset of naturalistic conversation\.Science advances9\(13\),pp\. eadf3197\.Cited by:[§4\.5](https://arxiv.org/html/2608.28630#S4.SS5.p1.1)\.
- S\. Roller, E\. Dinan, N\. Goyal, D\. Ju, M\. Williamson, Y\. Liu, J\. Xu, M\. Ott, E\. M\. Smith, Y\. Boureau,et al\.\(2021\)Recipes for building an open\-domain chatbot\.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume,pp\. 300–325\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- R\. Roy, J\. Raiman, S\. Lee, T\. Ene, R\. Kirby, S\. Kim, J\. Kim, and B\. Catanzaro \(2026\)PersonaPlex: voice and role control for full duplex conversational speech models\.arXiv preprint arXiv:2602\.06053\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- G\. K\. Sehun Lee \(2025\)Behavior\-sd: behaviorally aware spoken dialogue generation with large language models\.InProceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics,External Links:[Link](https://aclanthology.org/2025.naacl-long.484/)Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§2](https://arxiv.org/html/2608.28630#S2.p2.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p1.2),[§4\.1](https://arxiv.org/html/2608.28630#S4.SS1.p1.1),[§4\.5](https://arxiv.org/html/2608.28630#S4.SS5.p1.1),[Table 4](https://arxiv.org/html/2608.28630#S4.T4.1.4.1),[Table 4](https://arxiv.org/html/2608.28630#S4.T4.1.7.1)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.\(2025\)Openai gpt\-5 system card\.arXiv preprint arXiv:2601\.03267\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p5.1)\.
- S\. Team \(2024\)Silero vad: pre\-trained enterprise\-grade voice activity detector \(vad\), number detector and language classifier\.GitHub\.Note:[https://github\.com/snakers4/silero\-vad](https://github.com/snakers4/silero-vad)Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p2.1)\.
- W\. Tu, G\. Yang, R\. Yan, W\. Chen, Z\. Ma, Y\. Kang, K\. Yu, X\. Chen, and Z\. Zheng \(2025\)UltraVoice: scaling fine\-grained style\-controlled speech conversations for spoken dialogue models\.arXiv preprint arXiv:2510\.22588\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- A\. Viswanath and H\. Buschmeier \(2026\)Desirability of proactive robots: a user study on spoken interaction initiation\.InProceedings of the 2026 ACM/IEEE International Conference on Human\-Robot Interaction\. ACM, Edinburgh, UK,Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- J\. Wang, L\. Chen, A\. Khare, A\. Raju, P\. Dheram, D\. He, M\. Wu, A\. Stolcke, and V\. Ravichandran \(2024a\)Turn\-taking and backchannel prediction with acoustic and large language model fusion\.InICASSP 2024\-2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 12121–12125\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[§4\.4](https://arxiv.org/html/2608.28630#S4.SS4.p2.6)\.
- P\. Wang, S\. Lu, Y\. Tang, S\. Yan, W\. Xia, and Y\. Xiong \(2024b\)A full\-duplex speech dialogue scheme based on large language model\.Advances in Neural Information Processing Systems37,pp\. 13372–13403\.Cited by:[§4\.3](https://arxiv.org/html/2608.28630#S4.SS3.p2.1),[Table 3](https://arxiv.org/html/2608.28630#S4.T3)\.
- P\. Wang, E\. Han, A\. C\.M\. Queiroz, C\. DeVeaux, and J\. N\. Bailenson \(2025\)Predicting and understanding turn\-taking behavior in open\-ended group activities in virtual reality\.Proc\. ACM Hum\.\-Comput\. Interact\.9\(7\)\.External Links:[Link](https://doi.org/10.1145/3757498),[Document](https://dx.doi.org/10.1145/3757498)Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p3.1),[§3\.3](https://arxiv.org/html/2608.28630#S3.SS3.p6.1)\.
- \[55\]Q\. Wang, Z\. Meng, W\. Cui, Y\. Zhang, P\. Wu, B\. Wu, I\. King, L\. Chen, and P\. ZhaoNTPP: generative speech language modeling for dual\-channel spoken dialogue via next\-token\-pair prediction\.InForty\-second International Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- X\. Wang, Y\. Li, C\. Fu, Y\. Shen, L\. Xie, K\. Li, X\. Sun, and L\. Ma \(2024c\)Freeze\-omni: a smart and low latency speech\-to\-speech dialogue model with frozen llm\.arXiv preprint arXiv:2411\.00774\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§1](https://arxiv.org/html/2608.28630#S1.p2.1),[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[Table 2](https://arxiv.org/html/2608.28630#S4.T2.1.1.4.1),[Table 3](https://arxiv.org/html/2608.28630#S4.T3.20.16.20.1.1)\.
- B\. Wu, C\. Yan, C\. Hu, C\. Yi, C\. Feng, F\. Tian, F\. Shen, G\. Yu, H\. Zhang, J\. Li,et al\.\(2025a\)Step\-audio 2 technical report\.arXiv preprint arXiv:2507\.16632\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1),[§1](https://arxiv.org/html/2608.28630#S1.p2.1),[Table 3](https://arxiv.org/html/2608.28630#S4.T3.20.16.24.1.1)\.
- C\. Wu, S\. C\. Hoi, R\. Socher, and C\. Xiong \(2020\)TOD\-bert: pre\-trained natural language understanding for task\-oriented dialogue\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 917–929\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- W\. Wu, W\. Guan, K\. Wang, P\. Chen, Z\. Zha, J\. Li, J\. Fang, L\. Li, and Q\. Hong \(2025b\)Phoenix\-vad: streaming semantic endpoint detection for full\-duplex speech interaction\.arXiv preprint arXiv:2509\.20410\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- Z\. Xie and C\. Wu \(2024\)Mini\-omni2: towards open\-source gpt\-4o with vision, speech and duplex capabilities\.arXiv preprint arXiv:2410\.11190\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- J\. Xu, Z\. Guo, J\. He, H\. Hu, T\. He, S\. Bai, K\. Chen, J\. Wang, Y\. Fan, K\. Dang,et al\.\(2025\)Qwen2\. 5\-omni technical report\.arXiv preprint arXiv:2503\.20215\.Cited by:[Table 3](https://arxiv.org/html/2608.28630#S4.T3.20.16.27.1.1)\.
- K\. Xu, Y\. Jia, K\. Huang, J\. Chen, W\. Li, K\. Liu, F\. Xie, X\. Tang, and Y\. Hu \(2026\)FireRedASR2S: a state\-of\-the\-art industrial\-grade all\-in\-one automatic speech recognition system\.arXiv preprint arXiv:2603\.10420\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[§4\.3](https://arxiv.org/html/2608.28630#S4.SS3.p3.1),[Table 2](https://arxiv.org/html/2608.28630#S4.T2.1.1.8.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p2.1),[§1](https://arxiv.org/html/2608.28630#S1.p4.1),[§4\.1](https://arxiv.org/html/2608.28630#S4.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.28630#S4.T2.1.1.6.1)\.
- Y\. Yao, T\. Yu, A\. Zhang, C\. Wang, J\. Cui, H\. Zhu, T\. Cai, H\. Li, W\. Zhao, Z\. He,et al\.\(2024\)MiniCPM\-v: a gpt\-4v level mllm on your phone\.arXiv preprint arXiv:2408\.01800\.Cited by:[Table 3](https://arxiv.org/html/2608.28630#S4.T3.20.16.22.1)\.
- C\. Ye, L\. Liao, S\. Liu, and T\. Chua \(2022\)Reflecting on experiences for response generation\.InProceedings of the 30th ACM International Conference on Multimedia,pp\. 5265–5273\.Cited by:[§1](https://arxiv.org/html/2608.28630#S1.p1.1)\.
- F\. Yu, S\. Zhang, Y\. Fu, L\. Xie, S\. Zheng, Z\. Du, W\. Huang, P\. Guo, Z\. Yan, B\. Ma,et al\.\(2022\)M2MeT: the icassp 2022 multi\-channel multi\-party meeting transcription challenge\.InICASSP 2022\-2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 6167–6171\.Cited by:[§4\.5](https://arxiv.org/html/2608.28630#S4.SS5.p1.1)\.
- T\. Yu, Z\. Wang, C\. Wang, F\. Huang, W\. Ma, Z\. He, T\. Cai, W\. Chen, Y\. Huang, Y\. Zhao,et al\.\(2025\)Minicpm\-v 4\.5: cooking efficient mllms via architecture, data, and training recipe\.arXiv preprint arXiv:2509\.18154\.Cited by:[Table 3](https://arxiv.org/html/2608.28630#S4.T3.20.16.22.1)\.
- W\. Yu, S\. Wang, X\. Yang, X\. Chen, X\. Tian, J\. Zhang, G\. Sun, L\. Lu, Y\. Wang, and C\. Zhang \(2024\)Salmonn\-omni: a codec\-free llm for full\-duplex speech understanding and generation\.arXiv preprint arXiv:2411\.18138\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1),[§3\.2](https://arxiv.org/html/2608.28630#S3.SS2.p1.3)\.
- N\. Zargham, L\. Reicherts, M\. Bonfert, S\. T\. Voelkel, J\. Schoening, R\. Malaka, and Y\. Rogers \(2022\)Understanding circumstances for desirable proactive behaviour of voice assistants: the proactivity dilemma\.InProceedings of the 4th conference on conversational user interfaces,pp\. 1–14\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- A\. Zeng, Z\. Du, M\. Liu, K\. Wang, S\. Jiang, L\. Zhao, Y\. Dong, and J\. Tang \(2024\)Glm\-4\-voice: towards intelligent and human\-like end\-to\-end spoken chatbot\.arXiv preprint arXiv:2412\.02612\.Cited by:[§3\.2](https://arxiv.org/html/2608.28630#S3.SS2.p1.6),[Table 2](https://arxiv.org/html/2608.28630#S4.T2.1.1.5.1)\.
- C\. Zhang, K\. Yang, S\. Hu, Z\. Wang, G\. Li, Y\. Sun, C\. Zhang, Z\. Zhang, A\. Liu, S\. Zhu,et al\.\(2024\)Proagent: building proactive cooperative agents with large language models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 17591–17599\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
- H\. Zhang, W\. Li, R\. Chen, V\. Kothapally, M\. Yu, and D\. Yu \(2025a\)LLM\-enhanced dialogue management for full\-duplex spoken dialogue systems\.arXiv preprint arXiv:2502\.14145\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- Q\. Zhang, L\. Cheng, C\. Deng, Q\. Chen, W\. Wang, S\. Zheng, J\. Liu, H\. Yu, C\. Tan, Z\. Du,et al\.\(2025b\)Omniflatten: an end\-to\-end gpt model for seamless voice conversation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14570–14580\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p1.1)\.
- Z\. Zhou, Q\. Zhang, L\. Luo, J\. Liu, and R\. Zhou \(2025\)Open\-source full\-duplex conversational datasets for natural and interactive speech synthesis\.arXiv preprint arXiv:2509\.04093\.Cited by:[§2](https://arxiv.org/html/2608.28630#S2.p2.1)\.
## Appendix APrompt for Turn Prediction
You are an expert turn controller responsible for turn\-taking decisions\. Your task is to classify the user’s current audio into one of three actions\. Output Format:¡wait¿ OR ¡backchannel¿ OR ¡turn taking¿ Rule: Check if user’s INTENTION is complete 1\.¡wait¿\- Intention incomplete or unclear \- User is still formulating thoughts or mid\-sentence\. \- Examples: “I was thinking…” / “Could you…” \(trailing off\) 2\.¡backchannel¿\- Intention complete \(minimal feedback only\) \- User shares a simple fact, update, or statement\. \- No substantive reply or detailed engagement is needed\. \- Examples: “I’m done with my homework\.” / “It’s raining outside\.” 3\.¡turn\-taking¿\- Intention complete \(needs substantive reply\) \- User asks a question, requests help, or expects discussion\. \- Examples: “What time is it?” / “I had a terrible day today\.” Decision Logic \(Distinguish ¡backchannel¿ vs ¡turn taking¿\): \- Simple fact/statement→\\rightarrow¡backchannel¿ \- Needs answer/correction/emotional engagement→\\rightarrow¡turn taking¿ Key Criteria:Intention complete \+ \(needs response/action\) =¡turn taking¿
## Appendix BPrompt with VAD
You are an English real\-time conversational assistant managing turn\-taking\. Input: \- History: Previous conversation \(for context only\) \- Current user audio: Your focus for responding Response Rules: When current audio is ¡incomplete¿ \(user still speaking\): Choose ONE action based on the audio content: 1\.¡wait¿\- No response Example: User says “I was thinking about…” \(unclear intent\) Output: ¡wait¿ 2\.¡backchannel¿\- Brief acknowledgment Example: User says “So I went to the store and…” Output: ¡backchannel¿ Uh\-huh 3\.¡response¿\- Provide information Example: User says “What’s the capital of…” Output: ¡response¿ The capital of France is Paris\. When current audio is ¡complete¿ \(user finished\): Always respond: Output: ¡response¿ \[your complete answer\] Key Principles: \- Be concise for backchannels \(1\-3 words\) \- Be complete for responses \- Default to ¡wait¿ only if truly unclear
## Appendix CPrompt with Specific Judge Module
Now you are an English real\-time conversational assistant managing turn\-taking in conversations\. Your Role: You need to decide when and how to respond based on the current conversational state\. You have four possible actions: 1\. Take the Turn \(Full Response\) \- When: User has finished speaking OR there’s a natural opportunity \- Action: Provide a complete, substantive response \- Example: “The capital of France is Paris\. It’s known for…” 2\. Interrupt Turn\-Taking \(ITT\) \- When: User is still speaking, but you can predict their intent \- Action: Politely interrupt and provide a helpful response \- Example: User says “I was wondering about the…”→\\rightarrowYou respond “The capital of France?” 3\. BackChannel \- When: User is speaking and needs encouragement to continue \- Action: Give brief acknowledgment \(1\-3 words\) WITHOUT taking turn \- Examples: “Uh\-huh”, “I see”, “Right”, “Mm\-hmm”, “Go on” 4\. Wait \- When: No response is needed at this moment \- Action: Stay silent and wait for more information \- Output: ¡wait¿ Key Principles: \- Be context\-aware: Consider history and user’s speech completeness \- Be natural: Choose the most appropriate action for smooth flow \- Be concise: Keep backchannels short, make full responses informative
## Appendix DPrompt for Half\-duplex Model Itself
You are a real\-time English conversation assistant\. Output:¡wait¿ OR ¡backchannel¿ \[text\] OR ¡response¿ \[text\] Rule: Check if user’s INTENTION is complete ¡wait¿ \- Intention incomplete/unclear \- Don’t know what user wants yet \- Need more info to understand \- Examples: “I was thinking…” / “What’s the…” / “Can you…” ¡backchannel¿ \- Intention complete \(minimal acknowledgment\) \- User shares simple fact/update \- Brief, doesn’t invite conversation \- Examples: “I went shopping” \-¿ ¡backchannel¿ Nice ¡response¿ \- Intention complete \(needs substantive reply\) Use when: \- Direct question asked \- Factual error to correct \- Request for explanation/help \- Conversational engagement needed \- User invites discussion or expects your thoughts Examples: Questions: \- “What time is it?” \-¿ ¡response¿ It’s 3 PM\. \- “How does this work?” \-¿ ¡response¿ \[explanation\] Errors/corrections: \- “Paris is capital of Germany” \-¿ ¡response¿ Actually, Paris is France’s capital\. \- “Vaccines cause autism” \-¿ ¡response¿ That’s a common misconception\. Requests: \- “Can you explain X?” \-¿ ¡response¿ \[explanation\] \- “Help me understand this” \-¿ ¡response¿ \[help\] Conversational engagement: \- “I just got back from an amazing trip to Japan” \-¿ ¡response¿ Oh wow, … \- “I’m thinking about changing careers” \-¿ ¡response¿ That’s a big decision\. … \- “I had the worst day today” \-¿ ¡response¿ I’m sorry to hear that\. … \- “Guess what happened to me” \-¿ ¡response¿ What happened? Distinguish ¡backchannel¿ vs ¡response¿: \- “I made dinner” \-¿ ¡backchannel¿ Nice \(simple fact\) \- “I tried making sushi for the first time” \-¿ ¡response¿ Oh that’s cool\! … \- “It’s raining” \-¿ ¡backchannel¿ Yeah \(weather comment\) \- “It’s been raining for three days straight…” \-¿ ¡response¿ I can imagine that… Key: Intention complete \+ \(needs answer/correction/conversation\) = ¡response¿
## Appendix EPrompt for Interaction Evaluation
\[Evaluation Protocol for Full\-duplex Speech Interaction\] INPUT DATA:\{dialogue\_history\} CORE SYSTEM PRINCIPLES: \- Operational Mode: Real\-time streaming with low\-latency constraints\. \- Interjection Logic: The assistant is programmed for proactive responses\. Interruption is deemed valid if: \(i\) the user’s intent is sufficiently discernible for a complete reply, or \(ii\) immediate corrective feedback is required for factual or linguistic errors\. ASSESSMENT OBJECTIVES: 1\. Temporal Precision: Examine the final turn to determine if the assistant’s decision to preempt the user’s speech was justified\. Note that truncated user input results from the system’s cut\-off\. 2\. Semantic Alignment: For valid interruptions, evaluate whether the provided response maintains contextual coherence\. SCORING CRITERIA: Metric A \[Timing\]: Score 1 if the interjection was timely and well\-placed; otherwise 0\. Metric B \[Content\]: Score 1 if the response is contextually relevant and accurate; otherwise 0\. REQUIRED OUTPUT STRUCTURE: ”’ Analysis ¡detailed\_rationale\_for\_timing\_and\_coherence¿ Judge ¡timing\_binary\_score¿, ¡content\_binary\_score¿ ”’ EXECUTION: AnalysisSimilar Articles
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
DuplexGen introduces a method for adaptively synthesizing human-AI turn-taking dialogues, addressing the challenge of natural interaction timing in conversational AI.
A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism
This paper introduces TPA (Think, Plan, Ask), a proactive multi-agent dialogue framework using LLMs to systematically surface latent social language disorder traits in autism by selecting clinically grounded questioning strategies. It achieves 82.1% trait coverage, outperforming real clinical dialogues by clinicians.
Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue
This paper proposes a decoupled data approach to improve turn-taking in full-duplex dialogue by learning from real spoken dialogues while using text for semantics, leveraging a neural finite state machine framework to enhance naturalness and preserve semantic capabilities.
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
This paper introduces X2-Turn, a frame-synchronous dual-head model that jointly performs streaming ASR and turn state prediction on shared representations, improving turn-taking accuracy and latency in spoken dialogue systems.
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
TurnNat is a likelihood-based framework for automatically evaluating turn-taking naturalness in dyadic spoken dialogue, using a causal turn-taking prediction model trained on natural conversations to measure timing atypicality via negative log-likelihood.