FD-VAD:面向流式全双工语音的语义端点检测

arXiv cs.CL 论文

摘要

FD-VAD 提出一种无需 ASR、流式运行的语义端点检测方法,将因果音频窗口直接映射为 Continue/Stop 决策,适用于全双工语音智能体。该方法结合冻结的语音编码器与参数高效适配的语言模型,在 TurnBench 上取得了最先进(state-of-the-art)的 EOT 召回率。

arXiv:2609.35791v1 Announce Type: new Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ (at FP<=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.
查看原文
查看缓存全文

缓存时间: 2026/09/30 09:46

# FD-VAD: SEMANTIC ENDPOINT DETECTION FOR STREAMING FULL-DUPLEX SPEECH
Source: [https://arxiv.org/html/2609.35791](https://arxiv.org/html/2609.35791)
###### Abstract

Natural turn\-taking in full\-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent\. Acoustic voice activity detection lacks this semantic information, while cascaded ASR\-based endpointing introduces transcription dependence and additional processing stages\. We formulate semantic endpoint detection as a*causal audio\-language reasoning task*and introduceFD\-VAD, an ASR\-free streaming endpointer that maps bounded causal audio windows directly toContinue/Stopdecisions\. FD\-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter\-efficiently adapted language model, using a last\-chunk training objective for streaming inference\. We further introduce confidence\-gated endpoint commitment to control interruption versus delay and boundary\-focused hard\-negative sampling to improve decisions around ambiguous turn boundaries\. Across in\-domain and conversational evaluations, FD\-VAD outperforms strong streaming and non\-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set0\.8530\.853\(atFP≤0\.10\\mathrm\{FP\}\\leq 0\.10\) in a zero\-shot setting\. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking\.

###### Index Terms:

semantic endpoint detection, Voice Activity Detection, turn taking, full duplex, streaming speech

††address:University of Maryland, College Park## 1Introduction

Natural spoken interaction depends not only on*what*is said, but also on*when*each participant chooses to speak\. Humans continuously infer whether another speaker has completed a thought, is pausing to plan what to say next, or intends to continue after a hesitation\. This ability is fundamental to fluid turn\-taking: responding before a speaker has finished is disruptive, while waiting too long after a completed request makes an interaction feel unnatural\. The problem becomes particularly important for full\-duplex voice agents, which continuously listen and speak, and therefore cannot rely on the rigid listen–respond alternation used by conventional voice assistants\[[3](https://arxiv.org/html/2609.35791#bib.bib8),[22](https://arxiv.org/html/2609.35791#bib.bib9)\]\. A full\-duplex agent must instead maintain an evolving estimate of whether the user still holds the conversational floor\.

Voice activity detection \(VAD\) is commonly used in spoken systems to locate acoustic speech boundaries\. However, acoustic VAD is often insufficient because a period of silence does not necessarily indicate that a conversational turn has ended\. A speaker may pause while searching for a word, producing a disfluency, or planning the remainder of a request; conversely, a short utterance may already express a complete intent\. The decision is therefore inherently*semantic*: the system must determine whether the observed speech constitutes a complete conversational act, rather than merely detect whether acoustic energy has disappeared\. We refer to this problem as*semantic endpoint detection*\. In a streaming setting, it is especially challenging because the decision must be updated from partial audio without access to future speech\.

![Refer to caption](https://arxiv.org/html/2609.35791v1/figures/intro.png)Figure 1:Streaming semantic endpoint detection task, where 0 and 1 denoteContinueandStop, respectively\.FD\-VADuses linguistic as well as acoustic cues for end\-of\-turn detection\.Existing endpointing approaches expose a trade\-off between semantic reasoning, streaming operation, and architectural simplicity\. Acoustic methods such as Silero VAD\[[17](https://arxiv.org/html/2609.35791#bib.bib1)\]are lightweight and naturally streaming but cannot determine whether the preceding speech is semantically complete\. Cascaded approaches first transcribe speech and then use a language model or text classifier to determine whether the utterance is complete\[[18](https://arxiv.org/html/2609.35791#bib.bib3),[7](https://arxiv.org/html/2609.35791#bib.bib4)\]\. These methods provide access to lexical semantics, but introduce an explicit ASR dependency into the turn\-taking loop and inherit transcription errors and additional processing stages\. Semantic approaches such as Smart Turn\[[16](https://arxiv.org/html/2609.35791#bib.bib5)\], Easy Turn\[[10](https://arxiv.org/html/2609.35791#bib.bib18)\], FastTurn\[[21](https://arxiv.org/html/2609.35791#bib.bib19)\], Next\-Turn\[[19](https://arxiv.org/html/2609.35791#bib.bib20)\], and SpeculativeETD\[[12](https://arxiv.org/html/2609.35791#bib.bib21)\]incorporate semantics through utterance\-level classification, transcript/CTC pathways, timing supervision, or silence\-triggered semantic inference\. At the other end, full\-duplex systems such as Moshi and Freeze\-Omni\[[3](https://arxiv.org/html/2609.35791#bib.bib8),[22](https://arxiv.org/html/2609.35791#bib.bib9)\]learn turn\-taking jointly with response generation\. These methods improve semantic awareness, but typically introduce intermediate dependencies, triggering mechanisms or couple to particular dialogue architectures\.

We argue that semantic endpointing need not require these intermediate mechanisms and instead formulate it as a*streaming audio\-language reasoning task*: given only the speech observed so far, determine whether the speaker has expressed a complete conversational intent\. We introduceFD\-VAD, an ASR\-free semantic endpoint detector for streaming full\-duplex speech\. FD\-VAD maps short causal waveform windows directly toContinue/Stopdecisions\. We utilize a frozen self\-supervised speech encoder extracts acoustic and linguistic representations from the incoming waveform, while a lightweight modality adapter maps them into the embedding space of a small language model that performs the semantic endpoint decision\. Rather than processing the complete utterance, FD\-VAD repeatedly evaluates a bounded sliding window and supervises only the state at its most recent chunk\. The resulting model reasons about semantic completeness while preserving causal, fixed\-cost streaming inference\. Importantly, FD\-VAD requires neither transcript generation nor a separate acoustic\-to\-semantic cascade, and its endpointing policy remains independent of the downstream voice agents\.

Additionally, we address two properties particularly important for real\-time deployment\. First, endpoint errors are asymmetric: an earlyStopcan interrupt a user, whereas a conservative decision primarily adds delay\. We therefore introduce aconfidence\-gated decision mechanismthat explicitly controls this trade\-off through the predicted stop probability and the number of consecutive decisions required before committing an endpoint\. Second, ambiguous examples are concentrated near the true turn boundary rather than uniformly throughout an utterance\. To better model these cases, we introduceboundary\-focused hard\-negative sampling, which emphasizes difficult boundary regions during training while leaving the inference architecture unchanged\.Our main contributions are:

- •We reformulate semantic endpoint detection as acausal audio\-language reasoning task, where the model determines from partial speech whether the user should retain or yield the conversational floor\.
- •We introduceFD\-VAD111Code and model release:https://fd\-vad\.github\.io/, an ASR\-free audio\-to\-LLM architecture that performs streaming semantic endpoint detection for full\-duplex speech interaction\.
- •We proposeconfidence\-gated endpointingandboundary\-focused hard\-negative samplingto control the interruption–latency trade\-off and improve decisions near true endpoints\.
- •FD\-VAD achieves thehighest EOT recall on TurnBench devamong compared systems \(0\.8530\.853atFP≤0\.10\\mathrm\{FP\}\\leq 0\.10\) in a zero\-shot setting, and matches a strong non\-streaming turn classifier at the utterance level \(0\.9650\.965vs\.0\.9660\.966\)\.

## 2Related Work

Transcription\- and segmentation\-dependent endpointing\.Acoustic VAD detects speech boundaries from energy, spectral, or learned cues\[[13](https://arxiv.org/html/2609.35791#bib.bib2),[17](https://arxiv.org/html/2609.35791#bib.bib1)\], but provides no measure of semantic completion\. Cascaded systems such as TEN\[[18](https://arxiv.org/html/2609.35791#bib.bib3)\]obtain semantics through an ASR transcript, while Smart Turn\[[16](https://arxiv.org/html/2609.35791#bib.bib5)\]predicts completeness directly from an acoustically segmented utterance\. FD\-VAD instead operates continuously on causal audio windows without requiring either transcription or a completed speech segment\.

Streaming methods with intermediate signals\.Recent methods reduce this dependence while retaining auxiliary mechanisms\. Easy Turn\[[10](https://arxiv.org/html/2609.35791#bib.bib18)\]jointly predicts ASR and turn state, FastTurn\[[21](https://arxiv.org/html/2609.35791#bib.bib19)\]relies on streaming CTC representations, SpeculativeETD\[[12](https://arxiv.org/html/2609.35791#bib.bib21)\]triggers semantic inference after detecting silence, and Next\-Turn\[[19](https://arxiv.org/html/2609.35791#bib.bib20)\]learns endpoint timing through time\-to\-next\-speech supervision\. FD\-VAD removes these intermediate pathways and directly maps observed speech to a semanticContinue/Stopstate\.

Dialogue\-coupled and speech\-language models\.Full\-duplex models such as Moshi and Freeze\-Omni\[[3](https://arxiv.org/html/2609.35791#bib.bib8),[22](https://arxiv.org/html/2609.35791#bib.bib9)\]learn turn\-taking jointly with response generation, while TurnFSM\[[11](https://arxiv.org/html/2609.35791#bib.bib22)\]embeds semantic VAD within an LLM\-based dialogue\-state controller\. FD\-VAD instead remains independent of the downstream voice agent\. Architecturally, it follows the audio encoder–adapter–LLM paradigm used in speech\-language models\[[5](https://arxiv.org/html/2609.35791#bib.bib10),[2](https://arxiv.org/html/2609.35791#bib.bib11),[23](https://arxiv.org/html/2609.35791#bib.bib12)\], specializing it for causal endpoint prediction using bounded audio context and a last\-chunk objective\.

![Refer to caption](https://arxiv.org/html/2609.35791v1/figures/main.png)Figure 2:FD\-VAD architecture\.A causal audio window of incoming speech is encoded by frozen WavLM, projected into the embedding space of a LoRA\-adapted Qwen2\.5\-0\.5B language model forContinue/Stopendpoint prediction\. Training supervises only the final chunk of each window\.
## 3Method

Architecture: Figure[2](https://arxiv.org/html/2609.35791#S2.F2)shows FD\-VAD, an autoregressive LLM\-based architecture for semantic endpoint detection, containing three components:\(i\) WavLM Audio Encoder\.A frozen WavLM\-base\-plus encoder\[[1](https://arxiv.org/html/2609.35791#bib.bib13)\]maps raw waveform to frame\-level speech features\.\(ii\)Audio\-Text Modality Adapter\.To bridge the modality gap between audio and text, we utilize a modality adapter composed of two linear layers with a ReLU activation function that temporally downsamples and projects audio features into the embedding space of an LLM\.\(iii\) LLM Backbone\.We use Qwen2\.5\-0\.5B\-Instruct\[[25](https://arxiv.org/html/2609.35791#bib.bib14)\]as the LLM backbone, which takes the adapted audio features concatenated with text prompt embeddings as input to predict the semantic completeness state of the input speech segment\. LetAIA^\{I\}denote the WavLM feature sequence\. The adapter groups everykkconsecutive frames and appliesAP=W2​ReLU​\(W1​AI\+b1\)\+b2A^\{P\}=W\_\{2\}\\,\\mathrm\{ReLU\}\(W\_\{1\}A^\{I\}\+b\_\{1\}\)\+b\_\{2\}\. We usek=4k\{=\}4, so the 50 Hz speech representation becomes a 12\.5 Hz sequence of audio tokens\. The projected audio embeddings are concatenated with a short instruction prompt, and the LLM predicts a single state token,ContinueorStop, followed by<eos\>\. The speech encoder and base LLM remain frozen\. We jointly optimize the modality adapter and low\-rank LoRA parameters\[[8](https://arxiv.org/html/2609.35791#bib.bib15)\], for 10\.3M trainable parameters \(1\.7% of the total parameter count\)\. This keeps the endpoint detector modular and parameter\-efficient while reusing strong pretrained speech and language representations\.

Causal Sliding\-window Inference: To enable streaming inference, we employ a sliding window training strategy that makes predictions using only the audio within each window, reducing dependence on the entire input sequence, enabling incremental chunk\-wise predictions, and offering latency advantages\. FD\-VAD operates on a 10 ms waveform slicing grid\. For chunk indexcc, the model advances by a stride ofSf=320S\_\{f\}\{=\}320ms and retains at mostWf=2\.56W\_\{f\}\{=\}2\.56s of left context\. Thus, each decision is made every 320 ms from a 2\.56 s causal window which is encoded by WavLM into approximately 128 feature frames at 50 Hz\. The adapter downsamples these byk=4k\{=\}4to approximately 32 audio tokens at 12\.5 Hz before the LLM\.

Last\-chunk Training Objective: For each window, supervision is applied only to its terminal chunk\. Letyc∈\{0,1\}y\_\{c\}\\in\\\{0,1\\\}denoteContinueandStop, respectively, and define the target sequenceYc=\{yc,⟨eos⟩\}Y\_\{c\}\{=\}\\\{y\_\{c\},\\langle\\mathrm\{eos\}\\rangle\\\}\. We minimize the autoregressive cross\-entropy

ℒ\(θ\)=−∑t=1\|Yc\|logpθ\(Yc\(t\)∣Yc\(<t\),AcP,TP\),\\mathcal\{L\}\(\\theta\)=\-\\sum\_\{t=1\}^\{\|Y\_\{c\}\|\}\\log p\_\{\\theta\}\\\!\\left\(Y\_\{c\}^\{\(t\)\}\\mid Y\_\{c\}^\{\(<t\)\},A\_\{c\}^\{P\},T^\{P\}\\right\),\(1\)whereAcPA\_\{c\}^\{P\}denotes the adapted audio embeddings for the current window andTPT^\{P\}is the instruction prompt\.

Confidence Gating Commitment: A single erroneousStopcan prematurely interrupt a user\. At inference time we therefore optionally commitStoponly whenP⁡\(Stop\)≥τP\(\\textsc\{Stop\}\)\\geq\\tauforKKconsecutive decisions\. Increasingτ\\tauorKKreduces false interruptions but also delays the commit\. We treat this confidence gate as an operating\-point controller layered on top of the base model\.

Boundary\-focused Hard\-negative Sampling: Training errors concentrate near the true endpoint, where a one\-chunk timing shift changes the label\. We consider boundary\-focused hard\-negative sampling that oversamples windows whose terminal chunk lies within±2\\pm 2chunks of the endpoint\. This leaves inference unchanged but reallocates training mass towards ambiguous boundary cases\.

## 4Data and Experimental Setup

Training Data:We use the English subset of smart\-turn\-v3\.1\[[15](https://arxiv.org/html/2609.35791#bib.bib7)\], a 16 kHz conversational endpointing corpus with native*complete*and*incomplete*utterance labels\. We follow the official train/test partition\. The training set contains63,38663\{,\}386clips \(31,33831\{,\}338complete /32,04832\{,\}048incomplete;156156h\), comprising42,52942\{,\}529human and20,85720\{,\}857TTS recordings\. For complete utterances, theStoptransition is placed at the acoustic speech offset \(−40\-40dB relative to peak\), with subsequent chunks labeledStop; incomplete utterances remainContinuethroughout\. Incomplete examples are corpus\-native rather than artificially truncated complete utterances\.

Training Hyperparameters:We freeze the WavLM encoder and LLM backbone and train only the modality adapter and LoRA parameters \(r=16r\{=\}16\) for one epoch using AdamW with learning rate2×10−42\\times 10^\{\-4\}, cosine decay, and3%3\\%warmup\. Training uses an effective batch size of6464in bf16 on two A100 GPUs\.

Evaluation Data and Protocols:The held\-out in\-domain test set contains1,0001\{,\}000clips \(500500complete /500500incomplete;695695human /305305TTS;2\.52\.5h\), with median duration9\.19\.1s\. For synthetic\-to\-human transfer, we additionally train on TTS speech only and evaluate on unseen human recordings\. We evaluate conversational transfer zero\-shot on the TurnBench dev set\[[9](https://arxiv.org/html/2609.35791#bib.bib16)\], which contains real two\-channel human dialogue and uses event\-anchored end\-of\-turn \(EOT\) evaluation\. For this setting, Silero VAD\[[17](https://arxiv.org/html/2609.35791#bib.bib1)\]provides an acoustic gate and permits an EOT commit after at least11s of post\-speech silence; FD\-VAD receives no TurnBench training\. We report evaluation results as mean of three random seed runs\.

Evaluation Metrics:At the chunk level, we report precision, recall,F1F\_\{1\}, and accuracy withStopas the positive class\. We separately report the*chunk false\-stop rate*, defined as the fraction of chunks from incomplete utterances incorrectly predicted asStop\. At the utterance level, we report endpoint accuracy separately for complete and incomplete queries, together with95%95\\%bootstrap confidence intervals from20002000resamples\. We also report strict sentence accuracy \(Sent\-C/Sent\-I\), which requires all chunk decisions in an utterance to be correct, and near\-boundaryF1F\_\{1\}, computed over predictions within±2\\pm 2chunks of the true endpoint\. For confidence\-gated endpointing, the*false\-interruption rate*is the fraction of incomplete utterances for which the system commits any prematureStop, and endpointing delay isL=tcommit−tEOTL=t\_\{\\mathrm\{commit\}\}\-t\_\{\\mathrm\{EOT\}\}, whereL<0L<0denotes a commit before the annotated endpoint\. We report this delay jointly with false\-interruption rate\. On TurnBench, we follow its EOT protocol and report recall atFP≤0\.10\\mathrm\{FP\}\\leq 0\.10together with median commit latency\.

## 5Results

Table 1:Chunk\-level semantic endpoint detection on the held\-out smart\-turn\-v3\.1\[[14](https://arxiv.org/html/2609.35791#bib.bib6)\]test set \(1,000 utterances; 500 complete and 500 incomplete\)\.Stopis the positive class; false\-stop is the fraction of incomplete\-speech chunks incorrectly predicted asStop\. FD\-VAD outperforms streaming baselines with lowest false\-stop rate\.Table 2:Utterance\-level endpoint classification on the held\-out smart\-turn\-v3\.1\[[14](https://arxiv.org/html/2609.35791#bib.bib6)\]test set with 95% bootstrap CI over 2K samples\. FD\-VAD shows the strongest balanced performance across complete and incomplete turns in streaming mode \(Str\), comparable to the strongest non\-streaming semantic classifier\.In\-Domain Chunk and Utterance Performance:Tables[1](https://arxiv.org/html/2609.35791#S5.T1)and[2](https://arxiv.org/html/2609.35791#S5.T2)show that FD\-VAD reliably detects semantic endpoints while avoiding premature stops\. At the chunk level, FD\-VAD reachesF1=0\.832F\_\{1\}\{=\}0\.832with only a0\.6%0\.6\\%false\-stop rate on incomplete utterances\. Raw chunk accuracy is less informative becauseStopaccounts for only4\.8%4\.8\\%of the labels; indeed, the always\-Continuebaseline already achieves0\.9530\.953accuracy\. More importantly, the balanced utterance\-level evaluation shows that FD\-VAD performs consistently on both complete and incomplete turns \(0\.9720\.972/0\.9690\.969\), unlike acoustic VAD, whose performance drops sharply on incomplete speech \(0\.998→0\.6800\.998\\rightarrow 0\.680\)\. This gap confirms that acoustic termination alone is insufficient for semantic endpointing\. FD\-VAD also substantially outperforms the Whisper\-based semantic cascades and performs relatively better compared to Smart Turn \(0\.9650\.965vs\.0\.9760\.976overall\), despite operating streaming rather than on a completed utterances\.

CategorySystemRecallFPLat\. \(ms\)↓\\downarrowTurn\-takingVAP\[[4](https://arxiv.org/html/2609.35791#bib.bib17)\]0\.8410\.045463ESPnet turn\-taking0\.8360\.074895ESPnet TT \(per\-ch\.\)0\.6400\.100846Smart Turn\[[16](https://arxiv.org/html/2609.35791#bib.bib5)\]0\.7540\.1001010Moshi\[[3](https://arxiv.org/html/2609.35791#bib.bib8)\]0\.2120\.066771SoulX\-Duplug\[[24](https://arxiv.org/html/2609.35791#bib.bib23)\]0\.1940\.154†89SemanticKyutai semantic VAD0\.8030\.1001024Gemini Live0\.6650\.0471197OpenAI semantic VAD0\.3100\.037763OpenAI server VAD0\.9330\.563†281AcousticWavLM\-large anchor0\.8250\.1001017Mimi endpointer0\.7590\.047742WavLM\-base causal0\.4850\.100709WavLM\-large causal0\.4720\.100683RMS energy VAD0\.5950\.547†−98\-98OursFD\-VAD0\.8530\.0971019*Oracle \(annotator\)*1\.0000\.0000

Table 3:TurnBench\[[9](https://arxiv.org/html/2609.35791#bib.bib16)\]dev end\-of\-turn \(EOT\) performance under benchmark false\-positive constraint \(FP≤0\.10\\mathrm\{FP\}\\leq 0\.10\)\. FD\-VAD is evaluated zero\-shot and achieves the highest recall among qualifying systems\.†\\daggerdenotes systems that exceed the false\-positive budget\.Zero\-Shot Conversational EOT Transfer:The in\-domain evaluation uses isolated utterances, whereas practical endpointing operates within continuous dialogue\. To apply FD\-VAD to continuous audio, Silero VAD\[[17](https://arxiv.org/html/2609.35791#bib.bib1)\]provides an acoustic gate that determines*when*an endpoint is plausible, while FD\-VAD determines*whether*the preceding speech is semantically complete\. Table[3](https://arxiv.org/html/2609.35791#S5.T3)shows that under the TurnBench false\-positive constraint \(FP≤0\.10\\mathrm\{FP\}\\leq 0\.10\),FD\-VAD achieves the highest EOT recallamong qualifying systems \(0\.8530\.853\), despite being evaluated zero\-shot\. However, FD\-VAD’s10191019ms median commit latency is higher than VAP’s463463ms\. Fig\.[3](https://arxiv.org/html/2609.35791#S5.F3)further illustrates this trade\-off: FD\-VAD lies at the upper edge of the recall–false\-positive operating region, but not on the best recall–latency frontier\. The large separation between gated and ungated FD\-VAD additionally shows that semantic endpoint prediction benefits substantially from acoustic gating in continuous dialogue\.

0\.05\.10\.150\.20\.20\.40\.40\.60\.60\.80\.8VAPESPnetFD\-VADWavLMKyutaiungatedFP≤0\.10\\leq 0\.10false\-positive rateEOT recall\(a\) Recall vs\. false positives040080012000\.20\.20\.40\.40\.60\.60\.80\.8VAPFD\-VADWavLMungatedSoulXmedian commit latency \(ms\)EOT recall\(b\) Recall vs\. delayFigure 3:TurnBench\[[9](https://arxiv.org/html/2609.35791#bib.bib16)\]dev operating characteristics\. Zero\-shot gated FD\-VAD achieves the highest EOT recall within the benchmark false\-positive constraint \(FP≤0\.10\\mathrm\{FP\}\\leq 0\.10\)\. The latency view highlights the complementary trade\-off between endpoint reliability and commit speed; selected competitors are annotated and remaining baselines are shown in gray\.Streaming Latency Analysis:A key advantage of FD\-VAD over non\-streaming ASR\-based endpointing is that semantic inference proceeds concurrently with the user’s speech rather than beginning only after the utterance completion\. LetTspeechT\_\{\\mathrm\{speech\}\}denote the utterance duration,TstrideT\_\{\\mathrm\{stride\}\}is FD\-VAD’s inference stride, andtFDt\_\{\\mathrm\{FD\}\}is it’s per\-window processing time\. An offline ASR\-based cascade can produce its semantic endpoint decision only after the utterance has been observed, giving a endpoint decision time ofTspeech\+tASR\+tsemanticT\_\{\\mathrm\{speech\}\}\+t\_\{\\mathrm\{ASR\}\}\+t\_\{\\mathrm\{semantic\}\}, while FD\-VAD continuously updates its endpoint state everyTstrideT\_\{\\mathrm\{stride\}\}\(=320=320ms\) while speech is arriving\. On a single NVIDIA A100, FD\-VAD requires only45\.845\.8ms per2\.562\.56s causal window \(46\.846\.8ms p90\), and this cost remains constant with utterance duration\. At our operating point \(τ=0\.9\\tau\{=\}0\.9,K=1K\{=\}1\), the resulting median endpoint delay is190190ms with a2\.0%2\.0\\%false\-interruption rate\. Thus, longer utterances increase the latency of non\-streaming semantic endpointing, whereas FD\-VAD amortizes semantic reasoning through turns\.

Table 4:Ablation results on a matched 8k\-clip subset\. All variants use the same FD\-VAD architecture and differ only in the listed settings\. Near\-b\. isF1F\_\{1\}within±2\\pm 2chunks of the endpoint; Sent\-C is strict all\-chunk accuracy on complete utterances\. We study robustness to training data by training FD\-VAD on synthetic\-only data that transfers well per chunk but degrades at utterance\-level\.Ablations and Stability: Table[4](https://arxiv.org/html/2609.35791#S5.T4)reveals two useful design trends\. First, reducing the stride from320320ms to160160ms slightly raises aggregate chunk accuracy but substantially degrades complete\-turn and near\-boundaryF1F\_\{1\}\. Second, jointly adapting the modality interface and explicitly emphasizing boundary examples are important: freezing the adapter degrades all endpoint\-oriented metrics, whereas boundary\-focused sampling yields the strongest complete\-turn, near\-boundary, and strict sentence performance\. The three\-seed variation is small \(0\.9804±0\.00050\.9804\\pm 0\.0005chunk accuracy\), indicating that these trends are not driven by a particular initialization\.

Robustness to Training Data Source: FD\-VAD model is trained on a mixture of human and synthetic speech\. To assess its dependence on human training audio, we additionally train FD\-VAD using only synthetic speech and evaluates it on held\-out human recordings\. The synthetic\-only model retains similar per\-chunk performance on human speech, with a1\.11\.1\-point reduction relative to mixed\-data training, although the stricter all\-chunk utterance metrics degrade more substantially, emphasizing the importance of human training data for consistent endpoint decisions across complete utterances\.

## 6Conclusion

We introduced FD\-VAD, a streaming semantic endpoint detector that reformulates turn completion as causal audio\-language reasoning over partial speech\. By mapping bounded audio context directly toContinue/Stopdecisions, FD\-VAD removes the need for intermediate transcription while remaining modular with respect to the downstream voice agent\. Our evaluations show that FD\-VAD transfers effectively to continuous real\-time dialogue, where semantic reasoning complements conventional acoustic gating\. We establish that direct audio\-to\-language reasoning provides a simple and portable foundation for turn\-taking in full\-duplex spoken agents\.

## References

- \[1\]S\. Chen, C\. Wang, Z\. Chen, Y\. Wu, S\. Liu, Z\. Chen, J\. Li, N\. Kanda, T\. Yoshioka, X\. Xiao, J\. Wu, L\. Zhou, S\. Ren, Y\. Qian, Y\. Qian, J\. Wu, M\. Zeng, X\. Yu, and F\. Wei\(2022\)WavLM: large\-scale self\-supervised pre\-training for full stack speech processing\.IEEE Journal of Selected Topics in Signal Processing16\(6\),pp\.1505–1518\.External Links:[Document](https://dx.doi.org/10.1109/JSTSP.2022.3188113)Cited by:[§3](https://arxiv.org/html/2609.35791#S3.p1.1),[Table 1](https://arxiv.org/html/2609.35791#S5.T1.1.1.3.1),[Table 1](https://arxiv.org/html/2609.35791#S5.T1.1.1.4.1),[Table 1](https://arxiv.org/html/2609.35791#S5.T1.1.1.5.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.5.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.6.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.7.1)\.
- \[2\]W\. Chen, Z\. Ma, R\. Yan, Y\. Liang, X\. Li, R\. Xu, Z\. Niu, Y\. Zhu, Y\. Yang, Z\. Liu, K\. Yu, Y\. Hu, J\. Li, Y\. Lu, S\. Liu, and X\. Chen\(2025\)SLAM\-Omni: timbre\-controllable voice interaction system with single\-stage training\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\.2262–2282\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.115)Cited by:[§2](https://arxiv.org/html/2609.35791#S2.p3.1)\.
- \[3\]A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour\(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.Technical reportKyutai\.External Links:2410\.00037Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p1.1),[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p3.1),[Table 3](https://arxiv.org/html/2609.35791#S5.T3.1.1.6.1)\.
- \[4\]E\. Ekstedt and G\. Skantze\(2022\)Voice activity projection: self\-supervised learning of turn\-taking events\.InProc\. Interspeech,pp\.5190–5194\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2022-10955)Cited by:[Table 1](https://arxiv.org/html/2609.35791#S5.T1.1.1.7.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.4.2),[Table 3](https://arxiv.org/html/2609.35791#S5.T3.1.1.2.2)\.
- \[5\]Q\. Fang, S\. Guo, Y\. Zhou, Z\. Ma, S\. Zhang, and Y\. Feng\(2025\)LLaMA\-Omni: seamless speech interaction with large language models\.InProc\. International Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.35791#S2.p3.1)\.
- \[6\]K\. Fu, R\. Wen, A\. Lin, S\. Qin, R\. Gan, H\. Wang, and Q\. Wang\(2026\)X2\-turn: frame\-synchronous dual\-head modeling for joint streaming asr and turn state prediction\.arXiv preprint arXiv:2608\.10878\.Cited by:[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.9.1)\.
- \[7\]Z\. Gao, S\. Zhang, I\. McLoughlin, and Z\. Yan\(2022\)Paraformer: fast and accurate parallel transformer for non\-autoregressive end\-to\-end speech recognition\.InProc\. Interspeech,pp\.2063–2067\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2022-9996)Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1)\.
- \[8\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2022\)LoRA: low\-rank adaptation of large language models\.InProc\. International Conference on Learning Representations \(ICLR\),Cited by:[§3](https://arxiv.org/html/2609.35791#S3.p1.1)\.
- \[9\]F\. Jiang, R\. Sanabria, S\. Deshmukh, B\. Veluri, S\. M\. V\. Williams, E\. K\. Suen, G\. Lee, K\. Y\. Choi, T\. Umeki, R\. Kubo, S\. Udupa, C\. Huang, S\. S\. Kuan, Z\. Tao, S\. Krishna, S\. E\. Eskimez, Y\. Tsao, H\. Lee, and S\. Watanabe\(2026\)TurnBench: a multi\-domain benchmark for turn\-taking dynamics in spoken dialogue\.External Links:2608\.25218Cited by:[§4](https://arxiv.org/html/2609.35791#S4.p3.1),[Figure 3](https://arxiv.org/html/2609.35791#S5.F3.1),[Figure 3](https://arxiv.org/html/2609.35791#S5.F3.2),[Table 3](https://arxiv.org/html/2609.35791#S5.T3.2),[Table 3](https://arxiv.org/html/2609.35791#S5.T3.3)\.
- \[10\]G\. Li, C\. Wang, H\. Xue, S\. Wang, D\. Gao, Z\. Zhang, Y\. Lin, W\. Li, L\. Xiao, Z\. Fu, and L\. Xie\(2025\)Easy turn: integrating acoustic and linguistic modalities for robust turn\-taking in full\-duplex spoken dialogue systems\.arXiv preprint arXiv:2509\.23938\.Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p2.1)\.
- \[11\]Z\. Lin, T\. Du, Q\. Huang, Z\. Zhang, N\. Zheng, L\. Xiao, Y\. Lu, J\. Chen, and Z\. Wu\(2026\)TurnFSM for full\-duplex dialogue system: internalizing state\-machine logic for streaming semantic voice activity detection and utterance\-level rejection\.arXiv preprint arXiv:2609\.04240\.Cited by:[§2](https://arxiv.org/html/2609.35791#S2.p3.1)\.
- \[12\]H\. Ok, S\. Yoo, and J\. Lee\(2026\)Speculative end\-turn detector for efficient speech chatbot assistant\.InProc\. 64th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\.45184–45197\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.2094)Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p2.1)\.
- \[13\]Y\. Park and S\. Lee\(2012\)Voice activity detection using global speech absence probability based on teager energy for speech enhancement\.IEICE Transactions on Information and SystemsE95\-D\(10\),pp\.2568–2571\.External Links:[Document](https://dx.doi.org/10.1587/transinf.E95.D.2568)Cited by:[§2](https://arxiv.org/html/2609.35791#S2.p1.1)\.
- \[14\]Pipecat AI\(2025\)Smart Turn Data v3\.1\-test\.Note:Hugging Face DatasetTest split; accessed Sep\. 1, 2026External Links:[Link](https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.1-test)Cited by:[Table 1](https://arxiv.org/html/2609.35791#S5.T1.2),[Table 1](https://arxiv.org/html/2609.35791#S5.T1.3),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.2),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.3)\.
- \[15\]Pipecat AI\(2025\)Smart Turn Data v3\.1\-train\.Note:Hugging Face DatasetTrain split; accessed Sep\. 1, 2026External Links:[Link](https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.1-train)Cited by:[§4](https://arxiv.org/html/2609.35791#S4.p1.1)\.
- \[16\]Pipecat AI\(2025\)Smart Turn v3: native\-audio semantic turn detection\.Note:Open\-source model and softwareAvailable: https://github\.com/pipecat\-ai/smart\-turnCited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p1.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.13.1),[Table 3](https://arxiv.org/html/2609.35791#S5.T3.1.1.5.1)\.
- \[17\]Silero Team\(2024\)Silero VAD: pre\-trained enterprise\-grade voice activity detector \(vad\), number detector and language classifier\.GitHub\.Note:GitHub repositoryAvailable: https://github\.com/snakers4/silero\-vadCited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p1.1),[§4](https://arxiv.org/html/2609.35791#S4.p3.1),[Table 1](https://arxiv.org/html/2609.35791#S5.T1.1.1.6.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.3.1),[§5](https://arxiv.org/html/2609.35791#S5.p2.1)\.
- \[18\]TEN Team\(2025\)TEN Turn Detection: turn detection for full\-duplex dialogue communication\.Note:GitHub repositoryAvailable: https://github\.com/TEN\-framework/ten\-turn\-detectionCited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p1.1),[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.11.1)\.
- \[19\]T\. Tsoi, J\. Deng, Y\. Zhu, H\. Q\. Dang, T\. Cao, N\. Kuzmin, T\. Zhong, and S\. Lui\(2026\)Next\-turn: duration\-aware streaming endpoint detection via time\-to\-next\-speech\-onset prediction\.arXiv preprint arXiv:2606\.18094\.Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p2.1)\.
- \[20\]Ultravox\.ai Team\(2025\)UltraVAD: Context\-Aware Audio\-Native Neural Endpointing Model\.Hugging Face\.Note:https://huggingface\.co/fixie\-ai/ultraVADCited by:[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.12.1)\.
- \[21\]C\. Wang, H\. Xue, C\. He, J\. Hu, S\. Wang, B\. Wu, Y\. Ji, J\. Zheng, R\. Chen, Z\. Zhu, and L\. Xie\(2026\)FastTurn: unifying acoustic and streaming semantic cues for low\-latency and robust turn detection\.arXiv preprint arXiv:2604\.01897\.Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p2.1)\.
- \[22\]X\. Wang, Y\. Li, C\. Fu, Y\. Zhang, Y\. Shen, L\. Xie, K\. Li, X\. Sun, and L\. Ma\(2025\)Freeze\-omni: a smart and low latency speech\-to\-speech dialogue model with frozen LLM\.InProc\. International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.267,pp\.63345–63354\.Cited by:[§1](https://arxiv.org/html/2609.35791#S1.p1.1),[§1](https://arxiv.org/html/2609.35791#S1.p3.1),[§2](https://arxiv.org/html/2609.35791#S2.p3.1)\.
- \[23\]Y\. Wang, H\. Liu, Z\. Cheng, R\. Wu, Q\. Gu, Y\. Wang, and Y\. Wang\(2025\)VocalNet: speech LLMs with multi\-token prediction for faster and high\-quality generation\.InProc\. Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\.19584–19601\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.989)Cited by:[§2](https://arxiv.org/html/2609.35791#S2.p3.1)\.
- \[24\]R\. Yan, W\. Chen, Z\. Liu, Z\. Ma, H\. Lin, H\. Wen, H\. Xie, J\. Wu, Y\. Liang, Y\. Zhao, P\. Feng, J\. Qian, H\. Meng, Y\. Dai, S\. Yin, M\. Tao, L\. Xie, K\. Yu, X\. Wang, and X\. Chen\(2026\)SoulX\-duplug: plug\-and\-play streaming state prediction module for realtime full\-duplex speech conversation\.External Links:2603\.14877,[Link](https://arxiv.org/abs/2603.14877)Cited by:[Table 2](https://arxiv.org/html/2609.35791#S5.T2.1.1.8.1),[Table 3](https://arxiv.org/html/2609.35791#S5.T3.1.1.7.1)\.
- \[25\]A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. Qiu\(2024\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§3](https://arxiv.org/html/2609.35791#S3.p1.1)\.

相似文章

FastVideo/FastVideo-FastH3-四步-预览版-v1-VSA-无数据

Hugging Face Models Trending

FastVideo 发布 FastH3 预览版 v1,这是一个 AI 模型检查点,能够从文本生成同步视频和音频,使用四次 Transformer 前向传播,通过无数据 DMD2 和 VSA-H3 在 90% 稀疏度下训练。