沉默或重叠并非失败:全双工口语对话模型基于意图条件的话轮转换评估
摘要
论文介绍了TACT,一个用于评估全双工口语对话模型中话轮转换的基准,它使用基于意图条件的连续评分来替代二元规则,展示了与人类判断更好的一致性。
arXiv:2609.27372v1 Announce Type: new
Abstract: Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.
查看缓存全文
缓存时间: 2026/09/24 09:20
# Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models
Source: [https://arxiv.org/html/2609.27372](https://arxiv.org/html/2609.27372)
###### Abstract
Benchmarks for full\-duplex spoken dialogue models score turn\-taking with binary fixed\-window rules that reward immediate response or silence by completeness of the prior turn\. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker’s latent intent, identifiable only from that speaker’s behavior\. We introduce TACT, a benchmark of 9,728 episodes and 73\.2 hours from five dyadic corpora; each episode carries dialogue history, a per\-speaker memory profile, and an annotator\-derived posterior over six intent classes\. Scoring replaces binary windows with a strictly proper threshold\-weighted continuous ranked probability score whose weights are intent\-conditioned timing kernels fitted to human floor\-transfer\-offset distributions, proving boundedness, consistency, and binary reduction\. Across eleven systems the best model reaches 0\.47 against a human topline of 0\.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0\.81 versus 0\.46 for binary metrics\.
###### Index Terms:
full\-duplex spoken dialogue models, turn\-taking, benchmark evaluation, intent recognition, theory of mind, proper scoring rules
## IIntroduction
Full\-duplex spoken dialogue models \(SDMs\) listen and speak simultaneously, deciding continuously when to take the floor, yield it, backchannel, or remain silent\[[1](https://arxiv.org/html/2609.27372#bib.bib28),[2](https://arxiv.org/html/2609.27372#bib.bib29),[3](https://arxiv.org/html/2609.27372#bib.bib30),[4](https://arxiv.org/html/2609.27372#bib.bib26)\]\. Evaluation of these decisions has relied on binary event detection inside fixed look\-ahead windows: Full\-Duplex\-Bench scores whether a model takes over within a window after a pause or turn end and how fast\[[5](https://arxiv.org/html/2609.27372#bib.bib41)\], its successor extending the same logic to overlap\[[6](https://arxiv.org/html/2609.27372#bib.bib42)\]\. Under such metrics, a response inside the window is correct, silence after completion is a miss, and incoming speech before completion is an intrusion, whatever the interlocutor was doing with that silence or inviting with that incompleteness\.
This paper begins from a two\-sided observation well established in conversation analysis but absent from machine evaluation: silence after a completed turn is not intrinsically a failure, and overlapping talk before completion is not intrinsically a violation\. Human floor transfer offsets concentrate around 200 ms\[[7](https://arxiv.org/html/2609.27372#bib.bib4),[8](https://arxiv.org/html/2609.27372#bib.bib5),[9](https://arxiv.org/html/2609.27372#bib.bib7)\], yet long gaps are routinely tolerated when the prior turn was rhetorical, floor\-holding, or abandoned, or when the recipient is expected to reflect\[[10](https://arxiv.org/html/2609.27372#bib.bib1),[11](https://arxiv.org/html/2609.27372#bib.bib2),[12](https://arxiv.org/html/2609.27372#bib.bib10)\]; symmetrically, backchannels, collaborative completions, and urgent clarifications are routinely launched in overlap, before any turn end exists\[[13](https://arxiv.org/html/2609.27372#bib.bib3),[14](https://arxiv.org/html/2609.27372#bib.bib12)\]\. The value of a response offset therefore depends on the speaker’s latent intent: the same trailing\-off question followed by 1\.5 s of silence can demand an immediate answer from one speaker and forbid one from another, and the same mid\-clause incoming can be cooperative or disruptive\. Crucially, intent is often identifiable only from that speaker’s behavior across interactions, because speakers differ systematically in pause tolerance, backchannel solicitation, and floor\-holding style\[[15](https://arxiv.org/html/2609.27372#bib.bib56),[16](https://arxiv.org/html/2609.27372#bib.bib11)\]; an evaluation blind to intent and memory rewards the degenerate policy of answering everything quickly and never coming in early, on which current SDMs have converged\.
We operationalize this observation in*TACT*\(Theory\-of\-mind Aware Conversational Timing\), which supersedes fixed\-window evaluation in three ways\. First, every episode carries long multi\-turn history, a per\-speaker memory profile assembled from other episodes of the same speaker, and a six\-class latent intentzzannotated by at least three humans and fused with a calibrated large\-language\-model judge into a posteriorq^\(z∣e\)\\hat\{q\}\(z\\mid e\)\. Second, scoring is continuous, two\-sided, and strictly proper: a threshold\-weighted continuous ranked probability score whose nonnegative weight functionswz\(t\)w\_\{z\}\(t\)are derived from intent\-conditioned asymmetric timing kernelsuz\(t\)u\_\{z\}\(t\)over the offset domain𝒯\\mathcal\{T\}, so that negative offsets are anticipatory onsets overlapping the ongoing turn, fitted to human floor\-transfer\-offset distributions from CANDOR\[[15](https://arxiv.org/html/2609.27372#bib.bib56)\], SSSD\[[17](https://arxiv.org/html/2609.27372#bib.bib57)\], and otoSpeech\[[18](https://arxiv.org/html/2609.27372#bib.bib61)\]dyads, with analogous proper scores for length and prosody\-weighted overlap; we prove boundedness, strict propriety with the human conditional law as unique maximizer, reduction of Full\-Duplex\-Bench to a degenerate weight limit, and identifiability of the intent mixture\. Third, the benchmark is realistic in scale and provenance: 9,728 episodes and 73\.2 hours from five ecologically diverse public dyadic corpora, plus synthetic probes for rare conditions\.
Evaluating eleven frontier systems, we find that models appearing adequate under binary metrics collapse under intent conditioning\. The best attains a composite of 0\.47 against a human topline of 0\.86, the ranking reorders substantially relative to a reproduced Full\-Duplex\-Bench composite \(Spearman rank correlation 0\.55\), and all systems exhibit an*over\-eagerness*pathology whose rigidity is two\-sided: near\-uniform fast responding regardless of intent, paired with an inability to launch cooperative early onsets where humans routinely overlap\. A memory\-swap ablation shows that model behavior is essentially invariant to the speaker profile while held\-out human behavior is not, and the TACT composite correlates with held\-out human judgments at Spearmanρ=0\.81\\rho=0\.81versus0\.460\.46for the binary metrics it replaces\.
## IIRelated Work
Turn allocation was formalized by Sacks, Schegloff, and Jefferson\[[10](https://arxiv.org/html/2609.27372#bib.bib1)\]; smooth transfers with modal gaps near 200 ms are a cross\-linguistic universal\[[7](https://arxiv.org/html/2609.27372#bib.bib4)\]requiring predictive planning of turn ends\[[8](https://arxiv.org/html/2609.27372#bib.bib5),[19](https://arxiv.org/html/2609.27372#bib.bib8)\]from lexico\-syntactic and intonational completion cues\[[20](https://arxiv.org/html/2609.27372#bib.bib9),[16](https://arxiv.org/html/2609.27372#bib.bib11)\]\. The floor\-transfer\-offset \(FTO\) distribution, long studied on telephone corpora\[[21](https://arxiv.org/html/2609.27372#bib.bib60)\], is heavy\-tailed and well fitted by ex\-Gaussian forms tracking sequence organization and preference\[[22](https://arxiv.org/html/2609.27372#bib.bib6),[12](https://arxiv.org/html/2609.27372#bib.bib10),[9](https://arxiv.org/html/2609.27372#bib.bib7)\]; reflective responses are delayed, rhetorical or floor\-holding turns license long silence\[[11](https://arxiv.org/html/2609.27372#bib.bib2)\], and backchannels occupy an overlap\-permissive regime cued by prosody\[[13](https://arxiv.org/html/2609.27372#bib.bib3),[14](https://arxiv.org/html/2609.27372#bib.bib12)\]: timing is a conditional distribution given intent, the structure our scoring adopts, with a taxonomy motivated by\[[23](https://arxiv.org/html/2609.27372#bib.bib13),[24](https://arxiv.org/html/2609.27372#bib.bib14)\]\. Computational treatments evolved from finite\-state floor control\[[25](https://arxiv.org/html/2609.27372#bib.bib18),[26](https://arxiv.org/html/2609.27372#bib.bib20)\]to self\-supervised prediction: TurnGPT predicts turn shifts from lexical context\[[27](https://arxiv.org/html/2609.27372#bib.bib21)\], Voice Activity Projection \(VAP\) predicts joint future voice activity from stereo audio\[[28](https://arxiv.org/html/2609.27372#bib.bib22)\]with real\-time and multilingual extensions\[[29](https://arxiv.org/html/2609.27372#bib.bib23),[30](https://arxiv.org/html/2609.27372#bib.bib24)\], and production endpointers train jointly with recognition\[[31](https://arxiv.org/html/2609.27372#bib.bib25),[32](https://arxiv.org/html/2609.27372#bib.bib19)\]; these predictors supply machinery we repurpose for scoring\.
Full\-duplex behavior appeared commercially in XiaoIce and task\-oriented stacks\[[33](https://arxiv.org/html/2609.27372#bib.bib27),[4](https://arxiv.org/html/2609.27372#bib.bib26)\]; dGSLM introduced end\-to\-end two\-channel modeling\[[1](https://arxiv.org/html/2609.27372#bib.bib28)\], followed by Moshi’s parallel speech\-text streams\[[2](https://arxiv.org/html/2609.27372#bib.bib29)\], time\-multiplexing in SyncLLM\[[3](https://arxiv.org/html/2609.27372#bib.bib30)\], duplex fine\-tuning of text LLMs\[[34](https://arxiv.org/html/2609.27372#bib.bib31)\], listen\-while\-speaking objectives\[[35](https://arxiv.org/html/2609.27372#bib.bib32)\], a family of speech\-native assistants\[[36](https://arxiv.org/html/2609.27372#bib.bib40),[37](https://arxiv.org/html/2609.27372#bib.bib33),[38](https://arxiv.org/html/2609.27372#bib.bib39)\], persona\-controllable models such as PersonaPlex\[[39](https://arxiv.org/html/2609.27372#bib.bib37)\], and proprietary realtime stacks from GPT\-4o Realtime to gpt\-realtime\-2 and Gemini 3\.1 Flash Live\[[40](https://arxiv.org/html/2609.27372#bib.bib34),[41](https://arxiv.org/html/2609.27372#bib.bib35),[42](https://arxiv.org/html/2609.27372#bib.bib36)\]\. Benchmarks for such systems largely score semantic quality\[[43](https://arxiv.org/html/2609.27372#bib.bib48),[44](https://arxiv.org/html/2609.27372#bib.bib47)\]\. Interaction\-level evaluation began with Full\-Duplex\-Bench, which defined the four categories we retain via takeover rates and latencies inside fixed windows\[[5](https://arxiv.org/html/2609.27372#bib.bib41)\], and has since grown well beyond fixed local windows: v1\.5 probes four overlap scenarios with prosodic\-adaptation metrics\[[6](https://arxiv.org/html/2609.27372#bib.bib42)\], FD\-Bench scaled simulated interruptions\[[45](https://arxiv.org/html/2609.27372#bib.bib44)\], v2 adds a multi\-turn automated examiner\[[46](https://arxiv.org/html/2609.27372#bib.bib43)\], v3 uses entirely real human audio annotated for five disfluency categories under multi\-step tool use\[[47](https://arxiv.org/html/2609.27372#bib.bib45)\], and Talking Turns trains a supervised judge on human corpora\[[48](https://arxiv.org/html/2609.27372#bib.bib46)\]\. These increasingly cover overlap, multi\-turn structure, and real disfluent audio, yet each defines the appropriate behavior as a function of observed acoustic and dialogue context rather than the speaker’s latent intent\. TACT does not claim to be the first multi\-turn or overlap\-aware benchmark; its contribution is to condition the target on a latent, speaker\-specific, theory\-of\-mind\-and\-memory intent posterior and score continuous timing with a strictly proper, asymmetry\-aware rule\[[49](https://arxiv.org/html/2609.27372#bib.bib52),[50](https://arxiv.org/html/2609.27372#bib.bib51)\]rather than binary windows, calibrating the judge following\[[51](https://arxiv.org/html/2609.27372#bib.bib49),[52](https://arxiv.org/html/2609.27372#bib.bib50)\]\. Finally, attributing latent mental states is the classical theory\-of\-mind task\[[53](https://arxiv.org/html/2609.27372#bib.bib15)\], at which text probes show large models brittle\[[54](https://arxiv.org/html/2609.27372#bib.bib16),[55](https://arxiv.org/html/2609.27372#bib.bib17)\]; persona and memory help text dialogue\[[56](https://arxiv.org/html/2609.27372#bib.bib63),[57](https://arxiv.org/html/2609.27372#bib.bib64)\], yet no spoken\-dialogue evaluation tests whether a system adapts its timing to a specific speaker\.
## IIIThe TACT Benchmark
### III\-AEpisode Construction
A TACT episodeeeconsists of a dual\-channel audio contextAeA\_\{e\}ending at a decision point, a time\-aligned transcript historyHeH\_\{e\}averaging 3\.2 minutes and 11\.4 turns, a behavioral memory profileMs\(e\)M\_\{s\(e\)\}of the focal speakers\(e\)s\(e\), and annotations; episodes and profiles are mined from five public dyadic corpora and ratings from at least three annotators fuse with a calibrated LLM judge into the intent posterior conditioning the scoring\. The decision point is the end of a focal\-speaker turn \(for pause handling, an intra\-turn silence of at least 600 ms; for interruption, a region where the system holds the floor\)\. The system receives the focal\-speaker channel as streaming input with a rendered profile; its output is recorded over\[−Ta,Th\]\[\-T\_\{a\},T\_\{h\}\]with pre\-horizonTa=2T\_\{a\}=2s \(so anticipatory overlap onsets are observable\) and horizonTh=5T\_\{h\}=5s, and each episode is run three times, keeping the median\-scoring run\.
Episodes are mined from five public corpora chosen for ecological diversity\. CANDOR \(1,656 dyadic video calls with demographics and surveys\[[15](https://arxiv.org/html/2609.27372#bib.bib56)\]\) anchors the fairness audit \(2,560 episodes, 19\.8 h\); SSSD \(727 hours of crowdsourced spontaneous English dyads\[[17](https://arxiv.org/html/2609.27372#bib.bib57)\]\) supplies scale and speaker diversity \(2,304 episodes, 17\.1 h\); otoSpeech\-full\-duplex\-280h, raw 48 kHz stereo FLAC with one channel per speaker and a cleaned 141\-hour companion we use\[[18](https://arxiv.org/html/2609.27372#bib.bib61),[58](https://arxiv.org/html/2609.27372#bib.bib62)\], preserves the overlap ground truth mixed\-channel corpora destroy, essential for fitting anticipatory laws \(1,792 episodes, 13\.9 h\); Seamless Interaction\[[59](https://arxiv.org/html/2609.27372#bib.bib58)\]seeds interruption episodes \(1,536 episodes, 11\.5 h\); and the AMI individual\-headset subset, with characterized ASR baselines \(Whisper near 16–17% WER on AMI\-IHM\[[60](https://arxiv.org/html/2609.27372#bib.bib38)\]\), supplies the sole multiparty contrast mined as dyadic exchanges \(768 episodes, 5\.4 h\)\[[61](https://arxiv.org/html/2609.27372#bib.bib59)\]\. Synthetic TTS probes \(768 episodes, 5\.5 h\) cover conditions sparse in found data\. In total TACT comprises 9,728 episodes and 73\.2 hours over the four categories of\[[5](https://arxiv.org/html/2609.27372#bib.bib41)\]\(2,720 pause handling, 2,432 backchanneling, 2,592 smooth turn\-taking, 1,984 user interruption\), each crossed with the intent taxonomy\.
### III\-BSpeaker Memory Profiles
For every focal speaker we assemble a profileMsM\_\{s\}from episodes of the same speaker disjoint from the test episode \(mean 14\.6 minutes of evidence\), summarizing the FTO distribution, intra\-turn pause statistics, abstention tolerance, backchannel solicitation and production rates, articulation rate, and floor\-holding devices, rendered as structured features and a prompt synopsis\. Profiles are the only channel through which speaker\-specific expectations can flow; the swap ablation of Section[VI](https://arxiv.org/html/2609.27372#S6)exploits this\.
### III\-CIntent Taxonomy and Annotation Protocol
Each decision point is annotated with a latent intentzzover𝒵=\{\\mathcal\{Z\}=\\\{ans,ref,rhe,hold,bck,abn\}\\\}: immediate\-answer\-seeking, reflective, rhetorical, floor\-holding, cooperative\-incoming\-inviting, and turn\-abandonment, distilled from dialogue\-act inventories\[[23](https://arxiv.org/html/2609.27372#bib.bib13)\]and the timing literature\[[12](https://arxiv.org/html/2609.27372#bib.bib10),[11](https://arxiv.org/html/2609.27372#bib.bib2)\]so that classes induce distinct human timing on both sides of the turn end;bckcovers continuers, collaborative completions, and urgent clarifications, whose normative onset lies in overlap\. Three to five annotators \(median three\), shown full context, profile, and human continuation, label each episode, rate silence acceptability, and mark acceptable onset windows on both sides\. Agreement is Krippendorffα=0\.73\\alpha=0\.73for intent \(nominal\),0\.840\.84for onset windows, and0\.790\.79for overlap acceptability\[[62](https://arxiv.org/html/2609.27372#bib.bib55)\]; the empirical intent distribution is 31\.6%ans, 18\.2%hold, 16\.1%ref, 15\.3%bck, 9\.6%abn, 9\.2%rhe\. A temperature\-scaled LLM judge is fused in \(Section[IV](https://arxiv.org/html/2609.27372#S4)\); calibration cuts its expected calibration error from 8\.4% to 2\.1%, and the fusion predicts a held\-out annotator’s label better than either source alone\.
Because timing norms vary across populations\[[7](https://arxiv.org/html/2609.27372#bib.bib4)\], TACT is stratified by gender, age band, first\-language region, signal\-to\-noise ratio, and articulation\-rate tercile from CANDOR and Seamless metadata\[[15](https://arxiv.org/html/2609.27372#bib.bib56),[59](https://arxiv.org/html/2609.27372#bib.bib58)\], and kernels are fitted per slice where data permit, so the metric does not encode one population’s norm as universal\.
## IVIntent\-Conditioned Continuous Scoring
### IV\-ANotation and Intent Posterior
Leteedenote an episode with decision point at time origint=0t=0, focal speakers\(e\)s\(e\), historyHeH\_\{e\}, and profileMs\(e\)M\_\{s\(e\)\}\. The evaluated system’s action isye∈𝒯∪\{∅\}y\_\{e\}\\in\\mathcal\{T\}\\cup\\\{\\varnothing\\\}on the bounded offset domain𝒯=\[−Ta,Th\]\\mathcal\{T\}=\[\-T\_\{a\},T\_\{h\}\], with pre\-horizonTa=2T\_\{a\}=2s and evaluation horizonTh=5T\_\{h\}=5s:yey\_\{e\}is the onset offsetttof its first responsive vocalization within𝒯\\mathcal\{T\}\(negative offsets being anticipatory onsets in overlap with the ongoing turn\), or the no\-response\-within\-horizon outcome∅\\varnothing\(the atom we score explicitly\)\. WithKeK\_\{e\}annotator labelsz\(1\),…,z\(Ke\)z^\{\(1\)\},\\dots,z^\{\(K\_\{e\}\)\}, the empirical posterior isq^A\(z∣e\)=Ke−1∑k𝟏\[z\(k\)=z\]\\hat\{q\}\_\{\\mathrm\{A\}\}\(z\\mid e\)=K\_\{e\}^\{\-1\}\\sum\_\{k\}\\mathbf\{1\}\[z^\{\(k\)\}=z\], and the calibrated posterior fuses it with an LLM judgepJp\_\{\\mathrm\{J\}\}\[[51](https://arxiv.org/html/2609.27372#bib.bib49),[52](https://arxiv.org/html/2609.27372#bib.bib50)\], temperature\-scaled byT⋆T^\{\\star\}fitted on the development split,
q^\(z∣e\)=\(1−λ\)q^A\(z∣e\)\+λpJ\(z∣e\)1/T⋆∑z′pJ\(z′∣e\)1/T⋆,\\hat\{q\}\(z\\mid e\)\\;=\\;\(1\-\\lambda\)\\,\\hat\{q\}\_\{\\mathrm\{A\}\}\(z\\mid e\)\\;\+\\;\\lambda\\,\\frac\{p\_\{\\mathrm\{J\}\}\(z\\mid e\)^\{1/T^\{\\star\}\}\}\{\\sum\_\{z^\{\\prime\}\}p\_\{\\mathrm\{J\}\}\(z^\{\\prime\}\\mid e\)^\{1/T^\{\\star\}\}\},\(1\)withλ=0\.3\\lambda=0\.3chosen by held\-out annotator log\-likelihood\. Intuitively, \([1](https://arxiv.org/html/2609.27372#S4.E1)\) forms a Bayesian consensus belief over speaker intent by blending majority human annotations with a calibrated judge\. The posterior is a theory\-of\-mind estimate of the speaker’s state\[[53](https://arxiv.org/html/2609.27372#bib.bib15),[54](https://arxiv.org/html/2609.27372#bib.bib16)\]:HeH\_\{e\}alone leaves intent ambiguous in 38% of episodes, and conditioning onMs\(e\)M\_\{s\(e\)\}resolves over half of these\.
### IV\-BReference Laws and Weights for All Six Intents
For each of the six intentsz∈𝒵z\\in\\mathcal\{Z\}we specify a human reference lawhz=\(1−βz\)fz\+βzδ∅h\_\{z\}=\(1\-\\beta\_\{z\}\)\\,f\_\{z\}\+\\beta\_\{z\}\\,\\delta\_\{\\varnothing\}on𝒯=\[−Ta,Th\]\\mathcal\{T\}=\[\-T\_\{a\},T\_\{h\}\]with the no\-response atom∅\\varnothing\. Here,βz∈\[0,1\]\\beta\_\{z\}\\in\[0,1\]is the empirical human no\-response probability, estimated directly as the mean annotator silence acceptability on training dyads, whilefzf\_\{z\}is the conditional continuous FTO density fitted on responsive turns disjoint from test\. Becauseβz\\beta\_\{z\}andfzf\_\{z\}capture the discrete no\-response atom and response\-conditional timing respectively, they represent distinct, non\-conflicting components ofhzh\_\{z\}, with inter\-annotator label disagreements retained in the posteriorq^\(z∣e\)\\hat\{q\}\(z\\mid e\)\. Following\[[22](https://arxiv.org/html/2609.27372#bib.bib6)\], the responsive intentsans,bck, andabnuse the ex\-Gaussian density with parameters\(μz,σz,τz\)\(\\mu\_\{z\},\\sigma\_\{z\},\\tau\_\{z\}\)equal to\(0\.16,0\.11,0\.24\)\(0\.16,0\.11,0\.24\),\(−0\.32,0\.18,0\.20\)\(\-0\.32,0\.18,0\.20\), and\(1\.95,0\.50,0\.80\)\(1\.95,0\.50,0\.80\)s; thebckmode of−0\.19\-0\.19s places72%72\\%of onset mass inside the ongoing turn\. The reflective intentrefuses a shifted log\-normal of median1\.91\.9s\. Because the optimal action under a rhetorical \(rhe\) or floor\-holding \(hold\) turn is to withhold the floor,frhef\_\{\\textsc\{rhe\}\}andfholdf\_\{\\textsc\{hold\}\}are heavy\-tailed shifted log\-normals of median2\.62\.6s and3\.03\.0s, so the rare licensed response is late and most reference mass sits on∅\\varnothing\. The fitted no\-response masses, equal to the mean annotator acceptability of silence, areβans=0\.05\\beta\_\{\\textsc\{ans\}\}=0\.05,βref=0\.40\\beta\_\{\\textsc\{ref\}\}=0\.40,βrhe=0\.85\\beta\_\{\\textsc\{rhe\}\}=0\.85,βhold=0\.90\\beta\_\{\\textsc\{hold\}\}=0\.90,βbck=0\.45\\beta\_\{\\textsc\{bck\}\}=0\.45,βabn=0\.50\\beta\_\{\\textsc\{abn\}\}=0\.50\.
The asymmetric, intent\-conditioned importance of different offset regions is carried by a nonnegative weightwzw\_\{z\}derived from the same kernel shape,
uz\(t\)=\(fz\(t\)supt′fz\(t′\)\)γ,wz\(t\)=\(1−ϵ\)uz\(t\)\+ϵ,u\_\{z\}\(t\)=\\Big\(\\tfrac\{f\_\{z\}\(t\)\}\{\\sup\_\{t^\{\\prime\}\}f\_\{z\}\(t^\{\\prime\}\)\}\\Big\)^\{\\\!\\gamma\},\\quad w\_\{z\}\(t\)=\(1\-\\epsilon\)\\,u\_\{z\}\(t\)\+\\epsilon,\(2\)withγ=0\.7\\gamma=0\.7andϵ=0\.02\\epsilon=0\.02, sowz\>0w\_\{z\}\>0on𝒯\\mathcal\{T\}\(guaranteeing strict propriety below\) while preserving the kernel’s asymmetry up to the floorϵ\\epsilon: foranslate offsets are weighted above premature overlap, while forbcka well\-placed overlapping onset is weighted above any post\-completion one \(Fig\.[1](https://arxiv.org/html/2609.27372#S4.F1)\)\.
### IV\-CThreshold\-Weighted Continuous Ranked Probability Scoring
A model on episodes with posterior modezzemits a predictive lawPzP\_\{z\}overt∈𝒯t\\in\\mathcal\{T\}with a no\-response atom, summarized by its sub\-probability CDFPz\(t\)=Pr\(y≤t\)P\_\{z\}\(t\)=\\Pr\(y\\leq t\), whose defect1−Pz\(Th−\)=1−limt↑ThPz\(t\)1\-P\_\{z\}\(T\_\{h\}^\{\-\}\)=1\-\\lim\_\{t\\uparrow T\_\{h\}\}P\_\{z\}\(t\)represents the predicted no\-response mass on∅\\varnothingat horizonThT\_\{h\}\. We scorePzP\_\{z\}against a realized continuationyyby the threshold\-weighted continuous ranked probability score \(twCRPS\) of\[[49](https://arxiv.org/html/2609.27372#bib.bib52)\], built on the CRPS of\[[63](https://arxiv.org/html/2609.27372#bib.bib53),[50](https://arxiv.org/html/2609.27372#bib.bib51)\],
twCRPS\(Pz,y\)=∫𝒯wz\(t\)\(Pz\(t\)−𝟏\[y≤t\]\)2dt,\\mathrm\{twCRPS\}\(P\_\{z\},y\)\\;=\\;\\int\_\{\\mathcal\{T\}\}w\_\{z\}\(t\)\\,\\big\(P\_\{z\}\(t\)\-\\mathbf\{1\}\[\\,y\\leq t\\,\]\\big\)^\{2\}\\,dt,\(3\)where for the atom outcomey=∅y=\\varnothingthe indicator𝟏\[y≤t\]=0\\mathbf\{1\}\[\\,y\\leq t\\,\]=0for every finitet∈𝒯t\\in\\mathcal\{T\}, so withheld floor and post\-completion onset are scored by one object\. Intuitively, twCRPS measures the integrated squared discrepancy between the model’s predictive cumulative distribution and the empirical step function𝟏\[y≤t\]\\mathbf\{1\}\[y\\leq t\]of the realized event, weighted bywz\(t\)w\_\{z\}\(t\)so that timing errors are penalized strictly in proportion to their communicative severity under intentzz\. The negatively oriented twCRPS induces the proper divergencedz\(Pz∥hz\)=𝔼y∼hz\[twCRPS\(Pz,y\)\]−𝔼y∼hz\[twCRPS\(hz,y\)\]d\_\{z\}\(P\_\{z\}\\,\\\|\\,h\_\{z\}\)=\\mathbb\{E\}\_\{y\\sim h\_\{z\}\}\[\\mathrm\{twCRPS\}\(P\_\{z\},y\)\]\-\\mathbb\{E\}\_\{y\\sim h\_\{z\}\}\[\\mathrm\{twCRPS\}\(h\_\{z\},y\)\], and the population timing metric maps it to\[0,1\]\[0,1\]as
ST=∑z∈𝒵π\(z\)\(1−dz\(Pz∥hz\)Wz\),Wz=∫𝒯wz\(t\)𝑑t,S\_\{\\mathrm\{T\}\}\\;=\\;\\sum\_\{z\\in\\mathcal\{Z\}\}\\pi\(z\)\\left\(1\-\\frac\{d\_\{z\}\(P\_\{z\}\\,\\\|\\,h\_\{z\}\)\}\{W\_\{z\}\}\\right\),\\qquad W\_\{z\}=\\\!\\int\_\{\\mathcal\{T\}\}\\\!w\_\{z\}\(t\)\\,dt,\(4\)withπ\(z\)=Pr\(z\)\\pi\(z\)=\\Pr\(z\)the intent prior\. The per\-episode score uses the single\-sample twCRPS against the posterior\-mixed referencePq^=∑zq^\(z∣e\)hzP\_\{\\hat\{q\}\}=\\sum\_\{z\}\\hat\{q\}\(z\\mid e\)\\,h\_\{z\}, namelysT\(e\)=1−twCRPS\(Pq^,ye\)/W¯es\_\{\\mathrm\{T\}\}\(e\)=1\-\\mathrm\{twCRPS\}\(P\_\{\\hat\{q\}\},y\_\{e\}\)/\\overline\{W\}\_\{e\}withW¯e=∑zq^\(z∣e\)Wz\\overline\{W\}\_\{e\}=\\sum\_\{z\}\\hat\{q\}\(z\\mid e\)W\_\{z\}; its episode mean estimatesSTS\_\{\\mathrm\{T\}\}up to the irreducible single\-continuation variance reconciled below\.
Fig\. 1:Per\-intent human FTO densitiesfzf\_\{z\}\(top\) and induced twCRPS weightswz\(t\)w\_\{z\}\(t\)of \([2](https://arxiv.org/html/2609.27372#S4.E2)\) \(bottom\) on𝒯\\mathcal\{T\}: mass left of the turn end \(dashed\) is anticipatory overlap, thebckmode lies inside the ongoing turn, andrhe/holdplace most reference mass on the no\-response atomβz\\beta\_\{z\}\(dotted\) with a heavy late tail\.
### IV\-DLength, Overlap, and Composite
Length is scored by the same proper construction: withgzg\_\{z\}the human duration density given intent and context bucket andQzQ\_\{z\}the model’s predictive duration CDF, the aggregate is the bounded complement of the duration twCRPS divergence with the uniform weightw≡1w\\equiv 1on the duration support,
SL=∑zπ\(z\)\(1−dzL\(Qz∥gz\)WL\),S\_\{\\mathrm\{L\}\}\\;=\\;\\sum\_\{z\}\\pi\(z\)\\left\(1\-\\frac\{d\_\{z\}^\{\\mathrm\{L\}\}\(Q\_\{z\}\\,\\\|\\,g\_\{z\}\)\}\{W^\{\\mathrm\{L\}\}\}\\right\),\(5\)wheredzLd\_\{z\}^\{\\mathrm\{L\}\}is the CRPS divergence of\[[63](https://arxiv.org/html/2609.27372#bib.bib53)\]andWLW^\{\\mathrm\{L\}\}the length of the duration support, soSL∈\[0,1\]S\_\{\\mathrm\{L\}\}\\in\[0,1\]and equals11uniquely atQz=gzQ\_\{z\}=g\_\{z\}\. Overlap is scored with prosodic\-completion weighting: a VAP\-style estimator\[[28](https://arxiv.org/html/2609.27372#bib.bib22),[29](https://arxiv.org/html/2609.27372#bib.bib23)\], calibrated on annotator completion marks, assigns each instantτ\\tauof focal\-speaker speech a completion probabilityψ\(τ\)∈\[0,1\]\\psi\(\\tau\)\\in\[0,1\], and competitive \(non\-backchannel\) overlap onsetsoowith durationsdod\_\{o\}incur
sO\(e\)=exp\(−κ∑o∈Oe\(1−ψ\(τo\)\)do\),κ=1\.2s−1,s\_\{\\mathrm\{O\}\}\(e\)\\;=\\;\\exp\\\!\\Big\(\-\\kappa\\sum\_\{o\\in O\_\{e\}\}\\big\(1\-\\psi\(\\tau\_\{o\}\)\\big\)\\,d\_\{o\}\\Big\),\\qquad\\kappa=1\.2~\\mathrm\{s\}^\{\-1\},\(6\)so overlap at points of high prosodic completion, where competitive incoming is licensed\[[20](https://arxiv.org/html/2609.27372#bib.bib9),[16](https://arxiv.org/html/2609.27372#bib.bib11)\], is barely penalized while mid\-clause interruption is penalized in proportion to duration; backchannels are exempt\[[5](https://arxiv.org/html/2609.27372#bib.bib41)\]\. Paired with thebckreference law’s anticipatory mass, cooperative overlap is a high\-scoring action while disruptive interruption remains costly: neither silence nor overlap is a category, only an action under a posterior\. Each category aggregates its components as a fixed convex combination \(timing dominates pause handling and smooth turn\-taking, overlap dominates interruption, timing and length share backchanneling\); the compositeSSis the unweighted mean of the four category scores, with BCa bootstrap intervals over10410^\{4\}resamples\[[64](https://arxiv.org/html/2609.27372#bib.bib54)\]\.
### IV\-EFormal Guarantees
###### Proposition 1\(Boundedness and strict propriety\)\.
Fix an intentzzwithπ\(z\)\>0\\pi\(z\)\>0and weightwzw\_\{z\}strictly positive on the interior of𝒯\\mathcal\{T\}and integrable, withWz=∫𝒯wz<∞W\_\{z\}=\\int\_\{\\mathcal\{T\}\}w\_\{z\}<\\infty\. For every predictive lawPzP\_\{z\}on𝒯∪\{∅\}\\mathcal\{T\}\\cup\\\{\\varnothing\\\}and every realizedy∈𝒯∪\{∅\}y\\in\\mathcal\{T\}\\cup\\\{\\varnothing\\\},twCRPS\(Pz,y\)∈\[0,Wz\]\\mathrm\{twCRPS\}\(P\_\{z\},y\)\\in\[0,W\_\{z\}\], so the per\-intent term and henceSTS\_\{\\mathrm\{T\}\}of \([4](https://arxiv.org/html/2609.27372#S4.E4)\) lie in\[0,1\]\[0,1\]\. Moreover𝔼y∼hz\[twCRPS\(Pz,y\)\]\\mathbb\{E\}\_\{y\\sim h\_\{z\}\}\[\\mathrm\{twCRPS\}\(P\_\{z\},y\)\]is minimized over allPzP\_\{z\}uniquely atPz=hzP\_\{z\}=h\_\{z\}\(the twCRPS is strictly proper\), soST=1S\_\{\\mathrm\{T\}\}=1if and only ifPz=hzP\_\{z\}=h\_\{z\}for all suchzz; in particular every deterministic forecast and every law supported on the post\-completion half\-line incursST<1S\_\{\\mathrm\{T\}\}<1wheneverhzh\_\{z\}is dispersed or carries anticipatory mass\.
###### Proof\.
Boundedness: for eachtt,\(Pz\(t\)−𝟏\[y≤t\]\)2∈\[0,1\]\\big\(P\_\{z\}\(t\)\-\\mathbf\{1\}\[y\\leq t\]\\big\)^\{2\}\\in\[0,1\]sincePz\(t\)∈\[0,1\]P\_\{z\}\(t\)\\in\[0,1\]and the indicator is in\{0,1\}\\\{0,1\\\}; multiplying bywz\(t\)≥0w\_\{z\}\(t\)\\geq 0and integrating givestwCRPS∈\[0,Wz\]\\mathrm\{twCRPS\}\\in\[0,W\_\{z\}\], so1−twCRPS/Wz∈\[0,1\]1\-\\mathrm\{twCRPS\}/W\_\{z\}\\in\[0,1\]and the convex combination \([4](https://arxiv.org/html/2609.27372#S4.E4)\) stays in\[0,1\]\[0,1\]\.
Strict propriety: writeHz\(t\)=Pry∼hz\(y≤t\)H\_\{z\}\(t\)=\\Pr\_\{y\\sim h\_\{z\}\}\(y\\leq t\)for the proper CDF of the continuous part on𝒯\\mathcal\{T\}; the atom∅\\varnothingcontributes00to𝟏\[y≤t\]\\mathbf\{1\}\[y\\leq t\]at every finitett, hence𝔼y∼hz𝟏\[y≤t\]=Hz\(t\)\\mathbb\{E\}\_\{y\\sim h\_\{z\}\}\\mathbf\{1\}\[y\\leq t\]=H\_\{z\}\(t\)for allt∈𝒯t\\in\\mathcal\{T\}\. For fixedttthe indicator is Bernoulli with meanHz\(t\)H\_\{z\}\(t\), so the bias–variance identity gives𝔼\(Pz\(t\)−𝟏\[y≤t\]\)2=\(Pz\(t\)−Hz\(t\)\)2\+Hz\(t\)\(1−Hz\(t\)\)\\mathbb\{E\}\\big\(P\_\{z\}\(t\)\-\\mathbf\{1\}\[y\\leq t\]\\big\)^\{2\}=\\big\(P\_\{z\}\(t\)\-H\_\{z\}\(t\)\\big\)^\{2\}\+H\_\{z\}\(t\)\\big\(1\-H\_\{z\}\(t\)\\big\)\. By Tonelli,
𝔼y∼hz\[twCRPS\(Pz,y\)\]=∫𝒯wz\(Pz−Hz\)2𝑑t\+Rz,\\mathbb\{E\}\_\{y\\sim h\_\{z\}\}\[\\mathrm\{twCRPS\}\(P\_\{z\},y\)\]=\\int\_\{\\mathcal\{T\}\}w\_\{z\}\\,\(P\_\{z\}\-H\_\{z\}\)^\{2\}\\,dt\+R\_\{z\},withRz=∫𝒯wzHz\(1−Hz\)𝑑tR\_\{z\}=\\int\_\{\\mathcal\{T\}\}w\_\{z\}\\,H\_\{z\}\(1\-H\_\{z\}\)\\,dtindependent ofPzP\_\{z\}\. Hence the divergence isdz\(Pz∥hz\)=∫𝒯wz\(Pz−Hz\)2dt≥0d\_\{z\}\(P\_\{z\}\\\|h\_\{z\}\)=\\int\_\{\\mathcal\{T\}\}w\_\{z\}\\,\(P\_\{z\}\-H\_\{z\}\)^\{2\}\\,dt\\geq 0, with equality if and only ifPz\(t\)=Hz\(t\)P\_\{z\}\(t\)=H\_\{z\}\(t\)for Lebesgue\-almost everyttin\{wz\>0\}\\\{w\_\{z\}\>0\\\}\. Sincewz\>0w\_\{z\}\>0on the interior of𝒯\\mathcal\{T\}and bothPz,HzP\_\{z\},H\_\{z\}are right\-continuous,Pz=HzP\_\{z\}=H\_\{z\}everywhere on𝒯\\mathcal\{T\}; this pins downlimt↑ThPz\(t\)=limt↑ThHz\(t\)=1−βz\\lim\_\{t\\uparrow T\_\{h\}\}P\_\{z\}\(t\)=\\lim\_\{t\\uparrow T\_\{h\}\}H\_\{z\}\(t\)=1\-\\beta\_\{z\}, so the no\-response atom masses coincide andPz=hzP\_\{z\}=h\_\{z\}as mixed laws\. Therefore𝔼y∼hz\[twCRPS\]\\mathbb\{E\}\_\{y\\sim h\_\{z\}\}\[\\mathrm\{twCRPS\}\]is uniquely minimized athzh\_\{z\}, i\.e\. the score is strictly proper, and1−dz/Wz=11\-d\_\{z\}/W\_\{z\}=1exactly whenPz=hzP\_\{z\}=h\_\{z\}\. A deterministic forecast has a step CDF and a law on the post\-completion half\-line hasPz≡0P\_\{z\}\\equiv 0on\(−∞,0\)\(\-\\infty,0\); either differs from a dispersed or anticipatoryhzh\_\{z\}on a set of positivewzw\_\{z\}\-measure, forcingdz\>0d\_\{z\}\>0\. Summing overzzwithπ\(z\)\>0\\pi\(z\)\>0yieldsST∈\[0,1\]S\_\{\\mathrm\{T\}\}\\in\[0,1\]withST=1S\_\{\\mathrm\{T\}\}=1iffPz=hzP\_\{z\}=h\_\{z\}for all suchzz\. ∎
###### Corollary 1\(Composite strict propriety\)\.
The length aggregateSLS\_\{\\mathrm\{L\}\}of \([5](https://arxiv.org/html/2609.27372#S4.E5)\) is a CRPS divergence and the overlap score is a bounded penalty, both in\[0,1\]\[0,1\]; any fixed convex combination ofSTS\_\{\\mathrm\{T\}\},SLS\_\{\\mathrm\{L\}\}, andSOS\_\{\\mathrm\{O\}\}, and the unweighted mean of the four category scores, is a nonnegative weighted sum of strictly proper divergence terms, hence lies in\[0,1\]\[0,1\]and is maximized at the value11exactly when the model reproduces every human conditional timing, duration, and overlap law it combines\.
###### Proof\.
SLS\_\{\\mathrm\{L\}\}is the unweighted twCRPS of Proposition[1](https://arxiv.org/html/2609.27372#Thmproposition1)\(weightw≡1w\\equiv 1\) applied to durations, so it is strictly proper and in\[0,1\]\[0,1\]with unique maximizerQz=gzQ\_\{z\}=g\_\{z\};sO=exp\(−κ∑\(1−ψ\)d\)∈\(0,1\]s\_\{\\mathrm\{O\}\}=\\exp\(\-\\kappa\\sum\(1\-\\psi\)d\)\\in\(0,1\]equals11iff there is no low\-completion competitive overlap\. A convex combination∑cλcSc\\sum\_\{c\}\\lambda\_\{c\}S\_\{c\}withλc≥0\\lambda\_\{c\}\\geq 0,∑cλc=1\\sum\_\{c\}\\lambda\_\{c\}=1of terms each in\[0,1\]\[0,1\]lies in\[0,1\]\[0,1\], and since eachSc≤1S\_\{c\}\\leq 1with equality only at its own optimum, the combination attains11iff every active term does; the four\-category mean is the special case of equal weights\. ∎
###### Proposition 2\(Reduction to binary fixed\-window scoring\)\.
Let the posterior be degenerate at a single intentz0z\_\{0\}and the model emit the deterministic forecastPz0=𝟏\[t≥ye\]P\_\{z\_\{0\}\}=\\mathbf\{1\}\[t\\geq y\_\{e\}\]\. Replacing the smooth weightwz0w\_\{z\_\{0\}\}by the indicator weightw\(t\)=δ\-limit atWw\(t\)=\\delta\\text\{\-limit at \}W, i\.e\. takingwη\(t\)=η−1𝟏\[W<t≤W\+η\]w\_\{\\eta\}\(t\)=\\eta^\{\-1\}\\mathbf\{1\}\[W<t\\leq W\+\\eta\]and lettingη→0\+\\eta\\to 0^\{\+\}, the normalized twCRPS score1−twCRPS/Wz1\-\\mathrm\{twCRPS\}/W\_\{z\}converges to the response\-within\-window indicator𝟏\[0<ye≤W\]\\mathbf\{1\}\[\\,0<y\_\{e\}\\leq W\\,\]whose episode mean is exactly the takeover\-style statistic of Full\-Duplex\-Bench\[[5](https://arxiv.org/html/2609.27372#bib.bib41)\];W→∞W\\to\\inftyrecovers the unwindowed takeover rate, and the two\-sided indicator weight on\{−W′<t≤W\}\\\{\-W^\{\\prime\}<t\\leq W\\\}recovers the windowed overlap statistics of\[[6](https://arxiv.org/html/2609.27372#bib.bib42)\]\.
###### Proof\.
Concentratingwηw\_\{\\eta\}at the single thresholdt=Wt=Wand normalizing,1−twCRPS/Wz1\-\\mathrm\{twCRPS\}/W\_\{z\}tends to11when the realized onset precedesWW\(a response within the window\) and to00otherwise, that is to𝟏\[0<ye≤W\]\\mathbf\{1\}\[0<y\_\{e\}\\leq W\]; abstention givesPz0≡0P\_\{z\_\{0\}\}\\equiv 0and score00\. The episode mean is the windowed takeover rate,W→∞W\\to\\inftyremoves the upper edge, and the symmetric two\-sided indicator weight recovers the overlap window, so the binary fixed\-window metric is the degenerate indicator\-weight limit of the twCRPS score\. ∎
###### Proposition 3\(Identifiability of the intent mixture\)\.
If the per\-intent lawshzh\_\{z\}are pairwise distinct ex\-Gaussian/shifted\-log\-normal laws with an atom at∅\\varnothing, intents are drawn i\.i\.d\. from a speaker priorπs\\pi\_\{s\}, and the response law givenzzdoes not otherwise depend onss, thenπs\\pi\_\{s\}and\{hz\}\\\{h\_\{z\}\\\}are identifiable up to label permutation from the per\-speaker marginal response law\.
###### Proof\.
Atom masses separate from continuous parts by mutual singularity, so it suffices that the continuous components be linearly independent\. The ex\-Gaussian characteristic functionexp\(iμω−σ2ω2/2\)\(1−iτω\)−1\\exp\(i\\mu\\omega\-\\sigma^\{2\}\\omega^\{2\}/2\)\(1\-i\\tau\\omega\)^\{\-1\}has a unique simple pole atω=−i/τ\\omega=\-i/\\tau; ordering byτ\\tauand taking residues rules out nontrivial vanishing combinations, Gaussian\-mixture identifiability handles equal\-τ\\taugroups, and shifted log\-normals are separated by their support edge\. Linear independence makesπs\\pi\_\{s\}unique, licensing the memory profiles:πs\\pi\_\{s\}is recoverable from that speaker’s episodes and nothing less\. ∎
### IV\-FReconciling the Population Optimum with the Empirical Topline
Proposition[1](https://arxiv.org/html/2609.27372#Thmproposition1)fixes the population optimum atST=1S\_\{\\mathrm\{T\}\}=1, attained only byhzh\_\{z\}, yet the empirical human topline in Table[I](https://arxiv.org/html/2609.27372#S4.T1)is0\.860\.86because the per\-episode estimator scores a single held\-out continuationyey\_\{e\}against the reference\. Two finite\-sample effects separate it from11: each episode offers one realized continuation, so even a forecaster equal tohzh\_\{z\}pays the irreducible single\-draw termRz/Wz\>0R\_\{z\}/W\_\{z\}\>0from the proof above; and a held\-out speaker’s timing deviates from any pooledhzh\_\{z\}, a gap Proposition[3](https://arxiv.org/html/2609.27372#Thmproposition3)shows shrinks only with that speaker’s own evidence\. Together they account for the0\.140\.14shortfall, so the topline is a principled estimate of the population optimum\.
TABLE I:Main results on TACT by Full\-Duplex\-Bench scenario category \(Pause: pause handling, Backch\.: backchanneling, Smooth: smooth turn\-taking, Interr\.: user interruption\); the composite is the category mean with 95% BCa bootstrap intervals\. These are scenario types, distinct from the six latent intents of Fig\.[2](https://arxiv.org/html/2609.27372#S6.F2); in particular the Backch\. scenario is not thebckintent\. Ranks compare TACT and reproduced binary v1\-style orderings \(Spearman 0\.55\)\.
## VExperimental Setup
We evaluate eleven systems: open models dGSLM\[[1](https://arxiv.org/html/2609.27372#bib.bib28)\], Moshi\[[2](https://arxiv.org/html/2609.27372#bib.bib29)\], SyncLLM\[[3](https://arxiv.org/html/2609.27372#bib.bib30)\], Freeze\-Omni\[[34](https://arxiv.org/html/2609.27372#bib.bib31)\], Qwen2\.5\-Omni\[[37](https://arxiv.org/html/2609.27372#bib.bib33)\], MiniCPM\-o 2\.6\[[38](https://arxiv.org/html/2609.27372#bib.bib39)\], PersonaPlex\[[39](https://arxiv.org/html/2609.27372#bib.bib37)\], and proprietary points GPT\-4o Realtime\[[40](https://arxiv.org/html/2609.27372#bib.bib34)\], gpt\-realtime\-2\[[41](https://arxiv.org/html/2609.27372#bib.bib35)\], Gemini 3\.1 Flash Live\[[42](https://arxiv.org/html/2609.27372#bib.bib36)\]at minimal and high thinking\. Open models run locally in streaming mode, proprietary systems via realtime APIs\[[6](https://arxiv.org/html/2609.27372#bib.bib42)\]; prompt models receive synopses, dGSLM priming audio\. Reference baselines include a VAP\-gated cascade\[[28](https://arxiv.org/html/2609.27372#bib.bib22),[29](https://arxiv.org/html/2609.27372#bib.bib23)\], a never\-respond policy, and the human topline\. Reproduced binary v1 metrics\[[5](https://arxiv.org/html/2609.27372#bib.bib41)\]match published numbers \(Moshi pause takeover\>0\.98\>0\.98, median latency≈0\.26\\approx 0\.26s\)\. The validity study uses 1,280 held\-out episodes judged by five raters in 88 system\-condition cells\.
## VIResults
### VI\-AMain Comparison
Table[I](https://arxiv.org/html/2609.27372#S4.T1)reports main results\. The gap to humans is large: the best system, gpt\-realtime\-2, attains 0\.47 vs 0\.86 human topline, while binary v1 metrics make systems look adequate\. Rankings reorder substantially \(correlation 0\.55\): Moshi, second under binary scoring due to 0\.26 s latency, falls to tenth because instant takeover fails onref,rhe, andhold\. Gemini points expose a trade\-off invisible to binary scoring: minimal responds fastest \(0\.31 s\) and tops v1, while high waits median 0\.74 s, falling to seventh under v1 yet overtaking it on TACT\. The VAP cascade matches mid\-tier models \(0\.35\) and beats open models on timing, but degrades on backchannels\.
### VI\-BThe Over\-Eagerness Pathology
The scenario categories of Table[I](https://arxiv.org/html/2609.27372#S4.T1)and the intents of Fig\.[2](https://arxiv.org/html/2609.27372#S6.F2)are orthogonal cuts: each scenario composite is the prior\-weighted average of the Fig\.[2](https://arxiv.org/html/2609.27372#S6.F2)cells in its rows\. Over\-eagerness is traceable throughref,rhe, andhold\. All eleven systems do best onans\(0\.38–0\.78\) and collapse onref\(0\.09–0\.36\) andrhe\(0\.06–0\.30\)\. Humans respond within 1 s in 78% ofansepisodes but only 12% ofrefand 7% ofrhe; Moshi does so in 95–97% and conservative models in 62–76%, withholding the floor onrhein<4%<4\\%\(open\) to 24% \(best closed\) vs 74% for humans\. Rigidity is two\-sided: humans launch 63% of invited backchannels in overlap, vs 14% for PersonaPlex and<6%<6\\%elsewhere\. Models fail to be slow when reflection is warranted or early when invited, recurring as length miscalibration \(backchannels receive median 8\.4 s turns vs 1\.1 s human\)\.
Fig\. 2:Each cell is the TACT compositeSSof \([4](https://arxiv.org/html/2609.27372#S4.E4)\) restricted to episodes whose posterior mode is the column intent \(the intent\-conditioned composite; intent\-view, orthogonal to the scenario\-category axis of Table[I](https://arxiv.org/html/2609.27372#S4.T1)\)\. Systems approach humans only onansand degrade two\- to ten\-fold onref,rhe, andabn; the overlap\-optimalbckcolumn shows the anticipatory side of the rigidity\.
### VI\-CMemory Ablation: No Per\-Speaker Adaptation
Re\-running each system with profiles swapped across speakers of opposite behavioral type, we measureJSD2\\mathrm\{JSD\}\_\{2\}between onset distributions under true and swapped profiles\. Prompt\-capable models receive the prompt synopsis, Moshi receives profile features through its synchronized text stream, and dGSLM receives matched priming audio\. Held\-out human behavior shifts substantially \(JSD2=0\.23\\mathrm\{JSD\}\_\{2\}=0\.23\) while every system is essentially invariant \(ΔS≤0\.01\\Delta S\\leq 0\.01andJSD2≤0\.024\\mathrm\{JSD\}\_\{2\}\\leq 0\.024across all eleven systems, with 0\.024 for gpt\-realtime\-2 and 0\.003 for dGSLM being the swap extrema\), demonstrating that current models settle on a speaker\-independent policy where human behavior is speaker\-conditioned\[[54](https://arxiv.org/html/2609.27372#bib.bib16),[55](https://arxiv.org/html/2609.27372#bib.bib17)\]\.
### VI\-DMetric Validity
Over 88 cells judged by held\-out humans, TACT reaches Spearmanρ=0\.81\\rho=0\.81\(\[0\.77,0\.84\]\[0\.77,0\.84\]\), vs0\.460\.46for binary composite and0\.390\.39for latency\. An intent\-conditioned binary baseline attains onlyρ=0\.62\\rho=0\.62, showing that continuous twCRPS is essential: fixed windows cannot distinguish timing within windows or represent anticipatory, delayed, and absent onsets jointly\. Ablating memory, calibration, or prosody yieldsρ=0\.71\\rho=0\.71,0\.780\.78, and0\.740\.74\. Calibration reduces ECE from 8\.4% to 2\.1% \(λ=0\.3\\lambda=0\.3\)\.
### VI\-EFairness and Robustness Slices
Composites by slice show the human topline flat across gender, age, first\-language region, SNR, and articulation rate \(0\.84–0\.87\)\. Systems are not: the best loses 0\.07 between North American and Indian\-English speakers and 0\.06 between high\- and mid\-SNR conditions \(concentrated in timing\), with smaller gender \(0\.02\) and age \(0\.05\) gaps\.
### VI\-FLimitations
The six\-class taxonomy discretizes a continuum andα=0\.73\\alpha=0\.73leaves label noise; kernels are fit on English corpora, though the released recipe enables refitting where norms differ\[[7](https://arxiv.org/html/2609.27372#bib.bib4),[30](https://arxiv.org/html/2609.27372#bib.bib24)\]; the judge in \([1](https://arxiv.org/html/2609.27372#S4.E1)\) may share failure modes \(an annotator\-only mode ships\); and anticipatory scoring leans on VAP completion estimatesψ\\psithat released marks would replace\.
## VIIConclusion
TACT replaces binary fixed\-window turn\-taking with a strictly proper, intent\-conditioned twCRPS over human floor\-transfer\-offset laws, per\-speaker memory, and a calibrated posterior; data, kernels, and prompts are available in the Supplementary Material\. When to speak is an inference over a speaker’s latent intent, not a latency problem\.
## Acknowledgment
The authors disclose that Claude Opus 4\.8 was used for editing and rewriting all sections \(Abstract, Introduction, Related Work, Methods, Experiments, Results, Conclusion\) to improve the flow and presentation, and for assisting in implementation of the code\.
## References
- \[1\]T\. A\. Nguyen, E\. Kharitonov, J\. Copet, Y\. Adi, W\. Hsu, A\. Elkahky, P\. Tomasello, R\. Algayres, B\. Sagot, A\. Mohamed, and E\. Dupoux\(2023\)Generative spoken dialogue language modeling\.Transactions of the Association for Computational Linguistics11,pp\. 250–266\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p1.1),[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.2.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[2\]A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour\(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.Note:arXiv preprint arXiv:2410\.00037Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p1.1),[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.3.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[3\]B\. Veluri, B\. Peloquin, B\. Yu, H\. Gong, and S\. Gollakota\(2024\)Beyond turn\-based interfaces: synchronous LLMs as full\-duplex dialogue agents\.InProc\. Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p1.1),[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.4.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[4\]T\. Lin, Y\. Wu, F\. Huang, L\. Si, J\. Sun, and Y\. Li\(2022\)Duplex conversation: towards human\-like interaction in spoken dialogue systems\.InProc\. ACM SIGKDD Conference on Knowledge Discovery and Data Mining,Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p1.1),[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[5\]G\. Lin, J\. Lian, T\. Li, Q\. Wang, G\. Anumanchipalli, A\. H\. Liu, and H\. Lee\(2025\)Full\-duplex\-bench: a benchmark to evaluate full\-duplex spoken dialogue models on turn\-taking capabilities\.InProc\. IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\),Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p1.1),[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1),[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.3),[§V](https://arxiv.org/html/2609.27372#S5.p1.1),[Proposition 2](https://arxiv.org/html/2609.27372#Thmproposition2.p1.1.1)\.
- \[6\]G\. Lin, S\. S\. Kuan, Q\. Wang, J\. Lian, T\. Li, S\. Watanabe, and H\. Lee\(2026\)Full\-duplex\-bench v1\.5: evaluating overlap handling for full\-duplex speech models\.InProc\. IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p1.1),[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1),[Proposition 2](https://arxiv.org/html/2609.27372#Thmproposition2.p1.1.1)\.
- \[7\]T\. Stivers, N\. J\. Enfield, P\. Brown, C\. Englert, M\. Hayashi, T\. Heinemann, G\. Hoymann, F\. Rossano, J\. P\. de Ruiter, K\. Yoon, and S\. C\. Levinson\(2009\)Universals and cultural variation in turn\-taking in conversation\.Proceedings of the National Academy of Sciences106\(26\),pp\. 10587–10592\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p2.1),[§VI\-F](https://arxiv.org/html/2609.27372#S6.SS6.p1.1)\.
- \[8\]S\. C\. Levinson and F\. Torreira\(2015\)Timing in turn\-taking and its implications for processing models of language\.Frontiers in Psychology6,pp\. 731\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[9\]M\. Heldner and J\. Edlund\(2010\)Pauses, gaps and overlaps in conversations\.Journal of Phonetics38\(4\),pp\. 555–568\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[10\]H\. Sacks, E\. A\. Schegloff, and G\. Jefferson\(1974\)A simplest systematics for the organization of turn\-taking for conversation\.Language50\(4\),pp\. 696–735\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[11\]E\. A\. Schegloff\(2000\)Overlapping talk and the organization of turn\-taking for conversation\.Language in Society29\(1\),pp\. 1–63\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p1.1)\.
- \[12\]K\. H\. Kendrick and F\. Torreira\(2015\)The timing and construction of preference: a quantitative study\.Discourse Processes52\(4\),pp\. 255–289\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p1.1)\.
- \[13\]V\. H\. Yngve\(1970\)On getting a word in edgewise\.InPapers from the Sixth Regional Meeting of the Chicago Linguistic Society,pp\. 567–578\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[14\]N\. Ward and W\. Tsukahara\(2000\)Prosodic features which cue back\-channel responses in English and Japanese\.Journal of Pragmatics32\(8\),pp\. 1177–1207\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[15\]A\. Reece, G\. Cooney, P\. Bull, C\. Chung, B\. Dawson, C\. Fitzpatrick, T\. Glazer, D\. Knox, A\. Liebscher, and S\. Marin\(2023\)The CANDOR corpus: insights from a large multimodal dataset of naturalistic conversation\.Science Advances9\(13\),pp\. eadf3197\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§I](https://arxiv.org/html/2609.27372#S1.p3.1),[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1),[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p2.1)\.
- \[16\]A\. Gravano and J\. Hirschberg\(2011\)Turn\-taking cues in task\-oriented dialogue\.Computer Speech and Language25\(3\),pp\. 601–634\.Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p2.1),[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.3)\.
- \[17\]Z\. Sheikh, S\. Shimizu, S\. Arora, J\. Shi, S\. Cornell, X\. Li, and S\. Watanabe\(2025\)Scalable spontaneous speech dataset \(SSSD\): crowdsourcing data collection to promote dialogue research\.InProc\. Interspeech,Cited by:[§I](https://arxiv.org/html/2609.27372#S1.p3.1),[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1)\.
- \[18\]otoearth\(2025\)otoSpeech\-full\-duplex\-280h: raw English full\-duplex dyadic conversation corpus\.Note:Hugging Face dataset,https://huggingface\.co/datasets/otoearth/otoSpeech\-full\-duplex\-280hCited by:[§I](https://arxiv.org/html/2609.27372#S1.p3.1),[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1)\.
- \[19\]J\. P\. de Ruiter, H\. Mitterer, and N\. J\. Enfield\(2006\)Projecting the end of a speaker’s turn: a cognitive cornerstone of conversation\.Language82\(3\),pp\. 515–535\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[20\]S\. Bögels and F\. Torreira\(2015\)Listeners use intonational phrase boundaries to project turn ends in spoken interaction\.Journal of Phonetics52,pp\. 46–57\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.3)\.
- \[21\]J\. J\. Godfrey, E\. C\. Holliman, and J\. McDaniel\(1992\)SWITCHBOARD: telephone speech corpus for research and development\.InProc\. IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 517–520\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[22\]S\. G\. Roberts, F\. Torreira, and S\. C\. Levinson\(2015\)The effects of processing and sequence organization on the timing of turn taking: a corpus study\.Frontiers in Psychology6,pp\. 509\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§IV\-B](https://arxiv.org/html/2609.27372#S4.SS2.p1.1)\.
- \[23\]A\. Stolcke, K\. Ries, N\. Coccaro, E\. Shriberg, R\. Bates, D\. Jurafsky, P\. Taylor, R\. Martin, C\. Van Ess\-Dykema, and M\. Meteer\(2000\)Dialogue act modeling for automatic tagging and recognition of conversational speech\.Computational Linguistics26\(3\),pp\. 339–373\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p1.1)\.
- \[24\]H\. H\. Clark\(1996\)Using language\.Cambridge University Press,Cambridge, UK\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[25\]A\. Raux and M\. Eskenazi\(2009\)A finite\-state turn\-taking model for spoken dialog systems\.InProc\. Human Language Technologies: NAACL,pp\. 629–637\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[26\]D\. Schlangen and G\. Skantze\(2011\)A general, abstract model of incremental dialogue processing\.Dialogue and Discourse2\(1\),pp\. 83–111\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[27\]E\. Ekstedt and G\. Skantze\(2020\)TurnGPT: a transformer\-based language model for predicting turn\-taking in spoken dialog\.InFindings of the Association for Computational Linguistics: EMNLP,pp\. 2981–2990\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[28\]E\. Ekstedt and G\. Skantze\(2022\)Voice activity projection: self\-supervised learning of turn\-taking events\.InProc\. Interspeech,pp\. 5190–5194\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.2),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.13.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[29\]K\. Inoue, B\. Jiang, E\. Ekstedt, T\. Kawahara, and G\. Skantze\(2024\)Real\-time and continuous turn\-taking prediction using voice activity projection\.InProc\. International Workshop on Spoken Dialogue Systems Technology \(IWSDS\),Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.2),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.13.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[30\]K\. Inoue, B\. Jiang, E\. Ekstedt, T\. Kawahara, and G\. Skantze\(2024\)Multilingual turn\-taking prediction using voice activity projection\.InProc\. Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING\),pp\. 11873–11883\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1),[§VI\-F](https://arxiv.org/html/2609.27372#S6.SS6.p1.1)\.
- \[31\]S\. Chang, B\. Li, T\. Sainath, C\. Zhang, T\. Strohman, Q\. Liang, and Y\. He\(2022\)Turn\-taking prediction for natural conversational speech\.InProc\. Interspeech,pp\. 1821–1825\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[32\]G\. Skantze\(2021\)Turn\-taking in conversational systems and human\-robot interaction: a review\.Computer Speech and Language67,pp\. 101178\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p1.1)\.
- \[33\]L\. Zhou, J\. Gao, D\. Li, and H\. Shum\(2020\)The design and implementation of XiaoIce, an empathetic social chatbot\.Computational Linguistics46\(1\),pp\. 53–93\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[34\]X\. Wang, Y\. Li, C\. Fu,et al\.\(2025\)Freeze\-omni: a smart and low latency speech\-to\-speech dialogue model with frozen LLM\.InProc\. International Conference on Machine Learning \(ICML\),Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.5.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[35\]Z\. Ma, Y\. Song, C\. Du, J\. Cong, Z\. Chen, Y\. Wang, Y\. Wang, and X\. Chen\(2025\)Language model can listen while speaking\.InProc\. AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 24831–24839\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[36\]D\. Zhang, S\. Li, X\. Zhang, J\. Zhan, P\. Wang, Y\. Zhou, and X\. Qiu\(2023\)SpeechGPT: empowering large language models with intrinsic cross\-modal conversational abilities\.InFindings of the Association for Computational Linguistics: EMNLP,Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[37\]J\. Xu, Z\. Guo, J\. He,et al\.\(2025\)Qwen2\.5\-Omni technical report\.Note:arXiv preprint arXiv:2503\.20215Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.6.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[38\]OpenBMB Team\(2025\)MiniCPM\-o 2\.6: a GPT\-4o level MLLM for vision, speech, and multimodal live streaming on your phone\.Note:https://github\.com/OpenBMB/MiniCPM\-oCited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.7.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[39\]R\. Roy, J\. Raiman, S\. Lee, T\. Ene, R\. Kirby, S\. Kim, J\. Kim, and B\. Catanzaro\(2026\)PersonaPlex: voice and role control for full duplex conversational speech models\.InProc\. IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.8.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[40\]OpenAI\(2024\)GPT\-4o system card\.Note:arXiv preprint arXiv:2410\.21276Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.9.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[41\]OpenAI\(2026\)gpt\-realtime\-2: realtime API model documentation\.Note:https://developers\.openai\.com/api/docs/models/gpt\-realtime\-2Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.12.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[42\]Google DeepMind\(2026\)Gemini 3\.1 Flash Live: live API model documentation\.Note:https://ai\.google\.dev/gemini\-api/docs/models/gemini\-3\.1\-flash\-live\-previewCited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.10.1.1),[TABLE I](https://arxiv.org/html/2609.27372#S4.T1.3.11.1.1),[§V](https://arxiv.org/html/2609.27372#S5.p1.1)\.
- \[43\]Y\. Chen, X\. Yue, C\. Zhang, X\. Gao, R\. T\. Tan, and H\. Li\(2024\)VoiceBench: benchmarking LLM\-based voice assistants\.Note:arXiv preprint arXiv:2410\.17196Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[44\]J\. Ao, Y\. Wang, X\. Tian, D\. Chen, J\. Zhang, L\. Lu, Y\. Wang, H\. Li, and Z\. Wu\(2024\)SD\-Eval: a benchmark dataset for spoken dialogue understanding beyond words\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[45\]Y\. Peng, Y\. Chao, D\. Ng, Y\. Ma, C\. Ni, B\. Ma, and E\. S\. Chng\(2025\)FD\-Bench: a full\-duplex benchmarking pipeline designed for full duplex spoken dialogue systems\.InProc\. Interspeech,Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[46\]G\. Lin, S\. S\. Kuan, J\. Shi, K\. Chang, S\. Arora, S\. Watanabe, and H\. Lee\(2025\)Full\-duplex\-bench\-v2: a multi\-turn evaluation framework for duplex dialogue systems with an automated examiner\.Note:arXiv preprint arXiv:2510\.07838Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[47\]G\. Lin, C\. Chen, Z\. Chen, and H\. Lee\(2026\)Full\-duplex\-bench\-v3: benchmarking tool use for full\-duplex voice agents under real\-world disfluency\.Note:arXiv preprint arXiv:2604\.04847Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[48\]S\. Arora, Z\. Lu, C\. Chiu, R\. Pang, and S\. Watanabe\(2025\)Talking turns: benchmarking audio foundation models on turn\-taking dynamics\.InProc\. International Conference on Learning Representations \(ICLR\),Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[49\]T\. Gneiting and R\. Ranjan\(2011\)Comparing density forecasts using threshold\- and quantile\-weighted scoring rules\.Journal of Business & Economic Statistics29\(3\),pp\. 411–422\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§IV\-C](https://arxiv.org/html/2609.27372#S4.SS3.p1.1)\.
- \[50\]T\. Gneiting and A\. E\. Raftery\(2007\)Strictly proper scoring rules, prediction, and estimation\.Journal of the American Statistical Association102\(477\),pp\. 359–378\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§IV\-C](https://arxiv.org/html/2609.27372#S4.SS3.p1.1)\.
- \[51\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\(2023\)Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§IV\-A](https://arxiv.org/html/2609.27372#S4.SS1.p1.1)\.
- \[52\]Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu\(2023\)G\-Eval: NLG evaluation using GPT\-4 with better human alignment\.InProc\. Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 2511–2522\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§IV\-A](https://arxiv.org/html/2609.27372#S4.SS1.p1.1)\.
- \[53\]D\. Premack and G\. Woodruff\(1978\)Does the chimpanzee have a theory of mind?\.Behavioral and Brain Sciences1\(4\),pp\. 515–526\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§IV\-A](https://arxiv.org/html/2609.27372#S4.SS1.p1.2)\.
- \[54\]M\. Sap, R\. Le Bras, D\. Fried, and Y\. Choi\(2022\)Neural theory\-of\-mind? On the limits of social intelligence in large LMs\.InProc\. Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 3762–3780\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§IV\-A](https://arxiv.org/html/2609.27372#S4.SS1.p1.2),[§VI\-C](https://arxiv.org/html/2609.27372#S6.SS3.p1.1)\.
- \[55\]H\. Kim, M\. Sclar, X\. Zhou, R\. Le Bras, G\. Kim, Y\. Choi, and M\. Sap\(2023\)FANToM: a benchmark for stress\-testing machine theory of mind in interactions\.InProc\. Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1),[§VI\-C](https://arxiv.org/html/2609.27372#S6.SS3.p1.1)\.
- \[56\]S\. Zhang, E\. Dinan, J\. Urbanek, A\. Szlam, D\. Kiela, and J\. Weston\(2018\)Personalizing dialogue agents: I have a dog, do you have pets too?\.InProc\. Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 2204–2213\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[57\]J\. Xu, A\. Szlam, and J\. Weston\(2022\)Beyond goldfish memory: long\-term open\-domain conversation\.InProc\. Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 5180–5197\.Cited by:[§II](https://arxiv.org/html/2609.27372#S2.p2.1)\.
- \[58\]otoearth\(2025\)otoSpeech\-full\-duplex\-processed\-141h: cleaned full\-duplex dyadic conversation corpus\.Note:Hugging Face dataset,https://huggingface\.co/datasets/otoearth/otoSpeech\-full\-duplex\-processed\-141hCited by:[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1)\.
- \[59\]V\. Agrawal, A\. Akinyemi, K\. Alvero,et al\.\(2025\)Seamless interaction: dyadic audiovisual motion modeling and large\-scale dataset\.Note:arXiv preprint arXiv:2506\.22554Cited by:[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1),[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p2.1)\.
- \[60\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever\(2023\)Robust speech recognition via large\-scale weak supervision\.InProc\. International Conference on Machine Learning \(ICML\),pp\. 28492–28518\.Cited by:[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1)\.
- \[61\]J\. Carletta, S\. Ashby, S\. Bourban, M\. Flynn, M\. Guillemot, T\. Hain, J\. Kadlec, V\. Karaiskos, W\. Kraaij, M\. Kronenthal, G\. Lathoud, M\. Lincoln, A\. Lisowska, I\. McCowan, W\. Post, D\. Reidsma, and P\. Wellner\(2005\)The AMI meeting corpus: a pre\-announcement\.InProc\. International Workshop on Machine Learning for Multimodal Interaction \(MLMI\),pp\. 28–39\.Cited by:[§III\-A](https://arxiv.org/html/2609.27372#S3.SS1.p2.1)\.
- \[62\]K\. Krippendorff\(2004\)Content analysis: an introduction to its methodology\.2nd edition,Sage Publications,Thousand Oaks, CA\.Cited by:[§III\-C](https://arxiv.org/html/2609.27372#S3.SS3.p1.1)\.
- \[63\]J\. E\. Matheson and R\. L\. Winkler\(1976\)Scoring rules for continuous probability distributions\.Management Science22\(10\),pp\. 1087–1096\.Cited by:[§IV\-C](https://arxiv.org/html/2609.27372#S4.SS3.p1.1),[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.2)\.
- \[64\]B\. Efron and R\. J\. Tibshirani\(1993\)An introduction to the bootstrap\.Chapman and Hall,New York, NY\.Cited by:[§IV\-D](https://arxiv.org/html/2609.27372#S4.SS4.p1.3)\.相似文章
继续、适应或让步:全双工代理中对重叠语音的轮内适应
论文介绍了Duplex Cue,一个评估全双工语音代理轮内适应的框架,通过比较人类和模型对重叠语音的反应。
Enjoy Your Talk: 一个以人为中心的多轮对话基准,采用解耦的用户模拟、目标建模与评判
本文介绍了EYT-Bench,一个以人为中心的基准,用于评估多轮对话中的LLM,采用了解耦的用户模拟、目标建模和评判设计。它揭示了闭源和开源模型在客观意图追踪上差异显著,但在主观维度上相似;推理能力提升了客观追踪,而人物角色格式强烈影响轨迹分布。
TurnBench:一个用于口语对话中话轮转换动态的多领域基准测试(阅读时长25分钟)
TurnBench介绍了一个多领域基准测试,用于评估口语对话中的话轮转换动态,特点是拥有手工标注的语料库和标准化的评估协议,用于检测话轮结束和中断。
全双工语音对话模型中的同步与话轮转换
本文通过模拟两个Moshi模型实例之间的对话,利用CKA测量表征对齐并使用LSTM探针预测话轮边界,分析了全双工语音对话模型中的同步与话轮转换动态。
使用远距观看建模话轮转换:探究人类与AI生成话语中的沉默阈值
本文探讨了话轮转换中的沉默阈值如何在人类与AI生成的话语中有所不同,采用远距观看方法分析对话模式。