Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking
Summary
This paper proposes using semantic uncertainty derived from large language models to anticipate transition relevance places in spoken turn-taking, showing improved performance over baselines in dialogue systems.
View Cached Full Text
Cached at: 09/11/26, 08:20 AM
# Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking
Source: [https://arxiv.org/html/2609.10934](https://arxiv.org/html/2609.10934)
JP de RuiterAffiliation:Medford, Massachusetts, USAAffiliation:\{muhammad\.umair, jp\.deruiter\}@tufts\.edu
###### Abstract
Turn\-taking is a fundamental mechanism that governs when interlocutors speak and listen\. Although Spoken Dialogue Systems \(SDS\) exploit a range of linguistic, acoustic, and non\-verbal cues, they produce ill\-timed responses in unscripted interaction\. A central challenge is anticipating Transition Relevance Places \(TRPs\), or*opportunities*, not obligations, for a listener to take the floor\. Human listeners do not wait for turn endings; as an utterance unfolds, they use*expectations*about its developing meaning to anticipate TRPs and decide whether to take the floor\. We examine whether these evolving expectations can be modeled through*semantic uncertainty*—an LLM\-derived measure of how strongly a turn so far constrains what may plausibly come next\. To do so, we sample possible continuations of an ongoing turn and use changes in semantic dispersion to identify TRPs*within*turns\. We evaluate this account on a dataset with TRP labels derived from real\-time listener responses, rather than retrospective annotation\. Our approach substantially outperforms prompt\-based and fine\-tuned text\-only baselines, providing empirical support for the view that evolving semantic constraints inform perceived turn\-taking opportunities in unscripted interaction\.
## 1Introduction
Humans coordinate speaking and listening through a sophisticated turn\-taking mechanism\([Sacks et al\., 1974](https://arxiv.org/html/2609.10934#bib.bib50)\)\. Unlike in formal settings, where who speaks when may be predetermined, speaker selection in unscripted interaction is managed on a per\-turn basis\. This depends on conversationalists’ orientation to*Transition Relevance Places*\(TRPs\): points at which a turn may be treated as possibly complete and a listener may, but is not obligated to, respond\([Levinson, 1983](https://arxiv.org/html/2609.10934#bib.bib32);[Selting, 2000](https://arxiv.org/html/2609.10934#bib.bib54)\)\.
A key distinction concerns whether a TRP results in a speaker change\. If a listener takes the floor, the TRP occurs*between*turns and is directly observable\. Many TRPs, however, arise*within*turns, where a listener could have responded but did not\([Selting, 2000](https://arxiv.org/html/2609.10934#bib.bib54)\)\. These within\-turn TRPs are an integral part of the turn\-taking system but leave limited empirical trace in recorded data\([Umair et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib62);[Castillo\-López et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib11)\)\.
This creates a challenge for Spoken Dialogue Systems \(SDS\), which continue to produce ill\-timed responses and stilted feedback in unscripted interaction\([Algherairy and Ahmed, 2024](https://arxiv.org/html/2609.10934#bib.bib1);[Patamia et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib45);[Arora et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib4)\)\. Models trained on observable turn\-taking behavior \(speaker changes, backchannels, etc\.\) receive direct supervision for TRPs that listeners acted on, but not for the broader set of response opportunities they may have perceived\([Umair et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib62)\)\. This motivates a representational question: what information do humans use to anticipate TRPs, and can that be used to support their prediction in dialogue systems?
Human listeners do not wait for a turn to end; they recognize in advance where it may be complete\. This prospective recognition is known as*projection*\([Sacks et al\., 1974](https://arxiv.org/html/2609.10934#bib.bib50);[Schegloff, 1996](https://arxiv.org/html/2609.10934#bib.bib52)\)\. As each new word arrives, what can be projected changes, and listeners revise their*expectations*about how the turn may continue\([Magyari and de Ruiter, 2012](https://arxiv.org/html/2609.10934#bib.bib39);[Riest et al\., 2015](https://arxiv.org/html/2609.10934#bib.bib49)\)\.
We operationalize these evolving expectations as*semantic uncertainty*: an LLM\-derived measure of how an unfolding turn constrains what may plausibly come next\([Shorinwa et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib55)\)\. We estimate this signal by sampling possible continuations of the turn so far and measuring their semantic dispersion\([Nguyen et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib42)\)\. We ask whether changes in this uncertainty predict TRPs more accurately than direct baselines\. Our contribution is an empirical study of whether explicitly tracking evolving semantic constraint, derived from text alone, supports the prediction of perceived response opportunities in unscripted interaction\.
## 2Related Work
### 2\.1Turn\-Taking and the Role of Semantics
Transition Relevance Places \(TRPs\) are points in an utterance at which a listener could, but is not obligated to, initiate a response\([Sacks et al\., 1974](https://arxiv.org/html/2609.10934#bib.bib50)\)\. At these locations, a listener may take the floor, the current speaker may continue, or a listener may produce minimal contributions, such as backchannels \(e\.g\., hmm, uh\-huh;[Yngve](https://arxiv.org/html/2609.10934#bib.bib70)[1970](https://arxiv.org/html/2609.10934#bib.bib70)\) or continuers \(e\.g\., yeah, okay;[Schegloff](https://arxiv.org/html/2609.10934#bib.bib51)[1982](https://arxiv.org/html/2609.10934#bib.bib51)\)\. Depending on its timing, such a response may encourage the current speaker to continue or perform a specific conversational action\([Schegloff, 1982](https://arxiv.org/html/2609.10934#bib.bib51)\)\. Overlapping talk is likewise not necessarily interruptive\. Whether an entry is heard as competitive depends partly on where it begins relative to the developing turn and on how participants manage the overlap\([Schegloff, 2000](https://arxiv.org/html/2609.10934#bib.bib53);[Drew, 2009](https://arxiv.org/html/2609.10934#bib.bib14)\)\. While turn\-timing varies across cultures, the basic organization of rapid speaker transition is broadly shared across languages\([Stivers et al\., 2009](https://arxiv.org/html/2609.10934#bib.bib59)\)\. Given the importance of TRPs for coordinating turns, a central question emerges: how do listeners project TRPs as an utterance unfolds?
Experimental work suggests that, alongside prosodic and nonverbal cues, a turn’s developing lexico\-syntactic structure is particularly important for projecting possible completion\([de Ruiter et al\., 2006](https://arxiv.org/html/2609.10934#bib.bib13);[Levinson, 2016](https://arxiv.org/html/2609.10934#bib.bib33)\)\. Listeners accurately projected turn endings when lexico\-syntactic content was preserved but pitch was flattened to a monotone, whereas performance declined when words were made unrecognizable but intonation was retained\([de Ruiter et al\., 2006](https://arxiv.org/html/2609.10934#bib.bib13)\)\. They also anticipate upcoming words in a turn, the accuracy of which is closely linked to how well they estimate when a turn will end\([Magyari and de Ruiter, 2012](https://arxiv.org/html/2609.10934#bib.bib39)\)\.
Psycholinguistic accounts describe projection as an incremental process in which listeners form and revise expectations about how a turn may continue\([Riest et al\., 2015](https://arxiv.org/html/2609.10934#bib.bib49);[Levinson, 2016](https://arxiv.org/html/2609.10934#bib.bib33);[Magyari and de Ruiter, 2012](https://arxiv.org/html/2609.10934#bib.bib39)\)\. Each new word changes the lexico\-syntactic and semantic continuations that listeners may expect\([Tanenhaus et al\., 1995](https://arxiv.org/html/2609.10934#bib.bib60);[Altmann and Kamide, 1999](https://arxiv.org/html/2609.10934#bib.bib2);[Hale, 2001](https://arxiv.org/html/2609.10934#bib.bib20);[Levy, 2008](https://arxiv.org/html/2609.10934#bib.bib35)\)\. Some words narrow these possibilities, while extensions, repairs, qualifications, and redirections may introduce new ones and change how listeners expect the turn to develop\([Lerner, 1991](https://arxiv.org/html/2609.10934#bib.bib31);[Schegloff, 1996](https://arxiv.org/html/2609.10934#bib.bib52);[Selting, 2000](https://arxiv.org/html/2609.10934#bib.bib54);[Liddicoat, 2004](https://arxiv.org/html/2609.10934#bib.bib36)\)\.
Changes in semantic uncertainty may therefore reflect how the range of plausible meanings evolves over the course of a turn\. A decrease indicates that plausible continuations are becoming more similar in meaning, whereas an increase indicates greater variation\. Such changes may be relevant to within\-turn TRPs because semantic convergence may support anticipation of a continuation, while divergence may signal that an earlier expectation needs to be revised or deferred\([Magyari and de Ruiter, 2012](https://arxiv.org/html/2609.10934#bib.bib39);[Riest et al\., 2015](https://arxiv.org/html/2609.10934#bib.bib49)\)\.
Multimodal cues, such as prosody, gaze, and gesture, also play an important*disambiguating*role in TRP projection\. These cues are not interpreted independently of the unfolding semantics and pragmatics of a turn\([Kendrick et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib28)\)\. Rather, they help listeners distinguish possible completion from continuation and assess whether a response is relevant at a particular position\([Holler and Levinson, 2019](https://arxiv.org/html/2609.10934#bib.bib22)\)\. They are particularly informative in complex situations such as overlap, interruption, and miscommunication\([Skantze et al\., 2014](https://arxiv.org/html/2609.10934#bib.bib58);[Bögels and Torreira, 2015](https://arxiv.org/html/2609.10934#bib.bib8)\)\.
### 2\.2Computational Models of Turn\-Taking
Computational work has largely modeled turn\-taking through observable interactional outcomes\. TurnGPT predicts token\-level turn\-shift probabilities from text and speaker identity, while RC\-TurnGPT additionally conditions on a candidate system response\([Ekstedt and Skantze, 2020](https://arxiv.org/html/2609.10934#bib.bib15);[Jiang et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib26)\)\. Because both are trained on speaker changes, their targets are realized between\-turn transitions rather than within\-turn TRPs, which need not result in a listener taking the floor\([Threlkeld et al\., 2022](https://arxiv.org/html/2609.10934#bib.bib61)\)\.
Voice Activity Projection \(VAP\) models take a complementary approach by forecasting joint future speech activity from acoustic signals\. Turn holds, speaker switches, overlaps, and mutual silence can be derived from these forecasts\([Ekstedt and Skantze, 2022](https://arxiv.org/html/2609.10934#bib.bib16)\)\. VAP has been extended to multimodal, multi\-party, and multilingual settings and combined with text\-based approaches\([Onishi et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib44);[Inoue et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib25);[Leishman et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib30);[Wang et al\., 2024a](https://arxiv.org/html/2609.10934#bib.bib63);[Elmers et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib17)\)\.
Despite these advances, dialogue systems struggle to time responses\([Algherairy and Ahmed, 2024](https://arxiv.org/html/2609.10934#bib.bib1);[Patamia et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib45);[Arora et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib4)\)\. Predicting voice activity or speaker changes may therefore not fully capture when a response becomes interactionally relevant\([Liesenfeld et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib38)\)\.
This gap is difficult to study because many richly annotated interaction corpora organize turns around speaker changes\([Anderson et al\., 1991](https://arxiv.org/html/2609.10934#bib.bib3);[Jurafsky, 1997](https://arxiv.org/html/2609.10934#bib.bib27);[Calhoun et al\., 2010](https://arxiv.org/html/2609.10934#bib.bib9);[Reece et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib46)\), while task\-oriented corpora use system–user alternation\([Budzianowski et al\., 2018](https://arxiv.org/html/2609.10934#bib.bib7);[Si et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib56)\)\. These resources therefore primarily capture where speakers*did*respond rather than where they*could have*responded but did not\.
### 2\.3Semantic Uncertainty Quantification
Black\-box uncertainty quantification estimates variation in model outputs without requiring access to model output distributions, which may not be available for state\-of\-the\-art LLMs\([Liesenfeld and Dingemanse, 2024](https://arxiv.org/html/2609.10934#bib.bib37);[Shorinwa et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib55)\)\. Conventional measures such as normalized predictive entropy can conflate variation in meaning with variation in lexical form by treating semantically equivalent paraphrases as distinct outcomes\([Malinin and Gales, 2020](https://arxiv.org/html/2609.10934#bib.bib40);[Kuhn et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib29)\)\. Cluster\-based semantic uncertainty methods address this by grouping meaning\-equivalent outputs into semantic classes via entailment and computing entropy over those classes\([Kuhn et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib29);[Farquhar et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib18)\)\. More recent variants replace hard clustering with graded similarity through kernel\-based functions\([Nikitin et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib43)\)or pairwise sentence\-embedding similarities\([Nguyen et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib42)\)\.
These approaches have primarily been applied in tasks such as question answering, summarization, and translation\([Kuhn et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib29);[Farquhar et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib18);[Shorinwa et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib55)\)\. Less is known about whether these measures can characterize how an utterance becomes more or less semantically constrained as it unfolds\.
## 3Approach
### 3\.1Within\-Turn TRP Prediction Task
Following[Umair et al\. \(2024\)](https://arxiv.org/html/2609.10934#bib.bib62), we define our task as predicting*opportunities*, not obligations, for response within a single speaker’s unfolding turn\. As an utterance unfolds, the task is to predict whether it affords a possible listener response, regardless of whether a listener actually takes the floor\.
Formally, we define a single speaker’s turn as a*stimulus*S=⟨w1,…,wN⟩S=\\langle w\_\{1\},\\ldots,w\_\{N\}\\rangle, a sequence ofNNwords, where the number of words may vary across stimuli\. We further segment each stimulus into a series of*prefixes*, where each prefixPi=⟨w1,…,wi⟩P\_\{i\}=\\langle w\_\{1\},\\ldots,w\_\{i\}\\rangleconsists of the sequence of words from the first word throughwiw\_\{i\}\. We denote by𝒫S=⟨P1,…,PN⟩\\mathcal\{P\}\_\{S\}=\\langle P\_\{1\},\\ldots,P\_\{N\}\\ranglethe ordered set of all prefixes derived from a stimulusSS\.
For each prefixPiP\_\{i\}, we associate a binary*reference label*Ti∈\{0,1\}T\_\{i\}\\in\\\{0,1\\\}indicating whether the position immediately following its last wordwiw\_\{i\}is labeled as a TRP\. The resulting sequence𝐓S=⟨T1,…,TN⟩\\mathbf\{T\}\_\{S\}=\\langle T\_\{1\},\\ldots,T\_\{N\}\\rangleconstitutes a reference TRP labeling for a stimulus, withTNT\_\{N\}corresponding to the turn\-final position\. TRP predictions are therefore conditioned only on preceding linguistic material, reflecting the*causal*constraints under which listeners form TRP judgments\. Section[4\.1](https://arxiv.org/html/2609.10934#S4.SS1)describes the procedure by which the reference labels𝐓S\\mathbf\{T\}\_\{S\}are constructed from the listener\-response dataset\.
Algorithm 1Estimating Semantic Uncertainty for a Prefix1:Prefix
Pi=⟨w1,…,wi⟩P\_\{i\}=\\langle w\_\{1\},\\ldots,w\_\{i\}\\rangle; language model
ℳ\\mathcal\{M\}; number of continuations
K∈ℕ≥1K\\in\\mathbb\{N\}\_\{\\geq 1\}; maximum continuation length
L∈ℕ≥1L\\in\\mathbb\{N\}\_\{\\geq 1\}; embedding model
ℰ:Seq→ℝd\\mathcal\{E\}:\\text\{Seq\}\\rightarrow\\mathbb\{R\}^\{d\}; similarity scaling parameter
τ\>0\\tau\>0
2:Semantic uncertainty value
Ui∈ℝU\_\{i\}\\in\\mathbb\{R\}
3:Sample
KKstochastic continuations
Ai=\{ai,1,…,ai,K\}A\_\{i\}=\\\{a\_\{i,1\},\\ldots,a\_\{i,K\}\\\}from
ℳ\\mathcal\{M\}conditioned on
PiP\_\{i\}, each with maximum length
LLwords
4:for
j=1j=1to
KKdo
5:Compute continuation embedding
hi,j←ℰ\(ai,j\)h\_\{i,j\}\\leftarrow\\mathcal\{E\}\(a\_\{i,j\}\)
6:endfor
7:Construct semantic similarity matrix
S\(i\)∈ℝK×KS^\{\(i\)\}\\in\\mathbb\{R\}^\{K\\times K\}, where
Su,v\(i\)=cos\(hi,u,hi,v\)S^\{\(i\)\}\_\{u,v\}=\\cos\\\!\\left\(h\_\{i,u\},h\_\{i,v\}\\right\)
8:Compute semantic uncertainty
Ui←SNNE\(S\(i\)\)U\_\{i\}\\leftarrow\\mathrm\{SNNE\}\(S^\{\(i\)\}\)⊳\\trianglerightsee Eq\. \([1](https://arxiv.org/html/2609.10934#S3.E1)\)
9:return
UiU\_\{i\}
###### Definition 3\.1\(Within\-turn TRP Prediction\)
Given a stimulusSSand its associated prefix sequence𝒫S\\mathcal\{P\}\_\{S\}, produce a predicted binary label sequence𝐓^S=⟨T^1,…,T^N⟩\\widehat\{\\mathbf\{T\}\}\_\{S\}=\\langle\\widehat\{T\}\_\{1\},\\ldots,\\widehat\{T\}\_\{N\}\\rangle, where eachT^i∈\{0,1\}\\widehat\{T\}\_\{i\}\\in\\\{0,1\\\}indicates whether the position immediately followingwiw\_\{i\}is predicted to afford a possible transition, based solely on the prefixPiP\_\{i\}\. Predictions are evaluated against the corresponding reference labeling𝐓S\\mathbf\{T\}\_\{S\}\.
### 3\.2Predicting TRPs via Semantic Uncertainty
To support TRP prediction, we introduce*semantic uncertainty*as an intermediate representation\. For each prefixPiP\_\{i\}, semantic uncertainty is a scalar valueUi∈ℝU\_\{i\}\\in\\mathbb\{R\}that reflects how strongly the prefix constrains what may plausibly come next\. We use*semantic*to refer to variation in meaning among possible model outputs, rather than variation in surface form alone\([Kuhn et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib29);[Farquhar et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib18);[Nguyen et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib42)\)\. Intuitively,UiU\_\{i\}is low when the possible continuations ofPiP\_\{i\}are similar in meaning, and high when they differ more substantially\. Applied across the prefix sequence𝒫S\\mathcal\{P\}\_\{S\}, this yields a causal trajectory𝐔S=⟨U1,…,UN⟩\\mathbf\{U\}\_\{S\}=\\langle U\_\{1\},\\ldots,U\_\{N\}\\ranglethat tracks how the space of plausible continuations narrows or widens as the turn unfolds\. Section[3\.3](https://arxiv.org/html/2609.10934#S3.SS3)describes how we estimate this quantity\.
We do not associate TRPs with the absolute uncertainty value of a prefix\. Instead, we relate TRPs to*changes*in uncertainty as the turn unfolds \(Section[2\.1](https://arxiv.org/html/2609.10934#S2.SS1)\)\. These changes reflect shifts in how strongly the utterance constrains its plausible continuations, highlighting locations where a listener may, but is not obligated to, respond\.
We further define a TRP decision functionffthat maps a causal, variable\-length prefix of semantic uncertainty values to a binary label\. Fori=1,…,Ni=1,\\ldots,N, this label is given byT^i=f\(\(𝐔S\)1:i\)\\widehat\{T\}\_\{i\}=f\\\!\\left\(\\left\(\\mathbf\{U\}\_\{S\}\\right\)\_\{1:i\}\\right\), where\(𝐔S\)1:i=⟨U1,…,Ui⟩\\left\(\\mathbf\{U\}\_\{S\}\\right\)\_\{1:i\}=\\langle U\_\{1\},\\ldots,U\_\{i\}\\rangle\. This formulation is agnostic to the specific choice offf: it may operate over absolute uncertainty values, local changes, or aggregated statistics over time\. Section[3\.4](https://arxiv.org/html/2609.10934#S3.SS4)describes our concrete instantiation\.
### 3\.3Estimating Semantic Uncertainty
We estimateUiU\_\{i\}using the procedure summarized in Algorithm[1](https://arxiv.org/html/2609.10934#alg1)\. For each prefixPiP\_\{i\}, we sample continuations from a language model and compare their similarity in a sentence\-embedding space\. We then aggregate their pairwise similarities using*Semantic Nearest Neighbor Entropy*\([Nguyen et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib42), SNNE; see Equation[1](https://arxiv.org/html/2609.10934#S3.E1);\)\. SNNE is appropriate here because it uses sampled continuations rather than token\-level probability distributions, and because it compares continuations using pairwise semantic similarity instead of first grouping them into entailment\-based semantic clusters \([Kuhn et al\. 2023](https://arxiv.org/html/2609.10934#bib.bib29);[Farquhar et al\. 2024](https://arxiv.org/html/2609.10934#bib.bib18);[Nguyen et al\. 2025](https://arxiv.org/html/2609.10934#bib.bib42); see Section[2\.3](https://arxiv.org/html/2609.10934#S2.SS3)\)\. The resulting scalarUiU\_\{i\}reflects the semantic dispersion of the continuations\. We treat this embedding\-space dispersion as a model\-mediated proxy for semantic variation, without assuming independence from lexical form\. Appendix[E\.4](https://arxiv.org/html/2609.10934#A5.SS4)provides support for this treatment\.
SNNE\(S\(i\)\)\\displaystyle\\mathrm\{SNNE\}\(S^\{\(i\)\}\)=−1K∑u=1Klog\(∑v=1Kexp\(Su,v\(i\)τ\)\)\\displaystyle=\-\\frac\{1\}\{K\}\\sum\_\{u=1\}^\{K\}\\log\\left\(\\sum\_\{v=1\}^\{K\}\\exp\\\!\\left\(\\frac\{S^\{\(i\)\}\_\{u,v\}\}\{\\tau\}\\right\)\\right\)\(1\)
We measure semantic proximitySu,v\(i\)S^\{\(i\)\}\_\{u,v\}using cosine similarity, a standard measure for comparing sentence embeddings\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.10934#bib.bib47)\)\. The similarity scaling parameterτ\\taucontrols how strongly SNNE is influenced by nearest neighbors in the similarity matrix: smaller values emphasize the most similar continuation pairs, while larger values distribute weight more broadly across pairwise similarities\. Under this formulation, more negative SNNE values indicate more semantically constrained continuations, while less negative values indicate greater indeterminacy\. Applying Algorithm[1](https://arxiv.org/html/2609.10934#alg1)across all prefixes of a stimulus yields a prefix\-indexed uncertainty trajectory𝐔S=⟨U1,…,UN⟩\\mathbf\{U\}\_\{S\}=\\langle U\_\{1\},\\ldots,U\_\{N\}\\rangle, which tracks how semantic constraints evolve as the speaker’s turn unfolds\.
Algorithm 2Causal TRP decision rule for a PrefixT^i=f\(\(𝐔S\)1:i\)\\widehat\{T\}\_\{i\}=f\(\\left\(\\mathbf\{U\}\_\{S\}\\right\)\_\{1:i\}\)1:Index
i∈\{1,…,N\}i\\in\\\{1,\\ldots,N\\\}; Uncertainty prefix
\(𝐔S\)1:i=⟨U1,…,Ui⟩∈ℝi\\left\(\\mathbf\{U\}\_\{S\}\\right\)\_\{1:i\}=\\langle U\_\{1\},\\ldots,U\_\{i\}\\rangle\\in\\mathbb\{R\}^\{i\}; Window size
w∈ℕ≥1w\\in\\mathbb\{N\}\_\{\\geq 1\}; Causal smoothing operator
ℱ\\mathcal\{F\}; Threshold
θ∈ℝ\>0\\theta\\in\\mathbb\{R\}\_\{\>0\}; Constant
ε\>0\\varepsilon\>0
2:Predicted label
T^i∈\{0,1\}\\widehat\{T\}\_\{i\}\\in\\\{0,1\\\}
3:
j←max\(2,i−w\+1\)j\\leftarrow\\max\(2,i\-w\+1\)
4:if
j≥ij\\geq ithenreturn
00
5:endif
6:
Δ𝐔S←⟨0,U2−U1,…,Ui−Ui−1⟩\\Delta\\mathbf\{U\}\_\{S\}\\leftarrow\\langle 0,\\ U\_\{2\}\-U\_\{1\},\\ldots,U\_\{i\}\-U\_\{i\-1\}\\rangle
7:
Δ𝐔S~←ℱ\(Δ𝐔S\)\\Delta\\tilde\{\\mathbf\{U\}\_\{S\}\}\\leftarrow\\mathcal\{F\}\(\\Delta\\mathbf\{U\}\_\{S\}\)⊳\\trianglerightℱ\\mathcal\{F\}uses only indices≤i\\leq i
8:
Wi←⟨\(Δ𝐔S~\)j,…,\(Δ𝐔S~\)i−1⟩W\_\{i\}\\leftarrow\\langle\\left\(\\Delta\\tilde\{\\mathbf\{U\}\_\{S\}\}\\right\)\_\{j\},\\ldots,\\left\(\\Delta\\tilde\{\\mathbf\{U\}\_\{S\}\}\\right\)\_\{i\-1\}\\rangle
9:
b←median\(Wi\)b\\leftarrow\\mathrm\{median\}\(W\_\{i\}\),
s←MAD\(Wi\)s\\leftarrow\\mathrm\{MAD\}\(W\_\{i\}\)⊳\\trianglerightMAD\(X\)=median\(\|X−median\(X\)\|\)\\mathrm\{MAD\}\(X\)=\\mathrm\{median\}\(\|X\-\\mathrm\{median\}\(X\)\|\)
10:
z←\|\(Δ𝐔S~\)i−b\|s\+εz\\leftarrow\\dfrac\{\|\\left\(\\Delta\\tilde\{\\mathbf\{U\}\_\{S\}\}\\right\)\_\{i\}\-b\|\}\{s\+\\varepsilon\}
11:return
T^i←𝕀\[z\>θ\]\\widehat\{T\}\_\{i\}\\leftarrow\\mathbb\{I\}\[z\>\\theta\]
### 3\.4Rule over Uncertainty Dynamics
We define our decision functionffbased on the view that TRPs arise from local shifts in semantic constraint \(see Section[2\.1](https://arxiv.org/html/2609.10934#S2.SS1)\)\. Accordingly,ffidentifies candidate TRPs from both convergent shifts, where the utterance becomes more constrained, and divergent shifts, where unexpected extensions occur\. We adopt a deterministic rule to emphasize interpretability and leave more expressive realizations, such as learned models, for future work\.
Algorithm[2](https://arxiv.org/html/2609.10934#alg2)operationalizes this intuition\. It defines uncertainty change as discrete differences between successive prefixes\. Because uncertainty estimates are based on finite continuation samples, the change signal may contain sampling noise rather than genuine shifts\. The algorithm therefore attenuates this noise while preserving local change structure using a low\-pass filter\. It evaluates the current smoothed change relative to a trailing window of past changes to obtain a locally adaptive baseline\. It quantifies how exceptional the current change is using the median and Median Absolute Deviation \(MAD\) and predicts a TRP when the normalized value exceeds a fixed threshold\.
## 4Experimental Setup
### 4\.1Data: Empirical Within\-Turn TRPs
We use the participant\-response dataset111Stimulus and participant\-response audios are available at[https://osf\.io/k5pc9/overview?view\_only=5124d862448f4435b775d49a7b299d6d](https://osf.io/k5pc9/overview?view_only=5124d862448f4435b775d49a7b299d6d)\.released by[Umair et al\. \(2024\)](https://arxiv.org/html/2609.10934#bib.bib62), which consists of 55 single\-speaker stimuli and responses from 118 participants\. In the original study, the stimuli were organized into four presentation lists: two original lists and a reversed\-order version of each\. The reversed lists were used to counterbalance stimulus\-order effects, which refer to differences in response behavior caused by where a stimulus appeared in the sequence\([de Ruiter et al\., 2006](https://arxiv.org/html/2609.10934#bib.bib13);[Riest et al\., 2015](https://arxiv.org/html/2609.10934#bib.bib49)\)\.
The elicitation paradigm captured*within\-turn*response opportunities through real\-time listener responses rather than post\-hoc annotations or observed speaker changes\. Each participant was randomly assigned to one list and asked to verbalize brief backchannels \(e\.g\., ‘hmm’, ‘yes’\) when they perceived an opportunity for speech\.
Figure 1:Schematic illustration of empirical within\-turn TRP labeling\.SSdenotes the speaker’s word sequence,RRdenotes aggregated listener response distribution, and𝐓S\\mathbf\{T\}\_\{S\}denotes the resulting binary reference labels\. Participant response onsets are aligned to word\-adjacent intervalsIi,i\+1I\_\{i,i\+1\}, each corresponding to the position following prefixPiP\_\{i\}\. Intervals are labeled positive when their empirical response proportionIi,i\+1ProportionI\_\{i,i\+1\}^\{Proportion\}exceeds the stimulus\-specific threshold\. The example text is adapted from[Schegloff \(1982\)](https://arxiv.org/html/2609.10934#bib.bib51); the figure is inspired by[Umair et al\. \(2024\)](https://arxiv.org/html/2609.10934#bib.bib62)\.StatisticValueReleased dataset\([Umair et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib62)\)Stimuli55Participants118Participants per stimulus, median \(range\)58 \(51–60\)Reconstructed positions and labelsCandidate TRP positions \(total prefixes\)5,195TRP\-positive positions842Positive\-label prevalence \(842 / 5,195\)16\.2%DiagnosticLocal maxima among positive labels \(of 842\)72%Table 1:Dataset statistics\. The 16\.2% prevalence is computed over all candidate positions; the 72% local\-maximum statistic is computed over the 842 positive labels only and is diagnostic rather than definitional\.We derive the reference TRP labels𝐓S\\mathbf\{T\}\_\{S\}by aligning participant response onsets with the speaker’s word\-level timing\.[Umair et al\. \(2024\)](https://arxiv.org/html/2609.10934#bib.bib62)released stimulus and per\-participant response recordings, but not the orthographic transcripts, word\-level stimulus timings, or response\-onset timings required here\. We reconstructed these manually using ELAN and Praat\([Wittenburg et al\., 2006](https://arxiv.org/html/2609.10934#bib.bib67);[Boersma and Weenink, 2009](https://arxiv.org/html/2609.10934#bib.bib5)\)\. We excluded clear non\-response vocalizations, such as breaths, laughter, and throat clearings\. Because the original study reported no order effects, we pooled responses across original and reversed presentations of each stimulus, yielding a mean of 59 participants per stimulus \(median 58; see Table[1](https://arxiv.org/html/2609.10934#S4.T1)\)\.
For each stimulus, we define word\-adjacent intervals⟨I1,2,…,IN−1,N⟩\\langle I\_\{1,2\},\\ldots,I\_\{N\-1,N\}\\rangle\. Each intervalIi,i\+1I\_\{i,i\+1\}corresponds to the candidate TRP position immediately following prefixPiP\_\{i\}and its reference labelTiT\_\{i\}\. We compute its empirical response proportionIi,i\+1ProportionI\_\{i,i\+1\}^\{Proportion\}as the fraction of participants who responded within the interval, counting each participant at most once \(see Figure[1](https://arxiv.org/html/2609.10934#S4.F1)\)\.
To account for temporally dispersed response onsets and varying baseline responsiveness, we binarize these proportions using a stimulus\-specific threshold\. Specifically,Ti=1T\_\{i\}=1whenIi,i\+1Proportion\>μ\+0\.05σI\_\{i,i\+1\}^\{Proportion\}\>\\mu\+0\.05\\sigma, whereμ\\muandσ\\sigmaare the mean and standard deviation of the nonzero response proportions for that stimulus; otherwise,Ti=0T\_\{i\}=0\. Zero\-response intervals are excluded because they would pullμ\\mutoward zero and make the threshold uninformative\. The turn\-final positionTNT\_\{N\}is always labeled positive\. This yields 842 positive positions out of 5,195 \(16\.2%; see Table[1](https://arxiv.org/html/2609.10934#S4.T1)\)\. Among these, 72% are local maxima relative to adjacent intervals\. This statistic is diagnostic, not definitional: local maximality is not part of the labeling procedure\. Additional diagnostics are provided in Appendix[A](https://arxiv.org/html/2609.10934#A1); implications of the dataset size are discussed in Section[8](https://arxiv.org/html/2609.10934#S8)\.
### 4\.2Evaluation Metrics
The sparse nature of intervals in our data labeled as TRPs \(16\.2% of 5,195; see Table[1](https://arxiv.org/html/2609.10934#S4.T1)\) highlights that TRP prediction is highly imbalanced\. When predictions control system turn entry, a false positive may cause the system to interrupt the user’s ongoing turn, whereas a missed TRP may delay a response\([Skantze, 2021](https://arxiv.org/html/2609.10934#bib.bib57);[Arora et al\., 2025](https://arxiv.org/html/2609.10934#bib.bib4)\)\. For dialogue systems, false\-positive TRP predictions may therefore carry greater cost than missed ones\.
Accuracy is uninformative in this setting; majority\-class predictors can score well while failing to detect rare but meaningful events\([He and Garcia, 2009](https://arxiv.org/html/2609.10934#bib.bib21)\)\. Instead, we reportF0\.5F\_\{0\.5\}, which weighs precision more heavily than recall, and the True Negative Rate \(TNR\), which measures suppression of spurious predictions\. For comparability with prior work, we report balanced accuracy, which treats false positives and negatives symmetrically and therefore does not reflect the asymmetric interactional cost of mistimed turn entries\. Section[8](https://arxiv.org/html/2609.10934#S8)discusses the limitations of this approach\.
Across experiments, we report aggregated metrics using a two\-level macro\-average rather than pooling predictions across all stimuli\. Because stimuli vary in length and number of TRPs, pooling predictions would allow the most TRP\-dense stimuli to dominate\. We therefore compute each metric independently for each stimulus, average across stimuli within each list, and then average across lists\. This gives equal weight to each stimulus and each list in the final aggregate\.
## 5Experiments and Results
### 5\.1Prompt\-Based TRP Prediction
ExpertParticipantImaginedOracleModelF0\.5F\_\{0\.5\}PTNRBAF0\.5F\_\{0\.5\}PTNRBAF0\.5F\_\{0\.5\}PTNRBAF0\.5F\_\{0\.5\}PTNRBALLaMA\-3\.1\-8B0\.120\.120\.810\.50\.100\.120\.840\.490\.180\.160\.480\.500\.170\.160\.620\.49LLaMA\-3\.1\-70B0\.180\.220\.860\.530\.220\.210\.790\.540\.210\.240\.900\.540\.140\.170\.900\.53Mistral\-7B0\.100\.100\.790\.500\.120\.120\.770\.510\.200\.170\.350\.530\.190\.160\.010\.49Mixtral\-8x7B0\.060\.130\.970\.500\.040\.110\.990\.500\.100\.160\.910\.510\.170\.150\.530\.51Qwen2\.5\-7B0\.090\.110\.820\.500\.050\.090\.920\.500\.110\.150\.910\.510\.010\.020\.970\.50Qwen2\.5\-72B0\.080\.200\.990\.510\.080\.140\.970\.510\.210\.310\.930\.540\.160\.220\.880\.52Table 2:Prompt\-based within\-turn TRP prediction across instruction\-tuned models \(see Table[5](https://arxiv.org/html/2609.10934#A2.T5)\) and prompting conditions\. Expert and participant are direct prompting conditions; imagined and oracle are future\-context variants of the participant\-style prompt\. P denotes precision and BA denotes balanced accuracy\. Metrics are macro\-averaged as described in Section[4\.2](https://arxiv.org/html/2609.10934#S4.SS2)\. The highestF0\.5F\_\{0\.5\}value within each prompting condition is bolded\.We first test whether stronger instruction\-tuned models improve prompt\-based TRP prediction relative to[Umair et al\. \(2024\)](https://arxiv.org/html/2609.10934#bib.bib62)\. Given mixed evidence on whether detailed task\-specific prompts improve performance \(see[Reynolds and McDonell 2021](https://arxiv.org/html/2609.10934#bib.bib48);[Webson and Pavlick 2022](https://arxiv.org/html/2609.10934#bib.bib65);[Xu et al\. 2023](https://arxiv.org/html/2609.10934#bib.bib69)\), we compare four prompting conditions\. The*expert*condition provides an explicit, theory\-driven definition of TRPs, while the*participant*condition mirrors the listener instructions used to collect the dataset\. Inspired by response\-conditioned models \(see[Jiang et al\. 2023](https://arxiv.org/html/2609.10934#bib.bib26)\), we also test two future\-context variants: an*imagined*condition, where the model generates a plausible continuation before predicting, and an*oracle*condition, where the true upcoming same\-speaker words are provided\.
Table[2](https://arxiv.org/html/2609.10934#S5.T2)reports results for all six instruction\-tuned models \(see Appendix[B](https://arxiv.org/html/2609.10934#A2)\) across the four prompting conditions\. Across models, performance remains weak: balanced accuracy stays near chance, and both precision andF0\.5F\_\{0\.5\}are low\. Although the imagined condition yields the best within\-modelF0\.5F\_\{0\.5\}for four of six models, no prompting condition reliably identifies TRPs\. Oracle access to the true upcoming words also does not consistently improve performance\. Overall, direct prompting does not recover within\-turn TRPs, even when possible or actual same\-speaker future context is available\. Experimental details are in Appendix[C](https://arxiv.org/html/2609.10934#A3); prompt templates are in Appendix[E\.4](https://arxiv.org/html/2609.10934#A5.SS4)\. The prompts include no demonstrations from the evaluation dataset, making the setup zero\-shot with respect to the target distribution\([Brown et al\., 2020](https://arxiv.org/html/2609.10934#bib.bib6)\)\.
### 5\.2Supervised Fine\-Tuning \(SFT\) Based TRP Prediction
We next use supervised fine\-tuning \(SFT\), which, unlike prompting, updates model weights, to test whether explicit task adaptation improves within\-turn TRP prediction\. Because prompting shows no consistent benefit from model scale, we fine\-tune LLaMA\-3\.1\-8B\-Instruct with LoRA adapters rather than the strongest prompt\-based model, LLaMA\-3\.1\-70B\-Instruct \(Table[2](https://arxiv.org/html/2609.10934#S5.T2);[Hu et al\. 2021](https://arxiv.org/html/2609.10934#bib.bib24)\)\. Larger models may behave differently under SFT; we leave this comparison to future work\.
We construct training examples by pairing each prefixPiP\_\{i\}with its binary reference labelTiT\_\{i\}\. To avoid leakage across prefixes from the same turn, we split data by stimulus rather than prefix\. The train, validation, and test splits cover 60%, 20%, and 20% of the 55 stimuli, yielding 3,098, 1,203, and 894 prefixes, with 505, 188, and 149 TRP\-positive labels, respectively\. The held\-out test split preserves the original class imbalance\.
To assess whether performance is limited by sparse TRP\-positive training examples, we construct three nested supervision regimes—Low, Medium, and Full—from the training split\. For each regime, we select a target number of TRP\-positive prefixes and include adjacent negatives to preserve local context around TRP\-labeled positions\([He and Garcia, 2009](https://arxiv.org/html/2609.10934#bib.bib21)\)\. Validation and test splits are fixed across regimes i\.e\., the ablation varies only the training subset \(see Table[3](https://arxiv.org/html/2609.10934#S5.T3)\)\.
Table[3](https://arxiv.org/html/2609.10934#S5.T3)shows that supervision yields modest gains at the Medium regime, but improvements are limited and non\-monotonic\. Because results are based on a single stimulus\-level split, we interpret these trends qualitatively\. The pattern suggests that direct supervision provides some benefit, but performance does not consistently scale with additional data\. Full training details are in Appendix[D](https://arxiv.org/html/2609.10934#A4)\.
Training dataTest performanceRegimeSizeTRPsTRP %𝐅0\.5\\mathbf\{F\_\{0\.5\}\}PTNRBALow52815228\.790\.140\.170\.930\.51Medium1,01130329\.970\.270\.280\.890\.57Full1,52650533\.090\.150\.210\.910\.50Table 3:SFT training\-regime TRP prediction performance for LLaMA\-3\.1\-8B\-Instruct on a held\-out test set\. P denotes precision and BA denotes balanced accuracy\. Metrics are macro\-averaged \(Section[4\.2](https://arxiv.org/html/2609.10934#S4.SS2)\); the highest value in each performance column is bolded\.
### 5\.3Semantic\-Uncertainty\-Based TRP Prediction
Figure 2:Distribution ofF0\.5F\_\{0\.5\}across the fixed\-KKsemantic\-uncertainty sweep\. Each curve is an empirical cumulative distribution over configurations for one embedding model\. Mean±\\pmSD across configurations: SFR\-2R=0\.547±0\.033=0\.547\\pm 0\.033, all\-mpnet\-base\-v2=0\.540±0\.033=0\.540\\pm 0\.033, E5\-Mistral\-7B\-Instruct=0\.546±0\.031=0\.546\\pm 0\.031\.We now evaluate our proposed method \(see Section[3\.2](https://arxiv.org/html/2609.10934#S3.SS2)\)\. For the main evaluation, we fix sampling atK=25K=25continuations per prefix as well as the sampling model \(LLaMA\-3\.1\-8B\-Instruct\), and sweep the remaining configurations across qualitatively distinct settings within our compute budget\. The uncertainty\-estimation stage \(Algorithm[1](https://arxiv.org/html/2609.10934#alg1)\) spans 6 decoding settings, 3 embedding models \(110M–7B parameters\), and 4 SNNEτ\\tauvalues, yielding 72 uncertainty\-signal configurations\. The decision stage \(Algorithm[2](https://arxiv.org/html/2609.10934#alg2)\) applies 32 detector configurations to each trajectory, for72×32=2,30472\\times 32=2\{,\}304total configurations \(see Table[7](https://arxiv.org/html/2609.10934#A5.T7)\)\. Grid density reflects computational cost: we limit expensive continuation sampling to 6 settings but sweep the inexpensive decision stage over saved trajectories\. The grid was fixed before evaluation, and no configuration was selected using test performance; we report results across the full grid\. Full configuration details appear in Appendix[E\.1](https://arxiv.org/html/2609.10934#A5.SS1)\.
Figure[2](https://arxiv.org/html/2609.10934#S5.F2)summarizes the main fixed\-KKsweep\. Across configurations, the semantic\-uncertainty pipeline achieves a meanF0\.5F\_\{0\.5\}of 0\.545 \(SD=0\.032\\mathrm\{SD\}=0\.032\), ranging from 0\.385 to 0\.603\. The embedding\-specific distributions are closely aligned, with meanF0\.5F\_\{0\.5\}values ranging from 0\.540 to 0\.547\. Appendix[E\.4](https://arxiv.org/html/2609.10934#A5.SS4)complements this comparison by showing that the embedding space captures semantic similarity across lexical variation\. Results are also stable across continuation\-sampling regimes: more stochastic decoding increases continuation diversity and average SNNE, while the same prefixes remain in similar semantic regions across regimes \(see Appendix[E\.2](https://arxiv.org/html/2609.10934#A5.SS2)\)\.
A one\-wayη2\\eta^\{2\}decomposition shows that 75\.4% ofF0\.5F\_\{0\.5\}variance is associated with the decision ruleff; smoothing choices alone explain 52\.6%\. Upstream construction choices contribute much less: decoding setting, embedding model, andτ\\tauexplain 5\.5%, 0\.9%, and 0\.88%, respectively\. Although these one\-way effects are not additive, they indicate that SNNE preserves the TRP\-relevant signal across construction choices\. Varyingτ\\tauchanges the range of the uncertainty trajectory, but has only a small downstream effect \(see Appendix[E\.3](https://arxiv.org/html/2609.10934#A5.SS3)\)\.
The dominance of the decision rule and near\-identical performance across embedding models raise the possibility that the gains come from the rule \(Algorithm[2](https://arxiv.org/html/2609.10934#alg2)\) rather than the semantic information carried by SNNE \(Algorithm[1](https://arxiv.org/html/2609.10934#alg1)\)\. To test this, we replace the SNNE trajectory with two predictive\-entropy controls while holding the prefixes, decision rule, evaluation, and generation model \(LLaMA\-3\.1\-8B\-Instruct\) fixed\.*Next\-token entropy*\(NTE\) is the Shannon entropy of the model’s next\-token distribution at each prefix and tests whether the rule can exploit token\-level uncertainty without sampled continuations or embeddings\.*Normalized predictive entropy*\(NPE\) averages length\-normalized mean token surprisal over the sameK=25K=25continuations, testing whether their semantic dispersion adds information beyond their token probabilities\. NTE and NPE achieve meanF0\.5F\_\{0\.5\}scores of0\.3190\.319and0\.3280\.328, respectively \(see Table[4](https://arxiv.org/html/2609.10934#S5.T4)\)\. Their maxima,0\.3560\.356and0\.3760\.376, remain below the minimum semantic\-uncertainty \(F0\.5=0\.385F\_\{0\.5\}=0\.385\), showing that the decision rule is not signal\-agnostic and that the embedding\-based semantic dispersion captured by SNNE contributes information beyond token\-probability uncertainty\.
Approach𝐅0\.5\\mathbf\{F\}\_\{0\.5\}GPT\-4 Omni, participant\([Umair et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib62)\)†0\.15Prompt\-based prediction \(best\)0\.22Supervised fine\-tuning \(best\)‡0\.27Next\-token entropy \(NTE\)0\.319±0\.0700\.319\\pm 0\.070Normalized predictive entropy \(NPE\)0\.328±0\.0270\.328\\pm 0\.027Semantic uncertainty0\.545±0\.032\\mathbf\{0\.545\\pm 0\.032\}Table 4:SummaryF0\.5F\_\{0\.5\}comparison\. Baseline rows report the best result; uncertainty\-signal rows report mean±\\pmSD across configurations \(32 per alternative signal and 2,304 for semantic uncertainty\)\.†Approximated from the reported precision and recall under global\-threshold labels\.‡Evaluated on the held\-out split\.We next test whether performance using SNNE depends on continuation\-sample sizeKK\. Larger samples may improve dispersion estimates but increase sampling cost\. Holding the decision\-rule sweep fixed, we varyKKstarting at 2, the smallest nontrivial pairwise setting\. MeanF0\.5F\_\{0\.5\}reaches approximately 0\.54 atK=2K=2and remains near that level for larger samples, with a shallow optimum around four to five continuations \(see Figure[3](https://arxiv.org/html/2609.10934#S5.F3)\)\. This stability suggests that local uncertainty shifts are recoverable from small samples\. Moreover, SNNE atK=2K=2outperforms NPE atK=25K=25, indicating that its advantage is not explained by sample size alone\.
Figure 3:Continuation\-sample size ablation\. For eachKK, downstreamF0\.5F\_\{0\.5\}is collapsed over the decision\-rule sweep\. Lines show mean and median performance; shaded bands show interquartile range and±1\\pm 1std\.
## 6Discussion
Listeners in unscripted interaction project TRPs before a speaker’s turn has ended, using the turn\-so\-far to form expectations about what may plausibly come next\. This remains difficult for dialogue systems, partly because they are typically trained on observable outcomes rather than response opportunities that do not result in speaker change\.
We therefore ask whether evolving semantic constraints help identify within\-turn TRPs\. We estimate semantic uncertainty over sampled continuations and use local changes in its trajectory as the prediction signal\. Because this procedure is computationally costly, we leave its deployment in dialogue systems as an avenue for future work\.
Empirically, prompt\-based inference and SFT do not reliably predict within\-turn TRPs\. Prompting yields lowF0\.5F\_\{0\.5\}across models and conditions, while SFT produces only limited, non\-monotonic gains\. In contrast, SNNE\-based prediction performs substantially better\. Most performance variation is associated with the decision rule, but neither predictive\-entropy control reaches the SNNE performance range under the same rule, indicating that SNNE captures TRP\-relevant uncertainty beyond token\-probability uncertainty\. Performance remains stable from two sampled continuations onward, suggesting that informative local changes in SNNE are recoverable from small samples\.
Together, these findings suggest that evolving semantic structure provides a useful signal for within\-turn TRP prediction\. This is notable because our labels derive from real\-time listener responses rather than retrospective annotation, targeting opportunities that listeners perceived but did not necessarily act on\. Semantic uncertainty should nonetheless be read as a model\-mediated proxy for semantic variation, not a measurement of listener\-internal expectations\. For dialogue systems, this suggests that text\-only TRP prediction may benefit from intermediate representations of semantic constraint\.
## 7Conclusion
Even though spoken dialogue systems draw on linguistic, acoustic, and multimodal cues for turn\-taking, they continue to struggle to time responses appropriately in natural interaction\. One reason is that TRPs, or opportunities for speech, are not confined to turn endings\. Many occur within turns and leave no interactional trace when a listener does not respond, limiting their visibility in dialogue corpora\. Models trained on observable outcomes therefore receive evidence about where listeners did respond, but not necessarily about where they could have responded\. Our analyses show that this mismatch has measurable consequences\. Direct text\-to\-label prediction remains unreliable, even across diverse prompting conditions and with supervised fine\-tuning\. By contrast, semantic\-uncertainty\-based prediction provides a more reliable basis for identifying within\-turn TRPs by tracking how the space of plausible continuations becomes more or less semantically dispersed as an utterance unfolds\. The improvement is robust across the tested settings, but the broader contribution is representational\. Semantic uncertainty shows that changes in the space of plausible continuations carry information about when response opportunities arise\. Progress on turn\-taking may therefore require models that track how semantic constraints evolve over a turn, not only models trained on observable response outcomes\.
## 8Limitations
We acknowledge several limitations\. First, our empirical evaluation is based on a small, English\-only dataset of single\-speaker turns labeled through participant judgments of perceived TRPs \(see Section[4\.1](https://arxiv.org/html/2609.10934#S4.SS1)\)\. This dataset is well suited to studying within\-turn TRPs, which are rarely observable in standard corpora\. At the same time, it limits generalization\. Our findings have not yet been validated on widely used dialogue datasets such as Switchboard or SpokenWOZ\([Jurafsky, 1997](https://arxiv.org/html/2609.10934#bib.bib27);[Si et al\., 2023](https://arxiv.org/html/2609.10934#bib.bib56)\)\. This reflects a broader methodological challenge, since within\-turn TRPs are underrepresented in corpora derived from natural interaction\([Threlkeld et al\., 2022](https://arxiv.org/html/2609.10934#bib.bib61);[Umair et al\., 2024](https://arxiv.org/html/2609.10934#bib.bib62)\)\. Further validation across datasets and interactional settings is therefore required\.
Second, our evaluation isolates text\-derived semantic information and therefore does not capture the full multimodal structure of turn\-taking\. This choice allows us to test whether evolving semantic constraint contributes to within\-turn TRP prediction, but it does not show that semantic uncertainty alone is sufficient for response timing in deployed dialogue systems\. Acoustic, prosodic, and multimodal cues may interact with semantic uncertainty in ways that are not captured here\. Future work should therefore test whether the signal remains useful when integrated with models that incorporate these additional sources of information\.
Third, even within the text\-only setting, semantic uncertainty remains model\-mediated\. We do not claim that embedding\-space dispersion is independent of lexical form, only that it provides a proxy for semantic relatedness less tied to surface form than token\-level measures\. A diagnostic analysis supports this claim\. Duplicate\-continuation rates vary sharply across decoding settings, from roughly 67% to 1%, while prefix\-level mean embeddings remain stable, with pairwise cosine similarities of 0\.94–0\.97\. This reduces, but does not eliminate, the concern that the signal reflects surface form\. It also does not make the signal model\-independent; alternative language models, embedding spaces, or similarity functions could yield different uncertainty trajectories\.
Additionally, the semantic uncertainty measure we use is expensive relative to direct text to label prediction, specifically when samplingKKcontinuations per prefix \(see Section[3\.3](https://arxiv.org/html/2609.10934#S3.SS3)\)\. TheKKablation suggests that this cost is reducible\. MeanF0\.5F\_\{0\.5\}reaches 0\.54 atK=2K=2and remains near its best values for smallKK, indicating that the coarse uncertainty measures may still be informative enough to predict TRPs locally\. Regardless of smallerKKvalues, the per\-prefix cost of sampling and embedding remains higher than single\-pass prediction\. Future work is required to determine whether semantic uncertainty can be deployed in a real\-time system\.
Our analysis also relies on a deliberately simple decision mechanism\. The fixed, causal rule used here \(Section[3\.4](https://arxiv.org/html/2609.10934#S3.SS4)\) prioritizes interpretability over expressive capacity\. The variance decomposition in Section[5\.3](https://arxiv.org/html/2609.10934#S5.SS3)suggests this choice is consequential: the decision rule accounts for 75\.4% ofF0\.5F\_\{0\.5\}variance across configurations, with smoothing alone explaining 52\.6%\. More flexible models, which are capable of exploiting longer\-range dependencies or finer\-grained patterns in the uncertainty trajectory, could plausibly outperform a deterministic rule\. The reported performance should therefore not be read as a performance upper bound on uncertainty\-based TRP projection\.
Finally, our evaluation metrics reflect a particular interactional cost structure\. We prioritize precision\-weighted measures because false positive TRP predictions are often more disruptive than missed opportunities to respond \(Section[4\.2](https://arxiv.org/html/2609.10934#S4.SS2)\)\. This assumption may not hold in all settings\. More proactive or high\-initiative systems may require different trade\-offs between false positives and false negatives\. Our metrics also assess local prediction quality, not downstream outcomes such as conversational fluidity or user experience\. The results should therefore be interpreted as alignment with perceived TRPs under this cost structure, rather than as a complete measure of interactional success\.
## 9Ethical Considerations
We consider the ethical implications of this work in terms of how turn prediction signals might be used in deployed systems\. If such signals are treated as entitlements to speak rather than advisory cues, semantic uncertainty could lead to inappropriate interruptions or missed opportunities to respond\. We therefore emphasize that these signals are model\-mediated and should not be interpreted as evidence of user intent or readiness to yield the floor\. While turn\-taking is often described as universal, its realization varies across cultures and interactional settings, and systems that operationalize TRPs risk privileging particular norms if deployed without care\. Additionally, the participant\-judgment data used in this work were collected under ethical oversight, with informed consent, protections for participant anonymity, and approval from the relevant institutional review board\.
## 10Acknowledgments
We thank Vasanth Sarathy for early discussions on structuring our experiments, Bilal Ahmed for refining key ideas and providing feedback, and Julia Mertens for thoughtful discussions that strengthened our arguments\. We also acknowledge the Tufts University Department of Computer Science for institutional support, and Tufts Research Technology High Performance Computing for providing the computational resources used in this work\. AI assistants were used in this work to improve language consistency; all scientific content and results are the authors’ original work\.
## References
- Algherairy and Ahmed \(2024\)Atheer Algherairy and Moataz Ahmed\. 2024\.[A review of dialogue systems: current trends and future directions](https://doi.org/10.1007/s00521-023-09322-1)\.*Neural Computing and Applications*, 36\(12\):6325–6351\.
- Altmann and Kamide \(1999\)Gerry T\.M Altmann and Yuki Kamide\. 1999\.[Incremental interpretation at verbs: restricting the domain of subsequent reference](https://doi.org/10.1016/S0010-0277(99)00059-1)\.*Cognition*, 73\(3\):247–264\.
- Anderson et al\. \(1991\)Anne H Anderson, Miles Bader, Ellen Gurman Bard, Elizabeth Boyle, Gwyneth Doherty, Simon Garrod, Stephen Isard, Jacqueline Kowtko, Jan McAllister, Jim Miller, et al\. 1991\.The hcrc map task corpus\.*Language and speech*, 34\(4\):351–366\.
- Arora et al\. \(2025\)Siddhant Arora, Zhiyun Lu, Chung\-Cheng Chiu, Ruoming Pang, and Shinji Watanabe\. 2025\.[Talking turns: Benchmarking audio foundation models on turn\-taking dynamics](https://arxiv.org/abs/2503.01174)\.*Preprint*, arXiv:2503\.01174\.
- Boersma and Weenink \(2009\)Paul Boersma and David Weenink\. 2009\.[Praat: doing phonetics by computer \(version 5\.1\.13\)](http://www.praat.org/)\.
- Brown et al\. \(2020\)Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al\. 2020\.Language models are few\-shot learners\.*Advances in neural information processing systems*, 33:1877–1901\.
- Budzianowski et al\. \(2018\)Paweł Budzianowski, Tsung\-Hsien Wen, Bo\-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic\. 2018\.Multiwoz\-a large\-scale multi\-domain wizard\-of\-oz dataset for task\-oriented dialogue modelling\.In*Proceedings of the 2018 conference on empirical methods in natural language processing*, pages 5016–5026\.
- Bögels and Torreira \(2015\)Sara Bögels and Francisco Torreira\. 2015\.[Listeners use intonational phrase boundaries to project turn ends in spoken interaction](https://doi.org/10.1016/j.wocn.2015.04.004)\.*Journal of Phonetics*, 52:46–57\.
- Calhoun et al\. \(2010\)Sasha Calhoun, Jean Carletta, Jason M Brenier, Neil Mayo, Dan Jurafsky, Mark Steedman, and David Beaver\. 2010\.The nxt\-format switchboard corpus: a rich resource for investigating the syntax, semantics, pragmatics and prosody of dialogue\.*Language resources and evaluation*, 44\(4\):387–419\.
- Callison\-Burch et al\. \(2006\)Chris Callison\-Burch, Miles Osborne, and Philipp Koehn\. 2006\.Re\-evaluating the role of bleu in machine translation research\.In*11th conference of the european chapter of the association for computational linguistics*, pages 249–256\.
- Castillo\-López et al\. \(2025\)Galo Castillo\-López, Gael de Chalendar, and Nasredine Semmar\. 2025\.[A survey of recent advances on turn\-taking modeling in spoken dialogue systems](https://aclanthology.org/2025.iwsds-1.27/)\.In*Proceedings of the 15th International Workshop on Spoken Dialogue Systems Technology*, pages 254–271, Bilbao, Spain\. Association for Computational Linguistics\.
- Cer et al\. \(2017\)Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez\-Gazpio, and Lucia Specia\. 2017\.Semeval\-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation\.In*Proceedings of the 11th international workshop on semantic evaluation \(SemEval\-2017\)*, pages 1–14\.
- de Ruiter et al\. \(2006\)Jan P\. de Ruiter, H\. Mitterer, and N\. J\. Enfield\. 2006\.[Projecting the end of a speaker’s turn: A cognitive cornerstone of conversation](https://doi.org/10.1353/lan.2006.0130)\.*Language*, 82\(3\):515–535\.
- Drew \(2009\)Paul Drew\. 2009\."quit talking while i’m interrupting": A comparison between positions of overlap onset in conversation\.In M\. Haakana, M\. Laakso, and J\. Lindstrom, editors,*Talk in Interaction: Comparative Dimensions*, pages 70–93\. Finnish Literature Society\.
- Ekstedt and Skantze \(2020\)Erik Ekstedt and Gabriel Skantze\. 2020\.[TurnGPT: a transformer\-based language model for predicting turn\-taking in spoken dialog](https://doi.org/10.18653/v1/2020.findings-emnlp.268)\.In*Findings of the Association for Computational Linguistics: EMNLP 2020*, pages 2981–2990, Online\. Association for Computational Linguistics\.
- Ekstedt and Skantze \(2022\)Erik Ekstedt and Gabriel Skantze\. 2022\.[Voice Activity Projection: Self\-supervised Learning of Turn\-taking Events](https://doi.org/10.21437/Interspeech.2022-10955)\.In*Proc\. Interspeech 2022*, pages 5190–5194\.
- Elmers et al\. \(2025\)Mikey Elmers, Koji Inoue, Divesh Lala, and Tatsuya Kawahara\. 2025\.[Triadic Multi\-party Voice Activity Projection for Turn\-taking in Spoken Dialogue Systems](https://doi.org/10.21437/Interspeech.2025-2660)\.In*Interspeech 2025*, pages 3015–3019\.
- Farquhar et al\. \(2024\)Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal\. 2024\.Detecting hallucinations in large language models using semantic entropy\.*Nature*, 630\(8017\):625–630\.
- Fedus et al\. \(2022\)William Fedus, Barret Zoph, and Noam Shazeer\. 2022\.[Switch transformers: scaling to trillion parameter models with simple and efficient sparsity](https://dl.acm.org/doi/abs/10.5555/3586589.3586709)\.*J\. Mach\. Learn\. Res\.*, 23\(1\)\.
- Hale \(2001\)John Hale\. 2001\.[A probabilistic earley parser as a psycholinguistic model](https://doi.org/10.3115/1073336.1073357)\.In*Proceedings of the Second Meeting of the North American Chapter of the Association for Computational Linguistics on Language Technologies*, NAACL ’01, page 1–8, USA\. Association for Computational Linguistics\.
- He and Garcia \(2009\)Haibo He and Edwardo A\. Garcia\. 2009\.[Learning from imbalanced data](https://doi.org/10.1109/TKDE.2008.239)\.*IEEE Transactions on Knowledge and Data Engineering*, 21\(9\):1263–1284\.
- Holler and Levinson \(2019\)Judith Holler and Stephen C\. Levinson\. 2019\.[Multimodal language processing in human communication](https://doi.org/10.1016/j.tics.2019.05.006)\.*Trends in Cognitive Sciences*, 23\(8\):639–652\.
- Holtzman et al\. \(2019\)Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi\. 2019\.The curious case of neural text degeneration\.*arXiv preprint arXiv:1904\.09751*\.
- Hu et al\. \(2021\)Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen\-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen\. 2021\.Lora: Low\-rank adaptation of large language models\.*arXiv preprint arXiv:2106\.09685*\.
- Inoue et al\. \(2024\)Koji Inoue, Bing’er Jiang, Erik Ekstedt, Tatsuya Kawahara, and Gabriel Skantze\. 2024\.[Multilingual turn\-taking prediction using voice activity projection](https://aclanthology.org/2024.lrec-main.1036/)\.In*Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\)*, pages 11873–11883, Torino, Italia\. ELRA and ICCL\.
- Jiang et al\. \(2023\)Bing’er Jiang, Erik Ekstedt, and Gabriel Skantze\. 2023\.[Response\-conditioned turn\-taking prediction](https://doi.org/10.18653/v1/2023.findings-acl.776)\.In*Findings of the Association for Computational Linguistics: ACL 2023*, pages 12241–12248, Toronto, Canada\. Association for Computational Linguistics\.
- Jurafsky \(1997\)Dan Jurafsky\. 1997\.Switchboard swbd\-damsl shallow\-discourse\-function annotation coders manual\.*www\. dcs\. shef\. ac\. uk/nlp/amities/files/bib/ics\-tr\-97\-02\. pdf*\.
- Kendrick et al\. \(2023\)Kobin H\. Kendrick, Judith Holler, and Stephen C\. Levinson\. 2023\.[Turn\-taking in human face\-to\-face interaction is multimodal: gaze direction and manual gestures aid the coordination of turn transitions](https://doi.org/10.1098/rstb.2021.0473)\.*Philosophical Transactions of the Royal Society B: Biological Sciences*, 378\(1875\):20210473\.
- Kuhn et al\. \(2023\)Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar\. 2023\.Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation\.*arXiv preprint arXiv:2302\.09664*\.
- Leishman et al\. \(2024\)Sean Leishman, Peter Bell, and Sarenne Wallbridge\. 2024\.[Pairwiseturngpt: a multi\-stream turn prediction model for spoken dialogue](http://semdial.org/anthology/Z24-Leishman_semdial_0002.pdf)\.In*Proceedings of the 28th Workshop on the Semantics and Pragmatics of Dialogue \- Full Papers*, Trento, Italy\. SEMDIAL\.
- Lerner \(1991\)Gene H Lerner\. 1991\.[On the syntax of sentences\-in\-progress](https://doi.org/10.1017/S0047404500016572)\.*Language in Society*, 20\(3\):441–458\.
- Levinson \(1983\)Stephen C\. Levinson\. 1983\.[*Pragmatics*](https://www.cambridge.org/9780521294140)\.Cambridge Textbooks in Linguistics\. Cambridge University Press, Cambridge, U\.K\.
- Levinson \(2016\)Stephen C\. Levinson\. 2016\.[Turn\-taking in human communication – origins and implications for language processing](https://doi.org/10.1016/j.tics.2015.10.010)\.*Trends in Cognitive Sciences*, 20\(1\):6–14\.
- Levinson and Torreira \(2015\)Stephen C\. Levinson and Francisco Torreira\. 2015\.[Timing in turn\-taking and its implications for processing models of language](https://doi.org/10.3389/fpsyg.2015.00731)\.*Frontiers in Psychology*, 6:731\.
- Levy \(2008\)Roger Levy\. 2008\.[Expectation\-based syntactic comprehension](https://doi.org/10.1016/j.cognition.2007.05.006)\.*Cognition*, 106\(3\):1126–1177\.
- Liddicoat \(2004\)Anthony J\. Liddicoat\. 2004\.[The projectability of turn constructional units and the role of prediction in listening](https://doi.org/10.1177/1461445604046589)\.*Discourse Studies*, 6\(4\):449–469\.
- Liesenfeld and Dingemanse \(2024\)Andreas Liesenfeld and Mark Dingemanse\. 2024\.[Rethinking open source generative ai: open\-washing and the eu ai act](https://doi.org/10.1145/3630106.3659005)\.In*Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency*, FAccT ’24, page 1774–1787, New York, NY, USA\. Association for Computing Machinery\.
- Liesenfeld et al\. \(2023\)Andreas Liesenfeld, Alianda Lopez, and Mark Dingemanse\. 2023\.[The timing bottleneck: Why timing and overlap are mission\-critical for conversational user interfaces, speech recognition and dialogue systems](https://doi.org/10.18653/v1/2023.sigdial-1.45)\.In*Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue*, pages 482–495, Prague, Czechia\. Association for Computational Linguistics\.
- Magyari and de Ruiter \(2012\)Lilla Magyari and Jan P\. de Ruiter\. 2012\.[Prediction of turn\-ends based on anticipation of upcoming words](https://doi.org/10.3389/fpsyg.2012.00376)\.*Frontiers in Psychology*, 3:376\.
- Malinin and Gales \(2020\)Andrey Malinin and Mark Gales\. 2020\.Uncertainty estimation in autoregressive structured prediction\.*arXiv preprint arXiv:2002\.07650*\.
- MLC team \(2023\-2025\)MLC team\. 2023\-2025\.[MLC\-LLM](https://github.com/mlc-ai/mlc-llm)\.
- Nguyen et al\. \(2025\)Dang Nguyen, Ali Payani, and Baharan Mirzasoleiman\. 2025\.[Beyond semantic entropy: Boosting LLM uncertainty quantification with pairwise semantic similarity](https://doi.org/10.18653/v1/2025.findings-acl.234)\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 4530–4540, Vienna, Austria\. Association for Computational Linguistics\.
- Nikitin et al\. \(2024\)Alexander Nikitin, Jannik Kossen, Yarin Gal, and Pekka Marttinen\. 2024\.Kernel language entropy: Fine\-grained uncertainty quantification for llms from semantic similarities\.*Advances in Neural Information Processing Systems*, 37:8901–8929\.
- Onishi et al\. \(2023\)Kazuyo Onishi, Hiroki Tanaka, and Satoshi Nakamura\. 2023\.[Multimodal voice activity prediction: Turn\-taking events detection in expert\-novice conversation](https://doi.org/10.1145/3623809.3623837)\.In*Proceedings of the 11th International Conference on Human\-Agent Interaction*, HAI ’23, page 13–21, New York, NY, USA\. Association for Computing Machinery\.
- Patamia et al\. \(2025\)Rutherford Agbeshi Patamia, Ha Pham Thien Dinh, Ming Liu, and Akansel Cosgun\. 2025\.[Turn\-taking modelling in conversational systems: A review of recent advances](https://doi.org/10.3390/technologies13120591)\.*Technologies*, 13\(12\):591\.
- Reece et al\. \(2023\)Andrew Reece, Gus Cooney, Peter Bull, Christine Chung, Bryn Dawson, Casey Fitzpatrick, Tamara Glazer, Dean Knox, Alex Liebscher, and Sebastian Marin\. 2023\.The candor corpus: Insights from a large multimodal dataset of naturalistic conversation\.*Science advances*, 9\(13\):eadf3197\.
- Reimers and Gurevych \(2019\)Nils Reimers and Iryna Gurevych\. 2019\.[Sentence\-BERT: Sentence embeddings using Siamese BERT\-networks](https://aclanthology.org/D19-1410/)\.In*Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\)*, pages 3982–3992, Hong Kong, China\. Association for Computational Linguistics\.
- Reynolds and McDonell \(2021\)Laria Reynolds and Kyle McDonell\. 2021\.[Prompt programming for large language models: Beyond the few\-shot paradigm](https://doi.org/10.1145/3411763.3451760)\.In*Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems*, CHI EA ’21, New York, NY, USA\. Association for Computing Machinery\.
- Riest et al\. \(2015\)Carina Riest, Annett B\. Jorschick, and Jan P\. de Ruiter\. 2015\.[Anticipation in turn\-taking: mechanisms and information sources](https://doi.org/10.3389/fpsyg.2015.00089)\.*Frontiers in Psychology*, 6:89\.
- Sacks et al\. \(1974\)Harvey Sacks, Emanuel A\. Schegloff, and Gail Jefferson\. 1974\.[A simplest systematics for the organization of turn\-taking for conversation](https://doi.org/10.2307/412243)\.*Language*, 50\(4\):696–735\.
- Schegloff \(1982\)Emanuel A\. Schegloff\. 1982\.Discourse as an interactional achievement: Some uses of ‘uh huh’ and other things that come between sentences\.In Deborah Tannen, editor,*Analyzing Discourse: Text and Talk*, pages 71–93\. Georgetown University Press, Washington, DC\.
- Schegloff \(1996\)Emanuel A\. Schegloff\. 1996\.*Turn organization: one intersection of grammar and interaction*, page 52–133\.Studies in Interactional Sociolinguistics\. Cambridge University Press\.
- Schegloff \(2000\)Emanuel A\. Schegloff\. 2000\.[Overlapping talk and the organization of turn\-taking for conversation](http://www.jstor.org/stable/4168983)\.*Language in Society*, 29\(1\):1–63\.
- Selting \(2000\)Margret Selting\. 2000\.[The construction of units in conversational talk](http://www.jstor.org/stable/4169050)\.*Language in Society*, 29\(4\):477–517\.
- Shorinwa et al\. \(2025\)Ola Shorinwa, Zhiting Mei, Justin Lidard, Allen Z\. Ren, and Anirudha Majumdar\. 2025\.[A survey on uncertainty quantification of large language models: Taxonomy, open research challenges, and future directions](https://doi.org/10.1145/3744238)\.*ACM Comput\. Surv\.*, 58\(3\)\.
- Si et al\. \(2023\)Shuzheng Si, Wentao Ma, Haoyu Gao, Yuchuan Wu, Ting\-En Lin, Yinpei Dai, Hangyu Li, Rui Yan, Fei Huang, and Yongbin Li\. 2023\.Spokenwoz: A large\-scale speech\-text benchmark for spoken task\-oriented dialogue agents\.*Advances in Neural Information Processing Systems*, 36:39088–39118\.
- Skantze \(2021\)Gabriel Skantze\. 2021\.[Turn\-taking in conversational systems and human\-robot interaction: A review](https://doi.org/10.1016/j.csl.2020.101178)\.*Computer Speech & Language*, 67:101178\.
- Skantze et al\. \(2014\)Gabriel Skantze, Anna Hjalmarsson, and Catharine Oertel\. 2014\.[Turn\-taking, feedback and joint attention in situated human–robot interaction](https://doi.org/10.1016/j.specom.2014.05.005)\.*Speech Communication*, 65:50–66\.
- Stivers et al\. \(2009\)Tanya Stivers, N\. J\. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter de Ruiter, Kyung\-Eun Yoon, and Stephen C\. Levinson\. 2009\.[Universals and cultural variation in turn\-taking in conversation](https://doi.org/10.1073/pnas.0903616106)\.*Proceedings of the National Academy of Sciences*, 106\(26\):10587–10592\.
- Tanenhaus et al\. \(1995\)Michael K\. Tanenhaus, Michael J\. Spivey\-Knowlton, Kathleen M\. Eberhard, and Julie C\. Sedivy\. 1995\.[Integration of visual and linguistic information in spoken language comprehension](https://doi.org/10.1126/science.7777863)\.*Science*, 268\(5217\):1632–1634\.
- Threlkeld et al\. \(2022\)Charles Threlkeld, Muhammad Umair, and Jp de Ruiter\. 2022\.[Using transition duration to improve turn\-taking in conversational agents](https://aclanthology.org/2022.sigdial-1.20)\.In*Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue*, pages 193–203, Edinburgh, UK\. Association for Computational Linguistics\.
- Umair et al\. \(2024\)Muhammad Umair, Vasanth Sarathy, and Jan Ruiter\. 2024\.[Large language models know what to say but not when to speak](https://doi.org/10.18653/v1/2024.findings-emnlp.909)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 15503–15514, Miami, Florida, USA\. Association for Computational Linguistics\.
- Wang et al\. \(2024a\)Jinhan Wang, Long Chen, Aparna Khare, Anirudh Raju, Pranav Dheram, Di He, Minhua Wu, Andreas Stolcke, and Venkatesh Ravichandran\. 2024a\.[Turn\-taking and backchannel prediction with acoustic and large language model fusion](https://www.amazon.science/publications/turn-taking-and-backchannel-prediction-with-acoustic-and-large-language-model-fusion)\.In*ICASSP 2024\-2024 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\)*, pages 12121–12125\. IEEE\.
- Wang et al\. \(2024b\)Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei\. 2024b\.Improving text embeddings with large language models\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 11897–11916\.
- Webson and Pavlick \(2022\)Albert Webson and Ellie Pavlick\. 2022\.[Do prompt\-based models really understand the meaning of their prompts?](https://doi.org/10.18653/v1/2022.naacl-main.167)In*Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 2300–2344, Seattle, United States\. Association for Computational Linguistics\.
- Wei et al\. \(2022\)Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H\. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus\. 2022\.[Emergent abilities of large language models](https://openreview.net/forum?id=yzkSU5zdwD)\.*Transactions on Machine Learning Research*\.
- Wittenburg et al\. \(2006\)Peter Wittenburg, Hennie Brugman, Albert Russel, Alex Klassmann, and Han Sloetjes\. 2006\.[ELAN: a professional framework for multimodality research](http://www.lrec-conf.org/proceedings/lrec2006/pdf/153%5Fpdf.pdf)\.In*Proceedings of the Fifth International Conference on Language Resources and Evaluation \(LREC’06\)*, Genoa, Italy\. European Language Resources Association \(ELRA\)\.
- Wolf et al\. \(2020\)Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush\. 2020\.[Transformers: State\-of\-the\-art natural language processing](https://doi.org/10.18653/v1/2020.emnlp-demos.6)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, pages 38–45, Online\. Association for Computational Linguistics\.
- Xu et al\. \(2023\)Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao\. 2023\.Expertprompting: Instructing large language models to be distinguished experts\.*arXiv preprint arXiv:2305\.14688*\.
- Yngve \(1970\)Victor H Yngve\. 1970\.On getting a word in edgewise\.In*Papers from the sixth regional meeting Chicago Linguistic Society, April 16\-18, 1970, Chicago Linguistic Society, Chicago*, pages 567–578\.
- Zhang et al\. \(2026\)Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Guoyin Wang, and Fei Wu\. 2026\.[Instruction tuning for large language models: A survey](https://doi.org/10.1145/3777411)\.*ACM Comput\. Surv\.*, 58\(7\)\.
## Appendix AOperationalizing TRPs: Stimulus\-Specific Thresholding and Local Maximality
The nature of the listener\-response data permits several operationalizations of binary TRP labels\. A*global threshold*labels an interval positive when its response proportion exceeds a fixed cutoff across all stimuli\. A*stimulus\-specific threshold*labels an interval positive when its response proportion is elevated relative to the distribution for that stimulus\. A*local\-peak criterion*labels an interval positive when its response proportion exceeds those of its adjacent intervals\.
We use the stimulus\-specific threshold described in Section[4\.1](https://arxiv.org/html/2609.10934#S4.SS1)\. This choice reflects two properties of listener responses\. First, response onsets are temporally dispersed because listeners project TRPs and may begin responding before the stimulus turn ends\([Sacks et al\., 1974](https://arxiv.org/html/2609.10934#bib.bib50);[de Ruiter et al\., 2006](https://arxiv.org/html/2609.10934#bib.bib13);[Levinson and Torreira, 2015](https://arxiv.org/html/2609.10934#bib.bib34)\)\. High absolute agreement within any single interval is therefore unlikely\. Second, stimuli differ in baseline responsiveness, making a fixed global threshold overly conservative for low\-response stimuli and overly permissive for high\-response stimuli\. Stimulus\-specific thresholding accounts for this variability\.
The stimulus\-specific threshold identifies elevated response proportions but does not establish whether they form distinct peaks rather than broader regions of elevated responsiveness\. For each interval, we therefore compute a*local peak margin*: the difference betweenIi,i\+1ProportionI\_\{i,i\+1\}^\{Proportion\}and the larger response proportion of its immediate neighbors, where defined\. Positive margins indicate strict local maxima \(see Figure[4](https://arxiv.org/html/2609.10934#A1.F4)\)\. Among the 842 positive labels, 72% occur at strict local maxima\. This statistic is diagnostic, not definitional: local maximality is not part of the labeling rule\. It shows that the threshold generally selects positions that are locally prominent in the response distribution\.
Figure 4:Distribution of peak margins for detected TRPs and non\-TRP intervals\. Peak margin is defined as the difference between an interval’s response proportion and the maximum of its immediate neighbors\. Detected TRPs exhibit substantially larger positive margins, while non\-TRP intervals are centered near zero, indicating that the labeling procedure selects positions with locally elevated listener response rates\.
## Appendix BModel Selection and Experimental Infrastructure
### B\.1Model Selection
ModelTypeParams\.PrecisionLLaMA\-3\.1\-8B\-InstructDense8Bq4f16LLaMA\-3\.1\-70B\-InstructDense70Bq4f16Mistral\-7B\-InstructDense7Bq4f16Mixtral\-8×\\times7B\-InstructMoE46Bq4f16Qwen2\.5\-7B\-InstructDense7Bq4f16Qwen2\.5\-72B\-InstructDense72Bq4f16Table 5:Instruction\-tuned, open\-weight LLMs used throughout this work\. Dense models apply all parameters to each input, whereas Mixture\-of\-Experts \(MoE\) models route each input through a sparse subset of expert modules\. Parameter counts for MoE models report total parameters across experts\. q4f16 denotes 4\-bit weight quantization with FP16 compute\.Throughout this work, we use a set of instruction\-tuned, open\-weight LLMs spanning multiple architectural families and parameter scales \(see Table[5](https://arxiv.org/html/2609.10934#A2.T5)\)\. This selection tests whether observed behaviors vary systematically with model scale or architecture, rather than reflecting idiosyncrasies of a single model family\. The set includes both dense and mixture\-of\-experts architectures, which allocate representational capacity differently and exhibit distinct scaling and generalization properties\([Fedus et al\., 2022](https://arxiv.org/html/2609.10934#bib.bib19)\)\. Varying model scale probes whether limitations reflect insufficient capacity or the formulation of the TRP prediction task itself, since scaling can affect model capabilities unevenly across tasks\([Wei et al\., 2022](https://arxiv.org/html/2609.10934#bib.bib66)\)\. All models are instruction\-tuned, which improves adherence to task specifications\([Zhang et al\., 2026](https://arxiv.org/html/2609.10934#bib.bib71)\)\.
### B\.2Experimental Infrastructure
ExperimentStage\#ConfigsAggregateGPU\-hAggregateCPU\-hAvg wall/ unitPromptingbaselines\(§[5\.1](https://arxiv.org/html/2609.10934#S5.SS1)\)Expert67\.03728\.1480\.813 s/prefixParticipant610\.93243\.7281\.26 s/prefixImagined611\.75547\.0201\.36 s/prefixOracle620\.07880\.3122\.32 s/prefixSFT\(§[5\.2](https://arxiv.org/html/2609.10934#S5.SS2)\)Training30\.5912\.36611\.83 min/modelInference30\.3511\.4047\.02 min/modelSemanticuncertainty\(§[5\.3](https://arxiv.org/html/2609.10934#S5.SS3)\)Continuation sampling685\.335341\.3409\.86 s/prefixEmbedding184\.53018\.1200\.174 s/prefixDecision functionff2,304––<0\.001<0\.001s/prefixTotal––140\.609562\.438–Table 6:Compute summary for the main\-text experiments\. \# Configs reports the number of configurations or runs included in each stage\. Aggregate GPU\-h and CPU\-h report total compute for that stage\. Avg wall / unit reports average wall\-clock time for the natural unit of the stage: per prefix for prompting and semantic\-uncertainty stages, and per trained model for SFT stages\. The semantic\-uncertainty method is separated into continuation sampling, embedding, and decision\-rule evaluation because only continuation sampling requires new LLM generations\. The continuation\-sampling row aggregates the six decoding configurations used in the fixed\-K=25K\{=\}25sweep\. Embedding and decision\-rule configurations are computed from saved continuations, and decision\-rule evaluation is CPU\-only\.We implement efficient batched inference and consistent generation behavior across models using custom adaptations of the MLC\-LLM framework together with the HuggingFace ecosystem\([Wolf et al\., 2020](https://arxiv.org/html/2609.10934#bib.bib68);[MLC team, 2023\-2025](https://arxiv.org/html/2609.10934#bib.bib41)\)\. Model weights are stored with 4\-bit quantization and computation is performed in FP16, reducing memory usage and enabling large\-scale evaluation on limited hardware\. Parameter\-efficient fine\-tuning, which introduces low\-rank trainable parameters to a frozen base model, is implemented using LoRA adapters\([Hu et al\., 2021](https://arxiv.org/html/2609.10934#bib.bib24)\)\. Experiments run on NVIDIA A100 \(40 GB and 80 GB\) and H200 GPUs with a single accelerator per run and four CPU cores per GPU\. Per\-experiment aggregate compute usage is summarized in Table[6](https://arxiv.org/html/2609.10934#A2.T6)\. We do not include closed\-source LLMs because they do not provide the transparency required for controlled scientific evaluation, including access to model weights, training data, and decoding settings\([Liesenfeld and Dingemanse, 2024](https://arxiv.org/html/2609.10934#bib.bib37)\)\.
Across experiments, we use two decoding parameters, temperature andtoptop\-pp, to control LLM generation stochasticity\. Temperature rescales the shape of the full token distribution, whiletoptop\-ppgoverns its effective support by restricting sampling to the smallest set of tokens whose cumulative probability exceeds a threshold\([Holtzman et al\., 2019](https://arxiv.org/html/2609.10934#bib.bib23)\)\. Other decoding parameters \(e\.g\.,toptop\-kktruncation or repetition penalties\) use default values to limit interacting degrees of freedom and to preserve a consistent sampling regime\. Specific settings are reported in each experiment’s appendix\.
## Appendix CPrompt\-Based TRP Prediction Details
In Section[5\.1](https://arxiv.org/html/2609.10934#S5.SS1), we evaluate all LLMs \(see Table[5](https://arxiv.org/html/2609.10934#A2.T5)\) under the four prompting conditions using a fixed decoding configuration \(temperature = 0\.4,toptop\-pp= 1\.0;[Holtzman et al\., 2019](https://arxiv.org/html/2609.10934#bib.bib23)\)\. This allows us to balance diversity and determinism while treating generation stochasticity as a controlled source of variation rather than an object of analysis\. We generate a single prediction per interval; sampling multiple generations would substantially increase computational cost across the full model–condition grid\. As a result, reported metrics reflect a single sampled generation, and systematic analysis of decoding variability is left to future work\.
For each prefix, prompts instruct models to produce structured JSON outputs containing a binary TRP decision, a confidence score, and a brief justification\. Evaluation uses only the binary decision; confidence scores and justifications are recorded but excluded from reported metrics\. Outputs are parsed using a deterministic post\-processing routine, achieving a 99\.5% success rate over 124,680 generations\. The remaining failures primarily reflect minor deviations from the expected schema and are excluded prior to evaluation; given their low frequency, they are unlikely to affect comparative results\. Appendix[E\.4](https://arxiv.org/html/2609.10934#A5.SS4)shows the summarized prompt templates for each condition\.
## Appendix DSupervised Fine\-Tuning Details
In Section[5\.2](https://arxiv.org/html/2609.10934#S5.SS2), we fine\-tune LLaMA\-3\.1\-8B\-Instruct using LoRA adapters, rather than the strongest prompt\-based baseline model \(LLaMA\-3\.1\-70B\-Instruct; see Section[5\.1](https://arxiv.org/html/2609.10934#S5.SS1)\)\. Prompt\-based results indicate that larger models do not necessarily improve performance on the within\-turn TRP prediction task\. Fine\-tuning a larger model could in principle yield different results, as supervised adaptation and prompting exhibit distinct failure modes; we leave this for future work\.
We fine\-tune using a causal language modeling objective with a completion\-only loss, computing gradients only over the assistant completion corresponding to the binary TRP decision\. Restricting gradients to the completion avoids updating the model on prompt or context tokens that are not part of the decision target\. We do not introduce class\-weighted or precision\-aware losses, despite evaluating withF0\.5F\_\{0\.5\}\. Our aim is not to optimize downstream metrics but to test whether supervision alone improves TRP prediction under a standard SFT formulation\. Incorporating task\-specific loss shaping would confound this diagnostic\.
Training uses the AdamW optimizer with a base learning rate of2×10−52\\times 10^\{\-5\}, cosine scheduling with a warmup ratio of 0\.03, per\-device batch size of 4, and gradient accumulation over 8 steps\. Models are trained for up to 30 epochs with early stopping based on the completion\-only evaluation loss, and the best checkpoint is selected for evaluation\. At inference time, we generate TRP predictions using the same system prompt used during training, along with the same decoding parameters \(temperature = 0\.4,toptop\-pp= 1\.0\) as used for prompt\-based inference \(see Appendix[C](https://arxiv.org/html/2609.10934#A3)\)\.
## Appendix ESemantic Uncertainty Based TRP Prediction Details
StageHyperparameterValues\#Uncertaintyestimation\(Algorithm[1](https://arxiv.org/html/2609.10934#alg1)\)Generation modelLLaMA\-3\.1\-8B\-Instruct \(q4f16\)1Continuations per prefixK=25K=251Decoding setting\(T,top\-p\)\(T,\\mathrm\{top\}\\text\{\-\}p\)\(0\.50,0\.80\)\(0\.50,0\.80\);\(0\.80,0\.90\)\(0\.80,0\.90\);\(0\.85,0\.95\)\(0\.85,0\.95\);\(0\.90,0\.90\)\(0\.90,0\.90\);\(0\.90,1\.00\)\(0\.90,1\.00\);\(1\.10,1\.00\)\(1\.10,1\.00\)6Embedding modelSFR\-2R; all\-mpnet\-base\-v2; E5\-Mistral\-7B\-Instruct3SNNE scale \(τ\\tau\)\{0\.1,1,10,100\}\\\{0\.1,1,10,100\\\}4Decisionfunction\(Algorithm[2](https://arxiv.org/html/2609.10934#alg2)\)SmoothingEMA:α∈\{0\.3,0\.5,0\.7\}\\alpha\\in\\\{0\.3,0\.5,0\.7\\\};Boxcar: window∈\{3,5,7\}\\in\\\{3,5,7\\\};Gaussian:σ=1\.5\\sigma=1\.5, window∈\{5,10\}\\in\\\{5,10\\\}8Rise/drop windowsRise window∈\{5,10\}\\in\\\{5,10\\\}crossed with drop window∈\{5,10\}\\in\\\{5,10\\\}4Fixed detector settingsCausal operation; adaptive thresholding; threshold=1\.5=1\.5; minimum local points=3=3; derivative\-based rise/drop scoring; median centering; MAD\-to\-σ\\sigmafactor=1\.4826=1\.4826; minimum scaleε=10−8\\varepsilon=10^\{\-8\}1Table 7:Main fixed\-KKrobustness sweep, organized as a two\-phase pipeline\. In Phase 1, uncertainty estimation crosses 6 decoding settings, 3 embedding models, and 4 SNNEτ\\tauvalues, yielding 72 uncertainty\-signal configurations\. In Phase 2, the decision function converts each trajectory into TRP predictions using 32 detector configurations, formed by crossing 8 causal smoothing variants with 4 adaptive rise/drop window settings\. In total, the sweep evaluates72×32=2,30472\\times 32=2\{,\}304scored configurations\.### E\.1Experimental Setup
We evaluate the semantic uncertainty method \(see Section[3\.2](https://arxiv.org/html/2609.10934#S3.SS2)\) by sweeping three groups of parameters: continuation\-sampling settings, semantic\-uncertainty settings, and decision\-rule settings\. Table[7](https://arxiv.org/html/2609.10934#A5.T7)summarizes the configuration space\.
We use LLaMA\-3\.1\-8B\-Instruct as the continuation model\. This model provides a practical trade\-off between instruction\-following ability and computational cost\. Empirical runtimes indicate that LLaMA\-3\.1\-70B\-Instruct is between 6×\\timesand 20×\\timesslower than its smaller variant\. We therefore use the 8B model throughout and vary continuation diversity through six temperature and top\-ppsettings\.
We embed continuations using three different embedding models\. We includeSFR\-Embedding\-2\_Ras a general\-purpose embedding model for semantic similarity and retrieval, andall\-mpnet\-base\-v2as a smaller sentence\-transformer baseline\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.10934#bib.bib47)\)\. We also includeintfloat/e5\-mistral\-7b\-instructas a larger embedding model, allowing us to test whether the uncertainty signal depends on embedding\-model scale\([Wang et al\., 2024b](https://arxiv.org/html/2609.10934#bib.bib64)\)\.
For each model, we compute SNNE from pairwise cosine similarities between continuation embeddings\. While[Nguyen et al\. \(2025\)](https://arxiv.org/html/2609.10934#bib.bib42)report ROUGE\-L as the strongest similarity function for their task, we use embedding cosine similarity because our continuations are short \(≈\\approx5–10 tokens\) and often differ in surface form\. In this setting, lexical\-overlap measures may be sensitive to small wording differences rather than semantic similarity\([Callison\-Burch et al\., 2006](https://arxiv.org/html/2609.10934#bib.bib10)\)\.
### E\.2Semantic Uncertainty Robustness
To assess whether the semantic\-uncertainty signal is sensitive to continuation sampling, we summarize its behavior across the six decoding configurations \(see Table[7](https://arxiv.org/html/2609.10934#A5.T7)\)\. We holdK=25K=25andτ=0\.1\\tau=0\.1fixed to isolate the effect of temperature and top\-pp\.
First, we summarize the uncertainty trajectory for each decoding setting\. For each embedding model, we compute SNNE for every prefix, then measure the mean and standard deviation across prefixes \(see Table[8](https://arxiv.org/html/2609.10934#A5.T8)\)\. As decoding becomes more stochastic, mean SNNE becomes less negative, from−11\.74\-11\.74to−10\.61\-10\.61, indicating broader continuation sets\. The standard deviation across prefixes decreases from0\.560\.56to0\.210\.21, indicating compressed prefix\-level contrast\.
DecodingMetric, mean±\\pmSDTTppSNNEPref\. StdCos\.\.50\.80−11\.74±\.35\-11\.74\\pm\.35\.56±\.09\.56\\pm\.09\.72±\.17\.72\\pm\.17\.80\.90−10\.99±\.51\-10\.99\\pm\.51\.41±\.00\.41\\pm\.00\.62±\.22\.62\\pm\.22\.85\.95−10\.99±\.53\-10\.99\\pm\.53\.39±\.01\.39\\pm\.01\.62±\.23\.62\\pm\.23\.90\.90−10\.87±\.51\-10\.87\\pm\.51\.36±\.02\.36\\pm\.02\.60±\.23\.60\\pm\.23\.901\.00−10\.77±\.50\-10\.77\\pm\.50\.30±\.04\.30\\pm\.04\.58±\.25\.58\\pm\.251\.101\.00−10\.61±\.46\-10\.61\\pm\.46\.21±\.06\.21\\pm\.06\.54±\.27\.54\\pm\.27Table 8:Decoding diagnostic for fixedK=25K\{=\}25andτ=0\.1\\tau\{=\}0\.1\. SNNE is the mean uncertainty value across prefixes\. Pref\. Std is the standard deviation of SNNE across prefixes\. Cos\. is the average cosine similarity between distinct continuations sampled for the same prefix\. Values are mean±\\pmSD across the three embedding models\.Second, we measure semantic diversity within each prefix under the same decoding setting\. For each prefix, we compute the average cosine similarity between all distinct pairs of continuations sampled for that prefix, then average across prefixes and embedding models\. This continuation similarity decreases from0\.720\.72to0\.540\.54as decoding becomes more stochastic, indicating that higher stochasticity produces more semantically diverse continuations for the same prefix\.
Finally, we test whether the same prefix remains stable across decoding settings\. For each prefix and decoding setting, we average the embeddings of its sampled continuations\. We then compare these mean embeddings for the same prefix across every pair of the six decoding configurations\. Mean cosine similarity exceeds0\.940\.94, indicating that changing temperature and top\-ppincreases continuation diversity without moving the same prefix to a substantially different semantic region\.
These diagnostics show that changing temperature and top\-ppmakes the continuations broader and more diverse, but the same prefix still points to a similar semantic region across decoding settings\. The uncertainty trajectory is therefore not just an artifact of one sampling setup; it reflects information tied to the prefix itself\.
### E\.3SNNE Similarity Scaling and Discriminative Capacity
Our method predicts within\-turn TRPs from local changes in the semantic\-uncertainty trajectory\. The SNNE similarity scaling parameterτ\\taucontrols the numerical scale of this trajectory\. Smaller values weigh similar continuation pairs highly, while larger values distribute weight more evenly and compress differences among SNNE values\.
To quantify this effect, we compute a reference contrast for eachτ\\tauusing two syntheticK=25K=25similarity matrices\. One represents complete agreement, withS\(i\)u,v=1S^\{\(i\)\}\{u,v\}=1for allu,vu,v; the other is a positive\-similarity dispersion reference, withS\(i\)u,u=1S^\{\(i\)\}\{u,u\}=1andSu,v\(i\)=0S^\{\(i\)\}\_\{u,v\}=0foru≠vu\\neq v\. This contrast does not define the full SNNE range, since cosine similarities can be negative, but it provides a fixed calibration of how much numerical contrast SNNE expresses asτ\\tauvaries\. We compare it to the observed standard deviation of prefix\-level SNNE values from the fixed\-KKsweep, averaged over embedding models and decoding configurations\.
Table[9](https://arxiv.org/html/2609.10934#A5.T9)shows that the reference contrast shrinks rapidly asτ\\tauincreases\. This confirms that largeτ\\tauvalues compress SNNE scale\. Atτ=100\\tau=100, the prefix\-level standard deviation exceeds the reference contrast, meaning that this calibration no longer reflects the full empirical range of the signal\. The main sweep shows, however, that this compression has limited downstream effect\. The higher\-level implication is thatτ\\tauprimarily affects the scale of the uncertainty trajectory, while TRP prediction depends on how local changes in that trajectory are interpreted by the decision rule\.
τ\\tauRef\. ContrastPrefix StdStd / Contrast0\.13\.2180\.37111\.5%10\.9340\.0788\.4%100\.0960\.02223\.3%1000\.0100\.020206\.8%Table 9:Effect of the similarity scaling parameterτ\\tauon SNNE discriminative capacity forK=25K=25continuations\. Reference contrast is the SNNE difference between identical \(Su,v\(i\)=1S^\{\(i\)\}\_\{u,v\}=1for allu,vu,v\) and orthogonal \(Su,u\(i\)=1S^\{\(i\)\}\_\{u,u\}=1,Su,v\(i\)=0S^\{\(i\)\}\_\{u,v\}=0foru≠vu\\neq v\) similarity configurations\. Prefix Std denotes the observed standard deviation of SNNE across prefixes\.
### E\.4In\-Domain Semantic\-Similarity
Section[3\.3](https://arxiv.org/html/2609.10934#S3.SS3)treats embedding\-space dispersion as a model\-mediated proxy for semantic variation, without assuming independence from lexical form\. Prior work supports this interpretation: Sentence\-BERT is designed so that cosine similarity reflects semantic similarity\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.10934#bib.bib47)\), E5\-Mistral is evaluated on semantic\-similarity and retrieval tasks\([Wang et al\., 2024b](https://arxiv.org/html/2609.10934#bib.bib64)\), and Semantic Textual Similarity benchmarks compare model similarities with human judgments\([Cer et al\., 2017](https://arxiv.org/html/2609.10934#bib.bib12)\)\.
However, these evaluations largely concern complete sentences or passages, whereas Algorithm[1](https://arxiv.org/html/2609.10934#alg1)embeds LLM\-generated continuations of approximately 5–10 tokens conditioned on prefixes transcribed from spontaneous speech\. We therefore conduct a targeted in\-domain diagnostic to test whether the embedding space captures semantic similarity in this setting\.
We sample candidate continuations across all six decoding settings \(see Table[7](https://arxiv.org/html/2609.10934#A5.T7)\)\. For each continuation, the comparison reported here uses a triplet comprising the original, a similar\-length paraphrase preserving its meaning while changing its wording, and a similar\-length continuation from another stimulus with a different meaning\. We retain 78 cases after automatic checks for comparable length, distinct wording, and the absence of prefix copying, followed by manual verification of the intended semantic relationships\. For example, “was a really tough experience” is paired with the paraphrase “was incredibly challenging and difficult” and the unrelated continuation “from the main hallway suddenly\.” We embed the continuations without their prefixes usingSFR\-Embedding\-2\_R, one of the embedding models evaluated in the main experiment \(see Appendix[E\.1](https://arxiv.org/html/2609.10934#A5.SS1)\)\.
In 77 of 78 triplets, the original is more similar to its paraphrase than to the unrelated continuation \(98\.7%; 95% paired\-bootstrap CI: 96\.2–100%\)\. Mean cosine similarity is 0\.861 for original–paraphrase pairs and 0\.654 for original–unrelated pairs, with a mean paired difference of 0\.207 \(95% CI: 0\.190–0\.223\)\. Thus, for the short continuations used here, the embedding space captures semantic similarity across lexical variation, providing support for its use as a model\-mediated proxy\.
F Prompt TemplatesPrompt 1: Expert \(Theory\-Guided TRP Judgment\)\[ROLE\]You are a Conversation Analysis \(CA\) expert evaluating, incrementally after each prefix, whether a Transition Relevance Place \(TRP\) occursimmediately after the final word\. A TRP is an opportunity \(not an obligation\) for speaker transition or for the current speaker to begin a new Turn Construction Unit \(TCU\)\.\[SOURCES\]Note: For clarity of presentation, the academic sources originally included in the prompt have been omitted here\.\[TASK\]Given a sequence of words spoken by a single speaker \(<PREFIX\>\), decide whether it ends at a point where a listener could appropriately begin speaking or where the current speaker could transition to a new Turn Construction Unit\.\[DECISION RUBRIC\]A TRP is an*opportunity*, not an obligation, for speaker transition\.•Output 1 \(TRP\)when the utterance is syntactically and pragmatically complete \(e\.g\., an independent clause or completed social action\)\.•Output 0 \(no TRP\)when the utterance projects continuation \(e\.g\., open coordination or subordination, discourse markers, or continuation punctuation\)\.•If cues conflict, prefer0unless completion is decisively clear\.\[OUTPUT FORMAT\]```
{"does_trp_occur":0 or 1,"confidence":0.0-1.0,
"justification":"brief explanation (<=1000 chars)"}
```
The output must end with the token<END\>\. The confidence and justification fields are recorded for analysis but are not used for post\-processing or decision thresholding\.\[EXAMPLE \(SCHEMATIC\)\]Illustrative examples usefabricated text:•"I think we’re done\."→\\rightarrowdoes\_trp\_occur = 1•"I was thinking that"→\\rightarrowdoes\_trp\_occur = 0Prompt 2: Participant \(Intuitive TRP Judgment\)\[ROLE\]You are taking the role of a participant in a conversation study\. You imagine listening to a speaker and deciding whether you could naturally produce a brief encouraging response \(e\.g\., “yeah”, “mmhmm”\) at the end of what they just said\.\[TASK\]Given a single line of speech spoken by one speaker \(<PREFIX\>\), decide whether it ends at a point where you couldreasonably give a short listener response\. Always make a decision for the final word of the line\.\[DECISION RUBRIC\]A TRP is an*opportunity*, not an obligation, for speaker transition\.•Output 1 \(TRP\)when the utterance is syntactically and pragmatically complete \(e\.g\., an independent clause or completed social action\)\.•Output 0 \(no TRP\)when the utterance projects continuation \(e\.g\., open coordination or subordination, discourse markers, or continuation punctuation\)\.•If cues conflict, prefer0unless completion is decisively clear\.\[OUTPUT FORMAT\]```
{"does_trp_occur":0 or 1,"confidence":0.0-1.0,
"justification":"brief explanation (<=1000 chars)"}
```
The output must end with the token<END\>\. The confidence and justification fields are recorded for analysis but are not used for post\-processing or decision thresholding\.\[EXAMPLE \(SCHEMATIC\)\]Illustrative examples usefabricated text:•"I think we’re done\."→\\rightarrowdoes\_trp\_occur = 1•"I was thinking that"→\\rightarrowdoes\_trp\_occur = 0Prompt 3: Participant \(Imagined Future Continuation\)\[ROLE\]You are listening to a single speaker in a natural conversation and must judge whether the end of the speaker’s turn\-so\-far is a natural point for a brief listener response\.\[TASK\]First, imagine how the speaker is likely to continue next by writing a short continuation in the speaker’s voice \(1–2 clauses,≈\\approx10 words\)\. Then, decide whether the end of the given line is a point where a brief listener response would naturally fit\.\[DECISION GUIDELINES\]•Output 1 \(TRP\)if the turn\-so\-far sounds complete and the continuation would be optional\.•Output 0 \(no TRP\)if the turn\-so\-far leads directly into the provided continuation\.•Short items \(e\.g\., speaker acknowledgments\) may be complete on their own\.\[OUTPUT FORMAT\]```
{"does_trp_occur":0 or 1,
"projected_upcoming_turns":"imagined speaker continuation",
"confidence":0.0-1.0,
"justification":"brief explanation (<=1000 chars)"}
```
The output must end with the token<END\>\. The projected continuation is used only as part of the judgment and is not evaluated against ground truth\.\[EXAMPLE \(SCHEMATIC\)\]Illustrative examples usefabricated text:•"I think we’re done\."→\\rightarrowcontinuation:"We can leave whenever you’re ready\.",does\_trp\_occur = 1•"I was thinking that"→\\rightarrowcontinuation:"we should try something else\.",does\_trp\_occur = 0Prompt 4: Participant \(Oracle Future Continuation\)\[ROLE\]You are listening to a single speaker in a natural conversation and must judge whether the end of the speaker’s turn\-so\-far is a natural point for a brief listener response\.\[TASK\]You are given:1\.a line of speech spoken by one speaker \(<PREFIX\>\), and2\.theactual upcoming continuationof the same speaker\.Use the provided continuationexactly as writtento decide whether the end of the given line is a point where a brief listener response would naturally fit\.\[DECISION GUIDELINES\]•Output 1 \(TRP\)if the turn\-so\-far sounds complete and the continuation would be optional\.•Output 0 \(no TRP\)if the turn\-so\-far leads directly into the provided continuation\.•Short items \(e\.g\., speaker acknowledgments\) may be complete on their own\.\[OUTPUT FORMAT\]```
{"does_trp_occur":0 or 1,
"projected_upcoming_turns":"verbatim provided continuation",
"confidence":0.0-1.0,
"justification":"brief explanation (<=1000 chars)"}
```
The output must end with the token<END\>\. The provided continuation must be copied verbatim and is not generated by the model\. The confidence and justification fields are recorded for analysis but are not used for post\-processing or decision thresholding\.\[EXAMPLE \(SCHEMATIC\)\]Illustrative examples usefabricated text:•Turn\-so\-far:"I think we’re done\."
Continuation:"We can head out whenever you’re ready\."
→\\rightarrowdoes\_trp\_occur = 1•Turn\-so\-far:"I was thinking that"
Continuation:"we might need to try a different approach\."
→\\rightarrowdoes\_trp\_occur = 0Similar Articles
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation
This paper introduces HawkesLLM, a framework that models semantic uncertainty propagation in multi-step agentic text simulations by combining a multivariate Hawkes process for temporal influence and memory selection with a language model for text generation. Evaluation on a GDELT news-cascade case study shows improved late-stage semantic alignment under compact prompt-memory constraints.
Uncertainty-Aware Decision Making in Multimodal Large Language Models
This survey organizes research on uncertainty-aware decision making in multimodal large language models, covering sources of uncertainty, calibration methods, and actions to improve system reliability and safety.
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models
This paper proposes a training-free, uncertainty-aware inference framework for using large language models in operations research. The method uses short lookahead simulations and importance resampling to improve the coherence of mathematical formulations, outperforming standard baselines on OR benchmarks.
SKG-Eval: Stateful Evaluation of Multi-Turn Dialogue via Incremental Semantic Knowledge Graphs
Proposes SKG-Eval, a quasi-deterministic evaluation framework for multi-turn dialogue that uses incremental semantic knowledge graphs to detect cross-turn inconsistencies, contradiction, and topic drift, achieving higher correlation with human judgments.
Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue
This paper proposes a decoupled data approach to improve turn-taking in full-duplex dialogue by learning from real spoken dialogues while using text for semantics, leveraging a neural finite state machine framework to enhance naturalness and preserve semantic capabilities.