在流式全双工模型中控制反馈通道

arXiv cs.CL 论文

摘要

本文介绍了一种用于全双工语音对话模型的轻量级反馈通道头部,以预测和控制反馈时机,改善自然对话动态。

arXiv:2609.29418v1 Announce Type: new Abstract: Backchannels, brief acknowledgements like "uh-huh" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:21

# Controlling Backchannels in Streamable Full-duplex Models
Source: [https://arxiv.org/html/2609.29418](https://arxiv.org/html/2609.29418)
###### Abstract

Backchannels, brief acknowledgements like “uh\-huh” produced while the other party may still be talking, are central to natural conversation, but full\-duplex spoken dialogue models rarely model them explicitly\. We introduce a lightweight backchannel head that predicts, from a full\-duplex model’s own hidden states, when a backchannel should begin\. Once this probability crosses a tunable threshold, a backchannel is force\-decoded\. Attached to both a 7B \(PersonaPlex\) and a 1B \(F\-Actor\) model, it generalizes across scale\. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better\-timed backchannels\. Human raters judge the resulting backchannels on par with real ones\.

###### Index Terms:

full\-duplex, backchannel, turn\-taking

††address:1Karlsruhe Institute of Technology, Germany2AppTek, Germany3Charles University, Czech Republic4Sesame AI, USA5University of Edinburgh, UK[maike\.zuefle@kit\.edu](mailto:[email protected])## 1Introduction

Full\-duplex spoken dialogue models\[[1](https://arxiv.org/html/2609.29418#bib.bib2),[2](https://arxiv.org/html/2609.29418#bib.bib3),[3](https://arxiv.org/html/2609.29418#bib.bib10),[4](https://arxiv.org/html/2609.29418#bib.bib11)\]can listen and speak simultaneously, making them well suited for generating backchannels, i\.e\., short acknowledgements, like “uh\-huh”, a listener produces while the other party may still be talking\.

Backchannel modelling received attention before the advent of full\-duplex modelling, with earlier systems using explicit prediction heads to generate backchannels and improve rapport\[[5](https://arxiv.org/html/2609.29418#bib.bib15),[6](https://arxiv.org/html/2609.29418#bib.bib23)\]\. Voice Activity Projection \(VAP\) models predict backchanneling frame\-by\-frame as an external classifier in a cascaded pipeline\[[7](https://arxiv.org/html/2609.29418#bib.bib6),[8](https://arxiv.org/html/2609.29418#bib.bib5)\]\. Endpointing targets the more general problem of turn boundaries\[[9](https://arxiv.org/html/2609.29418#bib.bib8)\]\. Complementary work studies the lexico\-prosodic meaning of backchannels\[[10](https://arxiv.org/html/2609.29418#bib.bib13)\]\.

In full\-duplex dialogue modelling, recent models produce backchannels as a byproduct of training on conversational data\[[3](https://arxiv.org/html/2609.29418#bib.bib10),[4](https://arxiv.org/html/2609.29418#bib.bib11)\]\. Reinforcement learning can further shape this behaviour into more natural backchanneling\[[11](https://arxiv.org/html/2609.29418#bib.bib9)\]\. Though effective, this is expensive and yields an implicit policy with no controllable signal for when or why backchannels occur\. Closer to explicit control, one line of work first pretrains a full\-duplex model and then adds VAP heads for turn\-taking\[[12](https://arxiv.org/html/2609.29418#bib.bib12)\], while another probes a full\-duplex model’s representations for turn\-taking without pretraining\[[13](https://arxiv.org/html/2609.29418#bib.bib14)\]\.

Other existing evaluations of full\-duplex backchanneling focus on whether models backchannel at all\[[2](https://arxiv.org/html/2609.29418#bib.bib3),[4](https://arxiv.org/html/2609.29418#bib.bib11)\], and not whether those backchannels are appropriate\. Therefore, it is critical to evaluate backchannel placement against real human conversational timing, and to train a model that explicitly predicts when a backchannel should occur\.

We address this gap with a lightweight, model\-independent backchannel head\. It predicts, at each timestep, the probability that a backchannel should begin, and when this probability crosses a threshold, we force\-decode a backchannel token into the text stream, which in turn drives the corresponding audio output\. This is in the spirit of frame\-wise VAP\-based backchannel predictors\[[7](https://arxiv.org/html/2609.29418#bib.bib6)\], but attached directly to a full\-duplex model’s own hidden states rather than run as an external classifier\. We attach this head to two full\-duplex models at different scales, the 7B PersonaPlex\[[3](https://arxiv.org/html/2609.29418#bib.bib10)\]and the 1B F\-Actor\[[4](https://arxiv.org/html/2609.29418#bib.bib11)\], to show that the method generalizes\. Since F\-Actor does not support streaming inference, we additionally adapt it into a streaming model to enable this comparison\.

We evaluate the resulting models along three complementary axes\. First, we measure whether the model backchannels more often than the unmodified baseline\. Second, we probe the models’ hidden states to test whether they predict backchannel onset where real speakers actually backchanneled in human conversations, using human\-annotated timing as ground truth\. Third, we run a human evaluation in which participants rate how appropriate the backchannels in generated dialogues sound, finding that they are rated on par with human backchannels\.

Our contributions are as follows:

- •A model\-independent method for controlling backchannel timing, which we show generalizes across two architectures and two model scales\.
- •A human\-aligned evaluation of backchannel behaviour, combining a probing analysis against real human timing data with a human appropriateness study\.
- •A streaming version of F\-Actor to validate our method on models of multiple sizes\.

We release the code and models openly\.111[https://github\.com/MaikeZuefle/bcmore](https://github.com/MaikeZuefle/bcmore)

## 2Controllable Backchanneling

This section presents our proposed model\-independent backchannel head\. Our design fulfils two objectives: \(1\) At each frame, the model predicts whether the current point in the interlocutor’s speech is an appropriate moment for a backchannel, i\.e\., a moment at which a human would produce one\. \(2\) The system must provide an independent control mechanism to adjust how frequently backchannels are emitted without degrading the appropriateness of their placement\.

### 2\.1Backchannel Head

To predict backchannel timing without altering the backbone’s language modelling capabilities, we attach a lightweight MLP head directly to its hidden states\. Let𝐡t∈ℝd\\mathbf\{h\}\_\{t\}\\in\\mathbb\{R\}^\{d\}denote the hidden state of the backbone \(the temporal transformer\) at timesteptt\. The head maps𝐡t\\mathbf\{h\}\_\{t\}to a probability via a two\-layer MLP with a sigmoid outputσ⁡\(⋅\)\\sigma\(\\cdot\):

pt=σ⁡\(𝐖2​GELU⁡\(𝐖1​𝐡t\+𝐛1\)\+b2\),p\_\{t\}=\\sigma\(\\mathbf\{W\}\_\{2\}\\operatorname\{GELU\}\(\\mathbf\{W\}\_\{1\}\\mathbf\{h\}\_\{t\}\+\\mathbf\{b\}\_\{1\}\)\+b\_\{2\}\),\(1\)wherept∈\[0,1\]p\_\{t\}\\in\[0,1\]is the predicted probability of a backchannel onset, i\.e\., the first frame of a backchannel, att\+1t\+1, andb2b\_\{2\}is initialized to the empirical log\-odds of the class prior\.

To handle extreme class imbalance \(onsets account for∼\\sim1% of conversational frames\), the head is trained using Focal Loss\[[14](https://arxiv.org/html/2609.29418#bib.bib22)\]evaluated only over eligible frames𝒯valid\\mathcal\{T\}\_\{\\text\{valid\}\}where the partner is speaking and the agent is silent\. All frames during agent speech or mutual silence are excluded from𝒯valid\\mathcal\{T\}\_\{\\text\{valid\}\}\. Letyt∈\{0,1\}y\_\{t\}\\in\\\{0,1\\\}denote the frame\-level target backchannel onset label at timestept\+1t\+1and letp~t=yt​pt\+\(1−yt\)​\(1−pt\)\\tilde\{p\}\_\{t\}=y\_\{t\}p\_\{t\}\+\(1\-y\_\{t\}\)\(1\-p\_\{t\}\)denote the probability assigned to the target class using the predicted probabilitiesptp\_\{t\}\. The Focal Lossℒhead\\mathcal\{L\}\_\{\\text\{head\}\}is then computed as:

ℒhead=−1\|𝒯valid\|∑t∈𝒯validαt\(1−p~t\)γlog\(p~t\),\\mathcal\{L\}\_\{\\text\{head\}\}=\-\\frac\{1\}\{\|\\mathcal\{T\}\_\{\\text\{valid\}\}\|\}\\sum\_\{t\\in\\mathcal\{T\}\_\{\\text\{valid\}\}\}\\alpha\_\{t\}\(1\-\\tilde\{p\}\_\{t\}\)^\{\\gamma\}\\log\(\\tilde\{p\}\_\{t\}\),\(2\)whereαt=α​yt\+\(1−α\)​\(1−yt\)\\alpha\_\{t\}=\\alpha y\_\{t\}\+\(1\-\\alpha\)\(1\-y\_\{t\}\)balances class frequencies with hyperparametersα=0\.9\\alpha=0\.9andγ=2\.0\\gamma=2\.0\.

### 2\.2Thresholding and Conditioned Text Generation

During inference, we decouple placement detection from token emission via a thresholdτ∈\[0,1\]\\tau\\in\[0,1\]: a backchannel is triggered wheneverpt≥τp\_\{t\}\\geq\\tau, giving control over backchanneling frequency independently of placement quality\.

Upon triggering at timesteptt, we force\-decode the special\[BC\]token into the model’s text stream, replacing the new\-word marker\[EPAD\]\. The model then continues autoregressive generation to produce the backchannel’s lexical form \(e\.g\., “yeah”, “uh\-huh”, “right”\), conditioned on the injected token and surrounding context\.

Critically, we mask the loss on the\[BC\]token itself, so the backbone never learns to emit it spontaneously, leaving initiation entirely to the thresholded head, while retaining standard cross\-entropy loss on the subsequent transcript tokens that realize the backchannel’s content\.

## 3Experimental Setup

### 3\.1Models

We evaluate our approach on two full\-duplex models\. We trainPersonaPlex \(7B\)\[[3](https://arxiv.org/html/2609.29418#bib.bib10)\]with the backchannel mechanism from[Section2](https://arxiv.org/html/2609.29418#S2)on a single H100 80GB GPU for approximately 10 hours\. The backbone is fine\-tuned during training using LoRA\[[15](https://arxiv.org/html/2609.29418#bib.bib24)\], while the depth transformer is kept frozen\.

In addition, we adaptF\-Actor \(1B\)\[[4](https://arxiv.org/html/2609.29418#bib.bib11)\]\. Backchanneling requires listening and speaking in real time, since backchannels overlap with the interlocutor’s speech\. F\-Actor, however, does not support streaming inference, as it relies on NanoCodec\[[16](https://arxiv.org/html/2609.29418#bib.bib17)\]\. We therefore replace NanoCodec with the causal Mimi codec\[[2](https://arxiv.org/html/2609.29418#bib.bib3)\]and, following\[[17](https://arxiv.org/html/2609.29418#bib.bib16)\], add a depth transformer pretrained from[CSM 1B](https://huggingface.co/sesame/CSM-1B)\(frozen\) to model Mimi’s codebooks\. Speaker conditioning uses ECAPA\-TDNN\[[18](https://arxiv.org/html/2609.29418#bib.bib18)\]embeddings\. The temporal transformer backbone is[Llama 3\.2 1B Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)\[[19](https://arxiv.org/html/2609.29418#bib.bib25)\], with the BC head from[Section2\.1](https://arxiv.org/html/2609.29418#S2.SS1)attached\. Unlike PersonaPlex, we train it only to predict the system channel’s inner monologue and Mimi codes\. Training takes 13 hours on four A100 80GB GPUs\.

Models with the backchannel head carry the suffix \-BC\.

Figure 1:Correct backchannel timing F1 score per layer for predicting backchannel onset from frozen hidden states on TurnBench, for base and backchannel\-head\-trained \(BC\) models\. A prediction only counts as correct if it lands at the exact annotated frame\.Figure 2:Predicted backchannel frequency per layer, for base and backchannel\-head\-trained \(BC\) models on TurnBench\.
### 3\.2Data

#### Training\.

We use Fisher\[[20](https://arxiv.org/html/2609.29418#bib.bib20)\], which contains around 2000h of English conversations, restoring the 8kHz audio into 24kHz with Sidon\[[21](https://arxiv.org/html/2609.29418#bib.bib19)\]\. Each conversation is transcribed with[Parakeet TDT 0\.6B V3](https://huggingface.co/nvidia/Parakeet-TDT-0.6B-v3)\[[22](https://arxiv.org/html/2609.29418#bib.bib7)\], and split into 90s chunks at the nearest utterance end\. To condition the model to reliably produce backchannels upon receiving the\[BC\]token, we augment these dialogues by inserting additional backchannels from other occurrences at locations where the agent is silent for at least 1\.5 seconds, and the interlocutor keeps speaking for at least 1s, constraints chosen after pilot experiments\.

#### Evaluation\.

We use TurnBench\[[23](https://arxiv.org/html/2609.29418#bib.bib21)\], human\-human conversations with backchannel annotations from three annotators per speaker \(majority vote\)\. We restrict to the twelve conversations where all three segmented the same turns: 137 minutes, 574 backchannels out of 1,888 total utterances\. Each conversation is split into 60s windows \(plus 10s prior context\), snapped to the nearest silence gap and capped at 120s, and evaluated with leave\-one\-conversation\-out cross\-validation\. Since TurnBench gives only utterance\-level timestamps, we obtain word\-level ones via[Parakeet TDT 0\.6B V3](https://huggingface.co/nvidia/Parakeet-TDT-0.6B-v3)\[[22](https://arxiv.org/html/2609.29418#bib.bib7)\]on each turn’s own audio and transcript\.

## 4Evaluation

### 4\.1Probing

We test whether a model anticipates when a backchannel is appropriate by force\-decoding annotated human conversations and training a probe to predict from the hidden state at framettwhether a backchannel onset occurs att\+1t\+1\. We probe both the base models and their counterparts fine\-tuned with the backchannel head \(\-BC\): since the backbone is fine\-tuned jointly with the head \([Section3](https://arxiv.org/html/2609.29418#S3)\), training can change the hidden states, and the comparison tests whether it makes upcoming onsets easier to predict\. A separate probe, rather than the head itself, lets us measure the base models, which have no head, in the same way\.

#### Method\.

For each conversation, we force\-decode the ground\-truth audio of both speakers and the word\-aligned text stream, silencing the real backchannels so that predictions can also be tested during and after them, and extract the hidden states of every temporal\-transformer layer\. Probing all layers ensures a fair comparison with the base models, which may encode backchannel timing best at an intermediate layer, and shows whether training makes the signal available at the last layer, which the head reads at inference\. For each layer, we fit an MLP probe with class\-balanced weighting on TurnBench and evaluate it on held\-out conversations\.

#### Evaluation\.

Using the probe, we count how oftenp⁡\(yt=1∣ht\)p\(y\_\{t\}=1\\mid h\_\{t\}\)crosses a fixed threshold and compare this rate to the human rate\. We separately score whether the predicted timing itself is correct against the human annotations, reporting F1\.

Pause \(Synth\.\)Pause \(Candor\)BackchannelSmooth Turn TakingUser InterruptionModelTOR↓\\downarrowTOR↓\\downarrowTOR↓\\downarrowFreq↑\\uparrowJSD↓\\downarrowTOR↑\\uparrowLatency↓\\downarrowTOR↑\\uparrowLatency↓\\downarrowMoshi0\.4450\.5280\.2550\.0740\.8240\.7390\.1620\.9201\.377\+ RL\[[11](https://arxiv.org/html/2609.29418#bib.bib9)\]0\.2260\.4170\.0910\.0950\.7890\.9660\.1211\.0000\.461PersonaPlex0\.4820\.4440\.1820\.0460\.8410\.9580\.2190\.9400\.271\+ RL\[[11](https://arxiv.org/html/2609.29418#bib.bib9)\]0\.3280\.3610\.1270\.1220\.7830\.9500\.0791\.0000\.187\+ BClow\{\}\_\{\\text\{low\}\}\(ours\)0\.5690\.6200\.2910\.0810\.7800\.9580\.0980\.9350\.294\+ BCnormal\{\}\_\{\\text\{normal\}\}\(ours\)0\.5770\.6300\.2360\.0880\.7780\.9580\.0980\.9350\.295\+ BCmax\{\}\_\{\\text\{max\}\}\(ours\)0\.7660\.7730\.2180\.2590\.6720\.9830\.0150\.9450\.181F\-Actor0\.4090\.2960\.6730\.1020\.7430\.5130\.6750\.7852\.008\+ BClow\{\}\_\{\\text\{low\}\}\(ours\)0\.5330\.4630\.6000\.1180\.7210\.7310\.1240\.9251\.443\+ BCnormal\{\}\_\{\\text\{normal\}\}\(ours\)0\.3430\.2960\.6180\.1220\.7240\.5040\.9050\.8451\.613\+ BCmax\{\}\_\{\\text\{max\}\}\(ours\)0\.6500\.5690\.6000\.1510\.7090\.7650\.1870\.9201\.507Table 1:Comparison of Moshi, PersonaPlex, F\-Actor, and our BC variants on Full\-Duplex\-Bench v1\.

### 4\.2Evaluating the Generated Backchannels

We test whether the trained head produces correctly timed backchannels once it actually drives generation, and whether doing so comes at the cost of the model’s general full\-duplex conversational abilities\. We evaluate our BC\-head\-augmented models on Full\-Duplex\-Bench v1\[[24](https://arxiv.org/html/2609.29418#bib.bib1)\]alongside their base models, isolating the effect of the backchannel head\. The benchmark reports backchannel\-specific metrics alongside metrics for pause handling, turn\-taking, and interruption\. We use three thresholds:normalmatches the human TurnBench rate;low/maxyield fewer/more backchannels\.

### 4\.3Human Evaluation

In addition to the automatic evaluation, we assess whether human listeners judge the models’ backchannels as appropriately placed\. We randomly select 35 TurnBench samples of 10–20 seconds, each containing one backchannel, and force\-decode them with the PersonaPlex\-BC model, letting it predict the backchannel\. As a baseline, we also silence the real backchannels and reinsert them at random points\. Each sample thus has three versions, human, random, and model, each rated 1–5 for appropriate backchanneling by four annotators \(12 in total\) via the Pearmut platform\[[25](https://arxiv.org/html/2609.29418#bib.bib4)\]\.

## 5Results

We evaluate the backchannel\-head mechanism in three stages: probing its hidden states, generation, and human evaluation\.

### 5\.1Probing

#### Backchannel Frequency\.

[Figure2](https://arxiv.org/html/2609.29418#S3.F2)shows the probe’s predicted backchannel frequency per layer against the human rate\. PersonaPlex\-BC reaches this human rate when probing the last layer at a well\-calibrated threshold around 0\.5\. F\-Actor\-BC likewise improves over its own base model at the last layer, but to a smaller degree\. This suggests that the backchannel\-head mechanism successfully shifts the model’s hidden states toward human\-like backchannel frequency\.

#### Onset Timing Accuracy\.

F1 measures whether the probe’s predicted onset lands at the correct moment rather than merely how often it fires\.[Figure1](https://arxiv.org/html/2609.29418#S3.F1)shows PersonaPlex\-BC with substantially higher F1 than base PersonaPlex across nearly all layers, while F\-Actor\-BC’s large gain over base F\-Actor at middle layers narrows to a small improvement by the last layer\. F\-Actor\(\-BC\) also performs best at low decision thresholds\. A tolerance of two frames \(±\\pm0\.16s\) around the annotated onset improves both models’ scores\.

### 5\.2Evaluating the Generated Backchannels

We now evaluate the backchannels the model generates\.

#### Full\-Duplex\-Bench\.

[Table1](https://arxiv.org/html/2609.29418#S4.T1)compares the base models against our BC\-augmented versions and against RL\-based backchannel shaping\[[11](https://arxiv.org/html/2609.29418#bib.bib9)\]\. We find that BC backchannels more often and with better timing \(lower JSD\) than the base models\. A lower threshold produces more backchannels, as expected \(max vs\. low\)\. Compared to RL, PersonaPlex\-BCmax\{\}\_\{\\text\{max\}\}achieves higher backchannel frequency, better timing \(JSD\) and faster turn\-taking, but higher take\-over rates \(TOR\), especially on pauses\. F\-Actor’s pause TOR at the normal setting, in contrast, improves over its base model and is among the best in the table\.

#### Surface Forms\.

Beyond deciding when to backchannel, the model also has to choose what to say\. We test this in two ways: via free generation and via generation constrained to the training data’s own token paths\. Both produce similar lexical distributions, for example roughly doubling the human share ofyeah\. Despite this lexical similarity, listening to the audio, we find the constrained version sounds more natural\.

Figure 3:Results of the human evaluation\.

### 5\.3Human Evaluation

We ask human evaluators to judge backchannel appropriateness \([Figure3](https://arxiv.org/html/2609.29418#S5.F3)\): Human recordings score 4\.34/5, PersonaPlex\-BC 4\.37/5, and the random baseline 3\.77/5, with both human and PersonaPlex\-BC rated significantly higher than random \(p<0\.0001p<0\.0001\) but not significantly different from each other \(p=0\.47p=0\.47\)\. Inter\-annotator agreement \(Krippendorff’sα\\alpha, interval metric\) is0\.420\.42, moderate and consistent with the inherent subjectivity of naturalness judgments\.

## 6Discussion and Conclusion

We show that backchannel timing can be learned as a lightweight, pluggable signal, rather than left to emerge implicitly from conversational training\. Across two architectures, this explicit control improves backchannel frequency and timing, and in a human evaluation its backchannels are judged on par with human ones\. We release our models and code to make this control mechanism easy to build upon\.

## 7Acknowledgments

This work was supported by JSALT 2026 at Johns Hopkins University with funds from NSF CCRI Grant No\. 2120435, Google DeepMind, JHU HLTCOE, JHU AI2AI and ACL, and has received funding from the European Union’s Horizon research and innovation programme under grant agreement No 101135798, project Meetween \(My Personal AI Mediator for Virtual MEETtings BetWEEN People\)\.

Generative AI tools were used to assist with editing and grammar checking of the manuscript, as well as for coding and plotting\. All scientific content, analyses, and conclusions were developed and verified by the authors\.

## References

- \[1\]K\. Hu, E\. Hosseini\-Asl, C\. Chen, E\. Casanova, S\. Ghosh, P\. Żelasko, Z\. Chen, J\. Li, J\. Balam, and B\. Ginsburg\(2025\)Efficient and Direct Duplex Modeling for Speech\-to\-Speech Language Model\.InProc\. Interspeech,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p1.1)\.
- \[2\]A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour\(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.arXiv:2410\.00037\.Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p1.1),[§1](https://arxiv.org/html/2609.29418#S1.p4.1),[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p2.1)\.
- \[3\]R\. Roy, J\. Raiman, S\. Lee, T\. Ene, R\. Kirby, S\. Kim, J\. Kim, and B\. Catanzaro\(2026\)PersonaPlex: voice and role control for full duplex conversational speech models\.InProc\. ICASSP,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p1.1),[§1](https://arxiv.org/html/2609.29418#S1.p3.1),[§1](https://arxiv.org/html/2609.29418#S1.p5.1),[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p1.1)\.
- \[4\]M\. Züfle, O\. Klejch, N\. Sanders, J\. Niehues, A\. Birch, and T\. K\. Lam\(2026\)F\-Actor: controllable conversational behavior in full\-duplex models\.InACL Findings,External Links:ISBN 979\-8\-89176\-395\-1Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p1.1),[§1](https://arxiv.org/html/2609.29418#S1.p3.1),[§1](https://arxiv.org/html/2609.29418#S1.p4.1),[§1](https://arxiv.org/html/2609.29418#S1.p5.1),[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p2.1)\.
- \[5\]D\. Lala, P\. Milhorat, K\. Inoue, M\. Ishida, K\. Takanashi, and T\. Kawahara\(2017\)Attentive listening system with backchanneling, response generation and flexible turn\-taking\.InProc\. SIGDial,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p2.1)\.
- \[6\]R\. Ruede, M\. Müller, S\. Stüker, and A\. Waibel\(2018\)Yeah, right, uh\-huh: a deep learning backchannel predictor\.InProc\. IWSDS,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p2.1)\.
- \[7\]K\. Inoue, D\. Lala, G\. Skantze, and T\. Kawahara\(2025\)Yeah, un, oh: continuous and real\-time backchannel prediction with fine\-tuning of voice activity projection\.InProc\. NAACL\-HLT,External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.367),ISBN 979\-8\-89176\-189\-6Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p2.1),[§1](https://arxiv.org/html/2609.29418#S1.p5.1)\.
- \[8\]K\. Inoue, M\. Elmers, Y\. Fu, Z\. H\. Pang, T\. Mori, D\. Lala, K\. Ochi, and T\. Kawahara\(2026\)Multilingual and continuous backchannel prediction: a cross\-lingual study\.InProc\. IWSDS,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p2.1)\.
- \[9\]S\. Udupa, S\. Watanabe, P\. Schwarz, and J\. Cernocky\(2026\)Endpoint anticipation for low\-latency spoken dialogue\.InProc\. Interspeech,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p2.1)\.
- \[10\]L\. Qian and G\. Skantze\(2026\)Aligning backchannel and dialogue context representations via contrastive LLM fine\-tuning\.InProc\. ACL,Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p2.1)\.
- \[11\]A\. Ohashi, N\. Zeghidour, A\. Défossez, and E\. Kharitonov\(2026\)Multi\-faceted interactivity alignment in full\-duplex speech models\.arXiv:2606\.11167\.Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p3.1),[Table 1](https://arxiv.org/html/2609.29418#S4.T1.2.4.1.1),[Table 1](https://arxiv.org/html/2609.29418#S4.T1.2.6.1.1),[§5\.2](https://arxiv.org/html/2609.29418#S5.SS2.SSS0.Px1.p1.1)\.
- \[12\]S\. Rajaa\(2026\)DualTurn: learning turn\-taking from dual\-channel generative speech pretraining\.arXiv:2603\.08216\.Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p3.1)\.
- \[13\]P\. Riera, P\. Brusco, C\. Kuo, M\. Sancinetti, and S\. Branavan\(2026\)Synchronization and turn\-taking in full\-duplex speech dialogue models\.arXiv:2605\.20356\.Cited by:[§1](https://arxiv.org/html/2609.29418#S1.p3.1)\.
- \[14\]T\. Lin, P\. Goyal, R\. Girshick, K\. He, and P\. Dollar\(2017\)Focal loss for dense object detection\.InProceedings of the IEEE International Conference on Computer Vision \(ICCV\),Cited by:[§2\.1](https://arxiv.org/html/2609.29418#S2.SS1.p2.1)\.
- \[15\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2021\)LoRA: low\-rank adaptation of large language models\.arXiv:2106\.09685\.Cited by:[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p1.1)\.
- \[16\]E\. Casanova, P\. Neekhara, R\. Langman, S\. Hussain, S\. Ghosh, X\. Yang, A\. Jukic, J\. Li, and B\. Ginsburg\(2025\)NanoCodec: Towards High\-Quality Ultra Fast Speech LLM Inference\.InProc\. Interspeech,Cited by:[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p2.1)\.
- \[17\]N\. Torgashov, G\. E\. Henter, and G\. Skantze\(2026\)VoXtream2: full\-stream TTS with dynamic speaking rate control\.arXiv:2603\.13518\.Cited by:[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p2.1)\.
- \[18\]B\. Desplanques, J\. Thienpondt, and K\. Demuynck\(2020\)ECAPA\-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification\.InProc\. Interspeech,Cited by:[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p2.1)\.
- \[19\]A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma\(2024\)The Llama 3 herd of models\.arXiv:2407\.21783\.External Links:2407\.21783Cited by:[§3\.1](https://arxiv.org/html/2609.29418#S3.SS1.p2.1)\.
- \[20\]C\. Cieri, D\. Miller, and K\. Walker\(2004\)The Fisher corpus: a resource for the next generations of speech\-to\-text\.InLREC,Vol\.4,pp\. 69–71\.Cited by:[§3\.2](https://arxiv.org/html/2609.29418#S3.SS2.SSS0.Px1.p1.1)\.
- \[21\]W\. Nakata, Y\. Saito, Y\. Ueda, and H\. Saruwatari\(2026\)Sidon: fast and robust open\-source multilingual speech restoration for large\-scale dataset cleansing\.InProc\. ICASSP,Cited by:[§3\.2](https://arxiv.org/html/2609.29418#S3.SS2.SSS0.Px1.p1.1)\.
- \[22\]M\. Sekoyan, N\. R\. Koluguri, N\. Tadevosyan, P\. Zelasko, T\. Bartley, N\. Karpov, J\. Balam, and B\. Ginsburg\(2025\)Canary\-1B\-v2 & Parakeet\-TDT\-0\.6B\-v3: efficient and high\-performance models for multilingual ASR and AST\.arXiv:2509\.14128\.Cited by:[§3\.2](https://arxiv.org/html/2609.29418#S3.SS2.SSS0.Px1.p1.1),[§3\.2](https://arxiv.org/html/2609.29418#S3.SS2.SSS0.Px2.p1.1)\.
- \[23\]F\. Jiang, R\. Sanabria, S\. Deshmukh, B\. Veluri, S\. M\. V\. Williams, E\. K\. Suen, G\. Lee, K\. Y\. Choi, T\. Umeki, R\. Kubo,et al\.\(2026\)TurnBench: a multi\-domain benchmark for turn\-taking dynamics in spoken dialogue\.arXiv:2608\.25218\.Cited by:[§3\.2](https://arxiv.org/html/2609.29418#S3.SS2.SSS0.Px2.p1.1)\.
- \[24\]G\. Lin, J\. Lian, T\. Li, Q\. Wang, G\. Anumanchipalli, A\. H\. Liu, and H\. Lee\(2025\)Full\-Duplex\-Bench: a benchmark to evaluate full\-duplex spoken dialogue models on turn\-taking capabilities\.InProc\. ASRU Workshop,Cited by:[§4\.2](https://arxiv.org/html/2609.29418#S4.SS2.p1.1)\.
- \[25\]V\. Zouhar and T\. Kocmi\(2026\)Pearmut: human evaluation of translation made trivial\.arXiv:2601\.02933\.Cited by:[§4\.3](https://arxiv.org/html/2609.29418#S4.SS3.p1.1)\.

相似文章

基于对比 LLM 微调对齐对话附和信号与语境表征

arXiv cs.CL

KTH Royal Institute of Technology 的研究人员提出了一种两阶段框架,通过在对话转写文本上微调 LLMs,并结合对比学习构建联合嵌入空间,以实现对对话附和信号与语境的精准对齐。结果表明,相较于以往方法,该方案显著提升了语境与附和信号的匹配检索性能。

全双工语音模型中工具调用的前后端架构

arXiv cs.CL

本文提出了一种前后端架构,用于在全双工语音模型中实现工具调用。语音前端将任务委托给基于文本的LLM后端进行工具调用,从而保持低延迟和自然交互。