Tag
This paper introduces a lightweight backchannel head for full-duplex spoken dialogue models to predict and control the timing of backchannels, improving natural conversation dynamics.
A developer describes a failure mode in AI phone agents where the agent stalls mid-call without errors, causing dropped calls and turn-taking issues, and seeks diagnostic advice.
The paper introduces TACT, a benchmark for evaluating turn-taking in full-duplex spoken dialogue models that uses intent-conditioned continuous scoring to replace binary rules, demonstrating better alignment with human judgments.
This paper evaluates full-duplex speech models' ability to decide when to speak, finding that models like Moshi and PersonaPlex primarily respond to being addressed or silence rather than content-driven triggers such as false claims or hazards, identifying a gap in content understanding.
TurnBench introduces a multi-domain benchmark for assessing turn-taking dynamics in spoken dialogue, featuring a hand-labeled corpus and standardized evaluation protocols for end-of-turn and interruption detection.
This paper proposes using semantic uncertainty derived from large language models to anticipate transition relevance places in spoken turn-taking, showing improved performance over baselines in dialogue systems.
This paper proposes a decoupled data approach to improve turn-taking in full-duplex dialogue by learning from real spoken dialogues while using text for semantics, leveraging a neural finite state machine framework to enhance naturalness and preserve semantic capabilities.
This paper proposes a multi-task learning approach that leverages turn-taking dynamics to enhance intent recognition in multi-party conversations, outperforming existing methods that ignore interaction patterns.
An observation about how AI voice agents struggle with realistic customer behavior such as interruptions, self-corrections, and mid-sentence changes, suggesting turn-taking is a key challenge for enterprise voice AI.
This paper introduces X2-Turn, a frame-synchronous dual-head model that jointly performs streaming ASR and turn state prediction on shared representations, improving turn-taking accuracy and latency in spoken dialogue systems.
DuplexGen introduces a method for adaptively synthesizing human-AI turn-taking dialogues, addressing the challenge of natural interaction timing in conversational AI.
Introduces Instruct-FD, a benchmark for evaluating whether full-duplex speech systems can follow explicit turn-taking instructions. Results show the best model achieves only 64.4% adherence, highlighting a significant gap in instruction-following turn management.
This paper explores how silence thresholds in turn-taking differ between human and AI-generated discourse, using a distant viewing approach to analyze conversational patterns.
This paper reexamines addressee detection in multi-party dialogue, proposing continuous address levels over discrete labels and showing that address relates to gaze and backchannels beyond turn-taking, suggesting graded structure.
Presents a multimodal voice activity projection framework extending audio-only VAP to audio-visual inputs for turn-taking prediction in social robots, using pretrained backbones and low-rank adaptation. Achieves improvements on NoXi and Haru EDR corpora.
TurnNat is a likelihood-based framework for automatically evaluating turn-taking naturalness in dyadic spoken dialogue, using a causal turn-taking prediction model trained on natural conversations to measure timing atypicality via negative log-likelihood.
A developer describes building a multi-agent voice social-deduction game, solving turn-taking with a central conductor but struggling with shared memory and preserving social subtext when compressing conversation history into structured state.
This article shares hard-won lessons from building real-time voice AI agents, highlighting the importance of proper turn-taking, VAD handling, billing awareness, and avoiding echo loops.
This paper evaluates the abilities of large language models (LLMs) and multimodal LLMs for addressee detection, turn-change prediction, and next speaker prediction in multi-party meeting conversations. Results show text-based LLMs outperform supervised models and humans in next speaker prediction, while multimodal LLMs improve over text-only models in other tasks but remain below human performance.
BayLing-Duplex is a native full-duplex speech language model that enables a single autoregressive LLM to manage turn-taking and interruptions without external VAD modules, achieving high success rates and improved response quality over prior models.