full-duplex

Tag

Cards List
#full-duplex

OpenAI launches GPT-Live-1 for full-duplex voice agents (2 minute read)

TLDR AI · 14h ago Cached

OpenAI has launched GPT-Live-1, a full-duplex voice model for API that enables natural, bidirectional voice conversations for developers, reducing latency and improving turn-taking in voice agents.

0 favorites 0 likes
#full-duplex

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

arXiv cs.AI · 2026-09-04 Cached

The paper introduces DuplexSpeechBench-IFEval, a benchmark for evaluating implicit instruction following in full-duplex voice agents, with 1,038 test cases across eight roles and five protocols to assess real-time speech systems' adherence to explicit vs. persona-implied behaviors.

0 favorites 0 likes
#full-duplex

@simplifyinAI: Most voice assistants make you wait your turn, you talk, then it talks, back and forth. NVIDIA just built one that does…

X AI KOLs Timeline · 2026-08-22 Cached

NVIDIA has launched Nemotron 3 VoiceChat, an AI model enabling real-time, full-duplex voice interactions that allow natural interruptions, available via a live demo as part of its open-source NeMo Speech framework.

0 favorites 0 likes
#full-duplex

nvidia/NVIDIA-NemotronLabs-VoiceChat-11B · Hugging Face (full duplex)

Reddit r/LocalLLaMA · 2026-08-03 Cached

NVIDIA released NemotronLabs VoiceChat 11B, an open end-to-end full-duplex speech model enabling real-time conversational AI with ~450ms turn-taking latency, barge-in, and live tool calling, the first open full-duplex model to support tool calling.

0 favorites 0 likes
#full-duplex

OPEN AI: How we built a realtime system for responsive voice AI in six months

Reddit r/singularity · 2026-08-03 Cached

OpenAI describes how they built GPT-Live, a full-duplex realtime voice AI system that eliminates the turn detector, enabling natural continuous conversation. The article details architecture improvements in inference, context management, and media transport over six months.

0 favorites 0 likes
#full-duplex

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

arXiv cs.CL · 2026-08-03 Cached

M3-DuplexBench is a new multi-turn, multilingual, multidomain benchmark for evaluating full-duplex spoken dialogue systems, supporting English and Japanese across casual conversation and question answering domains.

0 favorites 0 likes
#full-duplex

Microsoft tests new MAI Realtime voice model (2 minute read)

TLDR AI · 2026-08-03 Cached

Microsoft is testing a new native real-time voice model, MAI Realtime, in early access on its MAI Playground. The full-duplex system supports multiple languages, low latency, and configurable turn-taking, positioning it as a competitor to OpenAI's GPT Live and Sesame.

0 favorites 0 likes
#full-duplex

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

arXiv cs.CL · 2026-07-24 Cached

Introduces Instruct-FD, a benchmark for evaluating whether full-duplex speech systems can follow explicit turn-taking instructions. Results show the best model achieves only 64.4% adherence, highlighting a significant gap in instruction-following turn management.

0 favorites 0 likes
#full-duplex

Agentic coding goes hands-free as OpenAI brings GPT-Live's full duplex voice control to Codex and ChatGPT on the desktop (5 minute read)

TLDR AI · 2026-07-24 Cached

OpenAI integrates GPT-Live's full-duplex voice control into Codex and ChatGPT desktop app, enabling hands-free agentic coding with multi-threaded task execution.

0 favorites 0 likes
#full-duplex

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

arXiv cs.CL · 2026-07-10 Cached

This paper evaluates the reliability of Gemini models as audio judges for scoring full-duplex voice agent conversations, finding that Gemini 2.5 Flash shows strong agreement with human raters on most dimensions, though model swaps require re-validation.

0 favorites 0 likes
#full-duplex

GPT‑Live

Hacker News Top · 2026-07-08 Cached

OpenAI announces GPT-Live, a new full-duplex voice model that enables more natural, real-time conversations by allowing simultaneous listening and speaking, with GPT-5.5 as the backend model.

0 favorites 0 likes
#full-duplex

OpenAI releases new voice models for more natural live conversations

TechCrunch AI · 2026-07-08 Cached

OpenAI released new full-duplex voice models GPT-Live-1 and GPT-Live-1 mini for more natural live conversations, allowing simultaneous speaking and listening, with improvements in turn-taking and context handling, and replacing Advanced Voice Mode in ChatGPT.

0 favorites 0 likes
#full-duplex

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

arXiv cs.CL · 2026-07-08 Cached

This paper introduces Lychee-FD, a native end-to-end full-duplex spoken language model that mitigates modality interference through a hierarchical parameter separation strategy, achieving significant improvements in speech intelligence and interaction fluidity.

0 favorites 0 likes
#full-duplex

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

arXiv cs.CL · 2026-07-03 Cached

TurnNat is a likelihood-based framework for automatically evaluating turn-taking naturalness in dyadic spoken dialogue, using a causal turn-taking prediction model trained on natural conversations to measure timing atypicality via negative log-likelihood.

0 favorites 0 likes
#full-duplex

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

arXiv cs.CL · 2026-06-15 Cached

BayLing-Duplex is a native full-duplex speech language model that enables a single autoregressive LLM to manage turn-taking and interruptions without external VAD modules, achieving high success rates and improved response quality over prior models.

0 favorites 0 likes
#full-duplex

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering

arXiv cs.CL · 2026-06-11 Cached

This paper identifies 'state inertia' in full-duplex spoken language models, where the model's internal predictive focus lags during user interruptions, and proposes a training-free activation steering method to improve interruption handling.

0 favorites 0 likes
#full-duplex

@kyutai_labs: New paper: Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models We use RL to post-train speech models (Mo…

X AI KOLs Following · 2026-06-10 Cached

Kyutai Labs released a new paper on using reinforcement learning to post-train speech models (Moshi and PersonaPlex) for more human-like interaction, including when to respond, wait, or give listening cues.

0 favorites 0 likes
#full-duplex

Full duplex vs half duplex - the spectrum of AI voice models [D]

Reddit r/MachineLearning · 2026-06-01

An analysis of half-duplex vs full-duplex architecture in AI voice models, discussing key features like overlap, backchannels, and barge-in that make voice agents sound robotic.

0 favorites 0 likes
#full-duplex

Raon-Speech Technical Report

arXiv cs.CL · 2026-05-26 Cached

Raon-Speech is a 9B-parameter speech language model for English and Korean, supporting understanding, answering, and generation, with a full-duplex extension Raon-SpeechChat for natural real-time conversation. It achieves strong performance across 42 benchmarks and is fully open-sourced.

0 favorites 0 likes
#full-duplex

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

arXiv cs.CL · 2026-05-21 Cached

This paper analyzes synchronization and turn-taking dynamics in full-duplex speech dialogue models by simulating conversations between two instances of the Moshi model, measuring representational alignment via CKA and predicting turn boundaries with LSTM probes.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback