speech-to-speech

Tag

Cards List
#speech-to-speech

@amitiitbhu: Design a Real-Time Voice AI Agent Read here:

X AI KOLs Timeline ↗ · 5d ago Cached

Outcome School 公开了一篇详细教程,讲解如何设计实时语音 AI 智能体,涵盖级联管道(STT→LLM→TTS)、端到端 Speech-to-Speech 模型与混合方案、延迟预算、打断处理(barge-in)、工具调用、记忆与上下文、电话系统集成、规模化部署、边缘情况、可观测性、安全与成本估算。

0 favorites 0 likes
#speech-to-speech

Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

arXiv cs.CL ↗ · 2026-09-28 Cached

This paper introduces randomized guidance as an efficient training method for tandem speech-to-speech models, enabling them to learn natural conversational behavior directly from real conversations without simulating backend LLM behavior.

0 favorites 0 likes
#speech-to-speech

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

arXiv cs.CL ↗ · 2026-09-28 Cached

This paper proposes a causality-aware framework for LLM-based simultaneous speech-to-speech translation, introducing a novel data pipeline and adaptive policy to improve quality-latency trade-off, achieving state-of-the-art results with reduced latency.

0 favorites 0 likes
#speech-to-speech

Gemini 3.8 Live with Live Avatar (1 minute read)

TLDR AI ↗ · 2026-09-25 Cached

Gemini 3.8 Live with Live Avatar brings real-time visual presence to conversational AI, allowing enterprises to create more natural and interactive digital experiences with dynamic avatars and multilingual support.

0 favorites 0 likes
#speech-to-speech

Gemini Live audio

Simon Willison's Blog ↗ · 2026-09-15 Cached

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, new speech-to-speech AI models, and the author built a web UI tool for voice conversations using these models.

0 favorites 0 likes
#speech-to-speech

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Google DeepMind Blog ↗ · 2026-09-15 Cached

Google introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, advanced AI models for real-time reasoning and voice agents, achieving top scores in various benchmarks.

0 favorites 0 likes
#speech-to-speech

GPT-Live-1 for phone agents - instruction following issues

Reddit r/AI_Agents ↗ · 2026-09-13

The article reports on real-world testing of OpenAI's GPT-Live-1 model for phone agents, highlighting issues with instruction adherence, alphanumeric errors, and language handling in production scenarios.

0 favorites 0 likes
#speech-to-speech

SpaceXAI launches Grok Voice Think Fast 2.0 on Agent Builder (2 minute read)

TLDR AI ↗ · 2026-07-30 Cached

xAI has launched Grok Voice Think Fast 2.0, an improved speech-to-speech model with higher transcription accuracy and lower latency, outperforming competitors on benchmarks.

0 favorites 0 likes
#speech-to-speech

@xiangxiang103: A bro from Bilibili made a cyber girlfriend, so immersive! A fully local AI girlfriend, no internet, no API key. VAD: Silero VAD v5 STT: Whisper LLM: local llama.cpp TTS: Qwen3-TTS All four models packed into 1…

X AI KOLs Timeline ↗ · 2026-07-25 Cached

Introduces a fully local AI girlfriend project made by a Bilibili developer, integrating four models: Silero VAD, Whisper, llama.cpp, and Qwen3-TTS, all packed into 15G VRAM with hot-swapping capability.

0 favorites 0 likes
#speech-to-speech

@googlegemma: Voice AI without the wait! Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the bra…

X AI KOLs Timeline ↗ · 2026-07-20 Cached

Google Gemma announces that developers can now use the Gemma 4 31B model as the brain for voice AI, enabled by Hugging Face and Cerebras for ultra-fast inference, as part of an open-source cascaded speech-to-speech stack.

0 favorites 0 likes
#speech-to-speech

@HuggingApps: NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words transcription, translation, sou…

X AI KOLs Following ↗ · 2026-07-20 Cached

NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.

0 favorites 0 likes
#speech-to-speech

@iluciddreaming: The four-layer architecture of voice assistants: VAD → Speech Recognition → LLM → Speech Synthesis. These four layers often come from different service providers. Wanting to replace one of them often means rewriting the entire pipeline. Hugging Face open-sourced speech-to-speech, modularizing these four layers so you can swap them out like components…

X AI KOLs Timeline ↗ · 2026-07-13 Cached

Hugging Face has open-sourced a modular speech-to-speech tool, which separates the four layers (VAD, speech recognition, LLM, and speech synthesis) and supports replacement, allowing users to flexibly swap components or run locally without rewriting the entire pipeline.

0 favorites 0 likes
#speech-to-speech

@Thom_Wolf: Most people should probably update their priors on the state of open-source speech-to-speech. It's honestly kind of min…

X AI KOLs Following ↗ · 2026-07-02 Cached

Thom Wolf and Cerebras released a fully open-source realtime voice demo with models and code, showcasing state-of-the-art speech-to-speech capabilities.

0 favorites 0 likes
#speech-to-speech

Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

arXiv cs.CL ↗ · 2026-07-01 Cached

This paper proposes a reference-based evaluation protocol for assessing prosody and rhythm in speech-to-speech AI systems, using matched human conversation data to provide interpretable behavioral plausibility checks.

0 favorites 0 likes
#speech-to-speech

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Hugging Face Blog ↗ · 2026-07-01 Cached

Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline combining open-source models (Nvidia's Parakeet, Gemma 4, Qwen3TTS) with Cerebras' fast inference, enabling natural conversational AI and powering robots like Reachy Mini.

0 favorites 0 likes
#speech-to-speech

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

arXiv cs.CL ↗ · 2026-06-25 Cached

This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.

0 favorites 0 likes
#speech-to-speech

Gemma 4 12B native encoder free voice input utilization suggest?

Reddit r/LocalLLaMA ↗ · 2026-06-14

Discusses leveraging Gemma 4 12B's encoder-free architecture for native voice input, seeking out-of-the-box solutions for low-latency streaming audio ingestion.

0 favorites 0 likes
#speech-to-speech

Gemini 3.5 Live Translate

Product Hunt ↗ · 2026-06-09

Gemini 3.5 Live Translate is a new audio model for real-time speech-to-speech translation.

0 favorites 0 likes
#speech-to-speech

@GoogleDeepMind: 3.5 Live Translate can convert speech into over 70 languages and processes it as it’s streamed - while keeping tone, pa…

X AI KOLs ↗ · 2026-06-09 Cached

Google DeepMind announces Live Translate, a feature that converts speech into over 70 languages in real-time while preserving tone, pace, and pitch for more natural conversations.

0 favorites 0 likes
#speech-to-speech

Fluid, natural voice translation with Gemini 3.5 Live Translate

Google DeepMind Blog ↗ · 2026-06-09 Cached

Google releases Gemini 3.5 Live Translate, an audio model for near real-time speech-to-speech translation in over 70 languages, preserving speaker intonation and pacing. It is rolling out across Google products including the Gemini Live API, Google Meet, and Google Translate.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback