@DanKornas: Real-time voice agents need more than an LLM call—they need transport, speech components, turn handling, and a path to …

X AI KOLs Timeline Tools

Summary

TEN is a framework for building real-time multimodal conversational AI agents, offering configurable STT, LLM, and TTS components, a visual designer, and deployment options including self-hosting and split deployment.

Real-time voice agents need more than an LLM call—they need transport, speech components, turn handling, and a path to deployment. TEN is a framework for real-time multimodal conversational AI for developers building voice and interactive agent applications. It helps you assemble and customize agents through extension-based STT, LLM, and TTS components, a visual TMAN Designer, and runnable examples. Key features: • Real-time voice assistant – supports both RTC and WebSocket connections. • Configurable agent components – edit STT, LLM, and TTS properties in TMAN Designer or property.json. • Practical examples – explore diarization, transcription, SIP calls, lip-sync avatars, and prompt-driven doodling. • Self-hosting path – build and run an example as a Docker image. • Split deployment option – host the backend on a container platform and the frontend on Vercel or Netlify. The root framework uses Apache 2.0 with additional restrictions; components in packages/ are released under Apache 2.0. Link in the reply
Original Article
View Cached Full Text

Cached at: 07/31/26, 10:52 AM

Real-time voice agents need more than an LLM call—they need transport, speech components, turn handling, and a path to deployment.

TEN is a framework for real-time multimodal conversational AI for developers building voice and interactive agent applications.

It helps you assemble and customize agents through extension-based STT, LLM, and TTS components, a visual TMAN Designer, and runnable examples.

Key features: • Real-time voice assistant – supports both RTC and WebSocket connections. • Configurable agent components – edit STT, LLM, and TTS properties in TMAN Designer or property.json. • Practical examples – explore diarization, transcription, SIP calls, lip-sync avatars, and prompt-driven doodling. • Self-hosting path – build and run an example as a Docker image. • Split deployment option – host the backend on a container platform and the frontend on Vercel or Netlify.

The root framework uses Apache 2.0 with additional restrictions; components in packages/ are released under Apache 2.0.

Link in the reply

Similar Articles

livekit/agents

GitHub Trending (daily)

LiveKit Agents is an open-source framework for building realtime, multimodal voice agents that can see, hear, and understand, with flexible STT/LLM/TTS integrations, job scheduling, telephony support, MCP compatibility, and a built-in test framework.

@AISuperDomain: Real-time voice agents are moving from "demo toys" to truly usable, and LiveKit Agents is arguably one of the most complete open-source frameworks. It doesn't just chain STT, LLM, and TTS; it also directly provides: • WebRTC real-time audio/video • Inbound and outbound phone calls • …

X AI KOLs Timeline

LiveKit Agents is an open-source real-time voice agent framework supporting WebRTC, telephony integration, semantic turn detection, MCP tool calling, and multi-agent handoff, helping developers quickly build real-time voice applications such as AI customer service and phone bots.

Your voice agent probably isn't slow because of the LLM.

Reddit r/AI_Agents

A developer debunks the common belief that LLM latency is the primary cause of slow voice agents, explaining that delays often stem from earlier stages like audio capture, VAD, and STT. They recommend logging specific latency metrics and testing various STT/TTS providers and orchestration frameworks to diagnose issues.

Building an AI voice agent from scratch: the parts that actually took our time

Reddit r/artificial

The author shares a postmortem on building a production phone-based AI voice agent, revealing that most engineering time was consumed by telephony infrastructure, turn detection, observability, and failure handling rather than core LLM behavior. They suggest using managed platforms like Vapi, Retell, or Dasha from the start to focus engineering effort on business logic.