@AISuperDomain: Real-time voice agents are moving from "demo toys" to truly usable, and LiveKit Agents is arguably one of the most complete open-source frameworks. It doesn't just chain STT, LLM, and TTS; it also directly provides: • WebRTC real-time audio/video • Inbound and outbound phone calls • …
Summary
LiveKit Agents is an open-source real-time voice agent framework supporting WebRTC, telephony integration, semantic turn detection, MCP tool calling, and multi-agent handoff, helping developers quickly build real-time voice applications such as AI customer service and phone bots.
View Cached Full Text
Cached at: 08/04/26, 08:15 PM
🎙️ Starter Agent
A starter agent optimized for voice conversations.
Code
🔄 Multi-user push to talk
Responds to multiple users in the room via push-to-talk.
Code
🎵 Background audio
Background ambient and thinking audio to improve realism.
Code
🛠️ Dynamic tool creation
Creating function tools dynamically.
Code
☎️ Outbound caller
Agent that makes outbound phone calls
Code
📋 Structured output
Using structured output from LLM to guide TTS tone.
Code
🔌 MCP support
Use tools from MCP servers
Code
💬 Text-only agent
Skip voice altogether and use the same code for text-only integrations
Code
📝 Multi-user transcriber
Produce transcriptions from all users in the room
Code
🎥 Video avatars
Add an AI avatar with Tavus, Bithuman, LemonSlice, and more
Code
🍽️ Restaurant ordering and reservations
Full example of an agent that handles calls for a restaurant.
Code
👁️ Gemini Live vision
Full example (including iOS app) of Gemini Live agent that can see.
Code
Similar Articles
livekit/agents
LiveKit Agents is an open-source framework for building realtime, multimodal voice agents that can see, hear, and understand, with flexible STT/LLM/TTS integrations, job scheduling, telephony support, MCP compatibility, and a built-in test framework.
@mylifcc: Voice agents are exploding, but most are still black boxes in production. Yesterday (July 21), LangSmith officially launched Python-side voice tracing support, covering the 4 most mainstream frameworks: Pipecat, LiveKit, OpenAI Real-time...
LangSmith officially launched Python-side voice tracing support, covering four mainstream frameworks: Pipecat, LiveKit, OpenAI Realtime, and Gemini Live. It brings voice conversations into the same observable, evaluable workflow as text agents, solving the pain point of poor debuggability in voice agents.
@DanKornas: Real-time voice agents need more than an LLM call—they need transport, speech components, turn handling, and a path to …
TEN is a framework for building real-time multimodal conversational AI agents, offering configurable STT, LLM, and TTS components, a visual designer, and deployment options including self-hosting and split deployment.
@DanKornas: Building a live voice agent requires coordinating audio streaming, turn detection, interruptions, model calls, and medi…
VideoSDK AI Agents is an open-source Python framework for building production-ready real-time voice and multimodal AI agents that join VideoSDK rooms as participants, with unified pipeline configuration and multiple execution modes.
@seclink: Recently this open-source tool has been quite popular. It looks like an open-source version of DingTalk Wukong and ByteDance Aily. You can use it to implement your own agent and integrate it into the aforementioned instant messaging platforms. Some guys tweaked it and used it to demo to investors, obtaining a considerable valuation. What makes investors remember...
CowAgent is an open-source AI assistant framework based on large language models. It supports autonomous task planning, long-term memory, knowledge base, multi-model switching, and multi-channel access (WeChat, Feishu, DingTalk, etc.), enabling rapid construction and deployment of personalized AI agents.