Tag
This survey paper evaluates real-time voice agents by proposing a taxonomy and a minimum reporting standard called TRG to address fragmentation across speech modeling, turn-taking, and agentic evaluation communities.
Sayble is an AI copilot that provides real-time suggestions during sales calls and meetings, along with automatic recap and email generation. It supports multiple platforms and languages for enhanced communication productivity.
An open-source voice AI agent with a live avatar can see, talk, think, and draw in real-time, and supports multilingual interactions, such as switching to Hindi mid-call.
A free web app that pairs animated pixel-art city nights with endless lofi music generated live in the browser using AI, offering a calming background for work or relaxation.
Google introduces Gemini 3.8 Live with Live Avatar, an update that integrates real-time visual avatars into conversational AI, enhancing interactions with features like lip-syncing, multilingual support, and asynchronous tool execution for enterprise applications.
OdysseyML introduces Agora-2, a next-generation multi-agent world model that supports up to 20 humans and agents interacting in real-time within a shared simulated environment, with a multiplayer research preview now available.
Nagi is a new open-source fast decision model that consistently outperforms other models like laya and semif in real-time gaming scenarios.
NVIDIA releases Nemotron 3 Diarization, an open-weight 100M-parameter model that achieves state-of-the-art speaker diarization with a 14.72% error rate, supporting real-time and offline processing for up to eight speakers.
Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.
OpenTrainDNN is an open-source, client-side web application that provides real-time visualization of deep neural network training, including backpropagation and weight updates, directly in the browser.
A prompt using Claude Opus 5.5 generates an interactive 3D scene of a pelican riding a bicycle, featuring physics simulations, achievements, and browser-based real-time sound effects.
NetEase Youdao's open-source AI models R2T2 and T3PO have topped Hugging Face leaderboards for speech recognition and translation, outperforming major competitors with impressive real-time performance and stability.
Speechka is a real-time voice translation tool that mimics the user's voice across 44 languages, available on macOS, Windows, and browsers.
PixVerse R2 introduces a unified scaling architecture for real-time audiovisual world models, leveraging block-sparse attention and continuous pretraining to enhance video generation and interactive control.
InstinctFlash is a high-performance serving framework for robotics models that enables real-time inference of 5B world-action models on Jetson Thor, with reported speedups of up to 33.78×.
Yohei Nakajima tweeted about a technology fast enough for live emotion detection, emphasizing its real-time performance.
The post compares AI models Jev and Laya, highlighting Laya's open-source, local, and low-latency capabilities for real-time decision-making in games.
This article discusses methods to prevent the babbling-idiot failure in time-triggered communication systems, enhancing reliability and safety.
Marcus Ng and Jamie Cai won $10k and became semi-finalists at HackTheNorth by building Adlib, a tool that creates flow charts, diagrams, and graphics in real-time from spoken words to replace traditional slides.
R2T2 is a low-latency and high-accuracy real-time speech recognition model that processes audio in small chunks and commits text without revision, suitable for applications like live captioning and translation.