Tag
ByteDance is developing a real-time spatial video model under founder Zhang Yiming's oversight, aiming to generate interactive virtual worlds for XR applications with a potential launch soon.
ZipDepth is a lightweight zero-shot monocular depth estimation model that achieves the best accuracy-efficiency trade-off, running in real time on any device from mobile phones to server GPUs, and has been accepted at ECCV 2026.
IBM and Confluent have partnered to offer time series foundation models for real-time intelligence on streaming data, now available in early access on Confluent Cloud.
An endless AI livestream called 'Infinite Slop' allows viewers to interact via chat, with AI generating the next video segment to form a continuous story, using MiniMax H3 and fal.ai.
Meta introduces Muse Voice Transcribe, a real-time audio perception model that excels in streaming ASR and diarization with multilingual support, topping public benchmarks.
Runway introduces Solaris, an Interface World Model that generates interactive interfaces frame by frame in real time without code, outperforming frontier LLMs.
Runway has released Solaris, its first Interface World Model, which can generate dynamic software interfaces in real-time based on user interactions, without the need for pre-set code, marking a fundamental shift in software operation.
Manzanas is a tool that allows controlling multiple iOS simulators across MacBooks in real time for faster app smoke testing, with optimizations for RAM usage and action execution speed.
Intel Mapper is a real-time interactive conflict map that uses AI to track and verify geolocated OSINT events in Ukraine and Syria, featuring territorial control and military flight tracking.
Ink-2 ranked #2 on the new VoiceCodeBench benchmark, demonstrating its suitability for real-time consumer apps, with GPT Live Transcribe taking the #1 spot.
Tavus introduces Sparrow-2, a state-of-the-art real-time conversational understanding AI model that helps voice AI systems decide when to listen, wait, speak, or keep speaking during conversations.
The world's first brain surgery with real-time AI assistance successfully removed a tumour, allowing surgeons to avoid vital vessels and nerves and enhance patient safety.
ABot-Recon enables real-time 3D reconstruction of large-scale environments from continuous video streams by using a fixed local context, achieving high efficiency and low memory usage.
EditaLive is a novel framework for real-time human-centric live-stream video editing that adapts image animation models using causal generation and distillation for efficient streaming inference.
Google introduces Gemini 3.5 Transcribe, a new AI model for precise and intelligent real-time speech-to-text transcription, available via APIs for developers.
This article argues that comparisons between WebSockets and Server-Sent Events should focus on event ordering and correctness to avoid inconsistent user interfaces, rather than just latency or simplicity.
S1 is a new foundation model that learns from a single example and can perform 10-minute tasks from video prompts without fine-tuning, showcased in real-time operation.
Breeze TTS 2 is an open-weight, real-time text-to-speech model that supports voice design and bilingual speech, ranking first among open-weight models on the Artificial Analysis TTS leaderboard.
A live translation application powered by Gemini Live API and LiveKit, enabling real-time broadcast translation for events with code and setup instructions provided.
PicoMQ is a system for durable real-time streams over HTTP, built on S3-compatible object storage, allowing for independent and scalable streams per use case.