Tag
YouTube announced new features including custom feeds and enhanced AI capabilities like Ask YouTube and Ask Music at the Made on YouTube event, set to roll out later this year.
Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.
NetEase Youdao has released two AI models: a streaming ASR model with 2B parameters and a Chinese-English simultaneous translation model with 14B parameters, both emphasizing real-time performance.
This article identifies and provides fixes for three bugs encountered when serving the MiMo-V2.6-Flash model with vLLM, including empty responses in streaming mode, reasoning loss in tool loops, and a hidden token output cap.
A developer built a small gateway to simplify multi-model agent workflows by handling streaming, tool calls, and provider-specific integration challenges across models like Claude and GPT.
This paper introduces SVEET, a framework for high-quality streaming video editing that leverages a pretrained video diffusion model to enable auto-regressive editing with real-time performance on a single GPU.
R2T2 is a low-latency and high-accuracy real-time speech recognition model that processes audio in small chunks and commits text without revision, suitable for applications like live captioning and translation.
The article discusses how streaming services and FAST channels are mimicking traditional cable models, reflecting cyclical trends in entertainment consumption and industry adaptations.
Max Verstappen, a four-time Formula 1 world champion, will attempt to race against 100 identical karts in a 30-lap event at Silverstone, streamed by Red Bull, Disney+, and ESPN.
The paper introduces Zing-0.5, a 5B parameter autoregressive world model for generating playable worlds with real-time user interaction through combined keyboard and text controls. It achieves high performance in navigation tasks and demonstrates low-cost real-time inference at 24 FPS.
Netflix announces three new Sega game adaptations: a Crazy Taxi film, a Sonic animated series for kids with attitude, and a live-action Stranger Than Heaven movie.
Omni-Streaming Thinking improves streaming omni-modal reasoning by deferring claims until cross-modal verification, reducing premature commitment and auditory hallucinations.
GPT6 Astra, an AI model, is playing Factorio Space Age and has reached the third planet, outperforming the previous model Fable 5.1. A stream is available, and it is expected to complete the game soon.
Confucius4-R2T2 is a low-latency, high-accuracy real-time speech recognition model developed by NetEase Youdao, featuring configurable chunking and stable output for applications like live captioning.
StreamAlign is a streaming text-aligned speech tokenization framework that enables real-time speech–text joint modeling, reducing latency and achieving state-of-the-art results on speech recognition and spoken language modeling tasks.
Prime Video has released the full trailer for Mike Flanagan's six-episode miniseries adaptation of Stephen King's Carrie, which aims to reinvent the story for the modern world.
NVIDIA announced expansions to NVIDIA AI for Media at IBC 2026, including advanced Synthetic Video Detector and 3D Body Pose technologies to enhance AI applications in broadcast, sports, and streaming.
Amazon Prime Video has introduced an AI-powered feature that synchronizes actors' lip movements with dubbed audio to enhance dubbing quality, initially available for the series Maxton Hall with plans for expansion.
Microsoft is imposing new monthly time limits on game streaming for Xbox Game Pass subscribers, with different tiers offering 5 to 15 hours per month, and additional streaming time available for purchase.
Microsoft has released VibeVoice-ASR-Streaming, a unified streaming ASR model that transcribes who said what with support for customized hotwords and 10 languages.