Tag
MOSS-VL is an open vision-language model family designed for real-time interaction, using gated cross-attention to process vision during generation and achieving top performance in streaming benchmarks among open-source models.
LingBot-World 2.0 is an advanced world model achieving unbounded interaction horizons, real-time 720p/60fps video streaming, diverse interactive elements, and an agentic harness integrating pilot and director agents. The model is released on Hugging Face with inference code and technical report.
AlayaWorld is an open-source framework for building interactive generative worlds that enables real-time user interaction and supports diverse actions. It unifies the complete development pipeline from data preparation to deployment.
New research from Thinking Machines critiques current single-threaded AI interaction models, arguing that they limit human-AI collaboration by forcing humans into clean input-output cycles. The lab proposes a new interaction model that supports continuous, multi-modal collaboration akin to real-time human conversation.
StepAudio 2.5 is a unified audio-language model that achieves state-of-the-art results across ASR, TTS, and real-time spoken interaction by leveraging task-tailored reinforcement learning from human feedback to optimize shared representations.
Reflecting on the 2013 film 'Her', this article examines how close current AI technology is to replicating the film's autonomous, real-time-interpreting AI, concluding that while progress has been made, full consciousness remains elusive.
Thinking Machines AI announces a research preview of interaction models, a new architecture designed for native, real-time human-AI collaboration across audio, video, and text. By replacing turn-based interfaces with a multi-stream, micro-turn design, the model aims to keep humans actively in the loop while delivering state-of-the-art intelligence and responsiveness.
Mira Murati's team showcased a preview of the new interaction model. Trained from scratch, it natively supports full-duplex real-time audio and video conversations, instant interruptions, multi-language translation, and dynamic multi-tasking. The demonstration verified its core capabilities in low-latency streaming interaction, multimodal perception, and concurrent task execution.
MiniCPM-o 4.5 is a 9B parameter multimodal model featuring Omni-Flow, a framework enabling real-time full-duplex interaction where the model can simultaneously perceive and respond proactively. It achieves state-of-the-art open-source performance comparable to Gemini 2.5 Flash and runs on edge devices with less than 12GB RAM.
OpenAI announces GPT-4o, a flagship multimodal model that processes audio, vision, text, and video in real-time with 232ms average audio response latency. The model matches GPT-4 Turbo on text/code while significantly improving multilingual, audio, and vision capabilities at 50% cheaper API costs.
The user shows two outfits in real-time via GPT-Live, and the AI gives specific clothing suggestions for the scenario of meeting parents, demonstrating multi-round image understanding and contextual recommendation capabilities.