real-time-interaction

Tag

Cards List
#real-time-interaction

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

Hugging Face Daily Papers ↗ · 2026-09-15 Cached

The paper introduces Zing-0.5, a 5B parameter autoregressive world model for generating playable worlds with real-time user interaction through combined keyboard and text controls. It achieves high performance in navigation tasks and demonstrates low-cost real-time inference at 24 FPS.

0 favorites 0 likes
#real-time-interaction

Omni Interaction Agent Technical Report

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper presents Gander, an end-to-end framework for omni interaction and agentic tasks, enabling real-time full-duplex interaction across multiple modalities with a Cerebellum-Brain architecture.

0 favorites 0 likes
#real-time-interaction

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

VoiceMem introduces a dual-brain streaming memory architecture for speech language models that improves retrieval accuracy, emotional personalization, and real-time efficiency.

0 favorites 0 likes
#real-time-interaction

MOSS-VL Technical Report

Hugging Face Daily Papers ↗ · 2026-08-15 Cached

MOSS-VL is an open vision-language model family designed for real-time interaction, using gated cross-attention to process vision during generation and achieving top performance in streaming benchmarks among open-source models.

0 favorites 0 likes
#real-time-interaction

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

Hugging Face Daily Papers ↗ · 2026-07-22 Cached

The paper proposes Super Star, a real-time framework for online co-speech gesture generation in digital humans using a causal multimodal autoregressive model with streaming speech and user feedback for continual adaptation.

0 favorites 0 likes
#real-time-interaction

robbyant/lingbot-world-v2-14b-causal-fast

Hugging Face Models Trending ↗ · 2026-07-08 Cached

LingBot-World 2.0 is an advanced world model achieving unbounded interaction horizons, real-time 720p/60fps video streaming, diverse interactive elements, and an agentic harness integrating pilot and director agents. The model is released on Hugging Face with inference code and technical report.

0 favorites 0 likes
#real-time-interaction

AlayaWorld: Long-Horizon and Playable Video World Generation

Hugging Face Daily Papers ↗ · 2026-07-07 Cached

AlayaWorld is an open-source framework for building interactive generative worlds that enables real-time user interaction and supports diverse actions. It unifies the complete development pipeline from data preparation to deployment.

0 favorites 0 likes
#real-time-interaction

Inside Thinking Machines' Interaction Models (17 minute read)

TLDR AI ↗ · 2026-07-01 Cached

New research from Thinking Machines critiques current single-threaded AI interaction models, arguing that they limit human-AI collaboration by forcing humans into clean input-output cycles. The lab proposes a new interaction model that supports continuous, multi-modal collaboration akin to real-time human conversation.

0 favorites 0 likes
#real-time-interaction

StepAudio 2.5 Technical Report

Hugging Face Daily Papers ↗ · 2026-05-22 Cached

StepAudio 2.5 is a unified audio-language model that achieves state-of-the-art results across ASR, TTS, and real-time spoken interaction by leveraging task-tailored reinforcement learning from human feedback to optimize shared representations.

0 favorites 0 likes
#real-time-interaction

How far from "Her"

Reddit r/ArtificialInteligence ↗ · 2026-05-14

Reflecting on the 2013 film 'Her', this article examines how close current AI technology is to replicating the film's autonomous, real-time-interpreting AI, concluding that while progress has been made, full consciousness remains elusive.

0 favorites 0 likes
#real-time-interaction

Interaction Models

Hacker News Top ↗ · 2026-05-11 Cached

Thinking Machines AI announces a research preview of interaction models, a new architecture designed for native, real-time human-AI collaboration across audio, video, and text. By replacing turn-based interfaces with a multi-stream, micro-turn design, the model aims to keep humans actively in the loop while delivering state-of-the-art intelligence and responsiveness.

0 favorites 0 likes
#real-time-interaction

@miramurati: Today we're sharing our work on interaction models. A new class of model trained from scratch to handle real-time inter…

X AI KOLs Following ↗ · 2026-05-11 Cached

Mira Murati's team showcased a preview of the new interaction model. Trained from scratch, it natively supports full-duplex real-time audio and video conversations, instant interruptions, multi-language translation, and dynamic multi-tasking. The demonstration verified its core capabilities in low-latency streaming interaction, multimodal perception, and concurrent task execution.

1 favorites 1 likes
#real-time-interaction

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Hugging Face Daily Papers ↗ · 2026-04-30 Cached

MiniCPM-o 4.5 is a 9B parameter multimodal model featuring Omni-Flow, a framework enabling real-time full-duplex interaction where the model can simultaneously perceive and respond proactively. It achieves state-of-the-art open-source performance comparable to Gemini 2.5 Flash and runs on edge devices with less than 12GB RAM.

0 favorites 0 likes
#real-time-interaction

Hello GPT-4o

OpenAI Blog ↗ · 2024-05-13 Cached

OpenAI announces GPT-4o, a flagship multimodal model that processes audio, vision, text, and video in real-time with 232ms average audio response latency. The model matches GPT-4 Turbo on text/code while significantly improving multilingual, audio, and vision capabilities at 50% cheaper API costs.

0 favorites 0 likes
#real-time-interaction

Image interaction with GPT-Live

YouTube AI Channels ↗ · 2026-07-09 Cached

The user shows two outfits in real-time via GPT-Live, and the AI gives specific clothing suggestions for the scenario of meeting parents, demonstrating multi-round image understanding and contextual recommendation capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback