Real-Time Omni-Modal Interaction Driven Whole-Body Mobile Manipulation
Summary
Unitree releases UnifoLM-OminiA-0.3, a single AI model for whole-body mobile manipulation with real-time omni-modal interaction, enabling autonomous home-care and wellness tasks.
View Cached Full Text
Cached at: 07/20/26, 05:32 PM
Similar Articles
Unitree’s UnifoLM-X2-1.0 reads the opponent, predicts the next move and reacts instantly. For the first time a world model controls a fully autonomous humanoid fight in real time.
UnifoLM-X2-1.0 is a world model that enables real-time autonomous combat for humanoid robots, breaking through bottlenecks in planning and decision-making and validating the feasibility of large-scale deployment.
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
ABot-M0.5 is a new World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency, achieving state-of-the-art results on long-horizon and fine-grained manipulation benchmarks.
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
This technical report introduces X-OmniClaw, a unified mobile agent system designed for multimodal understanding and interaction on Android devices. It details the architecture for perception, memory management, and action execution using on-device AI capabilities.
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
HiFi-UMI introduces a portable data-production system for robot-free UMI data that achieves high trajectory accuracy using stereo-inertial SLAM and wide-angle cameras. Training manipulation policies on this data alone enables zero-shot deployment on real robots, matching or exceeding teleoperation baselines across several model families, and the authors open-source a 2,000-hour high-fidelity dataset.
@rohanpaul_ai: Just a few days back, Thinking Machines Lab (TML), showcased a way of making AI interaction continuous instead of turn-…
Thinking Machines Lab and OpenBMB released MiniCPM-o 4.5, a 9B full-duplex omnimodal model with the Omni-Flow framework that enables continuous, time-aligned real-time video and voice interaction, surpassing previous models and available as open source.