Inertia-1: An Open Exploration to a Unified Motion Foundation Model
Summary
Inertia-1 is a research project that systematically explores the full lifecycle of motion models—data, sensing, objectives, and scale—to produce a unified representation that transfers across body placements, devices, and tasks without retraining, leveraging self-supervised pretraining on 18 million hours of accelerometry data.
View Cached Full Text
Cached at: 07/20/26, 03:31 PM
Similar Articles
Inertia-1: An Open Exploration of Wearable Motion Foundation Models
Inertia-1 is a fully open exploration of wearable motion foundation models using over 18.2 million hours of accelerometer data, systematically studying data, model, and training choices across 15 datasets for tasks like activity recognition and disease prediction.
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Embodied-R1.5 is a unified embodied foundation model that achieves state-of-the-art performance on 16 out of 24 embodied vision-language benchmarks using multi-task balanced reinforcement learning. It introduces a Planner-Grounder-Corrector closed-loop framework for long-horizon tasks and is open-sourced to facilitate future research.
Motion4Motion: Motion Transfer Across Subjects at Inference
Motion4Motion is a training-free motion transfer framework that models motion flow from a video instead of relying on skeleton structures, enabling motion transfer across diverse species (e.g., humans to animals) at inference time.
Minimalist Visual Inertial Odometry
A minimalist visual-inertial odometry approach uses four photodiodes with optical Gabor masks and a temporal convolutional network to achieve accurate planar motion estimation for differential-drive robots, validated across diverse indoor and outdoor terrains without real-world fine-tuning.
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
AnyMo is a geometry-aware framework for setup-agnostic human motion modeling using physics-grounded IMU simulation and graph encoding, achieving significant improvements in zero-shot activity recognition, cross-modal retrieval, and motion captioning across multiple datasets.