Tag
A robotic hand is trained using reinforcement learning to perform self-supported locomotion on its fingertips, along with tasks like fall recovery, keyboard pressing, and object pushing, as presented in a research paper from ETH Zurich.
Gen-HumanEgo is a new dataset released by GenrobotAI for robot learning, featuring over 1,800 hours of egocentric human data with 10K+ unique tasks and synchronized camera views including 3D hand keypoints and structured annotations.
DeltaWAM introduces delta-based world-action models for bimanual manipulation, enhancing efficiency and performance by predicting visual changes and actions. It demonstrates improved success rates and reduced computational overhead.
This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.
This survey paper reviews methods for embedding physics priors in robot learning, providing a unified taxonomy and discussing open challenges and future research directions in the field.
This paper proposes Time–Frequency Geometric Cross-Attention (TFGCA), a drop-in module for chunked vision-language-action models that improves action trajectory prediction by decomposing chunks into time-frequency representations and capturing geometric relationships, resulting in significant performance gains on benchmarks and real-robot tasks.
This study rigorously tests whether scaling web-video pre-training improves robot task performance, finding that larger models and more compute consistently yield better policies on real industrial tasks, with pre-training quality correlating with performance.
WuJi has open-sourced its MINT model and EgoPipeline for reconstructing 3D hand and camera motion from first-person video, along with 1,021 hours of egocentric data to aid robot learning.
Introduces Latent Energy Action Planning (LEAP), which optimizes action sequences through a frozen world model to improve goal-conditioned control, achieving a 17.3 percentage-point improvement in success rate across four control domains.
Scaling video pre-training to 120K hours boosts zero-shot success in World Action Models from 36.1% to 77.8% on real robots, enabling faster action prediction.
This paper compares the temporal robustness of expert and imitation-learned policies in dexterous manipulation tasks, finding that imitation-learned policies degrade more sharply with increased execution speeds, primarily due to insertion misalignments.
Introduces worldproof, an open-source tool for diagnosing world models, and finds that pixel metrics often cannot rank models on real robot video because evaluation horizons lack discriminative power; a usable window of 8-24 steps is identified.
Call for papers for the Scaling H2R workshop at CoRL 2026, focusing on scaling laws and diversity in human-to-robot learning. Submissions due Oct 7, 2026.
Light Origins announces a humanoid robot powered by a single perceptive whole-body policy that autonomously decides whether to walk, climb, or vault over unseen terrain, without skill labels, state machines, or motion generators at inference.
A survey paper organizing robot-learning techniques along an axis of frozen-weight policies (VLA models) versus agents that write their own executable skills as code, providing a taxonomy of self-improvement mechanisms and analyzing the emerging robot-skill economy.
Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.
Sergey Levine highlights the importance of learning from suboptimal robot data and promotes OopsieData, a community effort to collect and open-source the abundant but underused failed robot interaction clips.
HiFi-UMI introduces a portable data-production system for robot-free UMI data that achieves high trajectory accuracy using stereo-inertial SLAM and wide-angle cameras. Training manipulation policies on this data alone enables zero-shot deployment on real robots, matching or exceeding teleoperation baselines across several model families, and the authors open-source a 2,000-hour high-fidelity dataset.
This paper organizes embodied data sources into a five-level pyramid (real-robot, UMI, egocentric/exocentric, simulation, general vision-language), analyzing their trade-offs between scalability and robot alignment, and reviews recent embodied foundation models in terms of data recipes. It also discusses open challenges for building next-generation embodied systems.
The article discusses the potential benefits and challenges of using first-person video for robot learning, highlighting that while direct imitation is limited, the sequence of visual attention may transfer. It references LingBot-VLA 2.0 and calls for controlled evaluations to separate viewpoint effects from data volume.