Tag
A tweet highlighting that benchmarking robot policies is broken and sharing results from thousands of evaluations over 12 manipulation tasks to determine which policy to use.
HiFi-UMI introduces a portable data-production system for robot-free UMI data that achieves high trajectory accuracy using stereo-inertial SLAM and wide-angle cameras. Training manipulation policies on this data alone enables zero-shot deployment on real robots, matching or exceeding teleoperation baselines across several model families, and the authors open-source a 2,000-hour high-fidelity dataset.
MiniCPM launches a new series of robot models: MiniCPM-RobotManip and MiniCPM-RobotTrack, focusing on manipulation and tracking tasks respectively.
Xiaomi Robotics-1, a robot foundation model trained on 100,000 hours of real-world manipulation data, has been released on Hugging Face. The model can autonomously perform household tasks like folding laundry, loading a washer, and washing dishes.
RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.
Xiaomi introduces Xiaomi-Robotics-1, a vision-language-action foundation model trained on over 100,000 hours of real-world manipulation trajectories, demonstrating clear scaling laws and achieving high success rates on real-world tasks with minimal fine-tuning data.
Introduces Learning from Hindsight (LfH), a method that applies hindsight relabeling to RL post-training of vision-language-action models. By relabeling failed robot rollouts with the tasks they actually achieved, LfH achieves 5x improvement in sample efficiency on out-of-distribution manipulation tasks.
1X has developed new 25-DoF robotic hands for its NEO humanoid platform, achieving human-level dexterity, tactile sensing, and durability for real-world manipulation tasks.
Evaluation of 30+ embodied AI models finds that current generalist robot policies lack robustness for real-world manipulation, leading to the creation of RoboDojo.
AI developed at UMA_Robots demonstrates picking and scanning deformable, slippery, thin, and fragile items.
RoboDojo is a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies, featuring 42 simulation tasks and 18 real-world tasks across multiple evaluation dimensions.
OmniTacTune introduces a two-stage reinforcement learning pipeline for adapting tactile feedback to pretrained visual robot policies, achieving 85-100% success on contact-rich manipulation tasks within 40-80 minutes.
This paper introduces AgentCanvas, a typed-graph runtime for embodied agents, and KDLoop, a coding-agent search procedure, to automate the design of embodied agent architectures, evaluating across multiple embodied tasks and revealing challenges like rollout noise and local edit basins.
Discusses the potential of AI-generated personalized media populating social media feeds without consent, raising concerns about manipulation and attention economy.
ASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving significant improvements in manipulation and household tasks while enabling sim-to-real transfer.
SceneBot is a unified motion tracking framework that conditions a single humanoid policy on both reference motions and per-link contact labels, enabling free-space locomotion, terrain traversal, and whole-body manipulation in contact-rich environments.
SimFoundry is a modular system that automates real-to-sim scene construction from video, generating digital twins and affordance-preserving variations for zero-shot robot policy training, achieving strong transfer to real-world tasks and high simulation-to-real performance prediction.
NVIDIA's new chips enable running 500B parameter models locally, highlighting that AI safety measures are merely behavioral speed bumps that vanish offline, posing unprecedented risks for deception and manipulation at scale.
InSight presents a framework for autonomous skill acquisition in vision-language-action (VLA) models by enabling steerability at the primitive-action level and using a VLM-guided data flywheel to generate demonstrations, achieving manipulation tasks like block flipping and pouring without human demonstrations.
The paper presents World Value Model (WVM), a generalist robotic value model that combines world models with value estimation to accurately assess task progression and improve robotic policy learning from mixed-quality data, achieving state-of-the-art results on standard benchmarks and a new suboptimal data benchmark.