Tag
MobileVLA-R1 2.0 is an RL-enhanced vision-language-action framework that couples structured reasoning with mobile robot control, achieving improvements in long-horizon instruction following on real-world platforms.
An evaluation of 9 vision-language models on an expert mechanic's brake-repair exam shows they know the theory but fail in noticing and adapting during real-world deployment.
Google DeepMind released Gemini Robotics 2, an AI model that controls entire humanoid robots with whole-body motions and improved dexterity, along with updates to embodied reasoning and on-device models for multi-robot coordination and safety.
DeepMind introduces Gemini Robotics 2, a set of AI models (VLA, ER, On-Device) that enable robots to perform whole-body control, fine dexterity, and multi-robot collaboration, with local on-device execution and fast adaptation to new robot bodies.
SleepWalk is a three-tier benchmark for evaluating vision-language models' ability to predict spatially coherent trajectories in 3D environments from textual instructions and visual observations, revealing systematic failures in grounded spatial reasoning under occlusions and multi-step instructions.