Tag
NVIDIA Research introduces Hydra-0, a generalist world model that conditions on action flow to represent robot actions as pixel motion, enabling learning across diverse embodiments with significant error reductions and zero-shot capabilities.
SeededGrasp proposes a data-efficient framework that uses a vision-language model to predict a seed point for a lightweight grasp generator, enabling language-guided grasping in complex scenes with multiple robot embodiments. The method outperforms baselines with 72% simulation and 78% real-world success, and includes a new large-scale multi-embodiment grasping dataset.
CHORUS is a decentralized method enabling multiple robots with different embodiments to collaborate using a single Vision-Language-Action policy.