Tag
SeededGrasp proposes a data-efficient framework that uses a vision-language model to predict a seed point for a lightweight grasp generator, enabling language-guided grasping in complex scenes with multiple robot embodiments. The method outperforms baselines with 72% simulation and 78% real-world success, and includes a new large-scale multi-embodiment grasping dataset.
CHORUS is a decentralized method enabling multiple robots with different embodiments to collaborate using a single Vision-Language-Action policy.