Tag
SeededGrasp proposes a data-efficient framework that uses a vision-language model to predict a seed point for a lightweight grasp generator, enabling language-guided grasping in complex scenes with multiple robot embodiments. The method outperforms baselines with 72% simulation and 78% real-world success, and includes a new large-scale multi-embodiment grasping dataset.
This paper introduces UESF-Bench, a large-scale benchmark for unified embodied human seeking and following, and proposes SeekFollow-VLA, a vision-language-action framework that handles semantic-guided exploration and reliable behavior switching between searching and following.