Tag
MIT researchers developed VLASH, a method enabling vision-language-action models to predict future robot states, doubling speed and reducing lag in tasks like pick-and-place and table tennis.
SmolVLA is a compact vision-language-action model that achieves competitive robotic control performance at reduced computational cost, enabling deployment on consumer-grade hardware. It introduces asynchronous inference and leverages community-collected datasets.