Tag
Zero-WAM is a causal video-action model that enables zero-shot robotic manipulation of unseen tasks by conditioning on in-context human video guidance, with the HumanGen dataset and a future-chunk prediction objective to improve generalization.
LingBot-VA 2.0 is a video-action foundation model trained from scratch for robot control, achieving 225 Hz closed-loop execution with 13B parameters (1.9B active per token) and outperforming prior models on RoboTwin 2.0.
A robot using the LingBot-VA 2.0 video-action model picks objects off a moving conveyor belt in real-time at 1x speed, predicting future movements rather than reacting only to the current frame.