@heyshrutimishra: New video model just dropped. But this one isn't built for cinematic video. LingBot-Video is designed for embodied inte…
Summary
LingBot-Video, a 30B-parameter video model with sparse MoE, designed for embodied intelligence, is open-sourced. It outperforms existing models on RBench, trained on 70K+ hours of embodied data.
View Cached Full Text
Cached at: 07/09/26, 09:35 AM
New video model just dropped. But this one isn’t built for cinematic video.
LingBot-Video is designed for embodied intelligence. The training data is what sets it apart: 70,000+ hours of manipulation, navigation, and egocentric interaction. Internet video teaches how things look. This teaches how things change when you act on them.
30B parameters, only 3B active at inference. Sparse MoE, roughly 3x faster at long sequences. Already beats Wan2.6, Seedance 1.5 Pro, and Cosmos3 Super on RBench from Peking University and ByteDance.
It’s open source, which matters here. Researchers can actually build on it instead of just reading about it.
Robbyant (@robbyant_brain): Today we open-source LingBot-Video — the first MoE-based video foundation model built for embodied intelligence. 🔹30B params, only 3B active at inference. 🔹Augmented with 70K hours of embodied data on top of large-scale internet video pretraining. 🔹Already outperforming
Similar Articles
@_akhaliq: LingBot-Video is out on Hugging Face MoE-based video foundation model built for embodied intelligence 30B params, only …
LingBot-Video, a 30B parameter MoE-based video foundation model for embodied intelligence, has been released on Hugging Face with only 3B active parameters at inference, augmented with 70K hours of embodied data.
robbyant/lingbot-video-moe-30b-a3b
LingBot-Video is the first open-source large-scale MoE video generation model for embodied intelligence, featuring efficient MoE architecture, massive embodied data training, and multi-reward system for high aesthetics, physical rationality, and task completion.
@rohanpaul_ai: Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 fr…
LingBot-VA 2.0 is a video-action foundation model trained from scratch for robot control, achieving 225 Hz closed-loop execution with 13B parameters (1.9B active per token) and outperforming prior models on RoboTwin 2.0.
LingBot-Video: sparse-MoE video diffusion transformer (13B total, 1.4B active) post-trained as an action-conditioned world model[R]
LingBot-Video is a 13B sparse-MoE video diffusion transformer (1.4B active) post-trained with RL as an action-conditioned world model, open-sourced with weights and code. It includes a physical-plausibility reward graded by a VLM and frames itself as a policy evaluator and action planner, though closed-loop robot results are absent.
@robbyant_brain: LingBot-VLA 2.0 is now open-source — our next-gen embodied foundation model. 60,000 hours of high-quality pretraining d…
LingBot-VLA 2.0, an open-source embodied foundation model, has been released with 60,000 hours of pretraining data supporting 20 robot configurations across 17 brands, capable of sub-130ms inference on RTX 4090.