LingBot-Video: sparse-MoE video diffusion transformer (13B total, 1.4B active) post-trained as an action-conditioned world model[R]
Summary
LingBot-Video is a 13B sparse-MoE video diffusion transformer (1.4B active) post-trained with RL as an action-conditioned world model, open-sourced with weights and code. It includes a physical-plausibility reward graded by a VLM and frames itself as a policy evaluator and action planner, though closed-loop robot results are absent.
Similar Articles
robbyant/lingbot-video-moe-30b-a3b
LingBot-Video is the first open-source large-scale MoE video generation model for embodied intelligence, featuring efficient MoE architecture, massive embodied data training, and multi-reward system for high aesthetics, physical rationality, and task completion.
@_akhaliq: LingBot-Video is out on Hugging Face MoE-based video foundation model built for embodied intelligence 30B params, only …
LingBot-Video, a 30B parameter MoE-based video foundation model for embodied intelligence, has been released on Hugging Face with only 3B active parameters at inference, augmented with 70K hours of embodied data.
@heyshrutimishra: New video model just dropped. But this one isn't built for cinematic video. LingBot-Video is designed for embodied inte…
LingBot-Video, a 30B-parameter video model with sparse MoE, designed for embodied intelligence, is open-sourced. It outperforms existing models on RBench, trained on 70K+ hours of embodied data.
@rohanpaul_ai: Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 fr…
LingBot-VA 2.0 is a video-action foundation model trained from scratch for robot control, achieving 225 Hz closed-loop execution with 13B parameters (1.9B active per token) and outperforming prior models on RoboTwin 2.0.
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
LingBot-Video presents a DiT-based video pretraining framework with Mixture-of-Experts architecture, specialized data augmentation, and multi-dimensional reward system for embodied intelligence applications.