@robbyant_brain: LingBot-VLA 2.0 is now open-source — our next-gen embodied foundation model. 60,000 hours of high-quality pretraining d…
Summary
LingBot-VLA 2.0, an open-source embodied foundation model, has been released with 60,000 hours of pretraining data supporting 20 robot configurations across 17 brands, capable of sub-130ms inference on RTX 4090.
View Cached Full Text
Cached at: 07/07/26, 11:38 PM
LingBot-VLA 2.0 is now open-source — our next-gen embodied foundation model. 60,000 hours of high-quality pretraining data — combining curated robotic demonstrations and egocentric human operation videos 20 robot configurations across 17 brands — Astribot, Leju, Unitree, Franka, Fourier, Realman, and more Whole-body DoF: heads, waists, dexterous hands, and mobile bases — enabling far more complex task scenarios Inference under 130ms on RTX 4090 — developer events launching soon #EmbodiedAI #Robotics #OpenSource #VLA
Similar Articles
From Foundation to Application: Improving VLA Models in Practice
This paper presents LingBot-VLA 2.0, which enhances VLA foundation models for robotics by improving generalization across tasks and embodiments, expanding action space to whole-body degrees of freedom, and incorporating predictive dynamics modeling for better temporal reasoning.
@rohanpaul_ai: Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 fr…
LingBot-VA 2.0 is a video-action foundation model trained from scratch for robot control, achieving 225 Hz closed-loop execution with 13B parameters (1.9B active per token) and outperforming prior models on RoboTwin 2.0.
@AdinaYakup: Powered by the new LingBot Vision https://huggingface.co/collections/robbyant/lingbot-vision… - Apache 2.0 - 4 versions…
LingBot Vision is a new visual foundation model released under Apache 2.0, available in four sizes (small, base, large, giant) and pretrained with masked boundary modeling. It powers a depth estimation system that tops multiple benchmarks.
@_akhaliq: LingBot-Video is out on Hugging Face MoE-based video foundation model built for embodied intelligence 30B params, only …
LingBot-Video, a 30B parameter MoE-based video foundation model for embodied intelligence, has been released on Hugging Face with only 3B active parameters at inference, augmented with 70K hours of embodied data.
robbyant/lingbot-world-v2-14b-causal-fast
LingBot-World 2.0 is an advanced world model achieving unbounded interaction horizons, real-time 720p/60fps video streaming, diverse interactive elements, and an agentic harness integrating pilot and director agents. The model is released on Hugging Face with inference code and technical report.