@drfeifei: I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between @StanfordS…
Summary
Fei-Fei Li highlights a new test-time training approach for robotic learning, developed in collaboration between Stanford SVL and NVIDIA Robotics, which scales robot model context to 8000 timesteps with constant inference cost.
View Cached Full Text
Cached at: 07/16/26, 10:12 AM
I’m very excited by this test time training work for robotic learning! It’s an awesome collaboration between @StanfordSVL and @NVIDIARobotics !
Jim Fan (@DrJimFan): We scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot policies used to live their lives a few frames at a time (< 0.1 sec), instantly forgetting what just happened. We pushed to 3 orders of magnitude
Similar Articles
@drfeifei: 1/N Long horizon, complex tasks that truly matter in everyday life are not solved problems by today’s robotics, requiri…
Dr. Fei-Fei Li announces the second year of Stanford's BEHAVIOR Challenge, a robotics competition tackling long-horizon, complex everyday tasks, with new tasks, improved evaluation, and an $11,000 prize pool.
@svlevine: Today (June 3), I'll be speaking at CVPR at the Test-Time Scaling for Computer Vision WS (1:30 pm PT) about how we can …
Sergey Levine announces he will be speaking at CVPR workshops on test-time scaling for computer vision and robot policy generalization, as well as on deployment of foundation models.
@drfeifei: https://x.com/drfeifei/status/2062247238143996275
Fei-Fei Li and the World Labs team present a functional taxonomy of world models, distinguishing between renderers, physics engines, and other components within the reinforcement learning loop, and arguing that spatial intelligence is AI's next frontier.
@robbyant_brain: LingBot-VLA 2.0 is now open-source — our next-gen embodied foundation model. 60,000 hours of high-quality pretraining d…
LingBot-VLA 2.0, an open-source embodied foundation model, has been released with 60,000 hours of pretraining data supporting 20 robot configurations across 17 brands, capable of sub-130ms inference on RTX 4090.
@yukangchen_: Excited to share our new blog: Scaling Video Training with Parallelism https://research.nvidia.com/labs/eai/blogs/scali…
This blog from NVIDIA Research discusses how sequence parallelism can scale long-video training systems for both understanding and generation, addressing the challenge of fitting very long video sequences across multiple GPUs.