@zhijianliu_: Reasoning VLAs can think. They just can't think fast. Until now. Introducing FlashDrive 716 ms → 159 ms on RTX PRO 6000…
Summary
FlashDrive reduces reasoning vision-language-action model inference latency from 716 ms to 159 ms on RTX PRO 6000—up to 5.7× faster—with zero accuracy loss, enabling real-time autonomous applications.
View Cached Full Text
Cached at: 04/21/26, 09:00 AM
Reasoning VLAs can think. They just can’t think fast. Until now. Introducing FlashDrive 716 ms → 159 ms on RTX PRO 6000 (up to 5.7×) Zero accuracy loss FlashDrive = streaming inference + DFlash speculative reasoning + ParoQuant W4A8 Real-time reasoning for autonomous
Similar Articles
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
FlashDrive is an algorithm-system co-design framework that cuts the inference latency of vision-language-action models for autonomous driving by 4.7× (from 717 ms to 151 ms on a single GPU) using streaming KV-cache reuse, non-autoregressive diffusion drafting, and adaptive step caching, with negligible accuracy loss.
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
VisualThink-VLA introduces a visual intermediate reasoning framework for vision-language-action policies that preserves spatial precision and dramatically reduces latency compared to text-based reasoning, achieving sub-second inference and state-of-the-art success rates on robot manipulation benchmarks.
@songhan_mit: explore VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
MIT researchers developed VLASH, a method enabling vision-language-action models to predict future robot states, doubling speed and reducing lag in tasks like pick-and-place and table tennis.
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
TurboVLA introduces a new Vision-Language-Action paradigm that directly maps vision and language to action, achieving 97.7% success on LIBERO with only 0.2B parameters and real-time inference at 32 Hz on consumer GPUs, significantly reducing computational cost.
@AdinaYakup: Step-3.7-Flash New VL model from @StepFun_ai 198B / 11B active - MoE 256K context 3 reasoning level Up to 400 tokens/sec
StepFun releases Step-3.7-Flash, a new large vision-language MoE model with 198B parameters (11B active), 256K context, and up to 400 tokens/sec inference speed.