TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Hugging Face Daily Papers Papers

Summary

TurboVLA introduces a new Vision-Language-Action paradigm that directly maps vision and language to action, achieving 97.7% success on LIBERO with only 0.2B parameters and real-time inference at 32 Hz on consumer GPUs, significantly reducing computational cost.

Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation. In this work, we introduce TurboVLA, a new VLA paradigm that reformulates the conventional V to L to A pathway as a direct V + L to A mapping. Instead of using a large language model as the central interface between perception and action, TurboVLA independently encodes visual observations and language instructions, directly exchanges information between them through lightweight bidirectional vision-language interaction, and predicts continuous action chunks with a compact decoder. This simple design constructs task-conditioned representations directly from visual and linguistic features, significantly reducing the computational and memory costs of VLA inference. On LIBERO, TurboVLA achieves 97.7% average success with only 0.2B parameters, 31.2 ms inference latency, and 0.9 GB inference VRAM on a consumer-grade RTX 4090, matching or outperforming substantially larger VLA policies. These results establish TurboVLA as a simple and effective alternative to the prevailing LLM-centric VLA paradigm, offering a new perspective on how vision, language, and action can be connected for efficient robotic manipulation. Code is available at https://github.com/H-EmbodVis/TurboVLA.
Original Article
View Cached Full Text

Cached at: 07/30/26, 05:46 AM

Paper page - TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Source: https://huggingface.co/papers/2607.27205

Abstract

Vision-language-action(VLA)modelscommonlyadoptanLLM-centricVtoLtoApathway,wherevisualobservationsareprojectedintotherepresentationspaceofalargelanguagemodelbeforebeingdecodedintorobotactions.Althougheffective,thisdesignincurssubstantialcomputationandmemoryoverheadateverypolicyinvocation.Inthiswork,weintroduceTurboVLA,anewVLAparadigmthatreformulatestheconventionalVtoLtoApathwayasadirectV+LtoAmapping.Insteadofusingalargelanguagemodelasthecentralinterfacebetweenperceptionandaction,TurboVLAindependentlyencodesvisualobservationsandlanguageinstructions,directlyexchangesinformationbetweenthemthroughlightweightbidirectionalvision-languageinteraction,andpredictscontinuousactionchunkswithacompactdecoder.Thissimpledesignconstructstask-conditionedrepresentationsdirectlyfromvisualandlinguisticfeatures,significantlyreducingthecomputationalandmemorycostsofVLAinference.OnLIBERO,TurboVLAachieves97.7%averagesuccesswithonly0.2Bparameters,31.2msinferencelatency,and0.9GBinferenceVRAMonaconsumer-gradeRTX4090,matchingoroutperformingsubstantiallylargerVLApolicies.TheseresultsestablishTurboVLAasasimpleandeffectivealternativetotheprevailingLLM-centricVLAparadigm,offeringanewperspectiveonhowvision,language,andactioncanbeconnectedforefficientroboticmanipulation.Codeisavailableathttps://github.com/H-EmbodVis/TurboVLA.

View arXiv pageView PDFProject pageGitHub15Add to collection

Get this paper in your agent:

hf papers read 2607\.27205

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.27205 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.27205 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.27205 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Just open-sourced FastVLA

Reddit r/LocalLLaMA

FastVLA, an open-source Vision-Language-Action model, now runs 5 Hz robotics on an L4 GPU.

TBD-VLA: Temporal Block Diffusion Vision Language Action Model

Hugging Face Daily Papers

TBD-VLA introduces a discrete vision-language-action framework that combines block diffusion with autoregressive generation to achieve efficient temporal action modeling and faster inference, significantly outperforming prior VLA approaches in simulation and real-world manipulation tasks.