N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
Summary
Introduces N_0-VTLA, a vision-tactile-language-action foundation model for contact-rich manipulation, featuring large-scale tactile pretraining and advantage-conditioned offline policy improvement, with strong results on real-robot and simulation benchmarks.
View Cached Full Text
Cached at: 08/03/26, 05:30 AM
Paper page - N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
Source: https://huggingface.co/papers/2607.23782
Abstract
WepresentN_0-VTLA,avision-tactile-language-action(VTLA)foundationmodelcapableof(1)fine-grainedcontact-richmanipulationwithtactileperceptionandtactile-feedbackcontrol,and(2)offlinepolicyimprovementfromstoreddeploymentdata.Buildingoncurrentvision-basedbackbones,weproposeatrainingrecipefortactileintegrationconsistingofvisuo-tactilepre-training,stagedtactile-pathwayintegration,andadvantage-conditionedofflinepolicyimprovement.Duringpre-training,thepolicylearnsbroadcontactpriorsfromNeoData,ourlarge-scalevisuo-tactilerobotdataset;toourknowledge,N_0-VTLAisthefirstVTLAmodelpretrainedontactiledataatscale.Duringpost-training,weaugmentthepolicywithapredictivetactilepathwaythatdistillsthecontactpatternslearnedatscaleintothefinemotionadjustmentsrequiredbydownstreamtactile-centricmanipulation.Forofflinepolicyimprovement,weintroduceALTER,anadvantage-conditionedofflinereinforcementlearningmethodthatconvertsrelativeprogressandtrajectory-eventcomparisonsintobinaryadvantagelabelsforpolicytrainingonafixeddeploymentcorpus,furtherimprovingtask-specificlearningoncontact-richskillssuchasdeformableobjectmanipulation.Acrosscontact-richbenchmarks,N_0-VTLAoutperformsstrongbaselinesbywidemargins:itwinsallninereal-robotNeoRealtasksandreaches63.8%meansuccessonatwenty-tasksimulationsuite,against44.0%forthestrongestbaseline.N_0-VTLApoliciestrainedwithALTERreach75-95%successonthreelong-horizonreal-robottasks.Theseresultslayafoundationforversatiletactile-drivenmanipulationpolicies.
View arXiv pageView PDFProject pageGitHub31Add to collection
Get this paper in your agent:
hf papers read 2607\.23782
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.23782 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.23782 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.23782 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
N₀-TWAM is a tactile-native world-action model for contact-rich manipulation, trained at scale on visuo-tactile data from 6 embodiments and 450 tasks. The authors release code and pretrained checkpoints, positioning it as the first tactile world-action model trained at scale.
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
AffordanceVLA introduces a unified framework using structured affordance forecasting as an intermediate representation to improve perception-action mapping in robotic manipulation, leveraging vision-language models and a Mixture-of-Transformer architecture.
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
LaMem-VLA proposes a latent-memory-native framework that integrates short-term and long-term historical experience directly into Vision-Language-Action reasoning, enabling better performance on long-horizon robotic manipulation tasks.
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
VisualThink-VLA introduces a visual intermediate reasoning framework for vision-language-action policies that preserves spatial precision and dramatically reduces latency compared to text-based reasoning, achieving sub-second inference and state-of-the-art success rates on robot manipulation benchmarks.
LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
LabVLA is a vision-language-action model for scientific laboratory automation, trained with a two-stage approach combining action token pretraining and flow matching. It achieves state-of-the-art success rates on the LabUtopia benchmark by leveraging simulated data to bridge the gap between household demonstrations and lab-specific tasks.