RoboTTT: Context Scaling for Robot Policies
Summary
RoboTTT scales visuomotor context to 8K timesteps for robot policies, enabling one-shot imitation from human video demonstrations, on-the-fly policy improvement, and robustness to perturbations. It achieves an 87% improvement over baselines and completes a five-minute, ten-stage assembly task that no baseline could.
View Cached Full Text
Cached at: 07/20/26, 09:43 AM
Paper page - RoboTTT: Context Scaling for Robot Policies
Source: https://huggingface.co/papers/2607.15275 Authors:
,
,
,
,
,
,
,
,
,
Abstract
Recentrobotfoundationmodelsoperatewithsingle-steporshort-historyvisuomotorcontext.WeintroduceTest-Time-TrainingRobotPolicies(RoboTTT),arobotmodelandtrainingrecipethatscalevisuomotorcontextto8Ktimesteps,threeordersofmagnitudebeyondstate-of-the-artpolicies,withoutgrowinginferencelatency.Atthiscontextlength,weunlocknewrobotcapabilities:one-shotin-contextimitationfromhumanvideodemonstrations,on-the-flypolicyimprovement,robustnesstoperturbations,andstrongerperformanceonmulti-stage,long-horizontasks.Wealsoobserve,forthefirsttime,steadygainsinclosed-loopperformanceaspretrainingcontextlengthscales.Atitscore,RoboTTTintegratesTest-TimeTrainingintorobotfoundationmodelssuchasVision-Language-Actionpolicies,yieldingasequencemodelwhoserecurrentstateconsistsoffastweights,parametersupdatedbygradientdescentduringbothtrainingandinference,compressinghistoriesintoweightspaceandretrievingcontextualinformationforlong-contextconditioning.Toscaletrainingcontextlength,therecipecombinessequenceactionforcingwithtruncatedbackpropagationthroughtime.Onchallengingreal-robotmanipulationtasks,RoboTTTimprovesoverallperformanceby87%overthesingle-stepcontextbaselineandfullycompletesafive-minute,ten-stageassemblytask,whichnobaselineeverdoes.RoboTTTtrainedwith8K-timestepcontextoutperformsthesamemodelpretrainedwith1Ktimestepsby62%,suggestingcontextlengthasanewscalingaxisforrobotfoundationmodels.Videosareavailableathttps://research.nvidia.com/labs/gear/robottt/
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.15275
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.15275 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.15275 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.15275 in a Space README.md to link it from this page.
Collections including this paper3
Similar Articles
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
RoboTALES introduces a two-stage framework combining LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training, significantly outperforming existing methods on long-horizon manipulation tasks.
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
RoboLab is a high-fidelity simulation benchmarking framework for evaluating task-generalist robotic policies, introducing the RoboLab-120 benchmark with 120 tasks across visual, procedural, and relational competency axes. It enables scalable, realistic task generation and systematic analysis of policy behavior under controlled perturbations to assess true generalization capabilities.
LeRobot v0.5.0: Scaling Every Dimension
LeRobot v0.5.0 is a major release featuring support for Unitree G1 humanoid robots, new policy architectures (Pi0-FAST VLAs, Real-Time Chunking), streaming video encoding for 3x faster training, and EnvHub for loading simulation environments from Hugging Face Hub.
In-Context Robot Learning with VLM Agents
This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
OmniTacTune introduces a two-stage reinforcement learning pipeline for adapting tactile feedback to pretrained visual robot policies, achieving 85-100% success on contact-rich manipulation tasks within 40-80 minutes.