RoboTTT: Context Scaling for Robot Policies

Hugging Face Daily Papers Papers

Summary

RoboTTT scales visuomotor context to 8K timesteps for robot policies, enabling one-shot imitation from human video demonstrations, on-the-fly policy improvement, and robustness to perturbations. It achieves an 87% improvement over baselines and completes a five-minute, ten-stage assembly task that no baseline could.

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:43 AM

Paper page - RoboTTT: Context Scaling for Robot Policies

Source: https://huggingface.co/papers/2607.15275 Authors:

,

,

,

,

,

,

,

,

,

Abstract

Recentrobotfoundationmodelsoperatewithsingle-steporshort-historyvisuomotorcontext.WeintroduceTest-Time-TrainingRobotPolicies(RoboTTT),arobotmodelandtrainingrecipethatscalevisuomotorcontextto8Ktimesteps,threeordersofmagnitudebeyondstate-of-the-artpolicies,withoutgrowinginferencelatency.Atthiscontextlength,weunlocknewrobotcapabilities:one-shotin-contextimitationfromhumanvideodemonstrations,on-the-flypolicyimprovement,robustnesstoperturbations,andstrongerperformanceonmulti-stage,long-horizontasks.Wealsoobserve,forthefirsttime,steadygainsinclosed-loopperformanceaspretrainingcontextlengthscales.Atitscore,RoboTTTintegratesTest-TimeTrainingintorobotfoundationmodelssuchasVision-Language-Actionpolicies,yieldingasequencemodelwhoserecurrentstateconsistsoffastweights,parametersupdatedbygradientdescentduringbothtrainingandinference,compressinghistoriesintoweightspaceandretrievingcontextualinformationforlong-contextconditioning.Toscaletrainingcontextlength,therecipecombinessequenceactionforcingwithtruncatedbackpropagationthroughtime.Onchallengingreal-robotmanipulationtasks,RoboTTTimprovesoverallperformanceby87%overthesingle-stepcontextbaselineandfullycompletesafive-minute,ten-stageassemblytask,whichnobaselineeverdoes.RoboTTTtrainedwith8K-timestepcontextoutperformsthesamemodelpretrainedwith1Ktimestepsby62%,suggestingcontextlengthasanewscalingaxisforrobotfoundationmodels.Videosareavailableathttps://research.nvidia.com/labs/gear/robottt/

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2607\.15275

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.15275 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.15275 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.15275 in a Space README.md to link it from this page.

Collections including this paper3

Similar Articles

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

Hugging Face Daily Papers

RoboLab is a high-fidelity simulation benchmarking framework for evaluating task-generalist robotic policies, introducing the RoboLab-120 benchmark with 120 tasks across visual, procedural, and relational competency axes. It enables scalable, realistic task generation and systematic analysis of policy behavior under controlled perturbations to assess true generalization capabilities.

LeRobot v0.5.0: Scaling Every Dimension

Hugging Face Blog

LeRobot v0.5.0 is a major release featuring support for Unitree G1 humanoid robots, new policy architectures (Pi0-FAST VLAs, Real-Time Chunking), streaming video encoding for 3x faster training, and EnvHub for loading simulation environments from Hugging Face Hub.

In-Context Robot Learning with VLM Agents

Hugging Face Daily Papers

This paper introduces GPT-Policy, a framework for in-context robot learning using vision-language models, enabling robots to learn from demonstrations without gradient updates. It evaluates the framework in real-robot trials, showing improved task completion.