Tag
StreamPI introduces a streaming multimodal temporal modeling framework for vision-language-action models, improving robot manipulation through instruction-anchored attention and randomized interval training without additional parameters.