Tag
SimpleMemVLA introduces a simple memory mechanism for Vision-Language-Action models by feeding intact timestamped video history into a pretrained VLM backbone, achieving state-of-the-art results on long-horizon manipulation tasks without dedicated memory modules.
GROVE is a training-free framework that grows a temporally stratified memory from continuous video streams, supporting both reactive QA and proactive assistance. It achieves state-of-the-art results on benchmarks like MM-lifelong and EgoServe.