Tag
Presents ForeWAM, a direct-policy World Action Model that provides predictive context to action generation via hidden future-slot KV states, avoiding explicit future video decoding while achieving high success rates on LIBERO benchmarks.
RxBrain is an embodied cognition foundation model that jointly reasons with language and visual imagination to represent embodied plans, using a unified multimodal Mixture-of-Transformers architecture. It achieves promising real-robot performance without large-scale action data.