Tag
This paper introduces VideoRAE, a representation autoencoder that leverages frozen video foundation models to create compact, reconstruction-capable, and generation-friendly video latents. It achieves state-of-the-art results on UCF-101 with faster convergence than competing autoencoders.
Xiaomi Robotics introduces U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis, treating embodied generation as an extension of image and video generation. It achieves state-of-the-art results on multiple embodied tasks, outperforming GPT-Image-2.0 and improving real-world manipulation success rates.
Proposes MotionVLA, a vision-language-action model for humanoid motion generation using a dual-stream frequency tokenizer that separately encodes pose and physical dynamics, achieving better diversity and consistency.
ARM presents a unified autoregressive framework for image understanding, generation, and editing using discrete semantic tokenization and reinforcement learning optimization, showing cross-task synergy.