Tag
ControlNet author Min Shen has open-sourced the FramePack video generation model, which requires only 6GB of VRAM to run a 13B model, generates a 1-minute 30fps video, takes 1.5 seconds per frame on an RTX 4090, and comes with a one-click Windows package.
This paper introduces Identifiable Token Correspondence, a method that models token correspondence across time frames to improve temporal consistency in transformer-based world models for visual reinforcement learning, achieving state-of-the-art results on multiple benchmarks.