Tag
This paper explores the synergy between visual understanding and generation in unified multimodal models, showing that task-decoupled architectures and end-to-end optimization can enhance performance by turning coexistence into synergy.