Tag
This paper audits whether gradient-conflict metrics (cosine similarity, conflict rates) actually predict the understanding–generation trade-off in unified multimodal models, using a controlled testbed (GridUMM) where the true trade-off is computable. Across 63 configurations and 372 checkpoints, no directional conflict metric reliably correlates with the trade-off, and a dose-response intervention suppressing conflict leaves the trade-off flat, while eff_rank and training loss outperform conflict geometry as diagnostics.
This paper introduces UMM-Reflection, a reinforcement learning method for unified multimodal models that enables self-repair of generated images, improving performance on benchmarks like GenEval, WISE, and T2I-CompBench++ without external verifiers.
Introduces Semantic Generative Tuning (SGT), a paradigm that uses image segmentation as a generative proxy to align visual understanding and generation in unified multimodal models, improving both comprehension and fidelity.
This paper introduces SenseNova-U1, a unified multimodal architecture that integrates understanding and generation tasks, releasing two variants (8B and 30B) that perform competitively in both perception and image synthesis.
The paper introduces JoyAI-Image, a unified multimodal foundation model that integrates a spatially enhanced MLLM with MMDiT to achieve state-of-the-art performance in visual understanding, text-to-image generation, and instruction-guided editing.