unified-models

Tag

Cards List
#unified-models

Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models

arXiv cs.LG ↗ · 2d ago Cached

This paper audits whether gradient-conflict metrics (cosine similarity, conflict rates) actually predict the understanding–generation trade-off in unified multimodal models, using a controlled testbed (GridUMM) where the true trade-off is computable. Across 63 configurations and 372 checkpoints, no directional conflict metric reliably correlates with the trade-off, and a dose-response intervention suppressing conflict leaves the trade-off flat, while eff_rank and training loss outperform conflict geometry as diagnostics.

0 favorites 0 likes
#unified-models

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Hugging Face Daily Papers ↗ · 6d ago Cached

This paper introduces UMM-Reflection, a reinforcement learning method for unified multimodal models that enables self-repair of generated images, improving performance on benchmarks like GenEval, WISE, and T2I-CompBench++ without external verifiers.

0 favorites 0 likes
#unified-models

Semantic Generative Tuning for Unified Multimodal Models

Hugging Face Daily Papers ↗ · 2026-05-18 Cached

Introduces Semantic Generative Tuning (SGT), a paradigm that uses image segmentation as a generative proxy to align visual understanding and generation in unified multimodal models, improving both comprehension and fidelity.

0 favorites 0 likes
#unified-models

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

This paper introduces SenseNova-U1, a unified multimodal architecture that integrates understanding and generation tasks, releasing two variants (8B and 30B) that perform competitively in both perception and image synthesis.

0 favorites 0 likes
#unified-models

Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Hugging Face Daily Papers ↗ · 2026-05-05 Cached

The paper introduces JoyAI-Image, a unified multimodal foundation model that integrates a spatially enhanced MLLM with MMDiT to achieve state-of-the-art performance in visual understanding, text-to-image generation, and instruction-guided editing.

0 favorites 0 likes
← Back to home

Submit Feedback