unified-multimodal-models

Tag

Cards List
#unified-multimodal-models

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Hugging Face Daily Papers · 2026-06-11 Cached

HYDRA-X presents a unified multimodal model that integrates image and video tokenization within a single Vision Transformer, achieving strong performance across understanding and generation tasks.

0 favorites 0 likes
#unified-multimodal-models

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

arXiv cs.CL · 2026-06-02 Cached

This paper introduces UniKE, the first benchmark for cross-modal knowledge editing in unified multimodal models (UMMs), revealing a significant modality gap where text edits achieve 92% efficacy but only 18.5% transfer to image generation. It proposes Reasoning-augmented Parameter Editing to improve cross-modal transfer, with gains up to 18.6 percentage points.

0 favorites 0 likes
#unified-multimodal-models

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

Hugging Face Daily Papers · 2026-05-18 Cached

LatentUMM introduces dual latent alignment to improve cross-modal consistency in unified multimodal models by aligning transformations and stabilizing latent dynamics.

0 favorites 0 likes
#unified-multimodal-models

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning

Hugging Face Daily Papers · 2026-05-12 Cached

UniPath proposes a framework for adaptive coordination of understanding and generation in unified multimodal models, leveraging coordination-path diversity to improve performance over fixed strategies.

0 favorites 0 likes
← Back to home

Submit Feedback