Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
Summary
This paper proposes a novel approach that conditions diffusion models on Multimodal Large Language Models (MLLMs) for subject-driven image generation, using VAE-based identity conditioning and a Dual Layer Aggregation module to improve both semantic understanding and identity preservation while mitigating copy-paste artifacts.
View Cached Full Text
Cached at: 05/27/26, 02:47 AM
Paper page - Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
Source: https://huggingface.co/papers/2605.26111
Abstract
A novel approach conditions diffusion models on multimodal large language models for subject-driven image generation, combining text and reference image encoding with VAE-based identity conditioning to improve both semantic understanding and identity preservation.
Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. This limitscross-modal reasoningabilities and causescopy-paste artifacts. Recent frameworks that connect multimodal models anddiffusion modelsimprove instruction following, but largely overlook identity preservation. To address these limitations, we conditiondiffusion modelsonMultimodal Large Language Models(MLLMs) that jointly encode text and reference images, and augment it withVAE-based identity conditioning. A novelDual Layer Aggregation(DLA) module is designed to aggregate multi-level MLLM features for optimal conditioning, and amulti-stage denoising strategyis applied to progressively balance thesemantic informationfrom MLLM andfine-detail identityfrom VAE during inference. Extensive experiments demonstrate that our approach harmonizes multimodal understanding with identity preservation, mitigates copy-paste issues, and achieves superior performance regarding human preference on subject-driven image generation. Our project website is available at https://zsh2000.github.io/squeeze-mllm-subject-gen/.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.26111
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.26111 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.26111 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.26111 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
This paper investigates how training-free acceleration can silently change generated content in diffusion-based multimodal large language models, and proposes paired diagnostics and consistency-control methods to mitigate content drift.
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
MMCORE introduces a unified multimodal image generation and editing framework that aligns VLM semantic embeddings with diffusion conditioning, achieving state-of-the-art fidelity without costly fusion or from-scratch training.
Semantic DLM+: Improving Diffusion Language Models through Bias-variance Trade-off in Transition Kernel Design
This paper theoretically analyzes diffusion language models through a bias-variance lens, identifying trade-offs between masking and uniform diffusion kernels. It proposes SemDLM+, which adds a global transition and semantic-frequency penalty to overcome the semantic basin problem, achieving competitive generation quality on LM1B and OpenWebText benchmarks.
Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
Multimodal Flow presents a fully continuous generative architecture for language and vision using a shared chunk-causal flow backbone over continuous embedding hyperchunks, with MF-1 pretrained at 0.6B to 1.6B scales that competitively performs on GenEval, DPG-Bench, VQAv2, MMBench, and POPE using only 150B tokens.
Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation
This paper proposes a multimodal generative framework integrating fine-tuned Stable Diffusion XL with LLaMA for controllable and culturally faithful Ulos motif generation, validated through ablation studies and qualitative evaluations.