Tag
HOMIE is a framework for human-object centric video personalization that integrates MLLM features to improve subject fidelity and interaction patterns, achieving state-of-the-art performance.
This paper proposes a novel approach that conditions diffusion models on Multimodal Large Language Models (MLLMs) for subject-driven image generation, using VAE-based identity conditioning and a Dual Layer Aggregation module to improve both semantic understanding and identity preservation while mitigating copy-paste artifacts.