Tag
WithEveryone introduces a unified framework for generating group images with up to ten identities by grounding identities to layout plans, improving identity similarity and reducing artifacts compared to previous models like GPT-Image-2.
Wan-Animate-2 is a new end-to-end character animation framework that consumes driving videos directly in a redesigned Diffusion Transformer, achieving high-fidelity motion generation and identity preservation. It also introduces a lightweight variant for real-time streaming animation, with open-source weights released.
Mirage Avatar X by Captions creates a realistic digital twin from just 10 seconds of video, preserving identity, expressions, and mannerisms without quality degradation.
Mirage Avatar X is a new AI avatar model that claims industry-leading identity preservation, expressiveness, and support for vertical and horizontal video.
This paper proposes a novel approach that conditions diffusion models on Multimodal Large Language Models (MLLMs) for subject-driven image generation, using VAE-based identity conditioning and a Dual Layer Aggregation module to improve both semantic understanding and identity preservation while mitigating copy-paste artifacts.
Soap2Soap presents a multi-agent framework for long-horizon video-to-video generation that maintains narrative structure and character identity across extended sequences via a dual-bridge consistency mechanism using a semantic screenplay and visual reference anchors.