WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
Summary
WithEveryone introduces a unified framework for generating group images with up to ten identities by grounding identities to layout plans, improving identity similarity and reducing artifacts compared to previous models like GPT-Image-2.
View Cached Full Text
Cached at: 08/21/26, 04:09 AM
Paper page - WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
Source: https://huggingface.co/papers/2608.20336
Abstract
WithEveryone enables reliable identity-preserving group image generation for up to ten people by grounding identities to layout plans and using region-based identity losses.
Identity-preserving image generationbecomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as anaddressed token, predicts a structuredidentity--layout plan, and renders the plan as a visual condition. Its key objective,Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching;ID Representation Forcingadditionally trains a prediction for each identity before image synthesis. On anidentity-disjoint benchmark, WithEveryone achieves the highesttarget-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2608\.20336
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.20336 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.20336 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.20336 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Qwen-Image-Agent proposes a unified agentic framework that addresses the context gap in text-to-image generation by integrating planning, reasoning, searching, and memory mechanisms. It introduces IA-Bench for evaluation and achieves state-of-the-art performance.
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
This paper introduces a persona-based evaluation framework that uses synthetic cognitive profiles to represent diverse human perspectives for pluralistic alignment in generative AI, addressing the limitations of monolithic benchmarks.
Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
The paper introduces JoyAI-Image, a unified multimodal foundation model that integrates a spatially enhanced MLLM with MMDiT to achieve state-of-the-art performance in visual understanding, text-to-image generation, and instruction-guided editing.
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
UniAR presents a unified autoregressive framework that uses a single discrete visual tokenizer to bridge visual understanding and generation, achieving state-of-the-art results in image generation and editing.
Personalized Image Generation with Reasoning and Reflection
The paper introduces the first unified benchmark for personalized image generation from rich user histories and proposes PEARL, a method coupling a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop, achieving 15% average improvement on personalization metrics.