WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Hugging Face Daily Papers Papers

Summary

WithEveryone introduces a unified framework for generating group images with up to ten identities by grounding identities to layout plans, improving identity similarity and reducing artifacts compared to previous models like GPT-Image-2.

Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.
Original Article
View Cached Full Text

Cached at: 08/21/26, 04:09 AM

Paper page - WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Source: https://huggingface.co/papers/2608.20336

Abstract

WithEveryone enables reliable identity-preserving group image generation for up to ten people by grounding identities to layout plans and using region-based identity losses.

Identity-preserving image generationbecomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as anaddressed token, predicts a structuredidentity--layout plan, and renders the plan as a visual condition. Its key objective,Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching;ID Representation Forcingadditionally trains a prediction for each identity before image synthesis. On anidentity-disjoint benchmark, WithEveryone achieves the highesttarget-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.

View arXiv pageView PDFProject pageGitHub4Add to collection

Get this paper in your agent:

hf papers read 2608\.20336

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.20336 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2608.20336 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.20336 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Hugging Face Daily Papers

Qwen-Image-Agent proposes a unified agentic framework that addresses the context gap in text-to-image generation by integrating planning, reasoning, searching, and memory mechanisms. It introduces IA-Bench for evaluation and achieves state-of-the-art performance.

Personalized Image Generation with Reasoning and Reflection

Hugging Face Daily Papers

The paper introduces the first unified benchmark for personalized image generation from rich user histories and proposes PEARL, a method coupling a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop, achieving 15% average improvement on personalization metrics.