Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Summary
The paper introduces Image Bundle Composition (IBC) to shift image retrieval from atomic matching to dynamic composition of cohesive image bundles, addressing limitations in traditional approaches. It proposes a benchmark dataset IBCBench and an agentic framework BundleWeaver that leverages LLMs and VLMs for relational composition.
View Cached Full Text
Cached at: 09/01/26, 11:47 AM
Paper page - Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
Source: https://huggingface.co/papers/2608.28695 Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
Image Bundle Composition reframes retrieval as dynamic assembly of relationally coherent image groups, with a benchmark and agentic framework addressing combinatorial joint relevance.
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is scored in isolation. However, this atomic paradigm fails to capture the complexity of human search intent within personal photo collections, where users often seek compact visual stories bound by structural relations rather than isolated snapshots. To address this limitation, we introduce **Image Bundle Composition(IBC)**, a novel paradigm that shifts the objective from ranking individual images to dynamically composing cohesive image bundles from a massive, unstructured photo pool. Since target bundles are not predefined, IBC presents a severe combinatorial explosion challenge and demands modeling non-decomposable joint relevance. To establish this paradigm, we construct **IBCBench**, the first IBC benchmark dataset containing 109,467 images and 667 verified queries, built via a semi-automated verification pipeline. Furthermore, we propose **BundleWeaver**, an agentic framework that reformulates IBC as query-conditioned incrementalhyperedge discovery. By employing aLarge Language Modelto adaptively search for missing relational roles and utilizing aVision-Language Modelfor whole-bundle verification,BundleWeavereffectively navigates the combinatorial space. Extensive experiments demonstrate that while state-of-the-art embedding models and static decompose-and-rerank paradigms suffer from relational blindness,BundleWeaverachieves substantial performance gains, highlighting the necessity of shifting from atomic scoring to dynamicrelational composition. Our dataset and code are available.
View arXiv pageView PDFGitHub2Add to collection
Get this paper in your agent:
hf papers read 2608\.28695
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.28695 in a model README.md to link it from this page.
Datasets citing this paper1
#### CyberDancer/IBCBench Viewer• Updatedabout 6 hours ago • 667 • 51
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.28695 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Benchmarking Composed Image Retrieval for Applied Earth Observation
This paper presents a unified benchmark for composed image retrieval in Earth observation, evaluating vision-language backbones and introducing a change-centric dataset (xView2-CIR) for disaster monitoring, highlighting distinct challenges compared to attribute-based retrieval.
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
InterLV-Search is a new benchmark introduced in this paper to evaluate interleaved language-vision agentic search, highlighting limitations in current systems regarding visual evidence seeking and multimodal integration.
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Qwen-Image-Agent proposes a unified agentic framework that addresses the context gap in text-to-image generation by integrating planning, reasoning, searching, and memory mechanisms. It introduces IA-Bench for evaluation and achieves state-of-the-art performance.
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
IV-CoT decomposes visual conditioning into structural and semantic cascades for improved structure-aware image generation, using training-only sketch supervision to guide structural queries. It achieves state-of-the-art results on GenEval and T2I-CompBench.
Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
This paper introduces an agentic framework combining LLMs and VLMs for consistent multi-instruction video editing across multiple shots, and proposes the MMLVE task and benchmark to evaluate performance.