SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Hugging Face Daily Papers Papers

Summary

SnapBench introduces a paired corruption benchmark for robust snap-and-ask multimodal retrieval on mobile, revealing image noise as a key degradation factor and proposing an adaptive fusion method for modality reliability.

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.
Original Article
View Cached Full Text

Cached at: 09/03/26, 07:50 AM

Paper page - SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Source: https://huggingface.co/papers/2608.29607 Published on Aug 30

·

Submitted byhttps://huggingface.co/zrchen03

zrchenon Sep 3

Abstract

SnapBench introduces paired corruption benchmarks for mobile snap-and-ask retrieval, revealing that image noise severely degrades multimodal retrieval and proposing an adaptive fusion method to calibrate modality reliability.

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information.Snap-and-ask retrievalis now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness insnap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-askmultimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, coveringdual-tower encodersandembedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack ofcross-modal fallbackunder noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further proposeMOOR(Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need forreliability-aware modality calibrationinsnap-and-ask retrieval.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2608\.29607

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2608.29607 in a model README.md to link it from this page.

Datasets citing this paper1

#### yefd/SnapBench Viewer• Updated2 days ago • 10.2k • 40 • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.29607 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory

arXiv cs.CL

Introduces SMMBench, a benchmark to evaluate multimodal agents' ability to retrieve, align, and compose evidence scattered across independently originated sources like conversations, tables, and documents. Experiments show current systems struggle with this source-distributed memory composition task.

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Hugging Face Daily Papers

Introduces MulTaBench, a benchmark of 40 datasets for multimodal tabular learning with text and image modalities, demonstrating that task-specific embedding tuning improves performance over frozen pretrained embeddings, particularly when modalities provide complementary predictive signals.

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

arXiv cs.AI

Introduces Blind-Spots-Bench, a benchmark designed to expose persistent failures in modern multimodal AI models on tasks that are trivial for humans. Evaluates a range of models, revealing performance gaps and that no single model dominates across all task types.