SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
Summary
SnapBench introduces a paired corruption benchmark for robust snap-and-ask multimodal retrieval on mobile, revealing image noise as a key degradation factor and proposing an adaptive fusion method for modality reliability.
View Cached Full Text
Cached at: 09/03/26, 07:50 AM
Paper page - SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
Source: https://huggingface.co/papers/2608.29607 Published on Aug 30
·
Submitted byhttps://huggingface.co/zrchen03
zrchenon Sep 3
Abstract
SnapBench introduces paired corruption benchmarks for mobile snap-and-ask retrieval, revealing that image noise severely degrades multimodal retrieval and proposing an adaptive fusion method to calibrate modality reliability.
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information.Snap-and-ask retrievalis now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness insnap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-askmultimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, coveringdual-tower encodersandembedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack ofcross-modal fallbackunder noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further proposeMOOR(Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need forreliability-aware modality calibrationinsnap-and-ask retrieval.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2608\.29607
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.29607 in a model README.md to link it from this page.
Datasets citing this paper1
#### yefd/SnapBench Viewer• Updated2 days ago • 10.2k • 40 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.29607 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion
MuteBench is a benchmark for evaluating multimodal fusion models under modality missing and within-modality missing conditions across clinical datasets. It provides insights into architecture robustness and suggests that diffusion-based imputation can help.
SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
Introduces SMMBench, a benchmark to evaluate multimodal agents' ability to retrieve, align, and compose evidence scattered across independently originated sources like conversations, tables, and documents. Experiments show current systems struggle with this source-distributed memory composition task.
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
Introduces MulTaBench, a benchmark of 40 datasets for multimodal tabular learning with text and image modalities, demonstrating that task-specific embedding tuning improves performance over frozen pretrained embeddings, particularly when modalities provide complementary predictive signals.
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Introduces Blind-Spots-Bench, a benchmark designed to expose persistent failures in modern multimodal AI models on tasks that are trivial for humans. Evaluates a range of models, revealing performance gaps and that no single model dominates across all task types.
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
MMShopBench introduces a real-log benchmark for multimodal, multi-turn shopping agents, requiring joint inference from user images and dialogue, with an offline sandbox and training set to improve open-source model performance.