Tag
This paper identifies a specialized subset of attention heads called CoRe heads in multimodal LLMs that exhibit functional sparsity in cross-modal retrieval. Causal interventions show these heads are crucial for multimodal reasoning, and leveraging this sparsity can accelerate inference.
Urban-ImageNet is a large-scale multi-modal dataset and evaluation benchmark for urban space perception from social media imagery, supporting scene classification, cross-modal retrieval, and instance segmentation tasks across 61 urban sites in 24 Chinese cities.
SIMMER proposes a novel MLLM-based embedding approach for cross-modal food image-recipe retrieval, replacing traditional dual-encoder architectures with a unified encoder and achieving state-of-the-art results on the Recipe1M dataset with significant improvements over prior methods.