Tag
SIMMER proposes a novel MLLM-based embedding approach for cross-modal food image-recipe retrieval, replacing traditional dual-encoder architectures with a unified encoder and achieving state-of-the-art results on the Recipe1M dataset with significant improvements over prior methods.