Tag
The paper introduces MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes to evaluate vision-language models' understanding of cultural context and humor, demonstrating that providing cultural context improves model performance.
This paper proposes MAVEN, a hierarchical framework for evaluating multimodal content against macro-societal values, featuring a benchmark and optimized evaluators for scalable assessment.
Introduces PerceptionRubrics, a rubric-based evaluation framework for multimodal AI that shifts from holistic matching to atomic auditing using 1,038 images and over 10,000 instance-specific rubrics, with gated scoring to enforce strict perceptual fidelity.
Researchers introduce MM-JudgeBias, a benchmark that exposes systematic compositional biases in multimodal large language models when used as automatic judges, testing 26 SOTA MLLMs across 1,800 samples.