Tag
This paper introduces a hierarchical framework for multimodal sexism detection in memes that models annotator disagreement using frozen Gemini Embedding 2 representations and a lightweight gated MLP, achieving 1st place on fine-grained sexism categorization at EXIST 2026.
Proposes CIST, a method that assigns separate sample-wise adaptive temperatures to teacher and student in knowledge distillation, producing consistently informative soft labels and relaxing rigid logit-scale matching. Experiments on vision and language tasks show consistent improvements over standard KD.