Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

arXiv cs.AI Papers

Summary

This paper introduces MAR-12, a framework using Vision-Language Models and multi-angle reasoning to detect and explain harmful humor in memes, achieving state-of-the-art accuracy on PrideMM and Memotion datasets.

arXiv:2607.15442v1 Announce Type: new Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate classification and human interpretability. However, existing multimodal classifiers either overlook these interdependencies or provide only limited interpretability. In this paper, we introduce MAR-12, a novel framework that leverages Vision Language Models (VLMs) for meme detection and understanding in settings where humorous and hateful elements may coexist. The framework first interprets each meme through twelve structured perspectives derived from humor and hate theories. It then applies a role-aware soft-gated attention mechanism to learn how much each perspective should contribute, followed by a prototype-based classifier for the final prediction. Finally, explanations are synthesized using both perspective-specific reasoning and learned attention weights, ensuring transparent and context-grounded justifications. We evaluate MAR-12 on the PrideMM and Memotion datasets, where it achieves up to 80.3% accuracy for humor detection and 75.9% accuracy for hate detection, outperforming state-of-the-art approaches. Furthermore, both human and GPT-4-based evaluations confirm that MAR-12 produces coherent and persuasive explanations, particularly for memes in which humorous and harmful cues co-occur.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:21 AM

# Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
Source: [https://arxiv.org/html/2607.15442](https://arxiv.org/html/2607.15442)
###### Abstract

Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist\. These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate classification and human interpretability\. However, existing multimodal classifiers either overlook such intertwinings or provide only limited interpretability\. In this paper, we introduce MAR\-12, a novel framework that leverages Vision–Language Models \(VLMs\) for meme detection and understanding, where both humorous and hateful elements can coexist\. The framework first interprets each meme through twelve structured perspectives derived from humor and hate theory\. It then applies a role\-aware soft\-gated attention mechanism to learn how much each perspective should contribute, followed by a prototype\-based classifier for final prediction\. Finally, explanations were synthesized using both perspective\-specific reasoning and learned attention weights, ensuring transparent and context\-grounded justifications\. We evaluate MAR\-12 on the PrideMM and Memotion datasets, where it achieves up to 80\.3% accuracy for humor detection and 75\.9% accuracy for hate detection, outperforming state\-of\-the\-art approaches\. Furthermore, both human and GPT\-4\-based evaluations confirm that MAR\-12 produces coherent and persuasive explanations, particularly for memes in which humorous and harmful cues co\-occur\.

*Warning: Contains potentially offensive content\.*

CODE—https://anonymous\.4open\.science/r/MAR\-12

## Introduction

> There is a thin line that separates laughter and pain, comedy and tragedy, humor and hurt\. — Erma Bombeck

Internet memes that blend images with short text have become a pervasive mode of online communication\. While many are designed simply to entertain, this same expressive format can also be weaponized: humor often functions as a rhetorical shield for hateful or discriminatory content\. This creates a blurred and often deliberately manipulated boundary between what is perceived as “funny” and what is harmful\. During the COVID\-19 pandemic, for instance, humorous memes served as coping mechanisms for fear and uncertainty\(Yus[2023](https://arxiv.org/html/2607.15442#bib.bib8)\), yet similar formats were appropriated to spread racist, conspiratorial, or stigmatizing narratives\(Steffen[2025](https://arxiv.org/html/2607.15442#bib.bib140)\)\. Likewise, LGBTQ\+ memes have long supported community expression and identity\(Griffin[2021](https://arxiv.org/html/2607.15442#bib.bib14); He and Chang[2025](https://arxiv.org/html/2607.15442#bib.bib15)\), but the same visual\-textual patterns have also been used to mock or marginalize queer identities\(Bakeret al\.[2020](https://arxiv.org/html/2607.15442#bib.bib17); Crawfordet al\.[2021](https://arxiv.org/html/2607.15442#bib.bib141)\)\. These examples underscore the need for automated systems that can differentiate benign humor from harmful messaging, especially when the intended meaning is intentionally ambiguous\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/x1.png)Figure 1:Comparison of our proposed MAR\-12 \(bottom\) with traditional meme detection \(upper\)\.A central challenge in meme understanding is that humor is inherently subjective and multidimensional: what one viewer finds amusing may appear offensive, confusing, or harmful to another\(Attardo[2024](https://arxiv.org/html/2607.15442#bib.bib12)\)\. Moreover, humor in memes is not generated by a single cue\. It may emerge from visual incongruity, textual irony, emotional mismatch, cultural reference, exaggeration, or linguistic play, often intertwined with undertones of hostility or stereotyping\(Kalloniatis and Adamidis[2024](https://arxiv.org/html/2607.15442#bib.bib20)\)\. As illustrated in Figure[1](https://arxiv.org/html/2607.15442#Sx1.F1)\(a\), traditional learning frameworks, whether unimodal or multimodal, tend to compress these heterogeneous cues into a single embedding, making them prone to overlooking the subtle interactions that determine whether a meme is humorous, hateful, or both\(Liuet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib114); Shahet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib87); Pramanicket al\.[2021](https://arxiv.org/html/2607.15442#bib.bib90)\)\. Although recent work fine\-tunes vision\-language models \(VLMs\) to improve alignment between images and text for hateful meme detection\(Caoet al\.[2023](https://arxiv.org/html/2607.15442#bib.bib27); Heeet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib137)\), these models typically behave as black boxes, providing little insight into why a given meme is categorized as humorous or harmful\.

Recent interpretability\-oriented approaches generate rationales or stepwise explanations for meme classification\(Jiet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib84); Rizwanet al\.[2026](https://arxiv.org/html/2607.15442#bib.bib134); Liuet al\.[2026](https://arxiv.org/html/2607.15442#bib.bib133)\), marking an important shift toward transparency\. However, these methods generally assume that humor and hate are mutually exclusive or that a single explanatory pathway is sufficient\. In reality, many memes deliberately blend comedic devices with harmful insinuations, producing layered meanings that require reasoning across multiple, and often conflicting cues\. A meme may be humorous due to incongruity or exaggeration while simultaneously invoking stereotypes or derogatory comparisons\. Existing explanation frameworks tend to highlight only the dominant signal and overlook minority cues critical for identifying harmful intent\. Moreover, they rarely distinguish foundational content from higher\-level interpretive mechanisms, making the resulting explanations difficult to align with established theories of humor and harmful speech\. These limitations suggest that understanding humor and hate co\-occurrence demands multiple coordinated interpretive lenses rather than one monolithic rationale\.

To address this gap, we propose Multiple\-Angle Reasoning \(MAR\), which guides the vision\-language model \(VLM\) to interpret a meme through a coordinated set of twelve complementary perspectives\. Instead of collapsing all cues into a single embedding, MAR decomposes interpretation into foundational content \(e\.g\., image description, extracted text\), humor mechanisms \(e\.g\., irony, absurdity, affective contrast, linguistic play\), and social\-meaning cues \(e\.g\., harmfulness, intent\), forming a coverage set of reasoning pathways grounded in humor theory and harmful speech analysis\. These perspectives are not rigid components; the model is free to emphasize or suppress them depending on the meme\. To integrate these outputs, MAR employs a role\-aware soft\-attention module that learns how each perspective should contribute to the final decision by explicitly encoding its semantic function\. Unlike standard attention, which treats inputs as interchangeable, this mechanism preserves functional distinctions and captures how the relevance of different cues shifts across humorous, hateful, or mixed\-intent memes\. As illustrated in Figure[1](https://arxiv.org/html/2607.15442#Sx1.F1)\(b\), this yields an interpretable distribution over reasoning angles and a final prediction that transparently reflects how competing interpretations were weighted\. Our key contributions are:

- •Our proposed MAR is the first framework to jointly examine humor and hate in memes through a diverse set of interpretable angles, covering mechanisms such as visual/textual incongruity, cultural grounding, emotional contrast, absurdity, linguistic play, harmfulness, and inferred intent\.
- •We introduce a role\-aware soft\-attention mechanism that dynamically weights these perspectives based on the input meme, enabling flexible, human\-aligned interpretation rather than fixed or hard\-coded reasoning\.
- •Through extensive experiments across two Meme datasets containing both humorous and hateful content, we demonstrate that MAR consistently outperforms state\-of\-the\-art models on both tasks, particularly in ambiguous cases where humor is intentionally used to mask harmful messaging\.

## Related Works

### Meme Detection

Early work on humorous meme classification relied on multimodal fusion of visual and textual features\(Guoet al\.[2020](https://arxiv.org/html/2607.15442#bib.bib98)\)\. Later models incorporated attention mechanisms to improve cross\-modal alignment\(Liuet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib114)\)\. However, these approaches still treat humor as a flat label and do not explicitly consider irony, absurdity, affective contrast, exaggeration, linguistic play, that shape humorous intent\. They also overlook that humor may coexist with subtle hostility\. In parallel, hateful meme detection has advanced through large multimodal benchmark, such as the Hateful Memes Challenge\(Kielaet al\.[2020](https://arxiv.org/html/2607.15442#bib.bib22)\)and related datasets\(Caoet al\.[2022](https://arxiv.org/html/2607.15442#bib.bib53)\), and fine\-tuning CLIP and similar models to improve multimodal semantics\(Heeet al\.[2022](https://arxiv.org/html/2607.15442#bib.bib120); Thapaet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib138)\)\. These models added object\-, face\-, or text\-level cues to detect explicit harm, yet they still treat harmfulness as a single signal and often fail when hate is masked by humor, satire, or cultural nuance\.

On the other hand, we reframe meme understanding as a multiple\-angle reasoning problem\. Instead of compressing diverse cues into a single embedding, we decompose interpretation into twelve complementary perspectives, enabling the model to differentiate benign humor from ambiguous or veiled hostility, something prior methods struggle with\. Through a role\-aware soft\-attention mechanism, MAR\-12 learns which perspectives are most informative for humorous, hateful, or mixed\-intent memes, yielding flexible, transparent, and human\-aligned predictions\. This holistic treatment directly mitigates the shortcomings of unimodal, multimodal, and black\-box VLM\-based systems by capturing the nuanced interplay between humor and harm\.

### Reasoning\-Based Meme Understanding

Recent work increasingly uses LLM\- and VLM\-based approaches to enhance interpretability in harmful or ambiguous memes\(Hee and Lee[2025](https://arxiv.org/html/2607.15442#bib.bib19); Rizwanet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib136)\)\. Few\-shot and zero\-shot systems such as\(Rizwanet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib136)\)and MiND\(Liuet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib79)\), LoReHM\(Huanget al\.[2024](https://arxiv.org/html/2607.15442#bib.bib46)\), and related reasoning\-oriented methods\(Linet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib80)\)employ visual–textual entailment or chain\-of\-thought prompting to explain why a meme may be harmful\. Debate\-inspired frameworks like ExplainHM\(Linet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib76)\)simulate opposing viewpoints and use a judge model to consolidate arguments\. Although generative reasoning approaches show promise, they remain constrained by structural assumptions: debate\-style systems reduce interpretation to a binary contest, while single\- or dual\-chain explanations cannot capture the multidimensional space needed to model humor and hate co\-occurrence, affective contrast, cultural grounding, or linguistic play\.

Rather than relying on one or two reasoning trajectories, our MAR\-12 encourage the VLM to reflect on a meme through twelve complementary perspectives spanning foundational content, humor mechanisms, and social\-meaning cues\. This structured decomposition avoids the binary framing of debate\-style methods and better captures the diverse ways humor and harm interact\. A role\-aware soft\-attention mechanism further learns which perspectives matter most for each input, enabling flexible, context\-sensitive explanations that highlight ambiguity, competing cues, and mixed intent\. By modeling humorous and harmful signals in a unified and interpretable manner, our approach produces explanations that are more faithful, nuanced, and aligned with human interpretive processes than prior generative reasoning systems\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/x2.jpeg)Figure 2:Overview of the proposed Multiple\-Angle Reasoning \(MAR\) framework\. Stage 1 Perspective\-Conditioned Prompting\. Stage 2 Role\-Aware Soft\-Gated Attention\. Stage 3 Explanation via LLM\-based Judemeny\.

## Methodology

To model the complex and often conflicting interpretive space involved in meme understanding, we introduce MAR\-12, a framework that encourages a VLM to reflect on a meme through a structured set of complementary perspectives\. Rather than forcing the model into fixed behaviors, MAR\-12 acts as a form of structured cognitive scaffolding, analogous to how a human might be prompted to “look at it from another angle” or “put yourself in someone else’s shoes\.” These perspectives provide breadth, while a role\-aware soft\-gated attention mechanism determines which angles matter most for each meme, ensuring freedom in reasoning without hard\-coding outcomes\. As illustrated in Figure[2](https://arxiv.org/html/2607.15442#Sx2.F2), MAR\-12 consists of three stages: \(1\)*Perspective\-conditioned prompting*, \(2\)*Role\-aware soft\-gated attention*, and \(3\)*Explanation via LLM\-Based Judgment*\.

### Perspective\-Conditioned Prompting

Given a memem=\(I,T\)m=\(I,T\)consisting of an imageIIand associated textTT, MAR\-12 interprets it through a structured set of twelve perspectives𝒫=\{p1,…,p12\}\\mathcal\{P\}=\\\{p\_\{1\},\\dots,p\_\{12\}\\\}\. Each perspective corresponds to a distinct interpretive lens grounded in linguistic humor theory, affective contrast modeling, and harmful\-speech analysis\.

#### Role\-Conditioned Prompting\.

To elicit each reasoning angle, we prompt the VLM by assigning it a perspective\-specific role, analogous to asking a human annotator to temporarily adopt a certain interpretive mindset \(“consider the cultural context”, “analyze irony”, “evaluate harmful intent”\)\. Formally, the response for perspectivepip\_\{i\}generates a natural\-language reasoning sequence, i\.e\.,

ri=𝒬θ​\(I,T,pi\),r\_\{i\}=\\mathcal\{Q\}\_\{\\theta\}\(I,T,p\_\{i\}\),\(1\)whererir\_\{i\}is a natural\-language reasoning sequence specific to viewpointpip\_\{i\}\. Collecting all twelve responses yields:

ℛm=\{r1,r2,…,r12\}\.\\mathcal\{R\}\_\{m\}=\\\{r\_\{1\},r\_\{2\},\\dots,r\_\{12\}\\\}\.\(2\)Eachrir\_\{i\}provides a viewpoint\-specific interpretation, and together they form a diverse basis for downstream aggregation\. This stage does not impose constraints on which perspectives the model should rely on; instead, it encourages intellectual breadth while allowing Stage 2 to determine relevance\. Due to space constraints, the full prompts for all twelve perspectives are provided in Appendix[B](https://arxiv.org/html/2607.15442#A2)\(see Table[10](https://arxiv.org/html/2607.15442#A2.T10)\)\. The appendix describes the rationale, structure, and exact wording of each prompt, allowing readers to fully reproduce our perspective\-conditioned prompting setup\.

### Role\-Aware Multi\-Angle Fusion

To integrate the twelve reasoning traces, we employ a role\-aware soft\-gated attention mechanism that differentiates between the functional roles of each perspective \(e\.g\., humor vs\. safety\)\. Each reasoning sequencerir\_\{i\}is encoded into a fixed\-dimensional embedding𝐡i∈ℝd\\mathbf\{h\}\_\{i\}\\in\\mathbb\{R\}^\{d\}\. Unlike standard attention which treats inputs as interchangeable, we introduce a role embedding𝐞i∈ℝd\\mathbf\{e\}\_\{i\}\\in\\mathbb\{R\}^\{d\}that captures the specific conceptual function of perspectivepip\_\{i\}\. The relevance scoreuiu\_\{i\}is computed by integrating the reasoning content with its functional role:

ui=𝐰⊤​tanh⁡\(𝐖h​𝐡i\+𝐖e​𝐞i\+𝐛\),u\_\{i\}=\\mathbf\{w\}^\{\\top\}\\tanh\(\\mathbf\{W\}\_\{h\}\\mathbf\{h\}\_\{i\}\+\\mathbf\{W\}\_\{e\}\\mathbf\{e\}\_\{i\}\+\\mathbf\{b\}\),\(3\)where𝐖h\\mathbf\{W\}\_\{h\},𝐖e\\mathbf\{W\}\_\{e\},𝐰\\mathbf\{w\}, and𝐛\\mathbf\{b\}are learnable parameters\. The normalized weightsαi\\alpha\_\{i\}ensure the contribution of each perspective is context\-dependent, yielding an aggregated representation

𝐡agg=∑i=112αi​𝐡i\.\\mathbf\{h\}\_\{\\text\{agg\}\}=\\sum\_\{i=1\}^\{12\}\\alpha\_\{i\}\\mathbf\{h\}\_\{i\}\.\(4\)This representation is then concatenated with the visual embedding𝐯\\mathbf\{v\}to form the final multimodal vector𝐳=\[𝐡agg∥𝐯′\]\\mathbf\{z\}=\[\\mathbf\{h\}\_\{\\text\{agg\}\}\\parallel\\mathbf\{v\}^\{\\prime\}\], which is passed to a prototype\-based classifier\.

#### Prototype\-Based Classification\.

To promote geometric interpretability in decision\-making, We employ a cosine prototype\-based decision function, which is computed as

y^=arg⁡maxk⁡cos⁡\(𝐳,𝐜k\)\.\\hat\{y\}=\\arg\\max\_\{k\}\\ \\cos\(\\mathbf\{z\},\\mathbf\{c\}\_\{k\}\)\.\(5\)where each classk∈\{1,2\}k\\in\\\{1,2\\\}\(humor vs\. non\-humor or hate vs\. non\-hate\) is represented by a learnable prototype𝐜k\\mathbf\{c\}\_\{k\}\.

### Explanation Synthesizer

Beyond producing a binary prediction, MAR\-12 generates a final explanatory rationale that articulates*why*the model arrived at its decision\. This step is essential for transparency, especially in cases where humorous cues and harmful intent intersect in subtle or conflicting ways\. To construct this explanation, we repurpose the VLM as an explanation synthesizer\. The model is prompted with three key pieces of information: the predicted labely^\\hat\{y\}, the set of multi\-angle reasoning outputsℛm=\{r1,…,r12\}\\mathcal\{R\}\_\{m\}=\\\{r\_\{1\},\\dots,r\_\{12\}\\\}, and the learned attention weights\{αi\}i=112\\\{\\alpha\_\{i\}\\\}\_\{i=1\}^\{12\}that indicate the relevance of each perspective\.

To ensure that the explanation is grounded in the same reasoning process used during classification, MAR\-12 provides VLM with a structured instruction prompt that explicitly references both the content of each reasoning angle and its learned importance\. The model is guided using a template of the following form:

> “You are given a meme classification result\. The final prediction is: \[LABEL\]\. Below are the interpretations of the meme from multiple reasoning angles, each accompanied by its relevance weight: \[REASONING\_TEXT \+ WEIGHTS\]\. Based on these weighted perspectives, summarize the key factors that led to the decision\. Explain which perspectives were most influential, which were less relevant, and how they collectively justify the final classification\.”

Formally, the explanation is produced as

Explanation=𝒬θexp​\(y^,ℛm,\{αi\}\),\\text\{Explanation\}=\\mathcal\{Q\}^\{\\text\{exp\}\}\_\{\\theta\}\(\\hat\{y\},\\mathcal\{R\}\_\{m\},\\\{\\alpha\_\{i\}\\\}\),\(6\)where the concatenated, weight\-conditioned reasoning signals serve as inputs to the explanation module\. This design ensures that the explanation meaningfully reflects the same multi\-angle reasoning used by the classifier, rather than being a generic or post\-hoc justification\. The resulting explanation highlights the dominant interpretive angles driving the prediction, acknowledges conflicting or low\-weight cues when necessary, and provides a human\-readable account of how humor or harmful meaning was inferred\. This final stage ensures that MAR\-12 functions not merely as a classifier but as a transparent reasoning system whose decisions can be examined and understood\. By conditioning the synthesis on the same role\-aware attention used by the classifier, MAR\-12 achieves process\-level faithfulness, where the explanation remains a transparent account of the model’s internal evidence flow even in cases of misclassification\.

Table 1:Statistical distributions of datasets where “Hu” and “Ha” represent humor and hateful, “Non\-Hu” and “Non\-Ha” represent non\-humor and non\-hatefulTaskSplitLabelPrideMMMemotionHumorTrainHu29444272Non\-Hu13851321TestHu3181069Non\-Hu190330HateTrainHa21213423Non\-Ha22082170TestHa248856Non\-Ha260543

Table 2:Comparison of LLaVA and Qwen\-VL on humor classification with and without few\-shot reasoning\.ModelACC \(%\)AUROC \(%\)Δ\\DeltaACC \(%\)LLaVA67\.4665\.75–L\-Reasoning63\.1256\.69−\-4\.34Qwen\-VL48\.7257\.41–Q\-Reasoning58\.9760\.44\+10\.25

## Experiments Settings

Table 3:Performance comparison across humor and hate meme detection tasks for PrideMM and Memotion datasets\.TaskModelPrideMMMemotionACC \(%\)AUC \(%\)F1 \(%\)ACC \(%\)AUC \(%\)F1 \(%\)HumorVisual Only \(Resnet50 \+ MLP\)66\.0858\.0161\.6776\.4850\.5766\.88Text Only \(T5 \+ MLP\)67\.8562\.3866\.1076\.2750\.6467\.02MemeCLIP\(Shahet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib87)\)78\.3073\.2776\.9976\.3451\.3167\.77MOMENTA\(Pramanicket al\.[2021](https://arxiv.org/html/2607.15442#bib.bib90)\)73\.5765\.7969\.9276\.4850\.00†66\.50PromptHate\(Caoet al\.[2022](https://arxiv.org/html/2607.15442#bib.bib53)\)73\.7771\.0173\.4676\.4850\.00†66\.19LoReHM \(LLaVA\-34B\)\(Huanget al\.[2024](https://arxiv.org/html/2607.15442#bib.bib46)\)70\.0956\.7164\.0776\.1850\.9167\.32MiND \(Qwen2\.5\-VL\-32B \)\(Liuet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib79)\)54\.4551\.0050\.4360\.7252\.3953\.27MAR\-12 \(T5 Embedding\)79\.6875\.1178\.6379\.1074\.5177\.85MAR\-12 \(Clip Embedding\)80\.0877\.6479\.8279\.8577\.0478\.88MAR\-12 \(T5xClip Embedding\)80\.2876\.0079\.3779\.1575\.8478\.38HateVisual Only \(Resnet50 \+ MLP\)62\.7262\.8962\.5761\.4750\.00†46\.80Text Only \(T5 \+ MLP\)72\.7872\.7772\.7859\.6151\.9554\.46MemeCLIP\(Shahet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib87)\)75\.3575\.0975\.3561\.7650\.9949\.58MOMENTA\(Pramanicket al\.[2021](https://arxiv.org/html/2607.15442#bib.bib90)\)75\.1575\.2275\.1461\.4750\.00†46\.80PromptHate\(Caoet al\.[2022](https://arxiv.org/html/2607.15442#bib.bib53)\)71\.8071\.6671\.7061\.4750\.00†46\.81LoReHM \(LLaVA\-34B\)\(Huanget al\.[2024](https://arxiv.org/html/2607.15442#bib.bib46)\)65\.0565\.1664\.9553\.7252\.4354\.13MiND \(Qwen2\.5\-VL\-32B \)\(Liuet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib79)\)61\.7360\.8761\.6573\.0568\.4670\.59MAR\-12 \(T5 Embedding\)75\.3575\.3675\.3576\.1570\.2071\.12MAR\-12 \(Clip Embedding\)73\.3773\.8172\.6675\.5373\.4271\.05MAR\-12 \(T5xClip Embedding\)73\.1873\.4772\.8675\.8974\.0472\.00
†An AUC of 50\.00 alongside accuracy levels matching the majority class \(e\.g\., 76\.48% for humor or 61\.47% for hate\) indicates that the baseline collapsed to predicting the majority label for all samples, thus failing to learn discriminative rank\-ordering\.

### Evaluation Datasets

To evaluate MAR\-12 in settings where humor and harmful meaning frequently interact, we require datasets that contain both humor and hate annotations for the same meme\. However such datasets are extremely scarce: among existing meme benchmarks, only two publicly available resources, i\.e\., PrideMM\(Shahet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib87)\)and Memotion\(Sharmaet al\.[2020](https://arxiv.org/html/2607.15442#bib.bib52)\), provide dual labels that allow a model to study humor and hate co\-occurrence on the same image–text pair\. This makes them uniquely suited for evaluating multi\-angle reasoning approaches such as MAR\-12, where the goal is to disentangle humorous mechanisms from potentially harmful or offensive undertones\.

*PrideMM*focuses on LGBTQ related memes, many of which combine satire, wordplay, and cultural references with socially or politically charged commentary\. This mixture often produces content where humor can either soften or conceal hostile messaging, making it a strong testbed for evaluating MAR\-12’s ability to analyze competing interpretive cues\.

*Memotion*provides meme annotations across humor, sarcasm, and offensiveness\. Its broader thematic coverage introduces a diverse range of humorous styles, absurd, sarcastic, situational, and text\-dominant, while also capturing offensiveness that may emerge independently or in parallel with humor\. For experimentation, we binarize humor by grouping*funny*,*very funny*, and*hilarious*as humorous and*not funny*as non\-humorous\. Offensive labels are treated as hateful\.

Table 4:Summary of baseline models and our proposed MAR\-12 method, with total and trainable parameter counts\.ModelTotal ParametersTrainable ParametersResNet\-50 \+ MLP \(Visual Baseline\)25\.48M1\.97MT5 \+ MLP \(Text Baseline\)110\.02M110\.02MPromptHate355\.41M355\.41MMemeCLIP431\.29M3\.68MMOMENTA431\.29M3\.68MLoReHM34\.00B34\.00BMAR\-12 \(Ours\)1\.31M1\.31M

### Models

We compare MAR\-12 against three families of strong baselines to demonstrate its advantage over unimodal, multimodal, and reasoning\-based approaches\.\(i\) Unimodal baselinesevaluate the contribution of isolated modalities: a Visual\-Only model \(ResNet50 \+ MLP\) and a Text\-Only model \(T5 \+ MLP\)\.\(ii\) Multimodal CLIP\-style baselinesincludeMOMENTA\(Pramanicket al\.[2021](https://arxiv.org/html/2607.15442#bib.bib90)\), which enriches CLIP features with object\- and face\-level cues from VGG\-19 and textual features from DistilBERT,MemeCLIP\(Shahet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib87)\), which incorporates trainable adapters into CLIP for multi\-aspect meme classification while preserving its generalization capabilities, andPromptHate\(Caoet al\.[2022](https://arxiv.org/html/2607.15442#bib.bib53)\), which reformulates meme classification as masked language modeling with prompt templates using RoBERTa\.\(iii\) Reasoning\-based VLM baselinesincludeLoReHM\(Huanget al\.[2024](https://arxiv.org/html/2607.15442#bib.bib46)\), which uses a single LLM agent to generate reasoning chains through visual\-textual entailment and commonsense inference, andMiND\(Liuet al\.[2025](https://arxiv.org/html/2607.15442#bib.bib79)\), which incorporate prompting or chain\-of\-thought reasoning but do not structure reasoning across diverse interpretive angles as MAR\-12 does\. This grouping highlights that MAR\-12 improves over systems designed around single cues, multimodal matching, and generic reasoning alike\.

As summarized in Table[4](https://arxiv.org/html/2607.15442#Sx4.T4), MAR\-12 is designed to be an extremely lightweight framework, containing only 1\.31 million trainable parameters\. This is orders of magnitude smaller than competitive reasoning\-based baselines like LoReHM, which relies on a 34\.00 billion parameter Large Multimodal Model\. This efficiency is achieved by keeping the heavy VLM and feature encoders frozen, ocusing the learning process solely on the role\-aware attention mechanism and the prototype\-based classification head\. Consequently, MAR\-12 provides a high\-performance alternative to end\-to\-end fine\-tuning, offering significant depth in reasoning without the prohibitive computational costs and memory requirements associated with large\-scale model optimization\.

### Implementation Details

#### Vision Language Models\.

Before evaluating MAR\-12 directly, we first assess which VLM is most responsive to multi\-angle reasoning prompts\. We compare LLaVA\(Liuet al\.[2023](https://arxiv.org/html/2607.15442#bib.bib142)\)and Qwen\-VL\(Baiet al\.[2023](https://arxiv.org/html/2607.15442#bib.bib96)\)under two settings: \(1\) direct classification with a single prompt and \(2\) classification augmented with few\-shot reasoning exemplars\. In the few\-shot setting, each exemplar includes a meme image, its label, and a GPT\-4\(Achiamet al\.[2023](https://arxiv.org/html/2607.15442#bib.bib83)\)generated explanation, allowing us to examine whether the model can integrate external reasoning cues\. As shown in Table[2](https://arxiv.org/html/2607.15442#Sx3.T2), LLaVA performs best in the direct setting \(67\.46% ACC\) but degrades when reasoning exemplars are introduced, suggesting difficulty in leveraging structured interpretive signals\. Qwen\-VL, however, improves from 48\.72% to 58\.97% accuracy and from 57\.41% to 60\.44% AUC, demonstrating a better performance in internalizing structured reasoning\. These results support our choice of Qwen\-VL as the backbone for MAR\-12, which relies heavily on high\-quality reasoning across multiple interpretive perspectives\.

#### MAR\-12 Encoders\.

We use CLIP ViT‑B/32 as the vision encoder, with image inputs resized to 224×224 pixels\. The agent text embeddings were obtained using T5 \(or CLIP text encoder in some variants\) and projected into a 1024‑dimensional mapped space from the original 768‑dimensional embeddings\. The fusion module included 1 mapping layer and 1 pre‑output layer, with dropout probabilities of 0\.1, 0\.4, and 0\.2 applied at different stages\. The classification head employed a cosine‑similarity‑based prototype classifier with a scale of 30 and margin ratio of 0\.2\.

#### MAR\-12 Training and Efficiency\.

We trained MAR\-12 on a high\-performance computing server equipped with NVIDIA L40S GPUs \(46,068 MiB VRAM\) and dual AMD EPYC 7543 32\-Core Processors\. To ensure computational efficiency and scalability for large\-scale meme moderation, we utilized mixed precision training \(torch\.float16\), which significantly optimized memory utilization and reduced the overall training footprint\. We used 4 data‑loading workers and fixed the random seed to 42 for reproducibility\. We used a learning rate of1×10−41\\times 10^\{\-4\}, weight decay of1×10−41\\times 10^\{\-4\}, and a batch size of 32 to train MAR\-12 over 20 epochs\. Training the model for 20 epochs on the Memotion dataset took less than 30 minutes, with an average inference time of 3\.5 seconds per meme\. This inference latency accounts for the parallelized processing of the twelve VLM reasoning traces and the final synthesis by the explainer\. A comprehensive breakdown of the computing resources, software environment \(CUDA 12\.5\), and specific training configurations is available in the Appendix\. As for the evaluation of the models’ performance, we base our choices on the average of their Accuracy \(Acc\.\), Area Under the Receiver Operating Characteristics curve\(AUROC\), and F1\-score\. We optimize these models using Adam optimizer\(Kingma[2014](https://arxiv.org/html/2607.15442#bib.bib132)\)and are implemented in PyTorch using the Huggingface’sTransformers111https://huggingface\.co/docs/transformerslibrary\.

## Experiments

### Performance Comparison with Baselines

After selecting Qwen\-VL as the backbone, we evaluate MAR\-12 against unimodal, multimodal, and reasoning\-based baselines\. As shown in Table[3](https://arxiv.org/html/2607.15442#Sx4.T3), MAR\-12 achieves state\-of\-the\-art results across both datasets and tasks\. For humor detection on PrideMM, MAR\-12 \(CLIP\) attains 80\.08% ACC, 77\.64% AUC, and 79\.82% F1, with the T5\+CLIP fusion variant slightly improving accuracy further\. On Memotion, MAR\-12 again outperforms all baselines, reaching 79\.85% ACC and 78\.88% F1\. Similar improvements appear in the hate detection task: MAR\-12 \(T5\) matches or exceeds the strongest baselines on PrideMM with the highest AUC \(75\.36%\), and MAR\-12 \(T5xCLIP\) achieves the best overall performance on Memotion \(74\.04% AUC and 72\.00% F1\)\. These results demonstrate that MAR\-12 surpasses unimodal systems, multimodal CLIP\-style models, and reasoning\-based VLMs alike\. Notably, several baselines such as MOMENTA and PromptHate exhibit a performance collapse on the Memotion dataset, achieving high accuracy \( 76%\) but random\-level AUC \(50\.00\)\. Our audit confirmed this was due to the models defaulting to majority\-class prediction as a result of the dataset’s class imbalance \(see Table 1\)\. MAR\-12 successfully overcomes this limitation by leveraging multi\-angle reasoning to maintain high discriminative power \(77\.04% AUC on humor\)

### Ablation Studies

To better understand the contribution of multi\-angle reasoning and the role\-aware attention mechanism within MAR\-12, we conduct two ablation studies\. The first evaluates the importance of different perspective groups, while the second examines how various attention mechanisms influence the aggregation of perspective\-specific reasoning\.

#### Effect of Removing Perspective Groups\.

MAR\-12 organizes its twelve reasoning angles into three functional categories: a*humor perspective group*\(capturing irony, absurdity, wordplay, and comedic framing\), a*safety perspective group*\(capturing harmfulness, intent, stereotyping, and derogatory cues\), and a*foundational perspective group*\(capturing OCR text, literal description, and basic semantic grounding\)\. Table[5](https://arxiv.org/html/2607.15442#Sx5.T5)shows how removing the humor or safety groups affects performance\. For humor classification, MAR\-12 with all perspectives achieves 79\.85% accuracy and 77\.64% AUC\. However, removing either the humor or safety perspectives causes AUC to collapse to 50\.00%, equivalent to random guessing\. A similar pattern occurs in the hate task: full MAR\-12 obtains 75\.53% accuracy and 73\.42% AUC, but removing humor or safety perspectives again reduces AUC to approximately 50%\.

To ensure a rigorous evaluation of each functional category, we completely retrained the Stage 2 role\-aware attention module and prototype\-based classifier from scratch for each ablation setting\. We did not merely mask the inputs at inference time; instead, the model was forced to attempt to learn a discriminative decision boundary using only the remaining perspective groups\. The collapse of the AUROC to 50\.00 \(random guessing\) when either the humor or safety groups are removed demonstrates that these specialized reasoning signals are indispensable for the learning process\. This confirms that foundational perspectives alone are insufficient for capturing the nuanced interplay of humor and harm in memes\. These findings indicate that when key interpretive groups are removed, the model is forced to rely solely on the foundational perspectives, which provide surface\-level cues but lack the task\-specific discriminative information needed for complex reasoning about humor or harmful intent\. Without the humor or safety groups, the role\-aware attention cannot meaningfully differentiate between perspectives and thus defaults to majority\-class predictions\. The near\-identical drops observed when removing either group also highlight that humor and safety cues function complementarily in detecting humor and hate co\-occurrence, reinforcing the need for a multi\-angle interpretive structure\.

#### Effect of Attention Mechanisms\.

We next evaluate whether MAR\-12’s improvements stem from the use of role\-aware attention or whether simpler attention variants suffice\. Table[6](https://arxiv.org/html/2607.15442#Sx5.T6)compares three alternatives, standard soft attention, gated attention, and multi\-head attention, using both MLP and ResNet visual backbones\. Soft attention achieves moderate performance but shows instability across backbones, as it treats all reasoning embeddings as interchangeable and cannot account for the conceptual distinctions between humor, safety, and foundational perspectives\. Gated attention introduces additional nonlinearity but still struggles to encode differences tied to perspective roles, yielding only marginal improvements\. Multi\-head attention performs worst, particularly with the MLP backbone, as its distributed heads tend to diffuse focus across many perspectives, preventing stable alignment with role semantics\.

Table 5:Ablation study evaluating MAR\-12 when individual perspective groups are removed\.TaskSettingACC \(%\)AUC \(%\)HuMAR\-1279\.8577\.64w/o Humor Perspectives76\.4150\.00w/o Safety Perspectives76\.4150\.00HaMAR\-1275\.5373\.42w/o Humor Perspectives61\.4750\.03w/o Safety Perspectives61\.4750\.03Table 6:Ablation study on the role of attention and backbone architecture in MAR\-12\.MethodHumorHateACC\(%\)ACC\(%\)MLP\+Soft Attention75\.7470\.79ResNet\+Soft Attention74\.1667\.11MLP\+Gated Attention76\.7371\.9ResNet\+Gated Attention76\.9272\.58MLP\+Multi\-head Attention72\.9866\.90ResNet\+Multi\-head Attention74\.9570\.27MLP\+Role\-aware Attention \(ours\)79\.3675\.11ResNet\+Role\-aware Attention \(ours\)78\.7875\.27

In contrast, MAR\-12’s role\-aware attention achieves the highest and most consistent performance across both tasks and backbones\. By explicitly incorporating learned role embeddings for each perspective group, the mechanism preserves functional distinctions between different types of reasoning and enables the model to prioritize complementary cues\. This prevents the diffusion of attention seen in multi\-head variants and avoids the homogenization of embeddings seen in soft attention\. With these role distinctions encoded, MAR\-12 effectively leverages the diverse reasoning signals produced by the perspective\-conditioned prompts, leading to superior generalization and more robust interpretive alignment across datasets\.

Table 7:Automatic GPT\-4 evaluation of the explanation quality of humor and hate tasks on Memotion test sets\.TaskHaNon\-HaHuNon\-HuHuNon\-HuInformativeness3\.383\.463\.813\.79Readability3\.963\.974\.114\.05Soundness3\.643\.793\.763\.69Conciseness3\.903\.923\.923\.94Persuasiveness3\.793\.823\.863\.84Table 8:Human evaluation of the explanation quality of humor and hate tasks on Memotion test sets\.TaskHaNon\-HaHuNon\-HuHuNon\-HuInformativeness4\.034\.023\.883\.75Readability3\.783\.733\.653\.47Soundness3\.733\.673\.633\.48Conciseness3\.884\.124\.003\.85Persuasiveness3\.803\.733\.683\.77

## Empirical Analysis

Since MAR\-12’s core contribution is interpretable multi\-angle reasoning, we evaluate the quality of its explanations using both automatic and human judgments\. Explanations in this setting must integrate diverse reasoning perspectives, reflect the learned importance weights, and articulate why humor or hate was inferred from a multimodal meme\. We therefore assess not only linguistic fluency but also whether the explanations meaningfully reflect MAR1\-2’s structured interpretive process\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/error_analysis.png)Figure 3:Example memes from four scenarios: \(a\) hateful and humorous, \(b\) hateful but non\-humorous, \(c\) non\-hateful but humorous, and \(d\) non\-hateful and non\-humorous\. Correct \(inGreen\) and incorrect \(inRed\) predictions of our method\. Important information is inBold\.### Evaluation of Explainability

#### Automatic Evaluation\.

Following recent practice\(Linet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib76)\), we use GPT\-4 as an automatic evaluator and score each explanation along five criteria that relate directly to MAR\-12’s design principles\.*Informativeness*captures whether the explanation introduces additional contextual or cultural insight beyond surface\-level observations, reflecting the breadth of MAR\-12’s multi\-angle reasoning\.*Readability*assesses fluency and structural coherence\.*Soundness*evaluates whether the explanation presents a logically grounded interpretation consistent with the role\-aware attention mechanism\.*Conciseness*measures the model’s ability to summarize relevant reasoning without redundancy\. Finally,*Persuasiveness*reflects whether the explanation forms a compelling justification aligned with the aggregated reasoning signals\. Each criterion is rated on a 5\-point Likert scale\.

To examine how MAR\-12 handles different humor–hate interactions, we evaluate explanations across four interpretive regimes: hateful and humorous, hateful but non\-humorous, non\-hateful but humorous, and neither hateful nor humorous\. Table[7](https://arxiv.org/html/2607.15442#Sx5.T7)reports averaged GPT\-4 scores for each category\. We observe that explanations remain consistently strong across most criteria, particularly in Readability \(3\.96–4\.11\) and Conciseness \(3\.90–3\.94\), indicating that MAR\-12 produces structurally coherent and succinct justifications\. Soundness is similarly stable \(3\.64–3\.79\), suggesting that role\-aware attention effectively integrates multiple reasoning angles into a coherent interpretive thread\. This stability is further supported by our quantitative alignment analysis, which proves the synthesizer’s narrative is causally grounded in the model’s internal attention weights rather than being a post\-hoc justification of the final label\. The most challenging regime is hateful\-and\-humorous memes, which score lowest on Informativeness \(3\.38\)\. These cases often require additional cultural or contextual grounding that may not be explicit in the meme, reflecting the inherent difficulty of reconciling competing humorous and harmful cues\.

#### Human Evaluation\.

Because automatic evaluation cannot fully capture the subjectivity and cultural nuance involved in meme interpretation\(Linet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib76)\), we complement it with a human evaluation conducted by five English\-proficient university students using the same five criteria\. Table[8](https://arxiv.org/html/2607.15442#Sx5.T8)summarizes the results\. Inter\-annotator agreement is moderate \(ICC = 0\.618; Spearman’sρ\\rho= 0\.643\), consistent with the subjective nature of humor and harm perception\. Human judgments reveal patterns similar to the automatic evaluation: explanations for hateful\-but\-non\-humorous memes receive the highest Soundness and Conciseness scores, likely because harmful cues tend to be explicit\. Explanations for humorous\-but\-non\-hateful memes achieve the best Readability and Persuasiveness scores, suggesting that MAR\-12’s humor\-focused reasoning is perceived as fluent and coherent\. Across all criteria, memes containing both humor and hate exhibit slightly lower scores, further confirming the interpretive difficulty of cases involving ambiguous or dual communicative intent\.

Overall, both automatic and human evaluations indicate that MAR\-12 produces explanations that are coherent, grounded, and well\-aligned with its multi\-angle reasoning framework\. The model reliably synthesizes diverse interpretive cues into concise and persuasive justifications\. The most challenging cases involve humor and hate co\-occurrence, where cultural or contextual subtleties may exceed what is explicitly represented in the meme\. These findings highlight the importance of multi\-angle reasoning: even when faced with conflicting signals, MAR\-12 maintains stable explanatory quality while making explicit the interpretive tensions that shape the final prediction\.

Table 9:Pearson correlations\. Alignment captures faithfulness to agent reasoning/attention rather than correctness, whereas margin correlates with correctness as expected\.PairPearsonrralignhumour\\mathrm\{align\_\{humour\}\}vs\.humcorrect\\mathrm\{hum\_\{correct\}\}0\.0300\.030alignhate\\mathrm\{align\_\{hate\}\}vs\.hatecorrect\\mathrm\{hate\_\{correct\}\}−0\.026\-0\.026marginhumour\\mathrm\{margin\_\{humour\}\}vs\.humcorrect\\mathrm\{hum\_\{correct\}\}0\.2410\.241marginhate\\mathrm\{margin\_\{hate\}\}vs\.hatecorrect\\mathrm\{hate\_\{correct\}\}0\.2670\.267alignhumour\\mathrm\{align\_\{humour\}\}vs\.alignhate\\mathrm\{align\_\{hate\}\}0\.4250\.425marginhumour\\mathrm\{margin\_\{humour\}\}vs\.marginhate\\mathrm\{margin\_\{hate\}\}−0\.053\-0\.053alignhumour\\mathrm\{align\_\{humour\}\}vs\.marginhumour\\mathrm\{margin\_\{humour\}\}−0\.045\-0\.045alignhate\\mathrm\{align\_\{hate\}\}vs\.marginhate\\mathrm\{margin\_\{hate\}\}−0\.033\-0\.033

### Quantitative Faithfulness

To move beyond perceived quality, we conduct a quantitative analysis of explanation faithfulness with an Alignment metric that measures the textual grounding of the Explanation Synthesizer in the twelve weighted reasoning traces, weighted by the model’s learned attention\. As shown in Table[9](https://arxiv.org/html/2607.15442#Sx6.T9), our analysis reveals that faithfulness is uncorrelated with prediction accuracy \(r=0\.030r=0\.030for humor;r=−0\.026r=\-0\.026for hate\)\. The results in Table[9](https://arxiv.org/html/2607.15442#Sx6.T9)verifies that the synthesizer faithfully narrates the internal evidential path even when the classifier is incorrect\.

Furthermore, the consistency of alignment values across tasks \(r=0\.425r=0\.425\) indicates that MAR\-12 induces a stable, role\-aware explanatory pattern that tracks internal attention rather than inventing its own saliency to match a label\. Extended analysis is included in Appendix[E](https://arxiv.org/html/2607.15442#A5)\.

### Case Study Analysis

To illustrate how MAR\-12 navigates competing humorous and harmful cues, we analyze four representative memes spanning the full humor and hate interpretive space, as shown in Figure[3](https://arxiv.org/html/2607.15442#Sx6.F3)\. These examples highlight how multi\-angle reasoning and perspective\-specific attention weights\{αi\}\\\{\\alpha\_\{i\}\\\}shape the final predictions, particularly in cases involving ambiguous communicative intent\.

From the natural\-language explanations produced by MAR\-12, we observe two recurring patterns\.Firstly, the coexistence of humor and hate substantially increases classification difficulty\. MAR\-12 must determine whether humor is benign, whether it masks harmful messaging, or whether it amplifies derogatory undertones\. This tension is evident in Figure[3](https://arxiv.org/html/2607.15442#Sx6.F3)\(a\), where humor\-focused perspectives emphasize the absurd juxtaposition of Mr\. Mime, a Pokémon known for creating invisible walls, with the slogan “That’s my wall,” referencing Donald Trump’s political rhetoric\. In contrast, safety\-focused perspectives attend to the derogatory implications embedded in the comparison\. The learned weightsαi\\alpha\_\{i\}must reconcile these conflicting cues; when they lean toward humorous framing, the model favors a benign interpretation, but when safety cues dominate, it signals harmful intent\. A similar conflict occurs in Figure[3](https://arxiv.org/html/2607.15442#Sx6.F3)\(c\), where the humorous exaggeration of exhaustion is correctly identified by humor perspectives, yet safety perspectives overweight hardship\-related visual elements, leading the model to infer harmful undertones despite the meme being non\-hateful\.

Secondly, misclassifications often arise from the inherently subjective and culturally dependent nature of humor and hate annotations\. Memes involving sensitive themes, gender, race, religion, or political ideology, can elicit divergent interpretations across annotators\. This is reflected in Figure[3](https://arxiv.org/html/2607.15442#Sx6.F3)\(b\) and[3](https://arxiv.org/html/2607.15442#Sx6.F3)\(d\)\. Although annotators labeled both memes as humorous but non\-hateful, MAR\-12 predicted them as humorous but hateful\. In Figure[3](https://arxiv.org/html/2607.15442#Sx6.F3)\(b\), safety\-focused perspectives correctly detect the alignment between a historical dictator and a sensitive textual reference, capturing a legitimate basis for concern\. In Figure[3](https://arxiv.org/html/2607.15442#Sx6.F3)\(d\), the model recognizes the situational humor conveyed by facial expression and caption, yet elevated attention to socio\-political cues leads to a cautious interpretation that diverges from annotator consensus\. These cases demonstrate how small shifts inαi\\alpha\_\{i\}can alter interpretive emphasis, especially when memes rely on subtle social nuance or culturally specific humor\.

Taken together, these examples reveal the real\-world interpretive complexity of memes where humor and harmful meaning intersect\. They also demonstrate a multi\-angle, role\-aware reasoning framework like MAR\-12 can expose the sources of ambiguity, articulate competing interpretations, and justify its final decision in a transparent manner\.

## Conclusion

We present MAR\-12, a multi\-angle reasoning framework for the joint detection and explanation of humorous and hateful content in multimodal memes\. By decomposing meme understanding into twelve theory\-driven analytical perspectives, our approach addresses the intertwined challenges of performance and interpretability in meme moderation\. Through extensive experiments on the PrideMM and Memotion datasets, we demonstrate that MAR\-12 consistently outperforms existing multimodal and LLM\-based baselines in both detection accuracy and explanation quality\. In addition, our human and automatic evaluations show that MAR\-12 produces coherent, context\-grounded justifications, particularly for challenging cases where humor, sarcasm, and harmful intent co\-occur\.

## Limitations and Future Work

There are several directions in which this work can be further improved: 1\) MAR\-12 relies on predefined analytical perspectives derived from humor and hate theory\. While these perspectives capture many common cues, they may not fully represent the diversity of cultural references, regional slang, or evolving meme conventions found online\. Expanding the perspective set or incorporating culturally adaptive reasoning modules could improve coverage across different social and linguistic contexts\. 2\) The current design relies on single\-turn explanation synthesis\. In some challenging cases, particularly those involving subtle cultural references or sensitive attributes such as race or identity, the generated explanations may miss important contextual details\. Future research could explore multi\-turn reasoning strategies for explanation generation and incorporating iterative refinement to improve the completeness and reliability of explanations\. 3\) Although we employ both human evaluation and GPT\-4\-based automatic assessment to measure explanation quality, discrepancies remain between machine\-based and human judgments\. For instance, large language models may exhibit systematic biases when evaluating explanations generated by similar models\. More robust automatic evaluation protocols, along with larger\-scale and more diverse human studies, are needed to provide more reliable and unbiased assessment of explanation quality\.

## Ethics and Broader Impact

By providing interpretable justifications, MAR\-12 aims to empower moderators and researchers to better identify harmful intent, particularly when humor is used to mask hateful or harmful content\. Nevertheless, we acknowledge the risk that generated explanations could be misused to craft more persuasive harmful memes or to evade automated moderation systems\. We strongly discourage and condemn such actions\. We sincerely appreciate all the volunteer annotators and remain mindful of the risks associated with annotation tasks\. In response, we \(1\) informed annotators in advance about the potentially harmful nature of the content, \(2\) ensured their explicit acknowledgment before participation, and \(3\) advised them to stop if they felt overwhelmed\.

## References

- J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.\(2023\)Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[Vision Language Models\.](https://arxiv.org/html/2607.15442#Sx4.SSx3.SSS0.Px1.p1.1)\.
- S\. Attardo \(2024\)Linguistic theories of humor\.Vol\.1,Walter de Gruyter GmbH & Co KG\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1)\.
- J\. Bai, S\. Bai, S\. Yang, S\. Wang, S\. Tan, P\. Wang, J\. Lin, C\. Zhou, and J\. Zhou \(2023\)Qwen\-vl: a frontier large vision\-language model with versatile abilities\.arXiv preprint arXiv:2308\.129661\(2\),pp\. 3\.Cited by:[Vision Language Models\.](https://arxiv.org/html/2607.15442#Sx4.SSx3.SSS0.Px1.p1.1)\.
- J\. E\. Baker, K\. A\. Clancy, and B\. Clancy \(2020\)Putin as gay icon? memes as a tactic in russian lgbt\+ activism\.LGBTQ\+ activism in Central and Eastern Europe: Resistance, representation and identity,pp\. 209–233\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p2.1)\.
- R\. Cao, M\. S\. Hee, A\. Kuek, W\. Chong, R\. K\. Lee, and J\. Jiang \(2023\)Pro\-cap: leveraging a frozen vision\-language model for hateful meme detection\.InProceedings of the 31st ACM International Conference on Multimedia,pp\. 5244–5252\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1)\.
- R\. Cao, R\. K\. Lee, W\. Chong, and J\. Jiang \(2022\)Prompting for multimodal hateful meme classification\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 321–332\.Cited by:[Appendix A](https://arxiv.org/html/2607.15442#A1.SSx5.SSS0.Px3),[Meme Detection](https://arxiv.org/html/2607.15442#Sx2.SSx1.p1.1),[Models](https://arxiv.org/html/2607.15442#Sx4.SSx2.p1.1.5),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.2.2.2),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.5.2)\.
- B\. Crawford, F\. Keen, and G\. Suarez\-Tangil \(2021\)Memes, radicalisation, and the promotion of violence on chan sites\.InProceedings of the international AAAI conference on web and social media,Vol\.15,pp\. 982–991\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p2.1)\.
- H\. Griffin \(2021\)Living through it: anger, laughter, and internet memes in dark times\.International Journal of Cultural Studies24\(3\),pp\. 381–397\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p2.1)\.
- Y\. Guo, J\. Huang, Y\. Dong, and M\. Xu \(2020\)Guoym at semeval\-2020 task 8: ensemble\-based classification of visuo\-lingual metaphor in memes\.InProceedings of the Fourteenth Workshop on Semantic Evaluation,pp\. 1120–1125\.Cited by:[Meme Detection](https://arxiv.org/html/2607.15442#Sx2.SSx1.p1.1)\.
- R\. He and J\. Chang \(2025\)Chinese medicine as a cure for gayness: satire as countercultural resistance against heteronormative symbolic violence in digital public sphere\.International Journal of Cultural Studies,pp\. 13678779241308289\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p2.1)\.
- M\. S\. Hee, Z\. Gao, Y\. Wang, X\. Chu, R\. K\. Lee, and Z\. Qin \(2025\)Contrastive instruction fine\-tuning large multimodal model for hateful meme classification\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 760–773\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1)\.
- M\. S\. Hee, R\. K\. Lee, and W\. Chong \(2022\)On explaining multimodal hateful meme detection models\.InProceedings of the ACM web conference 2022,pp\. 3651–3655\.Cited by:[Meme Detection](https://arxiv.org/html/2607.15442#Sx2.SSx1.p1.1)\.
- M\. S\. Hee and R\. K\. Lee \(2025\)Demystifying hateful content: leveraging large multimodal models for hateful meme detection with explainable decisions\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 774–785\.Cited by:[Reasoning\-Based Meme Understanding](https://arxiv.org/html/2607.15442#Sx2.SSx2.p1.1)\.
- J\. Huang, H\. Lin, L\. Ziyan, Z\. Luo, G\. Chen, and J\. Ma \(2024\)Towards low\-resource harmful meme detection with lmm agents\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 2269–2293\.Cited by:[Appendix A](https://arxiv.org/html/2607.15442#A1.SSx5.SSS0.Px4),[Reasoning\-Based Meme Understanding](https://arxiv.org/html/2607.15442#Sx2.SSx2.p1.1),[Models](https://arxiv.org/html/2607.15442#Sx4.SSx2.p1.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.11.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.18.1)\.
- J\. Ji, X\. Lin, and U\. Naseem \(2024\)Capalign: improving cross modal alignment via informative captioning for harmful meme detection\.InProceedings of the ACM Web Conference 2024,pp\. 4585–4594\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p4.1)\.
- A\. Kalloniatis and P\. Adamidis \(2024\)Computational humor recognition: a systematic literature review\.Artificial Intelligence Review58\(2\),pp\. 43\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1)\.
- D\. Kiela, H\. Firooz, A\. Mohan, V\. Goswami, A\. Singh, P\. Ringshia, and D\. Testuggine \(2020\)The hateful memes challenge: detecting hate speech in multimodal memes\.Advances in neural information processing systems33,pp\. 2611–2624\.Cited by:[Meme Detection](https://arxiv.org/html/2607.15442#Sx2.SSx1.p1.1)\.
- D\. P\. Kingma \(2014\)Adam: a method for stochastic optimization\.arXiv preprint arXiv:1412\.6980\.Cited by:[MAR\-12 Training and Efficiency\.](https://arxiv.org/html/2607.15442#Sx4.SSx3.SSS0.Px3.p1.2)\.
- H\. Lin, Z\. Luo, W\. Gao, J\. Ma, B\. Wang, and R\. Yang \(2024\)Towards explainable harmful meme detection through multimodal debate between large language models\.InProceedings of the ACM Web Conference 2024,pp\. 2359–2370\.Cited by:[Reasoning\-Based Meme Understanding](https://arxiv.org/html/2607.15442#Sx2.SSx2.p1.1),[Automatic Evaluation\.](https://arxiv.org/html/2607.15442#Sx6.SSx1.SSS0.Px1.p1.1),[Human Evaluation\.](https://arxiv.org/html/2607.15442#Sx6.SSx1.SSS0.Px2.p1.1)\.
- X\. Lin, C\. Jia, J\. Ji, H\. Han, and U\. Naseem \(2025\)Ask, acquire, understand: a multimodal agent\-based framework for social abuse detection in memes\.InProceedings of the ACM on Web Conference 2025,pp\. 4734–4744\.Cited by:[Reasoning\-Based Meme Understanding](https://arxiv.org/html/2607.15442#Sx2.SSx2.p1.1)\.
- H\. Liu, R\. Wei, G\. Tu, J\. Lin, C\. Liu, and D\. Jiang \(2024\)Sarcasm driven by sentiment: a sentiment\-aware hierarchical fusion network for multimodal sarcasm detection\.Information Fusion108,pp\. 102353\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1),[Meme Detection](https://arxiv.org/html/2607.15442#Sx2.SSx1.p1.1)\.
- H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee \(2023\)Visual instruction tuning\.Advances in neural information processing systems36,pp\. 34892–34916\.Cited by:[Vision Language Models\.](https://arxiv.org/html/2607.15442#Sx4.SSx3.SSS0.Px1.p1.1)\.
- O\. S\. Liu, P\. C\. Ng, D\. W\. Soh, and K\. N\. Plataniotis \(2026\)Yes florence, i will do better next time\! agentic feedback reasoning for humorous meme detection\.arXiv preprint arXiv:2601\.07232\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p4.1)\.
- Z\. Liu, C\. Fan, H\. Lou, Y\. Wu, and K\. Deng \(2025\)MIND: a multi\-agent framework for zero\-shot harmful meme detection\.arXiv preprint arXiv:2507\.06908\.Cited by:[Reasoning\-Based Meme Understanding](https://arxiv.org/html/2607.15442#Sx2.SSx2.p1.1),[Models](https://arxiv.org/html/2607.15442#Sx4.SSx2.p1.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.12.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.19.1)\.
- S\. Pramanick, S\. Sharma, D\. Dimitrov, M\. S\. Akhtar, P\. Nakov, and T\. Chakraborty \(2021\)MOMENTA: a multimodal framework for detecting harmful memes and their targets\.InFindings of the Association for Computational Linguistics: EMNLP 2021,pp\. 4439–4455\.Cited by:[Appendix A](https://arxiv.org/html/2607.15442#A1.SSx5.SSS0.Px2),[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1),[Models](https://arxiv.org/html/2607.15442#Sx4.SSx2.p1.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.1.1.2),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.4.4.2)\.
- N\. Rizwan, P\. Bhaskar, M\. Das, S\. S\. Majhi, P\. Saha, and A\. Mukherjee \(2025\)Exploring the limits of zero shot vision language models for hate meme detection: the vulnerabilities and their interpretations\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 1669–1689\.Cited by:[Reasoning\-Based Meme Understanding](https://arxiv.org/html/2607.15442#Sx2.SSx2.p1.1)\.
- N\. Rizwan, S\. Swain, P\. Bhaskar, G\. Aryan, S\. S\. Khan, and A\. Mukherjee \(2026\)See, explain, and intervene: a few\-shot multimodal agent framework for hateful meme moderation\.arXiv preprint arXiv:2601\.04692\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p4.1)\.
- S\. B\. Shah, S\. Shiwakoti, M\. Chaudhary, and H\. Wang \(2024\)MemeCLIP: leveraging clip representations for multimodal meme classification\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Miami, Florida, USA,pp\. 17320–17332\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.959/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.959)Cited by:[Appendix A](https://arxiv.org/html/2607.15442#A1.SSx5.SSS0.Px1),[Introduction](https://arxiv.org/html/2607.15442#Sx1.p3.1),[Evaluation Datasets](https://arxiv.org/html/2607.15442#Sx4.SSx1.p1.1),[Models](https://arxiv.org/html/2607.15442#Sx4.SSx2.p1.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.10.1),[Table 3](https://arxiv.org/html/2607.15442#Sx4.T3.5.17.1)\.
- C\. Sharma, D\. Bhageria, W\. Scott, S\. Pykl, A\. Das, T\. Chakraborty, V\. Pulabaigari, and B\. Gamback \(2020\)SemEval\-2020 task 8: memotion analysis–the visuo\-lingual metaphor\!\.arXiv preprint arXiv:2008\.03781\.Cited by:[Evaluation Datasets](https://arxiv.org/html/2607.15442#Sx4.SSx1.p1.1)\.
- E\. Steffen \(2025\)More than memes: a multimodal topic modeling approach to conspiracy theories on telegram\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 1831–1844\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p2.1)\.
- S\. Thapa, H\. Veeramani, L\. Hu, Q\. Zhang, W\. Wang, and U\. Naseem \(2025\)A multimodal prompt\-based framework for analyzing code\-mixed and low\-resource memes\.InProceedings of the International AAAI Conference on Web and Social Media,Vol\.19,pp\. 1913–1923\.Cited by:[Meme Detection](https://arxiv.org/html/2607.15442#Sx2.SSx1.p1.1)\.
- F\. Yus \(2023\)Meme\-mediated humorous communication\.InPragmatics of Internet Humour,pp\. 245–307\.Cited by:[Introduction](https://arxiv.org/html/2607.15442#Sx1.p2.1)\.

## Paper Checklist

1. 1\.For most authors… 1. \(a\)Would answering this research question advance science without violating social contracts, such as violating privacy norms, perpetuating unfair profiling, exacerbating the socio\-economic divide, or implying disrespect to societies or cultures?Yes, our work primarily focuses on utilizing VLMs to analyze and generate interpretations of humorous and hateful memes\. While these generated interpretations may reflect social stereotypes, our goal is to enhance hateful meme detection systems and improve the understanding of such content\. 2. \(b\)Do your main claims in the abstract and introduction accurately reflect the paper’s contributions and scope?Yes\. 3. \(c\)Do you clarify how the proposed methodological approach is appropriate for the claims made?Yes\. 4. \(d\)Do you clarify what are possible artifacts in the data used, given population\-specific distributions?Yes\. 5. \(e\)Did you describe the limitations of your work?Yes\. You may find them under “Limitations and Future Work” section\. 6. \(f\)Did you discuss any potential negative societal impacts of your work?Yes\. You may find them under “Ethics and Broader Impact” section\. 7. \(g\)Did you discuss any potential misuse of your work?Yes\. You may find them under “Ethics and Broader Impact” section\. 8. \(h\)Did you describe steps taken to prevent or mitigate potential negative outcomes of the research, such as data and model documentation, data anonymization, responsible release, access control, and the reproducibility of findings?N/A 9. \(i\)Have you read the ethics review guidelines and ensured that your paper conforms to them?Yes\.
2. 2\.Additionally, if your study involves hypotheses testing… 1. \(a\)Did you clearly state the assumptions underlying all theoretical results?N/A 2. \(b\)Have you provided justifications for all theoretical results?N/A 3. \(c\)Did you discuss competing hypotheses or theories that might challenge or complement your theoretical results?N/A 4. \(d\)Have you considered alternative mechanisms or explanations that might account for the same outcomes observed in your study?N/A 5. \(e\)Did you address potential biases or limitations in your theoretical framework?N/A 6. \(f\)Have you related your theoretical results to the existing literature in social science?N/A 7. \(g\)Did you discuss the implications of your theoretical results for policy, practice, or further research in the social science domain?N/A
3. 3\.Additionally, if you are including theoretical proofs… 1. \(a\)Did you state the full set of assumptions of all theoretical results?N/A 2. \(b\)Did you include complete proofs of all theoretical results?N/A
4. 4\.Additionally, if you ran machine learning experiments… 1. \(a\)Did you include the code, data, and instructions needed to reproduce the main experimental results \(either in the supplemental material or as a URL\)?Yes\. The Anonymous GitHub link can be found in the paper’s abstract\. 2. \(b\)Did you specify all the training details \(e\.g\., data splits, hyperparameters, how they were chosen\)?Yes\. These information can be found under “Implementation Details” section\. 3. \(c\)Did you report error bars \(e\.g\., with respect to the random seed after running experiments multiple times\)?No\. The results reported in Tables 3, 4, and 5 are based on a single run using a fixed random seed \(42\) to ensure reproducibility\. 4. \(d\)Did you include the total amount of compute and the type of resources used \(e\.g\., type of GPUs, internal cluster, or cloud provider\)?Yes\. These information can be found under “Implementation Details” section\. 5. \(e\)Do you justify how the proposed evaluation is sufficient and appropriate to the claims made?Yes\. These information can be found under “Experiments” section\. 6. \(f\)Do you discuss what is “the cost“ of misclassification and fault \(in\)tolerance?N/A
5. 5\.Additionally, if you are using existing assets \(e\.g\., code, data, models\) or curating/releasing new assets,without compromising anonymity… 1. \(a\)If your work uses existing assets, did you cite the creators?Yes\. 2. \(b\)Did you mention the license of the assets?N/A 3. \(c\)Did you include any new assets in the supplemental material or as a URL?N/A 4. \(d\)Did you discuss whether and how consent was obtained from people whose data you’re using/curating?N/A 5. \(e\)Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content?N/A 6. \(f\)If you are curating or releasing new datasets, did you discuss how you intend to make your datasets FAIR?N/A 7. \(g\)If you are curating or releasing new datasets, did you create a Datasheet for the Dataset?N/A
6. 6\.Additionally, if you used crowd sourcing or conducted research with human subjects,without compromising anonymity… 1. \(a\)Did you include the full text of instructions given to participants and screenshots?No, we provided the annotators with instructions during meetings\. All annotators are expert inthis field and no explicit instruction set was required 2. \(b\)Did you describe any potential participant risks, with mentions of Institutional Review Board \(IRB\) approvals?N/A 3. \(c\)Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation?No, all annotators were from authors’ institute\. 4. \(d\)Did you discuss how data is stored, shared, and deidentified?N/A

## Appendix AAppendix

### Computing Resources

We conducted our experiments on a high\-performance computing server equipped with two AMD EPYC 7543 32‑Core Processors \(64 cores per socket with simultaneous multithreading enabled, totaling 128 logical CPUs\)\. The system has 514 GiB of DDR4 system memory, providing ample capacity for large‑scale data loading and preprocessing\. The machine is outfitted with seven NVIDIA L40S GPUs, each offering 46,068 MiB of VRAM, enabling large‑batch multimodal training and inference\. The software environment includes Ubuntu 20\.04, CUDA 12\.5, and PyTorch 2\.x with mixed precision training enabled viatorch\.float16\. The high\-performance environment, particularly the 46,068 MiB VRAM of the L40S GPUs, allowed for the parallelized processing of the twelve perspective\-conditioned reasoning traces during data preparation\. This hardware setup ensures that the system can handle the high\-throughput requirements of real\-time social media moderation\. The total training time for the entire MAR\-12 framework on the Memotion dataset \(approximately 5,600 memes\) was completed in under 30 minutes, demonstrating that the framework is not only high\-performing but also highly efficient to train\.

For MAR‑12, we used CLIP ViT‑B/32 as the vision encoder, with image inputs resized to 224×224 pixels\. The agent text embeddings were obtained using T5 \(or CLIP text encoder in some variants\) and projected into a 1024‑dimensional mapped space from the original 768‑dimensional embeddings\. The fusion module included 1 mapping layer and 1 pre‑output layer, with dropout probabilities of 0\.1, 0\.4, and 0\.2 applied at different stages\. The classification head employed a cosine‑similarity‑based prototype classifier with a scale of 30 and margin ratio of 0\.2\. Training was performed for 20 epochs using the AdamW optimizer with a learning rate of1×10−41\\times 10^\{\-4\}, weight decay of1×10−41\\times 10^\{\-4\}, and a batch size of 32\. We used 4 data‑loading workers and fixed the random seed to 42 for reproducibility\. Object‑level cues were extracted with a pretrained Faster R‑CNN \(ResNet‑50‑FPN\) detector \(IoU threshold=0\.5, score threshold=0\.5\)\. Object features were encoded with VGG‑16 \(final classification layer removed\), and attribute features were extracted using DistilBERT with uncased tokenization\. All image features were normalized using ImageNet statistics\.

### Training Configurations

We applied a consistent training setup across all models to ensure comparability\. Unless otherwise stated, We adopted a unified training strategy across all models:

- •Optimizer:AdamW
- •Learning Rate:1×10−41\\times 10^\{\-4\}
- •Weight Decay:1×10−41\\times 10^\{\-4\}
- •Batch Size:32
- •Epochs:5
- •Loss Function:Cross\-Entropy Loss for binary classification \(2 output classes\)
- •Seed:42 \(for reproducibility\)

All models were trained onimage\_size=224inputs \(when applicable\), withnum\_workers=4used for data loading\. Cross\-entropy loss was used as the objective function for all binary classification tasks\. Mixed precision training was enabled throughout to improve computational efficiency and memory utilization\. The rapid training time \(<<30 minutes for 20 epochs\) is a direct result of MAR\-12’s lightweight design\. Unlike competitive Large Multimodal Model \(LMM\) baselines like LoReHM \(34\.00B trainable parameters\), MAR\-12 utilizes only 1\.31M trainable parameters\. Because the heavy VLM \(Qwen\-VL\) and encoders \(CLIP/T5\) are kept frozen, the training process focuses solely on the role\-aware attention and the prototype\-based classification head\. This architectural choice significantly reduces the computational overhead and carbon footprint of model development compared to end\-to\-end fine\-tuning\.

### Inference Complexity and Scalability

Inference Complexity and Scalability\. At inference time, MAR\-12 processes a single meme in an average of 3\.5 seconds\. This duration accounts for the end\-to\-end pipeline: \(1\) twelve parallelized VLM forward passes to generate reasoning traces, \(2\) the soft\-gated attention aggregation, \(3\) the prototype\-based classification, and \(4\) the final LLM\-based explanation synthesis\. While the 12 VLM passes introduce a higher cost than a single black\-box classification, this cost is offset by the quality and interpretability of the results\. To improve scalability, our implementation utilizes batch processing for the 12 perspective prompts, reducing redundant vision\-encoding passes\. Compared to iterative multi\-agent debate frameworks which can take upwards of 10–15 seconds per meme, MAR\-12’s 3\.5\-second latency provides a superior balance between deep reasoning and the response times required for large\-scale content moderation\.

### Evaluation Metrics

We evaluate model performance using Accuracy \(ACC\), Area Under the Receiver Operating Characteristic Curve \(AUC\), and F1‑score\. Accuracy measures the proportion of correctly classified memes and serves as a straightforward indicator of overall predictive performance\. However, given that humour and hate meme detection tasks can exhibit class imbalance \(Table 2\) and involve ambiguous, borderline cases, we complement accuracy with additional metrics\. AUC evaluates the model’s ability to rank positive and negative samples correctly across varying classification thresholds, making it less sensitive to class distribution and more informative for imbalanced settings\. F1‑score, the harmonic mean of precision and recall, is particularly important in this context: a model that predicts most memes as “humour” or “hate” may achieve high accuracy but poor balance between false positives and false negatives\.

We followed the official train/validation/test splits for the PrideMM and Memotion datasets\. During evaluation, we computed accuracy, AUC, and F1‑score using the official labels, and additionally saved full classification reports, confusion matrices, and per‑meme prediction logs\. All results reported in the paper are based on a single run using a fixed random seed for reproducibility\.

### Baseline Model\-Specific Details

#### MemeCLIP\(Shahet al\.[2024](https://arxiv.org/html/2607.15442#bib.bib87)\)

This model builds upon the CLIP ViT\-L/14 encoder, with frozen parameters\. Image and text features are independently projected using linear mapping layers, followed by adapter modules for task\-specific refinement\. The projected and adapted embeddings are normalized and fused via element\-wise multiplication\. The final representation is passed through an optional pre\-output MLP and a cosine\-based classifier\. Key architectural parameters include:

- •unmapped\_dim=768,map\_dim=1024
- •drop\_probs=\[0\.1, 0\.4, 0\.2\]
- •ratio=0\.2for adapter residual weighting
- •Cosine classifier with scaling factorscale=30

#### MOMENTA\(Pramanicket al\.[2021](https://arxiv.org/html/2607.15442#bib.bib90)\)

MOMENTA leverages both object\-level and attribute\-level visual signals fused with CLIP image and text features\. Intra\-modal fusion is performed via a custom cross\-modal attention mechanism, which aligns CLIP embeddings with extracted features \(e\.g\., from object detection or attribute graphs\)\. The two streams are then aggregated via a self\-attention fusion block to form a unified multimodal representation\. The final prediction is produced by a linear classifier\. This model captures both intra\- and inter\-modal dependencies with fine\-grained attention\.

#### PromptHate\(Caoet al\.[2022](https://arxiv.org/html/2607.15442#bib.bib53)\)

We use a RoBERTa\-based prompting model for binary classification\. The model architecture is based onroberta\-large, trained in a masked language modeling setup\. Task\-specific verbalizers are used to map label tokens to the logits extracted at the masked position\. This approach aligns with prompt\-based learning paradigms where no classifier head is introduced — the prediction is directly computed based on the masked token output distribution\.

#### LoReHM\(Huanget al\.[2024](https://arxiv.org/html/2607.15442#bib.bib46)\)

We utilize LLaVA34B as the LMM agent from the the open\-source perspectives\. Specifically, we implement the “llava\-v1\.6\-34b” for LLaVA\-34B\. The frozen pretrained vision and text Transformer encoders are implemented as CLIP with the specific version “ViT\-L/14@336px”\.

#### MAR\-12 \(with Role\-aware Attention\-based Aggregator\)\.

Our proposed MAR\-12 model integrates vision\-language grounding with structured reasoning via twelve role\-specific agents\. Each agent produces an embedding projected into a shared feature space using agent\-specific MLPs\. Attention is computed across agent embeddings, yielding a weighted representation that captures the most salient reasoning cues\.

The visual input is encoded using CLIP ViT\-L/14, and both the agent and image embeddings are passed through shared mapping and adapter layers \(similar to MemeCLIP\)\. The final representation is computed by element\-wise multiplication and classified using a cosine\-based classifier\. This architecture supports interpretability and extensibility via role\-aware agent prompting\.

A fixed random seed \(42\) was used across all experiments to ensure reproducibility\.

TableLABEL:tab:model\_summarylists the parameter count for all models\.

## Appendix BAppendix: Prompting Strategy, Psychological Grounding, and Explainer Prompt

### Design Goals and Rationale

Our prompting strategy is built to \(i\)separate roles\(facts→\\rightarrowinterpretation\), \(ii\)cover complementary mechanismsof humour and harm with orthogonal lenses, and \(iii\)constrain outputsso that the attention module can weight agent evidence and the explainer can*faithfully*summarize that attention\-weighted reasoning\.

#### Role separation\.

We elicit*foundational facts*\(scene description, OCR\) with report\-only prompts that explicitly prohibit interpretation, then invite*interpretive*agents \(humour and safety/intent\) to reason on top of those facts\. This reduces leakage and makes attention assignments diagnostic\.

#### Mechanism coverage\.

Humour in memes is multifactorial\. We therefore cover visual/textual incongruity, affective contrast, cultural schema, absurdity, wordplay, timing/punchline structure, and image–text alignment\. For safety, we cover harmful impact \(hatefulness\) and authorial motive \(intent\), recognizing that harmful memes often co\-opt humour mechanics\.

#### Output constraints\.

Prompts use imperative, scope\-limited language \(e\.g\., “avoid interpretation”, “return text only”\) so agent outputs are short, typed rationales that are easy to weight \(by attention\) and cite \(by the explainer\)\. This supports process\-level faithfulness rather than label\-chasing\.

### Agent Prompts \(12 Roles\)

RoleAgent NamePromptFoundational AnalysisGeneral Description AgentProvide a detailed description of the image\. What objects, characters, facial expressions, and scene elements are visible? Avoid interpretation, just describe what is shown\.OCR & Text Extraction AgentExtract all text visible in the image, including captions, embedded text, and any signs or symbols\. Return the extracted text only\.Humour\-Oriented ReasoningVisual Irony AgentLook at the visual content of the image\. Is there any contradiction, mismatch, or unexpected element that creates irony or humour? Describe how the image visually subverts expectations\.Textual Irony AgentRead the text in the image\. Is there irony, sarcasm, or a contradiction between what is said and what is implied? Explain whether this contributes to humour\.Emotion Contrast AgentAnalyze the emotional tone of the image \(e\.g\., facial expressions, setting\) and compare it with the tone of the text\. Is there a humorous contrast or mismatch between the two?Cultural Reference AgentDoes the meme refer to any cultural, social, political, or internet\-related topic that contributes to its humour? Identify and explain any references that are key to understanding the joke\.Absurdity Check AgentDoes the meme contain elements that are absurd, exaggerated, or nonsensical in a way that is intended to be funny? If yes, describe them and how they contribute to humour\.Pun & Wordplay Check AgentCheck if the meme contains puns, double meanings, rhymes, or wordplay\. If so, explain how the text plays with language to create humour\.Punchline Check AgentDoes the meme follow a setup and punchline structure? Identify the setup and the punchline, and explain how the timing or surprise of the punchline contributes to humour\.Image\-Text Alignment AgentEvaluate the relationship between the image and the text\. Do they work together to create humour, or do they contradict each other? Explain how their alignment or misalignment affects the meme’s humour\.Safety & Intent EvaluationHatefulness Detection AgentAnalyze the meme to determine if it includes offensive, derogatory, or hateful content\. Could the meme harm or marginalize any group of people, even if it’s intended to be funny? Provide reasoning\.Intent Interpretation AgentBased on the image and text, infer what the creator intended\. Was it meant to entertain, criticize, mock, or provoke? Is the humour light\-hearted or could it be seen as targeted or aggressive? Explain your interpretation\.

Table 10:Overview of the 12 specialized agents used in MAR\-12, including their roles and associated prompts\.
### Explainer Prompt and Interface to Attention

The explainer receives three inputs: \(1\) the final predictiony^\\hat\{y\}, \(2\) the set of agent responses\{ri\}i=112\\\{r\_\{i\}\\\}\_\{i=1\}^\{12\}, and \(3\) attention weights\{αi\}i=112\\\{\\alpha\_\{i\}\\\}\_\{i=1\}^\{12\}from the soft aggregation module\. The explainer is instructed to summarize*why*the system decided ony^\\hat\{y\}by following the attention\-weighted agent evidence\.

#### Explainer Prompt\.

> “Given the prediction: \[LABEL\], and the following agent insightsweighted by relevance\(αi\\alpha\_\{i\}\), summarize the reasoning behind the classification\. Cite the most influential agents \(higherα\\alpha\) and explain how their evidence supports the decision\. If there are conflicting signals from lower\-weighted agents, briefly acknowledge them\.”

### Psychological Grounding of Prompt Families

Incongruity & resolution\(visual/textual irony; image–text alignment\) motivate prompts that explicitly test for mismatches and their resolution\.Affective contrast\(emotion vs\. content\) motivates prompts that ask for tonal mismatches\.Schema/culture dependencemotivates prompts for cultural references to supply background needed for comprehension\.Absurdist/nonsense humourmotivates prompts to flag exaggeration/surrealism\.Linguistic ambiguitymotivates puns/wordplay prompts\.Timing/surprise \(setup–punchline\)motivates explicit structure prompts\.Superiority/benign violation/social identity harmsmotivate hatefulness prompts \(impact on groups\)\.Theory of mind/Gricean pragmaticsmotivates intent prompts \(speaker meaning vs\. literal form\)\.

### Why This Prompt Design Supports Faithfulness and Novelty

- •Process alignment by construction\.Agents yield typed, local rationales; the attention head learns a sparse mixture over these; the explainer is instructed to mirror that mixture\. This produces explanations that track the system’s internal evidence flow \(role\-aware soft attention\) rather than post\-hoc label rationalization\.
- •Separation of facts and interpretation\.Report\-only foundational prompts reduce contamination and make later interpretive weights more meaningful\.
- •Dual\-use lenses for humour vs\. harm\.The same mechanisms that make content funny can also make it harmful; prompting both humour and safety over the*same*lenses lets the model explain how those mechanisms were deployed \(benign vs\. targeted\)\.
- •Extensibility\.New lenses can be added as new agents without disrupting others; the attention layer simply learns when they matter, and the explainer inherits them in explanation\.

### Implementation Notes \(for Reproducibility\)

- •Prompts use imperative phrasing \(“identify whether”, “explain how”\) and scope guards \(“avoid interpretation”, “return text only”\) to standardize outputs\.
- •Agent responses should be concise \(2–5 sentences or bullet points\), focusing on evidence, not conclusions; this improves attention calibration and explainer citability\.
- •The explainer should be passed\(ri,αi\)\(r\_\{i\},\\alpha\_\{i\}\)pairs in descendingαi\\alpha\_\{i\}order to bias summaries toward most influential evidence while still acknowledging conflicts\.

## Appendix CError Analysis

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/img_4074.png)Figure 4:Label: No Humour, No Hate; Prediction:Humour, Hate- •The image shows two animated characters in an office setting\. One character, a large man with a stern expression, is leaning forward with his fist raised, seemingly in a confrontational stance\. The other character, a smaller person with blonde hair, is seated across from him, looking up at the larger character\. The background features a window showing a sunset sky\. Text at the top reads ”Trans rights are human rights” and ”Black rights are human rights,” while text at the bottom says ”Human rights are human rights\!”
- •OCR & Text Extraction Agent: ”Trans rights are human rights” ”Black rights are human rights” Me: ”Human rights are human rights\!”
- •Visual Irony Agent: The image humorously subverts expectations by presenting a character who appears to be advocating for human rights in a dramatic manner, yet the text at the bottom states ”Human rights are human rights\!” This repetition and the character’s exaggerated expression create an ironic contrast, suggesting a playful dismissal of the serious topic through over\-the\-top enthusiasm\.
- •Textual Irony Agent: The text implies that all human rights are inherently equal, which contrasts with the implication of prioritizing certain groups over others\. This creates a humorous effect by highlighting the absurdity of the statement through its own logical inconsistency\.
- •Emotion Contrast Agent: The image conveys a sense of frustration and intensity through the character’s exaggerated facial expression and body language, while the text presents a straightforward statement about human rights\. This creates a humorous contrast, as the dramatic reaction seems disproportionate to the simple assertion made in the text\.
- •Cultural Reference Agent: The meme humorously addresses the intersection of various human rights movements by juxtaposing the phrases ”Trans rights are human rights” and ”Black rights are human rights,” suggesting that these are part of a broader, unified concept of human rights\. The character’s exaggerated reaction emphasizes the importance of this unity, making the joke resonate with those who support the idea that all human rights are interconnected and essential\.
- •Absurdity Check Agent: Yes, the meme uses exaggerated body language and a dramatic pose from an animated character to humorously emphasize the point that all human rights are equal, making it absurdly clear and funny\.
- •Pun & Wordplay Check Agent: The meme uses wordplay by repeating ”human rights” in the text, creating a humorous echo that emphasizes the importance of human rights for all groups, including trans and black individuals\.
- •Punchline Check Agent: Yes, the meme follows a setup and punchline structure\. The setup is the statement ”Trans rights are human rights” and ”Black rights are human rights,” leading to the punchline ”Human rights are human rights\!” The humor comes from the unexpected repetition of the phrase ”human rights,” which emphasizes the universality of human rights and creates a comedic effect by highlighting the redundancy\.
- •Image\-Text Alignment Agent: The image and text work together to create humor by using the exaggerated reaction of the character to emphasize the point that ”Human rights are human rights\!” This misalignment between the serious statement about human rights and the over\-the\-top reaction adds to the comedic effect, as it contrasts the expected calmness with the unexpected intensity, suggesting a humorous disbelief or frustration at the idea being presented
- •Hatefulness Detection Agent: The meme does not include offensive, derogatory, or hateful content\. It humorously emphasizes the universal nature of human rights by highlighting that both trans rights and black rights are part of the broader concept of human rights, without targeting any specific group negatively\.
- •Intent Interpretation Agent: The creator likely intended to entertain with a humorous take on the concept of human rights, using exaggerated body language for comedic effect\. The meme could also be seen as mildly provocative by highlighting the importance of recognizing all human rights equally\.

In this instance, the meme was misclassified by the MAR\-12 as both humorous and hateful, despite the agent responses suggesting it does not contain hateful content\. Specifically, the Hatefulness Detection Agent determined that the meme does not include offensive or derogatory elements, but instead humorously underscores the universality of human rights by affirming that both trans rights and Black rights are integral to human rights, without negatively targeting any group\. Meanwhile, the Intent Interpretation Agent inferred that the creator’s aim was to entertain by exaggerating body language for comedic effect, perhaps also prompting reflection on the importance of recognizing all human rights equally\.

The misclassification likely stems from the subjectivity inherent in human annotations, particularly in content addressing sensitive topics where interpretations of humor and offense can vary greatly among annotators\. While the meme’s exaggerated presentation might have led some annotators to perceive it as mocking or trivializing serious issues, the agents interpreted it as a humorous yet earnest commentary on equality\. This divergence underscores the challenge of aligning model outputs with subjective human judgments in nuanced, socially charged content\. Addressing this challenge will require future work on better modeling subjectivity in humour and hate detection, including training with datasets that capture annotator disagreement, and exploring uncertainty\-aware classification frameworks that can reflect ambiguous or borderline cases\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/img_3893.png)Figure 5:Label: Humour, Hate; Prediction: HumourNo Hate- •General Description Agent: The image shows a protest sign held by someone wearing a red jacket\. The sign reads: ”AS A GIRL, I JUST HOPE THAT 1 DAY I HAVE AS MANY RIGHTS AS A GUN\.” The text is written in black letters on a white background, with decorative elements like hearts and stars around the words\. The background includes buildings and other people, suggesting an urban setting during a public demonstration\.
- •OCR & Text Extraction Agent: ”AS A GIRL, I JUST HOPE THAT 1 DAY I HAVE AS MANY RIGHTS AS A GUN”
- •Visual Irony Agent: The sign humorously contrasts the serious topic of gender equality with a playful, nonsensical comparison to gun rights, creating an ironic and unexpected juxtaposition that subverts expectations by using a lighthearted tone to address a serious issue\.
- •Textual Irony Agent: The sign expresses a serious hope for equality in rights, but the use of ”gun” as a metaphor for rights creates an ironic and humorous contrast, playing on the absurdity of equating rights with a dangerous object\.
- •Emotion Contrast Agent: The image conveys a serious and urgent tone, likely from a protest or demonstration, while the text on the sign uses a sarcastic and ironic tone, suggesting a humorous contrast by comparing rights to gun ownership, which is typically associated with power and control rather than equality\.
- •Cultural Reference Agent: The meme humorously critiques gender inequality by playfully suggesting that women should have as many rights as guns, which is a hyperbolic exaggeration meant to highlight the absurdity of the situation\. The use of colorful stickers and playful fonts adds to the comedic effect by contrasting the serious issue with a lighthearted presentation\.
- •Absurdity Check Agent: Yes, the meme uses absurdity and exaggeration by comparing the rights of a girl to those of a gun, which is nonsensical and intended to be humorous by highlighting the irony and absurdity of the situation\.
- •Pun & Wordplay Check Agent: The meme uses a pun by playfully juxtaposing ””rights”” with ””gun,”” creating a humorous and ironic statement about gender equality\. The use of asterisks around ””gun”” adds to the absurdity, emphasizing the unexpected nature of the comparison\.
- •Punchline Check Agent: The meme follows a setup and punchline structure where the setup is the hopeful statement about having equal rights as a girl, and the punchline is the unexpected and ironic twist of equating rights with a gun\. The humor comes from the stark contrast between the innocent hope for equality and the absurdity of suggesting that guns are a measure of rights, which subverts expectations and creates a humorous effect through irony\.
- •Image\-Text Alignment Agent: The image and text work together to create humor by using a provocative and absurd juxtaposition\. The sign’s message about gender equality is paired with the word ””gun,”” which is typically associated with power and rights, but here it is used in a way that subverts expectations, making the statement both shocking and humorous\.
- •Hatefulness Detection Agent: The meme uses a provocative and insensitive metaphor comparing gender rights to gun ownership, which could be interpreted as a harmful or derogatory statement towards women’s rights\. It may marginalize women by suggesting that their rights are equivalent to those of a weapon, potentially perpetuating stereotypes about gender equality\.
- •Intent Interpretation Agent: The creator likely intended to provoke thought and highlight the absurdity of gender inequality by using a provocative statement that contrasts the desire for equal rights with the metaphorical ”right” to bear arms\. The humor is not light\-hearted but rather targeted at drawing attention to serious issues through a bold and unconventional approach\.

The conflicting perspectives in the safety agents’ responses highlight a core challenge in humor detection: the subjective nature of humor itself, especially when intertwined with sensitive social topics\. While human annotators may have perceived the meme as humorous, the agents interpreted the content as potentially harmful and therefore did not classify it as light\-hearted humour\. This divergence underscores how subjectivity in humour perception, both among human annotators and within the model’s reasoning process, can lead to classification disagreements, particularly in cases involving controversial or polarizing subjects\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/img_4099.png)Figure 6:Label: Humour, Hate; Prediction: Humour,No Hate- •General Description Agent: The image features Cookie Monster from Sesame Street sitting in a chair with a bookshelf in the background\. He has a surprised expression on his face\. Text overlaid on the image reads: ”what is trans right?”
- •OCR & Text Extraction Agent: ”what is trans right? I WOULD SOONER GIVE UP COOKIES THAN SUPPORT TRANS RIGHT and do you support trans left?”
- •Visual Irony Agent: The image humorously subverts expectations by presenting Cookie Monster as a character who would give up cookies rather than support trans rights, creating an ironic contrast between his typically cookie\-loving persona and his unexpected stance on the issue\.
- •Textual Irony Agent: The sign expresses a serious hope for equality in rights, but the use of ”gun” as a metaphor for rights creates an ironic and humorous contrast, playing on the absurdity of equating rights with a dangerous object\.
- •Emotion Contrast Agent: The image features Cookie Monster with a neutral expression, seated in a library\-like setting, which contrasts humorously with the text that presents a serious and potentially controversial topic about trans rights\. The juxtaposition creates a comedic effect by highlighting the absurdity of the statement\.
- •Cultural Reference Agent: The meme humorously references the complex and often contentious issue of trans rights, using a well\-known character’s exaggerated reluctance to support trans rights as a punchline\. The joke plays on the absurdity of prioritizing cookies over supporting a marginalized group, highlighting the absurdity of some people’s stance on the matter\.
- •Absurdity Check Agent: Yes, the meme uses absurd exaggeration by suggesting someone would give up cookies rather than support trans rights, which is meant to be humorous by highlighting the extreme nature of the statement\. The juxtaposition of this extreme position with the question about supporting trans rights on the left adds to the comedic effect\.
- •Pun & Wordplay Check Agent: The meme uses wordplay by playfully juxtaposing ”trans right” and ”trans left,” creating a humorous contrast that relies on the audience’s understanding of the term ”trans” in both political contexts and its literal meaning\. The phrase ”I would sooner give up cookies than support trans right” is a clever pun, as it humorously suggests a strong preference for cookies over supporting transgender rights, playing on the idea of giving something up for a cause\.
- •Punchline Check Agent: The meme follows a setup and punchline structure where the setup asks about trans rights, leading to the punchline that humorously suggests giving up cookies rather than supporting trans rights, which is unexpected and absurd, contributing to its comedic effect\.
- •Image\-Text Alignment Agent: The image of Cookie Monster, known for his love of cookies, paired with the text that humorously suggests he would give up cookies rather than support trans rights, creates a playful contradiction that adds to the humor by juxtaposing his character trait with the serious topic of trans rights\. The alignment of the text with the image’s absurdity enhances the meme’s comedic effect\.
- •Hatefulness Detection Agent: The meme uses a character known for its controversial statements about trans rights, which can be seen as offensive and potentially harmful to the trans community by promoting a harmful narrative that equates support for trans rights with a lack of support for cookies\. This could marginalize and alienate individuals who identify as trans or support their rights\.
- •Intent Interpretation Agent: The creator likely intended to provoke thought or criticism about the topic of trans rights, using a humorous yet potentially insensitive approach that might be seen as aggressive by some viewers\.

The Hatefulness Detection Agent identified that the meme features a character known for controversial statements about trans rights, framing the meme as potentially offensive and harmful by perpetuating a narrative that could marginalize or alienate trans individuals and their allies\. This agent clearly signaled the presence of hateful or harmful content\. However, the Intent Interpretation Agent suggested that the creator’s goal was to provoke thought or criticism around trans rights using humor, albeit in a manner that could be perceived as insensitive or aggressive by some audiences\. This interpretation leans toward reading the meme as provocative satire rather than explicit hate\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/img_4095.png)Figure 7:Label: No Humour, Hate; Prediction:Humour, No Hate- •General Description Agent: The image is a cartoon featuring a shepherd holding a sheep with a rainbow\-colored head\. The shepherd says, ””I know\. And I found her\. On the left side, a flock of sheep is gathered around another figure, who is saying, ””Whoa\! Whoa\! Whoa\! Hold it right there\! He wasn’t lost, we kicked him out\!””
- •OCR & Text Extraction Agent: ”WHOA\! WHOA\! WHOA\! HOLD IT RIGHT THERE\! HE WASN’T LOST, WE KICKED HIM OUT\! I KNOW\. AND I FOUND HER\.
- •Visual Irony Agent: The image humorously subverts expectations by depicting Jesus holding a sheep with a rainbow\-colored head, which is an unexpected and whimsical twist on the traditional imagery of Jesus as the shepherd\. The dialogue further adds to the irony, as it suggests that Jesus found someone who was not lost but was instead kicked out, creating a playful and absurd scenario\.
- •Textual Irony Agent: The text implies that Jesus found someone who was not lost but was kicked out, which creates a humorous contradiction with the typical narrative of Jesus finding lost sheep\. This irony adds to the humor by subverting expectations\.
- •Emotion Contrast Agent: The image conveys a lighthearted and somewhat ironic tone through its simple line art and the unexpected twist of Jesus holding a sheep, which contrasts humorously with the serious implication of ”kicking someone out\.”
- •Cultural Reference Agent: The text, while straightforward, adds to this contrast by emphasizing the absurdity of the situation, creating a comedic effect\.
- •Absurdity Check Agent: The meme humorously plays on the idea of ”kicking someone out” by referencing a common phrase used when someone is lost, suggesting that the person was never lost in the first place but was intentionally excluded\. The use of Jesus holding a sheep with a rainbow\-colored head adds a whimsical twist, enhancing the comedic effect\.
- •Pun & Wordplay Check Agent: Yes, the meme contains absurd and exaggerated elements intended to be humorous\. The depiction of Jesus holding a sheep with a beard and a speech bubble saying ”I know\. And I found her” alongside the flock’s reaction of disbelief and concern creates a comical contrast between the expected and the unexpected\. The meme uses wordplay by playfully misinterpreting ”kicked him out” as ”kicked him out” of a relationship, creating a humorous twist on the biblical story of Jesus finding the lost sheep\.
- •Punchline Check Agent: Yes, the meme follows a setup and punchline structure\. The setup is the flock of sheep calling out to Jesus, thinking he’s found someone lost, while the punchline is Jesus saying, ”I know\. And I found her\.” The humor comes from the unexpected twist that Jesus is referring to a female companion, which is a surprising and playful take on the traditional narrative of Jesus finding the lost sheep\.
- •Image\-Text Alignment Agent: The image and text work together to create humor by presenting a humorous contradiction\. The text suggests that someone was kicked out of the group, while the image depicts Jesus holding a sheep with a rainbow\-colored headband, implying he found someone who wasn’t lost but rather part of the group\. This misalignment of expectations adds to the comedic effect\.
- •Hatefulness Detection Agent: The meme does not include offensive, derogatory, or hateful content\. However, it could potentially be seen as insensitive by some viewers due to its reference to transgender individuals, which might be considered harmful or marginalizing to those who identify with the LGBTQ\+ community\.
- •Intent Interpretation Agent: The creator likely intended to entertain with a humorous take on a biblical story, using irony and exaggeration to create a comedic effect\. The use of ”kicked him out” instead of ”lost” adds a layer of absurdity that is meant to be lighthearted rather than aggressive or critical\.

The misclassification appears to stem from the dominant influence of the intent interpretation agent’s output, which emphasized the meme’s humorous intent and overshadowed the caution flagged by the hatefulness detection agent\. This case illustrates a key limitation of the current attention\-weighting mechanism, where insufficiently calibrated weights can cause certain agent perspectives to disproportionately sway the final decision\. It underscores the importance of carefully balancing the contributions of each agent to ensure that signals indicating potential harm or offensiveness are not inadvertently diluted by agents emphasizing benign interpretations\.

## Appendix DInterpreting Agent\-Level Attention and Its Implications

This appendix analyzes the learned attention over the twelve role\-specialized agents and demonstrates how the model allocates evidential focus differently for humour versus hate detection\. The results substantiate two central claims of our work: \(i\) the multi\-agent, soft\-attention architecture yields task\-sensitive and human\-interpretable evidence integration, and \(ii\) humour and hate are not mutually exclusive; rather, they rely on distinct but complementary evidential pathways that our framework can disentangle and recombine\.

#### Which agents matter most per task\.

Figures[8](https://arxiv.org/html/2607.15442#A4.F8)and[9](https://arxiv.org/html/2607.15442#A4.F9)present ranked mean attention with 95% confidence intervals for the humour and hate heads, respectively\. For the humour head, the model assigns the highest weight to*Absurdity Agent*\(mean≈0\.223\\approx 0\.223, CI\[0\.215,0\.232\]\[0\.215,0\.232\]\), followed by*OCR & Text Extraction*\(mean≈0\.125\\approx 0\.125, CI\[0\.123,0\.127\]\[0\.123,0\.127\]\), with a long tail of humour\-oriented agents such as*Textual Irony*,*Visual Irony*,*Emotion Contrast*, and*Cultural Reference*\. For the hate head,*Hatefulness Detection*\(mean≈0\.199\\approx 0\.199, narrow CI around\[0\.198,0\.199\]\[0\.198,0\.199\]\) and*Intent Interpretation*\(mean≈0\.121\\approx 0\.121\) dominate, followed by*Emotion Contrast*and*Image–Text Alignment*\. These patterns are consistent with the qualitative nature of the tasks: humour often hinges on textual setups and incongruity \(captured by OCR and absurdity/irony agents\), whereas hate classification prioritizes targeted derogation and intent\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/attn-visualization/agent_attn_ranked_humour.png)Figure 8:Ranked mean attention \(with 95% CI\) over agents for the humour head\. The model emphasizes*Absurdity*and*OCR & Text Extraction*, followed by irony\-, contrast\-, and reference\-focused agents\.![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/attn-visualization/agent_attn_ranked_hate.png)Figure 9:Ranked mean attention \(with 95% CI\) over agents for the hate head\.*Hatefulness Detection*and*Intent Interpretation*lead, followed by*Emotion Contrast*and*Image–Text Alignment*\.
#### Humour and hate attend to different evidence\.

To directly compare tasks, Figure[10](https://arxiv.org/html/2607.15442#A4.F10)shows side\-by\-side means \(with 95% CI\) for each agent under the two heads, and Figure[11](https://arxiv.org/html/2607.15442#A4.F11)plots the difference*\(humour−\-hate\)*\. The largest positive deltas appear for*Absurdity*and*OCR*, confirming that humour is more text\-anchored and playfully illogical\. Conversely, the largest negative deltas occur for*Hatefulness Detection*and*Intent Interpretation*, indicating substantially higher emphasis under the hate head\. These asymmetries provide quantitative support that the model relies on partially disjoint cues for the two tasks while still allowing overlap when a meme is simultaneously humorous and harmful\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/attn-visualization/agent_compare_humour_vs_hate.png)Figure 10:Side\-by\-side comparison of mean attention \(with 95% CI\) for humour vs\. hate across all agents\. Humour places more weight on*Absurdity*and*OCR*; hate emphasizes*Hatefulness Detection*and*Intent Interpretation*\.![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/attn-visualization/agent_delta_humour_minus_hate.png)Figure 11:Attention differences \(humour−\-hate\)\. Positive bars indicate higher weight under humour; negative bars indicate higher weight under hate\. The largest positive deltas are for*Absurdity*and*OCR*; the largest negative deltas are for*Hatefulness Detection*and*Intent*\.
#### Compact overview across agents and tasks\.

Figure[12](https://arxiv.org/html/2607.15442#A4.F12)provides a concise heatmap summarizing mean attention per agent across the two tasks\. The humour head concentrates on*Absurdity*,*OCR*, and irony/contrast agents, whereas the hate head concentrates on*Hatefulness Detection*,*Intent*, and*Alignment/Contrast*\. This global view complements the ranked plots and makes the task\-specific specialisation particularly transparent\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/attn-visualization/agent_attn_heatmap_by_task.png)Figure 12:Heatmap of mean attention for each agent across \{humour, hate\}\. The humour head is text\- and absurdity\-oriented; the hate head is safety\- and intent\-oriented\.
#### Group\-level sanity and interpretability\.

Finally, Figure[13](https://arxiv.org/html/2607.15442#A4.F13)aggregates attention into three interpretable groups \(*general*,*humour*,*safety*\)\. The humour head assigns the largest mass to the humour group, with non\-trivial contributions from general cues \(e\.g\., OCR and scene description\)\. The hate head markedly increases the safety group’s share while still leveraging general and alignment/contrast information\. This group\-level analysis confirms that the learned attention respects intuitive priors and yields an interpretable division of labour across roles, supporting error analysis and governance use\-cases\.

![Refer to caption](https://arxiv.org/html/2607.15442v1/AnonymousSubmission/figures/attn-visualization/grouped_attention_by_task.png)Figure 13:Grouped attention \(mean with approximate CI\) for*general*,*humour*, and*safety*roles by task\. The humour head emphasizes humour roles; the hate head emphasizes safety roles\.
#### Top\-55agents per task\.

For quick reference, Table[11](https://arxiv.org/html/2607.15442#A4.T11)lists the top five agents for each task together with their mean attention and 95% CI \(rounded\)\. These values are consistent with Figures[8](https://arxiv.org/html/2607.15442#A4.F8)–[13](https://arxiv.org/html/2607.15442#A4.F13)and further illustrate that the attention mechanism is learning task\-aligned, human\-interpretable signals\.

Table 11:Top\-5 agents by mean attention for each task \(mean±\\pm95% CI\)\.TaskAgentMean95% CIHumourAbsurdity Agent0\.223\[0\.215, 0\.232\]OCR & Text Extraction0\.125\[0\.123, 0\.127\]Hatefulness Detection0\.086\[0\.085, 0\.088\]Textual Irony0\.085\[0\.084, 0\.086\]Intent Interpretation0\.067\[0\.067, 0\.068\]HateHatefulness Detection0\.199\[0\.198, 0\.199\]Intent Interpretation0\.121\[0\.121, 0\.122\]Emotion Contrast0\.085\[0\.085, 0\.085\]Image–Text Alignment0\.077\[0\.077, 0\.077\]Absurdity Agent0\.076\[0\.076, 0\.076\]

#### Implications for novelty and research objectives\.

The attention distributions document that the model does not rely on a single undifferentiated feature space; instead, it adaptively emphasizes role\-specialized evidence that is semantically aligned with each task\. This is precisely the motivation for our multi\-agent design and soft\-attention aggregator\. Moreover, the comparative analyses \(Figures[10](https://arxiv.org/html/2607.15442#A4.F10)–[11](https://arxiv.org/html/2607.15442#A4.F11)\) empirically support the claim that humour and hate are not mutually exclusive but attend to different cues; the architecture can therefore capture co\-occurrence by simultaneously up\-weighting both humour\- and safety\-relevant agents when necessary\. These findings reinforce the interpretability and governance value of our approach and substantiate the methodological contribution beyond accuracy\-only baselines\.

## Appendix EExplanator–Agent Alignment and Faithfulness \(Extended Analysis\)

MAR\-12 introduces explanation synthesizer that consumes \(i\) the 12 reasonings, \(ii\) the learned attention weights, and \(iii\) the system outcome, and then*explains*the final decision*in terms of*the attention\-weighted agent evidence\. The explanator’s job is not to discover attention but to faithfully narrate the model’s internal evidential path\.

#### Metrics\.

Per image and task \(humour, hate\), we compute:

1. 1\.Alignment\(align\_humour,align\_hate\): textual alignment between the explanator’s explanation and the 12 agents’ reasonings, weighted by the model’s learned attention\.
2. 2\.Margin\(margin\_humour,margin\_hate\): confidence margin of the classifier’s chosen label\.
3. 3\.Correctness\(hum\_correct,hate\_correct\): indicator of prediction accuracy\.

Table 12:Pearson correlations\. Alignment captures faithfulness to agent reasoning/attention rather than correctness, whereas margin correlates with correctness as expected\.PairPearsonrralignhumour\\mathrm\{align\_\{humour\}\}vs\.humcorrect\\mathrm\{hum\_\{correct\}\}0\.0300\.030alignhate\\mathrm\{align\_\{hate\}\}vs\.hatecorrect\\mathrm\{hate\_\{correct\}\}−0\.026\-0\.026marginhumour\\mathrm\{margin\_\{humour\}\}vs\.humcorrect\\mathrm\{hum\_\{correct\}\}0\.2410\.241marginhate\\mathrm\{margin\_\{hate\}\}vs\.hatecorrect\\mathrm\{hate\_\{correct\}\}0\.2670\.267alignhumour\\mathrm\{align\_\{humour\}\}vs\.alignhate\\mathrm\{align\_\{hate\}\}0\.4250\.425marginhumour\\mathrm\{margin\_\{humour\}\}vs\.marginhate\\mathrm\{margin\_\{hate\}\}−0\.053\-0\.053alignhumour\\mathrm\{align\_\{humour\}\}vs\.marginhumour\\mathrm\{margin\_\{humour\}\}−0\.045\-0\.045alignhate\\mathrm\{align\_\{hate\}\}vs\.marginhate\\mathrm\{margin\_\{hate\}\}−0\.033\-0\.033
#### Key findings\.

- •Faithfulness is disentangled from correctness\.Alignment is essentially uncorrelated with accuracy \(r≈0r\\approx 0in both tasks\), whereas the margin has the expected positive association with correctness \(Table[12](https://arxiv.org/html/2607.15442#A5.T12)\)\. This is desirable: the explainer is evaluated on whether it*follows*the attention\-weighted agent evidence, not on whether the classifier is right\. In contrast, the margin, as a confidence proxy, tracks correctness\.
- •Consistent explanatory style across tasks\.The moderate correlation between humour and hate alignments \(r=0\.425r=0\.425\) shows that MAR\-12 induces a stable, role\-aware explanatory pattern: when attention highlights certain agent roles, the explainer highlights the same roles irrespective of task\.
- •Margins reflect task difficulty, not shared structure\.The near\-zero correlation between humour and hate margins \(r=−0\.053r=\-0\.053\) indicates decision difficulty varies idiosyncratically by task/image; this further argues against using margins to assess explanation quality\.
- •Cross\-task reuse of humour signals on hate\.On the hate/safety task, the attention module frequently allocates non\-trivial mass to humour heads \(harm in memes is often mediated via humour\)\. The explainer mirrors this by citing humour\-related cues as part of its hate/safety explanations, evidence that explanations reflect the*same role mix*the model relied on\.
- •“Negative evidence” behavior\.When humour is absent, the attention module still consults humour heads to*verify lack of signal*; the explainer correspondingly points to those heads to justify a “not humorous” outcome\. Explanations thus remain faithful even for negative conclusions, not only for positive attributions\.
- •Stability across label quadrants\.Attention distributions \(especially for hate\) remain similar across HUM/HATE label combinations, and the explainer’s emphasis follows this stability, ruling out label\-chasing and supporting that the explainer narrates the attention\-determined process rather than the label itself\.

Table 13:Descriptive statistics for alignment and margin \(per image\)\.MetricMeanStdMinMaxalign​\_​humour\\mathrm\{align\\\_humour\}0\.11230\.11230\.03750\.03750\.000190\.000190\.364400\.36440align​\_​hate\\mathrm\{align\\\_hate\}0\.12840\.12840\.03950\.03950\.000200\.000200\.245270\.24527margin​\_​humour\\mathrm\{margin\\\_humour\}0\.70180\.70180\.26420\.26420\.001300\.001300\.987380\.98738margin​\_​hate\\mathrm\{margin\\\_hate\}0\.61390\.61390\.27170\.27170\.000190\.000190\.974550\.97455
#### Interpretation of magnitudes\.

Alignment values are concentrated in a narrow band \(means≈\\approx0\.11–0\.13; std≈\\approx0\.04; Table[13](https://arxiv.org/html/2607.15442#A5.T13)\), indicating*consistent grounding*of the explainer’s text in attention\-weighted agent rationales across items\. By contrast, margins vary widely \(std≈\\approx0\.26–0\.27\), reflecting heterogeneous decision difficulty—another reason alignment should not be conflated with correctness\.

#### Why this is novel in MAR\-12\.

1. 1\.Decoupled selection–narration\.By learning attention*outside*the explainer and requiring the explainer to condition on those weights, MAR\-12 makes faithfulness*testable*\. Our results show the explainer’s narrative tracks the external, role\-aware attention rather than inventing its own saliency\.
2. 2\.Role\-aware, multi\-agent alignment\.Explanations reflect domain structure \(humour, safety, general, …\) and naturally capture cross\-task coupling \(humour cues shaping hate decisions\) when attention says they matter\.
3. 3\.Faithfulness for negative conclusions\.The system surfaces*absence of evidence*via the same roles that would have been diagnostic, a property often missing in post\-hoc LLM explanations\.
4. 4\.Separation of concerns in evaluation\.Alignment measures*faithfulness*\(process\-level grounding\), while margin measures*confidence/accuracy*\. The empirical orthogonality between them is a distinctive empirical signature of our design\.

#### Implications\.

MAR\-12 yields explanations that are*causally grounded*in the model’s internal evidence flow \(agents→\\rightarrowattention→\\rightarrowclassifier\), independent of outcome accuracy, while preserving the intuitive link between confidence and correctness\. This supports transparent, accountable meme understanding: reviewers can inspect*which*roles the model relied on \(attention\) and*how*the explainer’s text follows that evidence \(alignment\)\.

Similar Articles

CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

arXiv cs.CL

Researchers from Peking University introduce CFMS, the first fine-grained Chinese multimodal sarcasm detection benchmark with 2,796 image-text pairs and a triple-level annotation framework (sarcasm identification, target recognition, explanation generation), along with a novel RL-augmented in-context learning method (PGDS) that significantly outperforms existing baselines.