Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes

arXiv cs.CL Papers

Summary

This paper presents a qualitative analysis of vision-language models for detecting hate speech in memes, evaluating their performance and reasoning under zero-shot and few-shot prompting.

arXiv:2608.26143v1 Announce Type: new Abstract: Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate. It remains very difficult to identify such complex and context-dependent hate speech. Although they display excellent performance on multimodal tasks, vision-language models (VLMs) tend to ignore context, irony, and other subtle cues that play a key role in identifying hateful memes. In this work, we present a qualitative analysis of four state-of-the-art VLMs: LLaVA-7B, Qwen-VL, GPT-4o mini, and Claude 3 Haiku. We evaluate these models under zero-shot and few-shot prompting to examine how contextual framing influences their outputs. Our analysis goes beyond simple classification accuracy and focuses on a qualitative evaluation of the models' generated justifications, providing a more in-depth understanding of their thought processes and constraints when dealing with hateful memes.
Original Article
View Cached Full Text

Cached at: 08/28/26, 09:20 AM

# Taylor and Francis Book Chapter
Source: [https://arxiv.org/html/2608.26143](https://arxiv.org/html/2608.26143)
###### Contents

1. [0Beyond Accuracy: A Qualitative Analysis of Vision\-Language Models for Hate Speech Detection in Memes](https://arxiv.org/html/2608.26143#id2)1. [1Introduction](https://arxiv.org/html/2608.26143#Ch0.S1) 2. [2Methodology](https://arxiv.org/html/2608.26143#Ch0.S2)1. [1Dataset](https://arxiv.org/html/2608.26143#Ch0.S2.SS1) 2. [2Model and Prompting Configurations](https://arxiv.org/html/2608.26143#Ch0.S2.SS2) 3. [3Task Formulation and Evaluation](https://arxiv.org/html/2608.26143#Ch0.S2.SS3) 3. [3Empirical Analysis and Findings](https://arxiv.org/html/2608.26143#Ch0.S3)1. [1Qualitative Error Analysis](https://arxiv.org/html/2608.26143#Ch0.S3.SS1)1. [Prompt Sensitivity and Reasoning Failure](https://arxiv.org/html/2608.26143#Ch0.S3.SS1.SSSx1) 2. [Correct Classification But Flawed Reasoning](https://arxiv.org/html/2608.26143#Ch0.S3.SS1.SSSx2) 3. [Contextual Misinterpretation](https://arxiv.org/html/2608.26143#Ch0.S3.SS1.SSSx3) 4. [Over\-sensitivity to Keywords](https://arxiv.org/html/2608.26143#Ch0.S3.SS1.SSSx4) 5. [Failure to Detect Coded or Nuanced Hate](https://arxiv.org/html/2608.26143#Ch0.S3.SS1.SSSx5) 2. [2Quantitative Performance Overview](https://arxiv.org/html/2608.26143#Ch0.S3.SS2) 3. [3Discussion](https://arxiv.org/html/2608.26143#Ch0.S3.SS3) 4. [4Conclusion & Future Scopes](https://arxiv.org/html/2608.26143#Ch0.S4)
2. [References](https://arxiv.org/html/2608.26143#bib)

\\chapterauthor

Muhammad Jawad Chowdhury[https://orcid.org/0009-0009-4129-546X](https://orcid.org/0009-0009-4129-546X), Adiba Hasan[https://orcid.org/0009-0006-9072-4054](https://orcid.org/0009-0006-9072-4054), Ishrak Hossain[https://orcid.org/0009-0004-1995-6122](https://orcid.org/0009-0004-1995-6122), Shahriar Ivan✉[https://orcid.org/0009-0001-2801-3028](https://orcid.org/0009-0001-2801-3028), and Sabbir Ahmed[https://orcid.org/0000-0001-5928-4886](https://orcid.org/0000-0001-5928-4886)Department of Computer Science and Engineering, Islamic University of Technology, Gazipur, Bangladesh\. \{jawad, adibahasan, ishrakhossain, shahriarivan, sabbirahmed\}@iut\-dhaka\.edu

## Chapter 0Beyond Accuracy: A Qualitative Analysis of Vision\-Language Models for Hate Speech Detection in Memes

\\chaptermark

Beyond Accuracy: A Qualitative Analysis of Vision\-Language Models for Hate Speech Detection in Memes

\\chapterinitial

Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems\. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate\. It remains very difficult to identify such complex and context\-dependent hate speech\. Although they display excellent performance on multimodal tasks, vision\-language models \(VLMs\) tend to ignore context, irony, and other subtle cues that play a key role in identifying hateful memes\. In this work, we present a qualitative analysis of four state\-of\-the\-art VLMs: LLaVA\-7B, Qwen\-VL, GPT\-4o mini, and Claude 3 Haiku\. We evaluate these models under zero\-shot and few\-shot prompting to examine how contextual framing influences their outputs\. Our analysis goes beyond simple classification accuracy and focuses on a qualitative evaluation of the models’ generated justifications, providing a more in\-depth understanding of their thought processes and constraints when dealing with hateful memes\.

### 1Introduction

As a form of cultural expression, memes are usually image\-based content with brief and humorous text, extensively spread online\. Though they are commonly used for entertainment, they can also be used to portray hate speech against individuals or a community, which has unfortunately become a very common issue in recent times\[[24](https://arxiv.org/html/2608.26143#bib.bib9)\]\. The growing societal concern around online hate has led to increasing efforts from both industry and academia to address this challenge\[[50](https://arxiv.org/html/2608.26143#bib.bib70),[21](https://arxiv.org/html/2608.26143#bib.bib71),[31](https://arxiv.org/html/2608.26143#bib.bib72),[36](https://arxiv.org/html/2608.26143#bib.bib1)\]\. This task becomes more challenging when the inherent context requires both image and textual reference\[[29](https://arxiv.org/html/2608.26143#bib.bib10)\]\. In the case of multimodal data, hate speech may arise from the combined interpretation of text and image rather than from the individual modalities taken into consideration separately\. Each modal may seem innocuous separately, but when taken as a whole, it can convey an offensive or hateful message when considered together\[[38](https://arxiv.org/html/2608.26143#bib.bib52),[13](https://arxiv.org/html/2608.26143#bib.bib74)\]\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/chapter1/figures/96180.png)

![Refer to caption](https://arxiv.org/html/2608.26143v1/chapter1/figures/98543.png)

![Refer to caption](https://arxiv.org/html/2608.26143v1/chapter1/figures/10285.png)

Figure 1:Demonstration of sample meme images fromHateful Memes Challenge Dataset \(HMCD\)\[[36](https://arxiv.org/html/2608.26143#bib.bib1)\], where, without understanding the context, it is hard to capture the underlying hate speech towards a targeted community\. Here, the memes are to be considered multimodally while understanding the context\.Recent advances in deep learning have revolutionized the modern era by providing state\-of\-the\-art solutions across a wide range of downstream tasks in computer vision\[[6](https://arxiv.org/html/2608.26143#bib.bib5),[47](https://arxiv.org/html/2608.26143#bib.bib4),[23](https://arxiv.org/html/2608.26143#bib.bib16),[5](https://arxiv.org/html/2608.26143#bib.bib15)\], natural language processing\[[34](https://arxiv.org/html/2608.26143#bib.bib20),[42](https://arxiv.org/html/2608.26143#bib.bib2)\], and related domains\[[27](https://arxiv.org/html/2608.26143#bib.bib13),[15](https://arxiv.org/html/2608.26143#bib.bib12),[33](https://arxiv.org/html/2608.26143#bib.bib8),[14](https://arxiv.org/html/2608.26143#bib.bib11)\]\. Building on these advances, modern language models have shown significant potential in terms of large\-scale knowledge representation, contextual reasoning, and semantic structure capture\.\[[8](https://arxiv.org/html/2608.26143#bib.bib14),[35](https://arxiv.org/html/2608.26143#bib.bib3),[9](https://arxiv.org/html/2608.26143#bib.bib7),[7](https://arxiv.org/html/2608.26143#bib.bib6)\]\. This has led to the emergence of large language models \(LLMs\)\[[37](https://arxiv.org/html/2608.26143#bib.bib24)\]and vision\-language models \(VLMs\)\[[26](https://arxiv.org/html/2608.26143#bib.bib25)\], which extend these capabilities to multimodal understanding\. LLMs and VLMs have shown strong performance in generating human\-like text and understanding complex multimedia content\[[19](https://arxiv.org/html/2608.26143#bib.bib68),[12](https://arxiv.org/html/2608.26143#bib.bib23),[1](https://arxiv.org/html/2608.26143#bib.bib22),[20](https://arxiv.org/html/2608.26143#bib.bib69)\]\. Since memes are inherently multimodal, often relying on sarcasm, cultural nuances, and intricate visual–textual interactions, they provide a particularly challenging testbed for assessing the real\-world reasoning abilities of such models\.

Recent advancements in multimodal learning have spurred interest in applying VLMs to this complex task of meme analysis\[[10](https://arxiv.org/html/2608.26143#bib.bib80),[30](https://arxiv.org/html/2608.26143#bib.bib73)\]\. The Hateful Memes Challenge dataset proposed by Kielaet al\.\[[36](https://arxiv.org/html/2608.26143#bib.bib1)\]introduced a benchmark, where non\-hateful modalities may form hateful messages individually when interpreted together\. Mathiaset al\.\[[39](https://arxiv.org/html/2608.26143#bib.bib63)\]extended this dataset by adding two sub\-tasks where the objective was to identify the hatred categories in memes\. Early research in this direction included using CLIP\[[46](https://arxiv.org/html/2608.26143#bib.bib55)\], which maps images and text onto a shared embedding space\. However, as a critical limitation, these systems, like HateCLIPper\[[38](https://arxiv.org/html/2608.26143#bib.bib52)\]and MemeCLIP\[[17](https://arxiv.org/html/2608.26143#bib.bib53)\], often struggled with detecting covert or culturally nuanced hate, as they relied more on representation matching than deep contextual reasoning\.

Modern pre\-trained VLMs, like Flamingo\[[10](https://arxiv.org/html/2608.26143#bib.bib80)\], LLaVA\[[40](https://arxiv.org/html/2608.26143#bib.bib31)\], and GPT\-4\[[44](https://arxiv.org/html/2608.26143#bib.bib81)\]have recently gained significant popularity\. Frameworks likeMemeGuard\[[32](https://arxiv.org/html/2608.26143#bib.bib28)\]used a fine\-tuned VLM for interpreting meme contexts, while Heeet al\.\[[28](https://arxiv.org/html/2608.26143#bib.bib62)\]proposed theHatReDdataset containing annotations of underlying hateful contextual reasons\. Several recent research endeavors have focused on effective prompting strategies to guide VLM reasoning\. Gavitet al\.\[[25](https://arxiv.org/html/2608.26143#bib.bib27),[41](https://arxiv.org/html/2608.26143#bib.bib51)\]provided a structured analysis of VLM capabilities on various meme classification tasks, comparing methods like Zero\-Shot learning\[[48](https://arxiv.org/html/2608.26143#bib.bib21)\], Few\-Shot learning\[[4](https://arxiv.org/html/2608.26143#bib.bib19),[2](https://arxiv.org/html/2608.26143#bib.bib18)\], and Chain\-of\-Thought \(CoT\)\[[52](https://arxiv.org/html/2608.26143#bib.bib17)\]\. Similarly, Zhuanget al\.\[[55](https://arxiv.org/html/2608.26143#bib.bib58)\]used chain\-of\-thought prompting with GPT\-4 to achieve state\-of\-the\-art results, and Van and Wu\[[49](https://arxiv.org/html/2608.26143#bib.bib56)\]introduced a definition\-guided prompting strategy\.

In addition to benchmark performance, a good amount of research has been done on the common failure modes of these models, providing a foundation for our qualitative analysis\[[30](https://arxiv.org/html/2608.26143#bib.bib73)\]\. The central challenge of contextual misinterpretation, where models fail to synthesize multimodal cues, was a key motivator for the Hateful Memes Challenge itself\[[36](https://arxiv.org/html/2608.26143#bib.bib1)\]\. This is often compounded by models exhibiting an over\-sensitivity to keywords, a form of unintended bias where the presence of a sensitive term can trigger a misclassification regardless of the benign context\[[22](https://arxiv.org/html/2608.26143#bib.bib83)\]\. Furthermore, studies have highlighted a persistent weakness in detecting coded or nuanced hate, particularly when it involves irony or sarcasm, which models often interpret literally\[[51](https://arxiv.org/html/2608.26143#bib.bib84)\]\. Perhaps most insidiously, researchers have identified the phenomenon of models achieving correct classifications through entirely flawed reasoning— a ‘right for the wrong reasons’ problem that masks a true lack of understanding\[[43](https://arxiv.org/html/2608.26143#bib.bib45)\]\. These established challenges underscore the limitations of relying solely on accuracy metrics and motivate a deeper, more diagnostic approach to VLM evaluation\[[16](https://arxiv.org/html/2608.26143#bib.bib85)\]\.

While the literature has identified these critical failure modes, most empirical evaluations of modern VLMs still prioritize a narrow set of quantitative metrics or test a limited range of models\. A critical gap remains in understanding how different model architectures \(e\.g\., open\-source and API\-based\) compare in their qualitative reasoning and failure modes under varied prompting conditions\. Most of the existing works focus on classification accuracy, where the important factors in the decision\-making process of different models were underexamined\. To address this gap, our work provides the following key contributions:

- Rigorous empirical evaluation of four modern VLMs: LLaVA\-7B, Qwen\-VL, GPT\-4o mini, and Claude 3 Haiku, to compare their capabilities\.
- Systematic analysis on the impact of prompting strategies by comparing a zero\-shot approach against a context\-rich few\-shot approach for each model, revealing their sensitivity to prompt design\.
- Detailed qualitative analysis demonstrating how established failure modes, such as flawed reasoning and contextual misinterpretation, manifest in these models\.

### 2Methodology

To evaluate the effectiveness of modern VLMs in detecting hate speech within multimodal memes, we conducted an empirical study using four state\-of\-the\-art models under two prompting strategies\. Our focus was on qualitative analysis to capture the nuances of model reasoning and failure modes, rather than relying solely on quantitative metrics\. Figure[2](https://arxiv.org/html/2608.26143#Ch0.F2)provides an overview of the approach, consisting of three components: 1\) Dataset description, 2\) Model and prompt settings, and 3\) Task formulation with evaluation metrics supporting the analysis\.

#### 1Dataset

We conducted our experiments using theHateful Memes Challenge Dataset \(HMCD\)\[[36](https://arxiv.org/html/2608.26143#bib.bib1)\], introduced in the NeurIPS 2020 Hateful Memes competition\. As meme classification often depends on the joint interpretation of visual and textual content, this benchmark is specifically designed to evaluate complex multimodal reasoning\. A key challenge of the dataset is the inclusion ofbenign confounders, which are instances that appear hateful when viewed through a single modality but are non\-hateful when both modalities are considered together\.

The full dataset contains10,000image\-text pairs labeled as eitherhatefulornon\-hateful\. Each sample falls into one of five semantic categories\[[36](https://arxiv.org/html/2608.26143#bib.bib1)\]:

- Random Non\-Hateful: Neutral memes without harmful content\.
- Benign Image Confounders: Appear hateful based on the caption alone, but the image negates the hateful implication\.
- Benign Text Confounders: Appear hateful from the image alone, but the caption clarifies the intent\.
- Unimodal Hate: Meme is hateful based solely on either the image or the text\.
- Multimodal Hate: Hatefulness emerges only when image and text are interpreted together\.

We experimented with a subset of500 memesfrom thedev\_seen\.jsonsplit\. This subset isclass\-balanced, consisting of250 hatefuland250 non\-hatefulsamples\. The selected examples reflect the diversity and complexity of the original dataset, including instances from all five semantic categories\. This is important as in this context of multimodal hate detection, understanding arises from analyzing both visual and textual content jointly\. Using balanced data ensures fair comparisons across models and prompts\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/x1.png)Figure 2:Overview of the empirical evaluation framework, which flows from dataset sampling and prompt configuration \(zero\-shot vs\. few\-shot\) to evaluation across four VLMs, culminating in the analysis of their classification and justification outputs\.
#### 2Model and Prompting Configurations

We evaluated four large VLMs, chosen to provide a balanced perspective on their diverse reasoning capabilities\. Our selection includes models known for high\-precision visual grounding \(Qwen2\-7B\)\[[3](https://arxiv.org/html/2608.26143#bib.bib82)\], strong general\-purpose inference \(GPT\-4o mini\[[45](https://arxiv.org/html/2608.26143#bib.bib36)\], Claude 3 Haiku\[[11](https://arxiv.org/html/2608.26143#bib.bib32)\]\), and foundational instruction\-following abilities \(LLaVA\-7B\)\[[40](https://arxiv.org/html/2608.26143#bib.bib31)\]\. This selection allows us to compare three distinct open\-source models against a highly capable, proprietary API\-based model \(Claude 3 Haiku\), which serves as a state\-of\-the\-art performance benchmark\. To assess the models’ sensitivity to prompt design, we use two distinct prompting strategies:zero\-shotandfew\-shot, commonly used in evaluating reasoning capabilities of LLMs\[[18](https://arxiv.org/html/2608.26143#bib.bib38)\]\.

Table 1:Sample Zero\-Shot Prompt for Meme ClassificationIn thezero\-shotsetting, models are provided only with task instructions to classify a meme ashatefulornon\-hateful, accompanied by a justification, with no examples\. In contrast, thefew\-shotprompt presents three labeled meme examples with justifications before the target instance\. This design enables in\-context learning and evaluates whether exposure to examples enhances predictive accuracy and reasoning consistency\.

#### 3Task Formulation and Evaluation

The task is formulated as a binary classification problem\. Let our dataset beD=\{\(Mi,yi\)\}i=1ND=\\\{\(M\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}, where each sample consists of a memeMi=\(Ii,Ti\)M\_\{i\}=\(I\_\{i\},T\_\{i\}\)composed of an imageIiI\_\{i\}and its associated textTiT\_\{i\}\. The ground truth label isyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}\. A Vision Language Model, defined as a functionfθf\_\{\\theta\}, analyzes the meme according to the given promptPjP\_\{j\}\(wherej∈\{zero,few\}j\\in\\\{\\text\{zero\},\\text\{few\}\\\}\) to produce a prediction:

fθ​\(Ii,Ti,Pj\)→\(y^i​j,Ji​j\)f\_\{\\theta\}\(I\_\{i\},T\_\{i\},P\_\{j\}\)\\rightarrow\(\\hat\{y\}\_\{ij\},J\_\{ij\}\)wherey^i​j∈\{0,1\}\\hat\{y\}\_\{ij\}\\in\\\{0,1\\\}is the predicted label andJi​jJ\_\{ij\}is the generated text justification\. Each model\-prompt configuration is evaluated on the same set of 500 memes\.

We performed a deep qualitative analysis of the models’ reasoning capabilities, which involved a manual review of the misclassified samples for each model\-prompt configuration, with a specific focus on the textual justification \(Ji​jJ\_\{ij\}\) provided by the model\. This analysis moves beyond simply identifyingifa model was wrong to understandingwhyit was wrong\.

The primary goal of this analysis is to identify recurring error archetypes and diagnose the underlying failure modes in the models’ reasoning\. Our manual review focused on categorizing these failures, looking for patterns such as:

- Prompt Sensitivity and Reasoning Failure:Investigates how a model’s reasoning pathway and final decision are influenced by the prompt design\[[54](https://arxiv.org/html/2608.26143#bib.bib44)\]\.
- Correct Classification But Flawed Reasoning :Instances where the model produces the correct classification label, but its justification reveals a flawed, irrelevant, or logically unsound reasoning process\[[43](https://arxiv.org/html/2608.26143#bib.bib45)\]\.
- Contextual Misinterpretation:Instances where the model correctly identifies individual elements \(e\.g\., objects in an image, keywords in text\) but fails to grasp the contextual meaning that makes the meme hateful or benign\[[36](https://arxiv.org/html/2608.26143#bib.bib1)\]\.
- Over\-sensitivity to Keywords:Errors where a model incorrectly flags benign content as hateful due to the presence of sensitive keywords \(e\.g\., related to race, religion, or gender\), even when used in a non\-hateful context\[[22](https://arxiv.org/html/2608.26143#bib.bib83)\]\.
- Failure to Detect Coded or Nuanced Hate:Cases where models miss hate speech that relies on subtle stereotypes, cultural in\-jokes, or coded language, rather than explicit slurs\[[51](https://arxiv.org/html/2608.26143#bib.bib84)\]\. This is particularly relevant for the ‘benign confounder’ memes in the HMCD\.

### 3Empirical Analysis and Findings

Understanding model behavior inmultimodal classification tasksrequires more than just numerical metrics\. To this end, we conduct a multi\-faceted analysis of four VLMs on hateful meme understanding\. Through manual inspection of model outputs, we can find patterns in both correct and incorrect classifications, as well as recurring inconsistencies in generated justifications\. Finally, we do a quantitative evaluation, where we report standard performance metrics to see how well the overall classification accuracy and robustness hold up across all model\-prompt configurations\.

#### 1Qualitative Error Analysis

Although quantitative measurements offer a broad perspective, a qualitative examination of misclassified memes is crucial for comprehending the particular failure modes of these sophisticated VLMs\. By reviewing the justifications provided by the models for their incorrect predictions, we identified several recurring error archetypes\.

##### Prompt Sensitivity and Reasoning Failure

A predominant failure mode identified was contextual misunderstanding, when a model’s reasoning was significantly affected by the prompt\. An illustrative instance is LLaVA\-7B’s examination of the meme in Figure[3](https://arxiv.org/html/2608.26143#Ch0.F3), which presents a horrible humor alluding to the Holocaust and provoked radically conflicting reactions from the model and the model’s classification transitioned fromnon\-hatefultohatefulupon altering the prompt from a zero\-shot to a few\-shot instruction\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/x2.png)Figure 3:For certain models, non\-representative few\-shot examples can be more detrimental than a zero\-shot prompt, as they can suppress the model’s ability to apply its own knowledgeAnalysis: This drastic shift reveals a critical vulnerability we termExample\-Induced Bias, a phenomenon related to the known sensitivity of large models to in\-context examples\[[53](https://arxiv.org/html/2608.26143#bib.bib86)\]\. In the zero\-shot setting, the model’s predictiony^\\hat\{y\}for a memeMMis primarily conditioned on its vast internal knowledge base,𝒦\\mathcal\{K\}, followingP​\(y^\|M,𝒦\)P\(\\hat\{y\}\|M,\\mathcal\{K\}\)\. In the few\-shot setting, however, the decision becomes heavily conditioned on the small set of provided examples,ℰ\\mathcal\{E\}\. The model’s behavior suggests that it heavily discounts its internal knowledge\. Ideally, the model should integrate both sources of information, but the observed failure can be expressed as the model’s reasoning collapsing from the ideal to the biased state:

P​\(y^\|M,𝒦,ℰ\)→P​\(y^\|M,ℰ\)P\(\\hat\{y\}\|M,\\mathcal\{K\},\\mathcal\{E\}\)\\rightarrow P\(\\hat\{y\}\|M,\\mathcal\{E\}\)As the examples inℰ\\mathcal\{E\}did not contain this specific type of historical hate imagery, the model showed overfitting to their superficial patterns, defaulted to a naive visual description, and failed to recognize the unambiguous hate\. This demonstrates that for certain models, non\-representative few\-shot examples can be more detrimental than a zero\-shot prompt, as they can suppress the model’s ability to apply its own knowledge\.

##### Correct Classification But Flawed Reasoning

Instances where a model arrives at the correct classification but for the incorrect reasons are more revealing than plain mistakes\. This phenomenon, often described as ‘getting the right answer for the wrong reason’\. Figure[4](https://arxiv.org/html/2608.26143#Ch0.F4)shows such an example\. It exposes the superficial nature of the model’s reasoning process and its reliance on flawed heuristics or hallucinated evidence\. Such cases are particularly deceptive because a simple accuracy check would mark them as a success, masking the model’s actual contextual ignorance\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/chapter1/figures/Two_Example_memes.png)Figure 4:Hateful memes correctly classified based on flawed or hallucinated reasoning\. Models here gave the correct result, but analyzing their justifications reveals their decision is based on the wrong reasons, most commonly for their tendency to focus on a benign/non\-existent text template while ignoring the hateful theme\.Analysis:These two cases reveal a common underlying failure: the models are not performing genuine multimodal reasoning but are instead using flawed shortcuts\. We can formalize this by considering that a model’s decision functionfθf\_\{\\theta\}should operate on the true features of a memeMM, which is composed of its imageIIand its actual textTa​c​t​u​a​lT\_\{actual\}\. The correct reasoning pathway is:

y^=fθ​\(I,Ta​c​t​u​a​l\)→hateful\\hat\{y\}=f\_\{\\theta\}\(I,T\_\{actual\}\)\\rightarrow\\text\{hateful\}However, the models’ justifications show they followed incorrect pathways, operating on a fabricated or incomplete feature setM′M^\{\\prime\}whereM′≠MM^\{\\prime\}\\neq M\.

In the first case \(Figure[4](https://arxiv.org/html/2608.26143#Ch0.F4)a\), GPT\-4o\-mini correctly classified the racist meme ashatefulbut justified its decision by misinterpreting the neutral meme format ‘white people is this X’ as derogatory, completely ignoring the actual hateful phrase ‘is this a shooting range\.’ In the second case \(Figure[4](https://arxiv.org/html/2608.26143#Ch0.F4)b\), LLaVA also correctly identified the misogynistic meme ashatefulbut based its justification on a hallucinated phrase \(‘the most dangerous people on the planet’\) rather than the actual text present in the image\. In both instances, the model focused on a benign or non\-existent text template while ignoring the hateful payload\.

These cases highlight a critical challenge: models can be correct for the wrong reasons\. Analyzing their justifications is therefore essential to verify true understanding and rule out ‘success’ from flawed heuristics or hallucination\.

##### Contextual Misinterpretation

The incapacity of all models to infer the practical knowledge required to decode hateful subtext was a crucial flaw\. They took the meme in Figure[5](https://arxiv.org/html/2608.26143#Ch0.F5)literally, but they were unable to make the necessary logical leap to recognize its cruel mockery of a child who is afflicted with cancer\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/chapter1/figures/Example_Memes.png)Figure 5:Hateful Meme Requiring Real\-World Inference\. Models here fail to understand the context of the meme and identify the scenario as a ‘joke’, failing to access the necessary real\-world knowledge\.As shown in Figure[5](https://arxiv.org/html/2608.26143#Ch0.F5), every model classified the meme as non\-hateful, providing justifications that indicate a surface\-level reading\.

Analysis: The models’ justifications reveal a shared cognitive gap: they identify the scenario as a ‘joke’ but fail to access the necessary real\-world knowledge\. We can formalize this failure by defining the meme’s features as having both explicit \(FexplicitF\_\{\\text\{explicit\}\}\) and implicit \(FimplicitF\_\{\\text\{implicit\}\}\) components\.

The models’ decision\-making appears limited to only the explicit features, approximating:

fθ​\(M\)≈fθ​\(Fexplicit\)f\_\{\\theta\}\(M\)\\approx f\_\{\\theta\}\(F\_\{\\text\{explicit\}\}\)whereas a successful function must operate on both\. This inability to integrate commonsense knowledge is a critical limitation\.

##### Over\-sensitivity to Keywords

A key source of false positives wasKeyword Over\-Sensitivity, where models flagged benign content as hateful due to the presence of a sensitive term\. For instance, the neutral meme in Figure[6](https://arxiv.org/html/2608.26143#Ch0.F6)was misclassified by both LLaVA\-7B and Claude 3 Haiku on this basis\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/x3.png)Figure 6:Justifications for Benign Meme, models misinterpreted harmless content as hateful due to the presence of some sensitive terms involving race, religion, and others\.The justifications provided by the models, shown in Figure[6](https://arxiv.org/html/2608.26143#Ch0.F6), reveal that their decisions were triggered almost exclusively by the presence of the word ‘Jewish’\.

Analysis: This case demonstrates a critical failure where both models hallucinated derogatory language\. The models’ decision function,fθ​\(M\)f\_\{\\theta\}\(M\), appears to have been disproportionately weighted by a single sensitive keyword,wsw\_\{s\}, such that:

fθ​\(M\)≈fθ​\(ws\)f\_\{\\theta\}\(M\)\\approx f\_\{\\theta\}\(w\_\{s\}\)This suggests the models’ strong safety alignment for thekeywordoverrode the neutral context of the meme, causing a false positive\.

##### Failure to Detect Coded or Nuanced Hate

A key limitation observed was the models’ failure to detect nuanced hate that requires external, real\-world knowledge, a challenge central to the HMCD’s design\. As shown in Figure[7](https://arxiv.org/html/2608.26143#Ch0.F7), models performed a literal analysis but missed the hateful subtext\. The cruel irony of pairing the text ‘family trip in Mexico’ with a distressing image of an immigrant family was lost on both Qwen2 and Claude 3 Haiku, who incorrectly classified the meme as non\-hateful\.

![Refer to caption](https://arxiv.org/html/2608.26143v1/x4.png)Figure 7:Failure Case: Inter\-Model Comparison on Coded or Nuanced Hate\. Both Qwen2\-7B and Claude 3 Haiku failed to detect the subtle hateful intent in this meme, highlighting limitations in interpreting implicit or ironic content\.Analysis: A notable cognitive gap is revealed by this case study, where models carried out a literal analysis, recognizing the term ‘family trip’ but not deducing from the visual cues the emotional context of distress\. This failure can be formalized by defining the meme’s features as having explicit \(FexplicitF\_\{\\text\{explicit\}\}\) and implicit \(𝒞context\\mathcal\{C\}\_\{\\text\{context\}\}\) components\. Models’ decision\-making was limited to the explicit features, approximating:

fθ​\(M\)≈fθ​\(Fexplicit\)f\_\{\\theta\}\(M\)\\approx f\_\{\\theta\}\(F\_\{\\text\{explicit\}\}\)whereas a successful function must recognize the ironic dissonance between them\. This inability to move beyond literal interpretation to understand that the meme’s ‘humor’ is a form of dehumanization highlights a key frontier for VLM development\.

#### 2Quantitative Performance Overview

To form an empirical baseline, we evaluated the four VLMs on our balanced test set using both zero\-shot and few\-shot prompting strategies\. We used standard classification metrics, such as Precision, Recall, and F1\-Score, to evaluate how well each model worked, for bothhatefulandnon\-hatefulclasses shown in Table[2](https://arxiv.org/html/2608.26143#Ch0.T2)\.

Table 2:Performance comparison of VLMs on Hateful Meme Detection under zero\-shot and few\-shot conditions\. For the critical ‘hateful’ class, the highest and lowest scores in each column are marked in bold and underlined, respectively\.Analysis of Quantitative Results: The quantitative data reveals significant heterogeneity in the performance profiles of the evaluated models\. This may be because of being trained on different sets of data, having different architectural designs, and using different alignment procedures\.

1. 1\.Overall Performance and Reasoning Quality:The community version of GPT\-4o\-mini consistently was the top\-performing model, achieving the highest F1\-Score for the ‘hateful’ class in both prompt configurations \(0\.59 and 0\.62\)\. Its superior performance showed a proper balance between precision and recall\. This suggests that its architecture is better suited for complex and inferential reasoning\. This finding is consistent with recent studies that show the official GPT\-4o model has demonstrated state\-of\-the\-art accuracy at detecting multimodal hate speech\[[49](https://arxiv.org/html/2608.26143#bib.bib56)\]\. We posit that its pre\-training on a large and varied set of internet data, which contains a lot of different cultural settings and meme formats, is a good fit for the specific problems that the HMCD dataset presents\.
2. 2\.The Precision\-Recall Trade\-off:A clear trade\-off emerges among the open\-source models\.Qwen2\-7Bachieved the highest precision \(0\.70 and 0\.71\), likely due to strong OCR and factual grounding, making it reliable for text\-based hate\. However, its recall was the lowest \(0\.34 and 0\.32\), missing roughly 70% of hateful content\. This statistical profile is directly supported by our qualitative findings\. For example, in our analysis of nuanced hate \(Figure[7](https://arxiv.org/html/2608.26143#Ch0.F7)\), Qwen2\-7B performed a perfect literal analysis of the ‘family trip in Mexico’ meme, but completely failed to grasp the cruel, ironic subtext required to identify it as hateful\. This tendency to miss non\-explicit hate in ambiguous cases is the primary driver of its low recall score\.

#### 3Discussion

Our results highlight a clear gap between how well the models seem to understand the task and how they actually perform\. Even the well\-performed model, GPT\-4o mini, did not always rely on solid reasoning to achieve its high scores\. Instead, all models frequently repeated the same types of mistakes—often using flawed logic to land on the right answer or failing to incorporate real\-world knowledge\. This pattern points to a dependence on shallow strategies rather than genuine understanding\. For example, Qwen2\-7B showed high precision but very low recall, meaning it avoided errors by taking an overly literal approach, but in doing so, missed most subtle cases of hate\. These findings suggest that accuracy alone is not a reliable measure of a model’s true ability to moderate content\.

### 4Conclusion & Future Scopes

Our quantitative findings show that, despite recent VLMs demonstrating impressive progress, performance discrepancies persist across architectures\. GPT\-4o\-mini exhibited the most balanced performance\. More importantly, our qualitative analysis reveals that all tested models have significant vulnerabilities, limiting reliability for real\-world content moderation\. A key limitation is their consistent failure to infer unstated, real\-world knowledge, with reasoning largely relying on the literal meaning of images and text and missing essential common\-sense and emotional understanding\. While VLMs are promising tools, they are not yet capable of reliably identifying complex multimodal hate speech, lacking the deep inferential skills necessary to interpret subtle expressions of hate and remaining fragile under contextual variations\.

Future research can focus on targeted evaluations and enhancing reasoning capabilities\. Extending studies to alternative prompting strategies, larger model families, and multilingual datasets can reveal whether observed failure modes are language\-specific or universal\. Assessing performance on specialized datasets, such as MAMI, can help determine whether models can identify precise targets and sub\-categories of hate\. Additionally, developing mitigation strategies can directly address observed error archetypes, such as ‘Example\-Induced Bias’ and ‘Flawed Reasoning’\. Collectively, these directions offer a path toward more reliable VLMs for content moderation and robust multimodal reasoning\.

## References

- \[1\]\(2024\)Performance evaluation of large language models in bangla consumer health query summarization\.In2024 27th International Conference on Computer and Information Technology \(ICCIT\),Vol\.,pp\. 2748–2753\.External Links:[Document](https://dx.doi.org/10.1109/ICCIT64611.2024.11022034),[Link](https://ieeexplore.ieee.org/abstract/document/11022034)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[2\]M\. Ahamed, R\. B\. Kabir, T\. T\. Dipto, M\. Al Mushabbir, S\. Ahmed, and Md\. H\. Kabir\(2024\)Performance analysis of few\-shot learning approaches for bangla handwritten character and digit recognition\.In2024 6th International Conference on Sustainable Technologies for Industry 5\.0 \(STI\),Vol\.,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/STI64222.2024.10951048),[Link](https://ieeexplore.ieee.org/abstract/document/10951048)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[3\]I\. Ahmed, S\. Islam, P\. P\. Datta, I\. Kabir, N\. U\. R\. Chowdhury, and A\. Haque\(2025\)Qwen 2\.5: a comprehensive review of the leading resource\-efficient llm with potentioal to surpass all competitors\.Authorea Preprints\.Cited by:[§2](https://arxiv.org/html/2608.26143#Ch0.S2.SS2.p1.1)\.
- \[4\]S\. Ahmed, Md\. B\. Hasan, T\. Ahmed, and Md\. H\. Kabir\(2025\)DExNet: combining observations of domain adapted critics for leaf disease classification with limited data\.External Links:2506\.18173,[Link](https://arxiv.org/abs/2506.18173)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[5\]S\. Ahmed, Md\. B\. Hasan, T\. Ahmed, Md\. R\. K\. Sony, and Md\. H\. Kabir\(2022\)Less is more: lighter and faster deep neural architecture for tomato leaf disease classification\.IEEE Access10\(\),pp\. 68868–68884\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2022.3187203),[Link](https://ieeexplore.ieee.org/document/9810234)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[6\]T\. Ahmed, Md\. B\. Hasan, S\. Ahmed, and Md\. H\. Kabir\(2024\)ExE\-Net: explainable ensemble network for potato leaf disease classification\.In2024 IEEE Canadian Conference on Electrical and Computer Engineering \(CCECE\),Vol\.,pp\. 335–339\.External Links:[Document](https://dx.doi.org/10.1109/CCECE59415.2024.10667205),[Link](https://ieeexplore.ieee.org/abstract/document/10667205)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[7\]T\. Ahmed, S\. Ivan, M\. Kabir, H\. Mahmud, and K\. Hasan\(2022\-08\-02\)Performance analysis of transformer\-based architectures and their ensembles to detect trait\-based cyberbullying\.Social Network Analysis and Mining12\(1\),pp\. 99\.External Links:ISSN 1869\-5469,[Document](https://dx.doi.org/10.1007/s13278-022-00934-4),[Link](https://doi.org/10.1007/s13278-022-00934-4)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[8\]T\. Ahmed, S\. Ivan, A\. Munir, and S\. Ahmed\(2024\)Decoding depression: analyzing social network insights for depression severity assessment with transformers and explainable ai\.Natural Language Processing Journal7,pp\. 100079\.External Links:ISSN 2949\-7191,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.nlp.2024.100079),[Link](https://www.sciencedirect.com/science/article/pii/S294971912400027X)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[9\]T\. Ahmed, M\. Kabir, S\. Ivan, H\. Mahmud, and K\. Hasan\(2021\)Am i being bullied on social media? an ensemble approach to categorize cyberbullying\.In2021 IEEE International Conference on Big Data \(Big Data\),Vol\.,pp\. 2442–2453\.External Links:[Document](https://dx.doi.org/10.1109/BigData52589.2021.9671594)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[10\]J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds, and e\. al\. Ring\(2022\)Flamingo: a visual language model for few\-shot learning\.Vol\.35,pp\. 23716–23731\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/edc2a8c8c4d2b96c9212e26cc9b894f6-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[11\]Anthropic\(2024\)Claude 3 haiku: our fastest model yet\.Note:Accessed: 2025\-07\-13External Links:[Link](https://www.anthropic.com/news/claude-3-haiku)Cited by:[§2](https://arxiv.org/html/2608.26143#Ch0.S2.SS2.p1.1)\.
- \[12\]N\. H\. Arif, S\. Rabby, M\. H\. H\. Papon, and S\. Ahmed\(2025\)Preemptive hallucination reduction: an input\-level approach for multimodal language model\.External Links:2505\.24007,[Link](https://arxiv.org/abs/2505.24007)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[13\]G\. Arya, M\. K\. Hasan, A\. Bagwari, N\. Safie, S\. Islam, F\. R\. A\. Ahmed, A\. De, M\. A\. Khan, and T\. M\. Ghazal\(2024\)Multimodal hate speech detection in memes using contrastive language\-image pre\-training\.IEEE Access12\(\),pp\. 22359–22375\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2024.3361322)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1)\.
- \[14\]A\. B\. M\. Ashikur Rahman, Md\. B\. Hasan, S\. Ahmed, T\. Ahmed, Md\. H\. Ashmafee, M\. R\. Kabir, and Md\. H\. Kabir\(2022\)Two decades of bengali handwritten digit recognition: a survey\.IEEE Access10\(\),pp\. 92597–92632\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2022.3202893),[Link](https://ieeexplore.ieee.org/document/9869842)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[15\]S\. Aziz, N\. H\. Arif, S\. Ahbab, S\. Ahmed, T\. Ahmed, and Md\. H\. Kabir\(2023\)Improved speech emotion recognition in bengali language using deep learning\.In2023 26th International Conference on Computer and Information Technology \(ICCIT\),Vol\.,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/ICCIT60459.2023.10441053),[Link](https://ieeexplore.ieee.org/abstract/document/10441053)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[16\]E\. M\. Bender, T\. Gebru, A\. McMillan\-Major, and S\. Shmitchell\(2021\)On the dangers of stochastic parrots: can language models be too big?\.InProceedings of the 2021 ACM conference on fairness, accountability, and transparency,pp\. 610–623\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p5.1)\.
- \[17\]S\. Bikram Shah, S\. Shiwakoti, M\. Chaudhary, and H\. Wang\(2024\)MemeCLIP: leveraging clip representations for multimodal meme classification\.arXiv e\-prints,pp\. arXiv–2409\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1)\.
- \[18\]T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.Cited by:[§2](https://arxiv.org/html/2608.26143#Ch0.S2.SS2.p1.1)\.
- \[19\]L\. Chen, J\. Li, X\. Dong, P\. Zhang, C\. He, J\. Wang, F\. Zhao, and D\. Lin\(2023\)ShareGPT4V: improving large multi\-modal models with better captions\.External Links:2311\.12793,[Link](https://arxiv.org/abs/2311.12793)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[20\]W\. Dai, J\. Li, D\. Li, A\. Tiong, J\. Zhao, W\. Wang, B\. Li, P\. N\. Fung, and S\. Hoi\(2023\)Instructblip: towards general\-purpose vision\-language models with instruction tuning\.Advances in neural information processing systems36,pp\. 49250–49267\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[21\]A\. Das, J\. S\. Wahi, and S\. Li\(2020\)Detecting hate speech in multi\-modal memes\.arXiv preprint arXiv:2012\.14891\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1)\.
- \[22\]L\. Dixon, J\. Li, J\. Sorensen, N\. Thain, and L\. Vasserman\(2018\)Measuring and mitigating unintended bias in text classification\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,pp\. 67–73\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p5.1),[4th item](https://arxiv.org/html/2608.26143#Ch0.S2.I2.i4.p1.1)\.
- \[23\]T\. R\. Fuad, S\. Ahmed, and S\. Ivan\(2025\)AQUA20: a benchmark dataset for underwater species classification under challenging conditions\.External Links:2506\.17455,[Link](https://arxiv.org/abs/2506.17455)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[24\]A\. Gandhi, P\. Ahir, K\. Adhvaryu, P\. Shah, R\. Lohiya, E\. Cambria, S\. Poria, and A\. Hussain\(2024\)Hate speech detection: a comprehensive review of recent works\.Expert Systems41\(8\),pp\. e13562\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/exsy.13562),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1111/exsy.13562),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1111/exsy\.13562Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1)\.
- \[25\]D\. Gavit, D\. Mazumder, S\. Das, and J\. Patro\(2025\)On vlms for diverse tasks in multimodal meme classification\.arXiv preprint arXiv:2505\.20937\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[26\]A\. Ghosh, A\. Acharya, S\. Saha, V\. Jain, and A\. Chadha\(2024\)Exploring the frontier of vision\-language models: a survey of current methodologies and future directions\.arXiv preprint arXiv:2404\.07214\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[27\]Md\. B\. Hasan, T\. Ahmed, S\. Ahmed, and Md\. H\. Kabir\(2023\)GaitGCN\+\+: improving gcn\-based gait recognition with part\-wise attention and dropgraph\.Journal of King Saud University \- Computer and Information Sciences35\(7\),pp\. 101641\.External Links:ISSN 1319\-1578,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jksuci.2023.101641),[Link](https://www.sciencedirect.com/science/article/pii/S1319157823001957)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[28\]M\. S\. Hee, W\. Chong, and R\. K\. W\. Lee\(2023\)Decoding the underlying meaning of multimodal hateful memes\.arXiv preprint arXiv:2305\.17678\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[29\]P\. C\. d\. Q\. Hermida and E\. M\. d\. Santos\(2023\-11\-01\)Detecting hate speech in memes: a review\.Artificial Intelligence Review56\(11\),pp\. 12833–12851\.External Links:ISSN 1573\-7462,[Document](https://dx.doi.org/10.1007/s10462-023-10459-7),[Link](https://doi.org/10.1007/s10462-023-10459-7)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1)\.
- \[30\]S\. Ivan, T\. Ahmed, S\. Ahmed, and M\. H\. Kabir\(2024\)A vision\-language multimodal framework for detecting hate speech in memes\.In2024 IEEE Canadian Conference on Electrical and Computer Engineering \(CCECE\),pp\. 464–468\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p5.1)\.
- \[31\]S\. Ivan\(2024\)Hate speech detection from multimodal memes using vision\-language transformer models\.Ph\.D\. Thesis,Department of Computer Science and Engineering \(CSE\), Islamic University of …\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1)\.
- \[32\]P\. Jha, R\. Jain, K\. Mandal, A\. Chadha, S\. Saha, and P\. Bhattacharyya\(2024\)Memeguard: an llm and vlm\-based framework for advancing content moderation via meme intervention\.arXiv preprint arXiv:2406\.05344\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[33\]A\. M\. Khan, A\. Ashrafee, R\. Sayera, S\. Ivan, and S\. Ahmed\(2022\)Rethinking cooking state recognition with vision transformers\.In25th International Conference on Computer and Information Technology \(ICCIT\),Vol\.,pp\. 170–175\.External Links:[Document](https://dx.doi.org/10.1109/ICCIT57492.2022.10055869),[Link](https://ieeexplore.ieee.org/abstract/document/10055869)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[34\]A\. Khan, F\. Kamal, M\. A\. Chowdhury, T\. Ahmed, M\. T\. R\. Laskar, and S\. Ahmed\(2023\-12\)BanglaCHQ\-summ: an abstractive summarization dataset for medical queries in Bangla conversational speech\.InProceedings of the First Workshop on Bangla Language Processing \(BLP\-2023\),Singapore,pp\. 85–93\.External Links:[Link](https://aclanthology.org/2023.banglalp-1.10),[Document](https://dx.doi.org/10.18653/v1/2023.banglalp-1.10)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[35\]A\. Khan, F\. Kamal, N\. Nower, T\. Ahmed, S\. Ahmed, and T\. Chowdhury\(2023\-12\)NERvous about my health: constructing a Bengali medical named entity recognition dataset\.InFindings of the Association for Computational Linguistics: EMNLP 2023,Singapore,pp\. 5768–5774\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.383),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.383)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[36\]D\. Kiela, H\. Firooz, A\. Mohan, V\. Goswami, A\. Singh, P\. Ringshia, and D\. Testuggine\(2020\)The hateful memes challenge: detecting hate speech in multimodal memes\.Advances in neural information processing systems33,pp\. 2611–2624\.Cited by:[Figure 1](https://arxiv.org/html/2608.26143#Ch0.F1),[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p5.1),[3rd item](https://arxiv.org/html/2608.26143#Ch0.S2.I2.i3.p1.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S2.SS1.p1.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S2.SS1.p2.1)\.
- \[37\]T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. Iwasawa\(2022\)Large language models are zero\-shot reasoners\.Advances in neural information processing systems35,pp\. 22199–22213\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[38\]G\. K\. Kumar and K\. Nandakumar\(2022\)Hate\-clipper: multimodal hateful meme classification based on cross\-modal interaction of clip features\.arXiv preprint arXiv:2210\.05916\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1),[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1)\.
- \[39\]M\. Lambert, S\. Nie, A\. Mostafazadeh Davani, D\. Kiela, V\. Prabhakaran, B\. Vidgen, and Z\. Waseem\(2021\)Findings of the woah 5 shared task on fine\-grained hateful memes detection\.InProceedings of the 5th Workshop on Online Abuse and Harms \(WOAH 2021\),pp\. 201–206\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1)\.
- \[40\]H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee\(2023\)Visual instruction tuning\.Advances in neural information processing systems36,pp\. 34892–34916\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1),[§2](https://arxiv.org/html/2608.26143#Ch0.S2.SS2.p1.1)\.
- \[41\]J\. Liu, R\. Tong, A\. Shen, S\. Li, C\. Yang, and L\. Xu\(2025\)MemeBLIP2: a novel lightweight multimodal system to detect harmful memes\.arXiv preprint arXiv:2504\.21226\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[42\]R\. Mahbub, I\. Khan, S\. Anuva, M\. S\. Shahriar, M\. T\. R\. Laskar, and S\. Ahmed\(2023\-12\)Unveiling the essence of poetry: introducing a comprehensive dataset and benchmark for poem summarization\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 14878–14886\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.920),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.920)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[43\]T\. Niven and H\. Kao\(2019\)Probing neural network comprehension of natural language arguments\.InProceedings of ACL,pp\. 4658–4664\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p5.1),[2nd item](https://arxiv.org/html/2608.26143#Ch0.S2.I2.i2.p1.1)\.
- \[44\]OpenAI\(2023\)GPT\-4 technical report\.CoRRabs/2303\.08774\.External Links:[Link](https://arxiv.org/abs/2303.08774),2303\.08774Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[45\]OpenAI\(2024\)GPT\-4o technical report\.External Links:[Link](https://openai.com/research/gpt-4o)Cited by:[§2](https://arxiv.org/html/2608.26143#Ch0.S2.SS2.p1.1)\.
- \[46\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p3.1)\.
- \[47\]S\. R\. Raiyan, Z\. Z\. Amio, and S\. Ahmed\(2025\)HaSPer: an image repository for hand shadow puppet recognition\.InICCV 2025 Workshop on Cultural Continuity of Artists,External Links:[Link](https://openreview.net/forum?id=AAGhroilqL)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p2.1)\.
- \[48\]T\. Tajwar, M\. Rahman, T\. A\. Chowdhury, S\. Ahmed, M\. Farazi, and Md\. H\. Kabir\(2023\)Improving zero\-shot semantic segmentation using dynamic kernels\.In2023 International Conference on Digital Image Computing: Techniques and Applications \(DICTA\),Vol\.,pp\. 395–402\.External Links:[Document](https://dx.doi.org/10.1109/DICTA60407.2023.00061),[Link](https://ieeexplore.ieee.org/abstract/document/10410952)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[49\]M\. Van and X\. Wu\(2025\)Detecting and mitigating hateful content in multimodal memes with vision\-language models\.arXiv preprint arXiv:2505\.00150\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1),[item 1](https://arxiv.org/html/2608.26143#Ch0.S3.I1.i1.p1.1)\.
- \[50\]R\. Velioglu and J\. Rose\(2020\)Detecting hate speech in memes using multimodal deep learning approaches: prize\-winning solution to hateful memes challenge\.arXiv preprint arXiv:2012\.12975\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p1.1)\.
- \[51\]B\. Vidgen, A\. Harris, D\. Nguyen, R\. Tromble, S\. Hale, and H\. Margetts\(2019\)Challenges and frontiers in abusive content detection\.InProceedings of the third workshop on abusive language online,Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p5.1),[5th item](https://arxiv.org/html/2608.26143#Ch0.S2.I2.i5.p1.1)\.
- \[52\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.
- \[53\]W\. M\. I\. L\. WorkRethinking the role of demonstrations: what makes in\-context learning work?\.Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S3.SS1.SSSx1.p2.5)\.
- \[54\]Z\. Zhao, E\. Wallace, S\. F\. Wang, S\. Singh, and M\. Gardner\(2021\)Calibrate before use: improving few\-shot performance of language models\.InProceedings of ICML,Cited by:[1st item](https://arxiv.org/html/2608.26143#Ch0.S2.I2.i1.p1.1)\.
- \[55\]Y\. Zhuang, K\. Guo, J\. Wang, Y\. Jing, X\. Xu, W\. Yi, M\. Yang, B\. Zhao, and H\. Hu\(2025\-01\)I know what you meme\! understanding and detecting harmful memes with multimodal large language models\.pp\.\.External Links:[Document](https://dx.doi.org/10.14722/ndss.2025.240415)Cited by:[§1](https://arxiv.org/html/2608.26143#Ch0.S1.p4.1)\.

Similar Articles