PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
Summary
Introduces PHANTOM, a large-scale open-source dataset of pre-generated adversarial attacks for vision-language models, covering 1010 high-level categories and 55 subcategories of harmful intents with 47,524 adversarial samples. The dataset aims to lower the barrier for adversarial research and enable systematic evaluation of VLM robustness and safety.
View Cached Full Text
Cached at: 06/24/26, 07:46 AM
# PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
Source: [https://arxiv.org/html/2606.24388](https://arxiv.org/html/2606.24388)
Hossein KhodadadiThe Italian Institute of Artificial Intelligence \(AI4I\), Turin, ItalyMauro DoreHikmaAI S\.r\.l\., Pula, ItalyMauro MeddaHikmaAI S\.r\.l\., Pula, ItalyNicola FrancoThe Italian Institute of Artificial Intelligence \(AI4I\), Turin, Italy
###### Abstract
We introduce a large‑scale, open‑source dataset of pre‑generated adversarial attacks for vision–language models \(VLMs\)\. The dataset is designed to be diverse, representative, and practical, extending existing benchmarks by covering1010high‑level categories and5555subcategories of harmful intents\. Our primary goal is to make adversarial data accessible to the research community, given the computational cost and complexity of generating large numbers of attacks\. The dataset comprises47 52447\\,524adversarial samples, generated using state‑of‑the‑art attack strategies from recent literature\. Our work complements existing efforts by consolidating and extending prior benchmarks from multiple established sources, resulting in7 8267\\,826intents, and introduce an additional category to broaden coverage\. This provides realistic evaluation resources for studying model robustness and alignment\. Our dataset intends to enable researchers and practitioners to systematically evaluate the robustness and safety of VLMs, fine‑tune attack‑generation models, and develop or stress‑test defensive guardrails under diverse adversarial conditions\. By releasing this resource, we aim to lower the barrier to adversarial research and foster more reproducible, comprehensive, and comparable evaluations of VLM safety\. The dataset has been released at:[https://huggingface\.co/datasets/it4lia/PHANTOM](https://huggingface.co/datasets/it4lia/PHANTOM)
Disclaimer: This paper and dataset contain content that may be disturbing or offensive, included solely for research purposes\.
11footnotetext:Equal contribution\.††footnotetext:Corresponding author:simone\.gallivanone@ai4i\.it## 1Introduction
With the rapid public deployment of vision–language models \(VLMs\) in both open\- and closed\-source settings, including safety\-critical and user\-facing applications, their robustness against adversarial prompting has become an increasingly important research concern \(see e\.g\.,\[[1](https://arxiv.org/html/2606.24388#bib.bib1),[2](https://arxiv.org/html/2606.24388#bib.bib2),[3](https://arxiv.org/html/2606.24388#bib.bib3),[4](https://arxiv.org/html/2606.24388#bib.bib4),[5](https://arxiv.org/html/2606.24388#bib.bib5)\]\)\. Recent studies consistently show that, despite improved alignment and scaling, state\-of\-the\-art multimodal models remain vulnerable to carefully crafted jailbreak attacks, particularly when harmful intents are distributed across visual and textual modalities \(see e\.g\.,\[[6](https://arxiv.org/html/2606.24388#bib.bib6),[7](https://arxiv.org/html/2606.24388#bib.bib7),[8](https://arxiv.org/html/2606.24388#bib.bib8),[9](https://arxiv.org/html/2606.24388#bib.bib9)\]\)\.
Unlike unimodal settings, multimodal safety violations often exploit cross\-modal reasoning and semantic alignment, significantly expanding the attack surface and complicating both detection and defense\. As a consequence, evaluating the robustness of VLMs requires large and diverse collections of adversarial image–text pairs\. This cost particularly affects resource\-constrained research groups and practitioners, for whom reproducing large\-scale multimodal attack generation may be impractical\. This is particularly true for vision\-language models, where attack generation is typically more resource\-consuming than in unimodal settings\. Unlike image\-only or text\-only attacks, multimodal attacks may require optimizing perturbations across multiple input spaces while preserving or exploiting their semantic alignment\. As a result, each attack iteration can involve forward and backward passes through multiple modality\-specific encoders and the cross\-modal alignment module, and the overall search space becomes larger and more constrained\. Although the exact overhead is model and attack dependent, the computational cost can be approximated as scaling with the combined cost of the involved modalities\. This makes systematic adversarial attack generation especially demanding for resource\-constrained actors\. While many existing open‑source benchmarks \(e\.g\.,\[[4](https://arxiv.org/html/2606.24388#bib.bib4),[10](https://arxiv.org/html/2606.24388#bib.bib10),[11](https://arxiv.org/html/2606.24388#bib.bib11),[12](https://arxiv.org/html/2606.24388#bib.bib12),[13](https://arxiv.org/html/2606.24388#bib.bib13),[14](https://arxiv.org/html/2606.24388#bib.bib14),[15](https://arxiv.org/html/2606.24388#bib.bib15),[1](https://arxiv.org/html/2606.24388#bib.bib1),[16](https://arxiv.org/html/2606.24388#bib.bib16),[17](https://arxiv.org/html/2606.24388#bib.bib17),[3](https://arxiv.org/html/2606.24388#bib.bib3),[2](https://arxiv.org/html/2606.24388#bib.bib2),[7](https://arxiv.org/html/2606.24388#bib.bib7),[18](https://arxiv.org/html/2606.24388#bib.bib18)\]\) provide tools and pipelines to generate and evaluate adversarial attacks, they typically do not release large collections of ready‑to‑use adversarial samples\. Only a limited number of datasets offer such pre‑generated attacks \(e\.g\.,\[[6](https://arxiv.org/html/2606.24388#bib.bib6),[19](https://arxiv.org/html/2606.24388#bib.bib19),[7](https://arxiv.org/html/2606.24388#bib.bib7),[2](https://arxiv.org/html/2606.24388#bib.bib2),[12](https://arxiv.org/html/2606.24388#bib.bib12)\]\), often focusing on specific attack types, categories, or linguistic settings\.
AEthical & SocialBPrivacy & DataCSafety & Physical HarmDCriminal & EconomicECybersecurity ThreatsFInfo & PoliticalGContent & CulturalHIP & OwnershipIDecision & CognitiveJChild Safety*\(new\)*RISK TAXONOMY10 categories⋅\\cdot55 subcategories⋅\\cdot7 826intentsBAPIDEATORMMLFC ATTACKCSDJATTACKSTRATEGIESAttack generation⊕\\oplus47 524pairsDeepSeek\-VL2GLM\-4\.6V\-FlashKimi\-VL\-A3BQwen3\-VL\-30BQwen3\.5\-27BQwen3\.6\-27BGemma\-4\-26BLLaVA\-v1\.6\-13BMinistral\-3\-14BOPEN\-SOURCEwhite\-boxGPT\-5\.4GPT\-5\.5Gemini 3\.1 ProClaude Opus 4\.6Claude Opus 4\.7Claude Opus 4\.8CLOSED\-SOURCEblack\-boxJudge⇒\\RightarrowASRper categoryAttack samplesPerturbedFlowchartFlippedGeneratedCollageintent selectionattackselectiontested againstTRANSFER
Figure 1:Overview of PHANTOM\. The risk taxonomy \(1010categories,5555subcategories,7 8267\\,826intents\) defines the dataset; the multimodal attacks \(BAP, IDEATOR, MML, FC ATTACK, CSDJ\) turn each intent into a multimodal adversarial sample \(harmful text prompt\+\+image\), giving47 52447\\,524\(prompt, image\) pairs\. Each pair is given white\-box to nine open\-source VLMs and transferred to six closed\-source, black\-box models; the judge scores every response to obtain per\-category ASR\. Bottom\-left: representative samples from different attack families\.In this work, we aim to complement existing efforts by releasing a large‑scale collection of ready‑to‑use multimodal adversarial samples, covering a broader range of attack strategies and safety categories\. Our goal is not to replace prior benchmarks, but to provide a practical resource that lowers the barrier to safety evaluation and enables reproducible and comprehensive robustness testing of multimodal models\. With this in mind, we designed and produced thePHANTOMdataset, a dataset of adversarial attacks for vision‑language models, which aims to fill this gap, and thus lower the barrier to systematic robustness evaluation\. The dataset contains attack samples in the form of image–text pairs for both single‑turn and conversational attacks\.
For a more detailed discussion on the design and content of the dataset we refer the reader to[section˜3](https://arxiv.org/html/2606.24388#S3)\. The samples have been generated against a variety of different open\-source models, from the following families: Qwen3\-VL\[[22](https://arxiv.org/html/2606.24388#bib.bib22)\], DeepSeek\-VL22\[[23](https://arxiv.org/html/2606.24388#bib.bib23)\], GLM\-4\.6V\[[24](https://arxiv.org/html/2606.24388#bib.bib24)\], Kimi\-VL\[[25](https://arxiv.org/html/2606.24388#bib.bib25)\], Qwen3\.5\[[26](https://arxiv.org/html/2606.24388#bib.bib26)\], Qwen3\.6\[[27](https://arxiv.org/html/2606.24388#bib.bib27)\]\. The generated samples were subsequently evaluated against state‑of‑the‑art proprietary models, including Claude Opus 4\.6 \- 4\.7 \- 4\.8, GPT‑5\.4 \- 5\.5, Gemini\-3\.1\-pro\. The results, which highlight cross‑model transferability and robustness trends, are presented in[section˜3\.4](https://arxiv.org/html/2606.24388#S3.SS4)\.
Our main contributions are:
- •PHANTOM, a large\-scale open\-source dataset of multimodal adversarial attacks for VLM safety evaluation\.
- •A curated taxonomy of7 8267\\,826harmful intents spanning1010categories and5555subcategories\.
- •47 52447\\,524adversarial samples generated using four attack strategies: BAP\[[8](https://arxiv.org/html/2606.24388#bib.bib8)\], IDEATOR\[[6](https://arxiv.org/html/2606.24388#bib.bib6)\], MML\[[9](https://arxiv.org/html/2606.24388#bib.bib9)\], FC ATTACK\[[20](https://arxiv.org/html/2606.24388#bib.bib20)\]and CSDJ\[[21](https://arxiv.org/html/2606.24388#bib.bib21)\]\.
- •A transferability analysis across both open\-source and proprietary VLMs\.
- •Structured metadata designed to support reproducibility, benchmarking, and downstream safety research\.
The paper is organized as follows\.[Section˜2](https://arxiv.org/html/2606.24388#S2)reviews related work\.[Section˜3](https://arxiv.org/html/2606.24388#S3)describes the dataset design, generation pipeline, and evaluation protocol\.[Section˜4](https://arxiv.org/html/2606.24388#S4)discusses the limitations of the current release, and[section˜5](https://arxiv.org/html/2606.24388#S5)addresses ethical considerations\.
## 2Related Works
In the landscape of adversarial attacks and model robustness, numerous efforts have benchmarked sensitive categories and created attack datasets across vision\-language models \(VLMs\) to evaluate vulnerabilities and establish foundations for model alignment\. To facilitate this review, we formalize the evaluation framework as a tupleℰ=\(𝒞,ℬ,𝒜,𝒥\)\\mathcal\{E\}=\(\\mathcal\{C\},\\mathcal\{B\},\\mathcal\{A\},\\mathcal\{J\}\)\. Letℳ\\mathcal\{M\}denote the target model which generates a responser∈ℛr\\in\\mathcal\{R\}from an image\-text input pair\(I,T\)\(I,T\)\.
- •Categories \(𝒞\\mathcal\{C\}\):A set ofnnsensitive domains𝒞=\{c1,…,cn\}\\mathcal\{C\}=\\\{c\_\{1\},\\dots,c\_\{n\}\\\}where model output must be constrained to ensure safety\.
- •Intents / Behaviors \(ℬ\\mathcal\{B\}\):A set of specific harmful intentsℬ=⋃c∈𝒞ℬc\\mathcal\{B\}=\\bigcup\_\{c\\in\\mathcal\{C\}\}\\mathcal\{B\}\_\{c\}, where eachb∈ℬcb\\in\\mathcal\{B\}\_\{c\}represents a concrete instance of a harmful objective within categorycc\.
- •Adversarial Attacks \(𝒜\\mathcal\{A\}\):A set of functionsf∈𝒜f\\in\\mathcal\{A\}that map a benign input to an adversarial input\(I′,T′\)\(I^\{\\prime\},T^\{\\prime\}\), optimized to exploit model misalignment such thatℳ\(I′,T′\)\\mathcal\{M\}\(I^\{\\prime\},T^\{\\prime\}\)aligns with a target behaviorbb\.
- •Judge \(𝒥\\mathcal\{J\}\):A classifier𝒥:ℛ×ℬ→\{0,1\}\\mathcal\{J\}:\\mathcal\{R\}\\times\\mathcal\{B\}\\rightarrow\\\{0,1\\\}that maps a model responserrand an intentbbto a binary success metric, where𝒥\(r,b\)=1\\mathcal\{J\}\(r,b\)=1indicates a successful adversarial exploit\.
[App\.˜D](https://arxiv.org/html/2606.24388#A4)summarizes the chronological evolution of adversarial attack benchmarks\. Detailed below, we review how different adversarial attack datasets included in these benchmarks or independently released, have evolved\.
### 2\.1Evolution of Early Multimodal Adversarial Attack Datasets
The study of adversarial attacks on language and multimodal models has evolved through a series of increasingly comprehensive datasets\. Early work by VAJM\[[16](https://arxiv.org/html/2606.24388#bib.bib16)\]introduced a dataset of32 22632\\,226samples, focusing on degradations related to gender, race, and human identity\. These included visual adversarial examples derived from4040behavioral categories, with attacks primarily generated through prompt tuning techniques\.
Subsequent efforts expanded both the scale and diversity of attacks\. The JailBreakV\-28K\[[19](https://arxiv.org/html/2606.24388#bib.bib19)\]dataset applied attacks not only to initial harmful prompts but also to broader behavioral patterns\. It includes20 00020\\,000text\-based jailbreak prompts and8 0008\\,000image\-based examples\. These attacks are derived from the RedTeam2K\[[19](https://arxiv.org/html/2606.24388#bib.bib19)\]benchmark, which covers approximately2 0002\\,000behaviors across1616categories\. The textual attacks were generated using methods such as GCG\[[17](https://arxiv.org/html/2606.24388#bib.bib17)\], Cognitive Overload, real\-world jailbreak prompt templates, and PAP\[[28](https://arxiv.org/html/2606.24388#bib.bib28)\], while the visual attacks leverage Stable Diffusion and typographic image techniques\.
The MM\-SafetyBench\[[2](https://arxiv.org/html/2606.24388#bib.bib2)\]dataset further advances multimodal evaluation by introducing5 0405\\,040text–image pairs derived from1 6801\\,680behaviors across1313categories\. In parallel, the Multiturn Human Jailbreaks\[[14](https://arxiv.org/html/2606.24388#bib.bib14)\]dataset explores iterative attack strategies, comprising2 9122\\,912attacks generated using a combination of automated methods, including AutoDAN\[[29](https://arxiv.org/html/2606.24388#bib.bib29)\], AutoPrompt\[[30](https://arxiv.org/html/2606.24388#bib.bib30)\], GCG\[[17](https://arxiv.org/html/2606.24388#bib.bib17)\], GPTFuzzer\[[31](https://arxiv.org/html/2606.24388#bib.bib31)\], and PAIR\[[32](https://arxiv.org/html/2606.24388#bib.bib32)\]\.
SafeBench\[[33](https://arxiv.org/html/2606.24388#bib.bib33)\]extends the evaluation setting by incorporating9 2009\\,200samples, including2 3002\\,300multimodal pairs, and introduces the audio modality\. It evaluates models under both adversarial and non\-adversarial conditions, using attack strategies such as LPT\[[34](https://arxiv.org/html/2606.24388#bib.bib34)\], PAP\[[28](https://arxiv.org/html/2606.24388#bib.bib28)\], and BAP\[[8](https://arxiv.org/html/2606.24388#bib.bib8)\]\. Notably, it is designed to assess safety risks even in the absence of explicit attacks\.
The MMJ dataset, derived from the MMJ benchmark\[[3](https://arxiv.org/html/2606.24388#bib.bib3)\], includes1 0001\\,000adversarial examples generated using methods such as FigStep\[[7](https://arxiv.org/html/2606.24388#bib.bib7)\], MM\-SafetyBench\[[2](https://arxiv.org/html/2606.24388#bib.bib2)\], HADES\[[18](https://arxiv.org/html/2606.24388#bib.bib18)\], ADV\-16\[[35](https://arxiv.org/html/2606.24388#bib.bib35)\], ADV\-64\[[35](https://arxiv.org/html/2606.24388#bib.bib35)\], ADV\-inf\[[35](https://arxiv.org/html/2606.24388#bib.bib35)\], ImgJP\[[36](https://arxiv.org/html/2606.24388#bib.bib36)\], and AttackVLM\[[37](https://arxiv.org/html/2606.24388#bib.bib37)\]\. This work highlights a critical limitation of overly conservative defenses, arguing that a system that refuses all prompts is not practically useful\.
BAVI\-Bench\[[12](https://arxiv.org/html/2606.24388#bib.bib12)\]significantly scales adversarial evaluation, containing316316k adversarial visual\-instruction samples\. It includes four types of image\-based B\-AVIs, ten types of text\-based B\-AVIs, and nine categories of content bias \(e\.g\., gender, violence, cultural, and racial biases\)\. The benchmark evaluates robustness using attacks such as PAR\[[38](https://arxiv.org/html/2606.24388#bib.bib38)\], Boundary\[[39](https://arxiv.org/html/2606.24388#bib.bib39)\], and SurFree\[[40](https://arxiv.org/html/2606.24388#bib.bib40)\]\.
The VLJailBreak benchmark\[[6](https://arxiv.org/html/2606.24388#bib.bib6)\]provides3 6543\\,654samples spanning1212safety topics and4646subcategories, offering a highly structured categorization\. It evaluates model vulnerabilities using GCG\[[17](https://arxiv.org/html/2606.24388#bib.bib17)\]and UMK\[[41](https://arxiv.org/html/2606.24388#bib.bib41)\]attack methods\.
Finally, the Adversarial Humanities Benchmark\[[42](https://arxiv.org/html/2606.24388#bib.bib42)\]investigates whether safety mechanisms generalize beyond familiar harmful prompt patterns\. It includes3 6003\\,600attack samples and shows that current safety techniques exhibit limited generalization, suggesting that a deeper understanding of non\-maleficence remains an open challenge in frontier model safety\. Complementarily, MultiBreak\[[5](https://arxiv.org/html/2606.24388#bib.bib5)\]focuses on realistic multi\-turn jailbreak scenarios, where harmful intents are progressively elicited through conversation rather than expressed in a single prompt\. It introduces10 38910\\,389multi\-turn adversarial prompts spanning2 6652\\,665harmful intents, and shows that diverse multi\-turn attacks can reveal fine\-grained vulnerabilities that may remain hidden under single\-turn evaluation\. Together, these works highlight the need for safety benchmarks that go beyond static or template\-based attacks, covering both broader semantic generalization and more realistic conversational adversarial settings\.
PHANTOM, on the other hand, identifies three key limitations in existing datasets\. First, the number of intents used to generate attacks is too limited to adequately represent the full range of categories; accordingly, PHANTOM expands this to7 8267\\,826intents, as detailed in[table˜2](https://arxiv.org/html/2606.24388#S3.T2)\. Second, it examines how state\-of\-the\-art attacks perform on more recent models that are commonly used in industrial and research settings as shown in[table˜3](https://arxiv.org/html/2606.24388#S3.T3)\. Third, it investigates how vulnerabilities discovered through prior attacks and models can be transferred to other, primarily black\-box, models\.
## 3Dataset design and production
In this section, we describe the dataset design and the process used to generate its samples\. Specifically, we outline the category and subcategory structure, as well as the JSON\-based intent specification employed during dataset construction\.
#### Settings\.
All experiments were conducted on a cluster, using NVIDIA A100 GPUs with 64GB of memory and Intel®Xeon®Platinum 8358 CPUs operating at 2\.60GHz\. In addition to the selected attack strategy, generated samples must be evaluated to determine whether they constitute successful attacks\. For the sake of reproducibility and to avoid bias stemming from ad hoc judgment criteria, we opted to rely on a publicly available automated judge as a common baseline\. In particular, we adopted the Abel\-24\-HarmClassifier proposed in\[[43](https://arxiv.org/html/2606.24388#bib.bib43)\]\. This choice was motivated by the increasing difficulty of fully and cleanly jailbreaking recent models, which often respond with partially harmful or evasive outputs\. Consequently, we selected a recent and aligned classifier to provide a more reliable assessment of attack harmfulness\.
### 3\.1Categories structure
Table 1:Number of intents per benchmarkAs mentioned in the introduction, we began our work by analyzing existing benchmarks for adversarial attacks\. The landscape of adversarial attack benchmarks is quite extensive; however, many of these benchmarks build upon previous work, in the sense that they naturally extend earlier foundations\. Two of the most influential works in this area are HarmBench\[[1](https://arxiv.org/html/2606.24388#bib.bib1)\]and AdvBench\[[17](https://arxiv.org/html/2606.24388#bib.bib17)\], which were designed as comprehensive collections of harmful intents and attack strategies for evaluating model safety\.
With the evolution of AI models, their expanding range of use cases, and their increasing accessibility to the general public, these early benchmarks have gradually become insufficient in terms of coverage\. As a result, several subsequent benchmarks have been proposed to address these limitations and extend prior efforts\. In this work, we studied1616benchmarks, including OmniSafeBench\-MM\[[4](https://arxiv.org/html/2606.24388#bib.bib4)\], VLJailbreakBench\[[6](https://arxiv.org/html/2606.24388#bib.bib6)\], Sorry\-Bench\[[11](https://arxiv.org/html/2606.24388#bib.bib11)\], B\-AVIBench\[[12](https://arxiv.org/html/2606.24388#bib.bib12)\], JailbreakBench\[[10](https://arxiv.org/html/2606.24388#bib.bib10)\], SafeBench\[[13](https://arxiv.org/html/2606.24388#bib.bib13)\], Multi\-Turn Human Jailbreaks\[[14](https://arxiv.org/html/2606.24388#bib.bib14)\], StrongREJECT\[[15](https://arxiv.org/html/2606.24388#bib.bib15)\], HarmBench\[[1](https://arxiv.org/html/2606.24388#bib.bib1)\], VAJM\[[16](https://arxiv.org/html/2606.24388#bib.bib16)\], AdvBench\[[17](https://arxiv.org/html/2606.24388#bib.bib17)\], MMJ\-Bench\[[3](https://arxiv.org/html/2606.24388#bib.bib3)\], MM\-SafetyBench\[[2](https://arxiv.org/html/2606.24388#bib.bib2)\], JailBreakV\-28K\[[19](https://arxiv.org/html/2606.24388#bib.bib19)\], FigStep\[[7](https://arxiv.org/html/2606.24388#bib.bib7)\], and HADES\[[18](https://arxiv.org/html/2606.24388#bib.bib18)\]\. A gap we identified across these benchmarks is the limited number of behaviors, which makes it difficult to disentangle whether observed vulnerabilities stem from insufficient robustness to specific categories or from the strength of the attacks themselves\. To address this, we incorporated7 8267\\,826behaviors into our study\.
Table 2:Number of intents per categoryBased on this analysis, we constructed a root dataset of harmful intents \(also referred to asbehaviorsorgoals\) by merging recent benchmarks that were not designed as direct extensions of one another\. The number of intents contributed by each benchmark and their distribution across categories are reported in[table˜1](https://arxiv.org/html/2606.24388#S3.T1)and[table˜2](https://arxiv.org/html/2606.24388#S3.T2), respectively\. In[App\.˜A](https://arxiv.org/html/2606.24388#A1)the reader can find a table listing all categories, subcategories together with their alphanumeric reference\. Since\[[4](https://arxiv.org/html/2606.24388#bib.bib4)\]carried out an extensive effort to reorganize and categorize harmful intents, while also providing a detailed taxonomy, we mapped all merged intents onto the classification proposed there\. However, we identified a gap in its coverage:*child safety*\. The motivation for introducing this category is twofold\. First, the widespread adoption of LLMs across users of all ages, including minors, raises the risk of exposure to content that may be harmful to their safety, such as methods to circumvent parental controls or age\-verification systems\. Second, given the particular vulnerability of minors, malicious actors could potentially exploit LLMs to obtain information on how to deceive or manipulate children in harmful situations\. For these reasons, we believe that including this category contributes to a more comprehensive understanding of the risks associated with LLMs, and may support the development of more effective safeguards and training strategies\.
Figure 2:PHANTOM intents dataset\. The horizontal axis shows the distribution of categories across the dataset, while the vertical axis shows the distribution of intents within each category, broken down by subcategory\.As a result, our dataset is organized into 10 high\-level categories, further divided into a total of 55 subcategories\. The main categories are: Ethical and Social Risks, Privacy and Data Risks, Safety and Physical Harm, Criminal and Economic Risks, Cybersecurity Threats, Information and Political Manipulation, Content and Cultural Safety, Intellectual Property and Ownership, Decision and Cognitive Risks, and Child Safety\. The full subdivision into subcategories is illustrated in[fig\.˜2](https://arxiv.org/html/2606.24388#S3.F2); we refer the reader to[App\.˜A](https://arxiv.org/html/2606.24388#A1)for an overview of the names of subcategories with respect to their reference code\.
We collected more than7 0007\\,000harmful intents from the following benchmarks: JailBreakV\_28K\[[19](https://arxiv.org/html/2606.24388#bib.bib19)\], MM\-SafetyBench\[[2](https://arxiv.org/html/2606.24388#bib.bib2)\], OmniSafeBench\-MM\[[4](https://arxiv.org/html/2606.24388#bib.bib4)\], and SafeBench\[[33](https://arxiv.org/html/2606.24388#bib.bib33)\]\. Moreover, we added747747intents related to the new Child Safety category, generated with the assistance of OpenAI GPT\-5\.4, accessed via API\. Since these benchmarks, together with our additions, may contain overlapping or semantically similar intents, we performed a cosine similarity analysis across all collected samples using embeddings generated by the sentence\-transformersall\-MiniLM\-L6\-v2model \(see\[[44](https://arxiv.org/html/2606.24388#bib.bib44)\]\)\.
We found that the overlap was not negligible, reaching values as high as90%90\\%\. Therefore, we decided to clean the dataset using this threshold, which left no pair of intents with a cosine similarity above90%90\\%\. At lower thresholds the number of intents flagged as similar grows:376376\(4\.7%4\.7\\%\) at85%85\\%and661661\(8\.2%8\.2\\%\) at80%80\\%\. We decided not to apply more aggressive cleaning, as we observed that an80%80\\%threshold often groups together semantically different intents, and we did not want to remove meaningful content\.
### 3\.2Adversarial attacks
To benchmark the performance of state\-of\-the\-art adversarial attacks against recent vision–language models, we initially selected a set of established attack strategies that had already proven effective on open\-source vision–language systems\. As shown in[fig\.˜3](https://arxiv.org/html/2606.24388#S3.F3), we analyze the attack success rate \(ASR\) as a function of generation time\. To enable large\-scale dataset generation, we ultimately focused on a limited subset of attacks offering the most favorable trade\-off between ASR and computational cost\. Specifically, we selected one single\-turn attack strategy, the Bi\-modal Adversarial Prompt \(BAP\) attack proposed in\[[8](https://arxiv.org/html/2606.24388#bib.bib8)\]; one multi\-turn strategy, IDEATOR, introduced in\[[6](https://arxiv.org/html/2606.24388#bib.bib6)\]; and two more typographic oriented attacks the Multi\-Modal Linkage Attack, proposed in\[[9](https://arxiv.org/html/2606.24388#bib.bib9)\]and the Flowchart attack in\[[20](https://arxiv.org/html/2606.24388#bib.bib20)\]\. A review of the attack strategies can be found in[App\.˜F](https://arxiv.org/html/2606.24388#A6)\.
Figure 3:Evaluation of ASR based on Attack strategy and delay per attackThe main motivation for selecting these strategies was their strong empirical performance\. Before converging on this subset, however, we experimented with additional attack strategies, namely QR–attack\[[2](https://arxiv.org/html/2606.24388#bib.bib2)\], JOOD\[[45](https://arxiv.org/html/2606.24388#bib.bib45)\], CS\-DJ\[[21](https://arxiv.org/html/2606.24388#bib.bib21)\], FigStep\[[7](https://arxiv.org/html/2606.24388#bib.bib7)\], HADES\[[18](https://arxiv.org/html/2606.24388#bib.bib18)\], HIMRD\[[46](https://arxiv.org/html/2606.24388#bib.bib46)\], VISCARA\[[47](https://arxiv.org/html/2606.24388#bib.bib47)\], MIDAS\[[48](https://arxiv.org/html/2606.24388#bib.bib48)\], ACZ attack\[[49](https://arxiv.org/html/2606.24388#bib.bib49)\]\. The results of this preliminary analysis are reported in[fig\.˜3](https://arxiv.org/html/2606.24388#S3.F3), based on5050samples generated against Qwen3\.5\-27B and evaluated with Abel\-24\-HarmClassifier\[[43](https://arxiv.org/html/2606.24388#bib.bib43)\]\. Additional considerations that informed our final choice are discussed below\.
While preserving the original attack pipelines, we introduced several minor modifications to better suit our dataset\-generation process\. In the following, we briefly describe how these methods were adapted and employed\. In future releases, we intend to expand the range of attack strategies and target models in order to cover a broader spectrum of vulnerabilities\.
### 3\.3Data structure and distribution
The dataset is structured according to the attack strategy and the target model used during the generation process\. The target models considered are DeepSeek\-VL22, GLM\-4\.6V\-Flash, Kimi\-VL\-A3B\-Instruct, Qwen3\-VL\-30B\-A3B\-Instruct, Qwen3\.5\-27B and Qwen3\.6\-27B\. Since the attacks target vision–language models, each sample consists of an image–prompt pair\. Each folder contains the generated image along with a structured metadata file, which enables the correct association between images and prompts and ensures full reproducibility of the attacks\. The current release contains a total of47 52447\\,524generated attacks\. Below, we discuss their distribution across target models, attack strategies, and categories\. Regarding the distribution of attacks across target models, we initially generated attacks uniformly for all models\. After conducting preliminary cross\-model evaluations, we focused further generation on the models that exhibited higher attack transferability, namely GLM\-4\.6V\-Flash and Qwen3\-VL\-30B\-A3B\-Instruct, Qwen3\.5\-27B and Qwen3\.6\-27B\. We refer the reader to[table˜3](https://arxiv.org/html/2606.24388#S3.T3)for a detailed breakdown\.
Table 3:Number of generated attacks per target model and attack strategy\.From a categorical perspective, we selected intents randomly and uniformly from our dataset of intents, discussed in[section˜3\.1](https://arxiv.org/html/2606.24388#S3.SS1)\. The distribution over categories and subcategories is shown in[fig\.˜4](https://arxiv.org/html/2606.24388#S3.F4)\.
Figure 4:Category coverage: overall and in\-category
### 3\.4Dataset tests and statistics
To validate our adversarial data generation pipeline, we conducted a series of experiments using the generated attacks\. In particular, we considered two evaluation settings:*white\-box*and*black\-box*testing\.
In the black\-box setting, we evaluated the attacks against several state\-of\-the\-art proprietary models accessed via API, namely Gemini 3\.1 Pro Preview, GPT\-5\.4, GPT\-5\.5, Claude Opus 4\.6, Claude Opus 4\.7, and Claude Opus 4\.8\. Model responses were collected and assessed using an automated judge\. Cases in which a model returned an empty response were treated as*hard refusals*and therefore counted as failed jailbreak attempts\.
In the white\-box setting, we focused primarily on the models used during adversarial generation, in order to evaluate the cross\-model transferability of the attacks\. Specifically, we tested DeepSeek\-VL22, GLM\-4\.6V\-Flash, Kimi\-VL\-A3B\-Instruct, Qwen3\-VL\-30B\-A3B\-Instruct, Qwen3\.5\-27B, Qwen3\.6\-27B, Gemma\-4\-26B\-A4B\-it, Llava\-v1\.6\-vicuna\-13b\-hf and Ministral\-3\-14B\-Instruct\-2512 all executed locally on NVIDIA A100 GPUs with 64 GB of memory\. To keep the evaluation compact, we consider a fixed subset of 1,100 attacks for each attack strategy\. This subset is selected once and reused across all models\. The choice of an odd number is motivated by the need to evenly cover all subcategories associated with each attack strategy; specifically, we select 20 attacks per subcategory\.
As in the generation phase, we employed Abel\-24\-HarmClassifier\[[43](https://arxiv.org/html/2606.24388#bib.bib43)\]as the baseline judge for evaluating model responses\. As evaluation metric, we relied on the widely used Attack Success Rate \(ASR\), defined as the percentage of successful jailbreak attempts, as determined by the judge, over the total number of attacks\. Part of the result, relative to black\-box and the most recent withe\-box models are reported inLABEL:tbl:asr\_x\_model\_x\_cat\_corpus\. The rest of the results can be found in[App\.˜C](https://arxiv.org/html/2606.24388#A3)\. For each model and category, we report the corresponding ASR, i\.e\., the ratio of successful attacks to the number of attacks sampled in that category\. This breakdown makes it possible to identify the categories for which each model is most vulnerable or most robust\.We emphasize once more that these results should be interpreted as a reference relative to the chosen baseline for assessing harmfulness, rather than as an absolute ground truth\.
Table 4:Table of ASR \(%\) per model and per category across all attacks, for category names refer to[App\.˜A](https://arxiv.org/html/2606.24388#A1)ModelABCDEFGHIJAverageBAPGemma\-4\-26B23\.7516\.0010\.0018\.0027\.1420\.0022\.5022\.5026\.8834\.0022\.00Qwen3\.6\-27B42\.5040\.0030\.7142\.0038\.5735\.8333\.7540\.0045\.0051\.0039\.82GPT\-5\.413\.7515\.0025\.0023\.0011\.4318\.3310\.006\.2510\.6223\.0015\.91GPT\-5\.547\.5041\.0059\.2968\.0047\.8642\.5046\.2535\.0048\.1250\.0049\.09Claude Opus 4\.66\.2511\.000\.7110\.008\.575\.838\.757\.506\.8817\.007\.91Claude Opus 4\.751\.2544\.0020\.7161\.0040\.0039\.1735\.0056\.2550\.0049\.0043\.64Claude Opus 4\.847\.5040\.0012\.1465\.0039\.2934\.1721\.2541\.2543\.1251\.0038\.73Gemini 3\.1 Pro23\.7539\.0035\.7145\.0074\.2928\.3330\.0027\.5025\.6244\.0038\.36IDEATORGemma\-4\-26B10\.0030\.0020\.7134\.0033\.5720\.0015\.0021\.2535\.6243\.0027\.36Qwen3\.6\-27B7\.507\.0011\.4314\.0015\.0016\.6718\.757\.5023\.1265\.0018\.82GPT\-5\.425\.0014\.0027\.1449\.0044\.2936\.6726\.2512\.5027\.5033\.0030\.45GPT\-5\.521\.2528\.0026\.4342\.0048\.5744\.1752\.5031\.2537\.5044\.0037\.82Claude Opus 4\.66\.256\.005\.7117\.0013\.575\.833\.755\.0013\.127\.008\.82Claude Opus 4\.725\.0039\.0012\.1453\.0030\.7130\.0032\.5045\.0034\.3833\.0032\.55Claude Opus 4\.816\.2524\.0012\.1447\.0032\.8624\.1717\.5033\.7531\.2520\.0026\.09Gemini 3\.1 Pro28\.7543\.0042\.1456\.0071\.4358\.3356\.2542\.5054\.3762\.0052\.64MMLGemma\-4\-26B95\.56100\.0090\.3498\.5199\.3796\.8894\.87100\.0096\.5595\.2696\.09Qwen3\.6\-27B80\.0085\.7180\.9789\.6394\.5587\.5071\.7986\.2183\.1488\.4286\.10GPT\-5\.471\.7673\.4574\.1367\.9270\.0072\.7978\.7580\.4680\.0095\.0076\.03GPT\-5\.591\.2584\.0090\.7192\.0090\.7189\.1782\.5080\.0081\.2561\.0084\.64Claude Opus 4\.671\.9561\.3232\.1460\.7863\.9563\.7841\.2548\.8162\.8654\.5556\.19Claude Opus 4\.71\.2510\.000\.716\.002\.860\.830\.005\.001\.881\.002\.82Claude Opus 4\.812\.5021\.003\.5726\.0010\.0011\.6711\.2521\.2513\.7523\.0014\.64Gemini 3\.1 Pro3\.530\.884\.200\.9440\.6281\.6266\.2593\.1093\.7129\.0043\.38FC ATTACKGemma\-4\-26B37\.5029\.002\.147\.0022\.1411\.6733\.7528\.7518\.7514\.0018\.91Qwen3\.6\-27B86\.2581\.0055\.7185\.0092\.8682\.5072\.5076\.2578\.1264\.0077\.27GPT\-5\.452\.5043\.0060\.7171\.0077\.1474\.1746\.2547\.5059\.3863\.0061\.00GPT\-5\.557\.5057\.0068\.5786\.0089\.2976\.6767\.5063\.7565\.0074\.0071\.36Claude Opus 4\.677\.5085\.0041\.4368\.0087\.8670\.8363\.7582\.5075\.6260\.0070\.82Claude Opus 4\.770\.0072\.0038\.5775\.0068\.5765\.0037\.5082\.5068\.7561\.0063\.45Claude Opus 4\.863\.7583\.0048\.5777\.0079\.2975\.0036\.2578\.7562\.5070\.0067\.45Gemini 3\.1 Pro35\.0046\.0015\.0028\.0067\.1426\.6735\.0055\.0040\.6221\.0037\.00CSDJGemma\-4\-26B53\.7575\.0027\.8653\.0085\.7157\.5027\.5041\.2536\.8855\.0051\.35Qwen3\.6\-27B71\.2568\.0064\.2974\.0084\.2970\.0070\.0060\.0061\.8870\.0069\.37GPT\-5\.478\.7578\.0091\.4387\.0087\.1486\.6766\.2561\.2563\.7581\.0078\.82GPT\-5\.582\.5080\.0088\.5791\.0095\.0085\.0081\.2570\.0069\.3887\.0083\.16Claude Opus 4\.673\.7588\.0067\.8685\.0093\.5784\.1748\.7571\.2573\.7583\.0077\.82Claude Opus 4\.770\.0090\.0057\.1478\.0095\.0088\.3350\.0071\.2559\.3888\.0074\.82Claude Opus 4\.880\.0089\.0071\.4391\.0095\.0087\.5050\.0071\.2559\.3889\.0078\.45Gemini 3\.1 Pro65\.0086\.0039\.2972\.0086\.4367\.5046\.2568\.7563\.1268\.0066\.18While the weakness of a model with respect to a given attack gives an interesting insight in how to chose the attack strategy, one may be interested in global weaknesses of a model\. To approximate such information we averaged the results across attack strategies and reported the results in[fig\.˜5](https://arxiv.org/html/2606.24388#S3.F5)\. As discussed, the7 8267\\,826intents in the PHANTOM benchmark provide sufficient coverage across categories to confidently analyze model vulnerabilities with respect to these specific domains\.
The radar diagrams provide an immediate, at\-a\-glance understanding of model robustness: a larger colored area corresponds to a greater amount of harmful content produced during testing\. However, an important clarification is needed\. In modern models, it is difficult to observe “pure” jailbreaks, i\.e\., cases in which the model responds directly and fully to a harmful request\. Instead, harmful content is more often embedded within longer responses that include benign context and argumentation\. Therefore, these diagrams should be interpreted as indicating a higher tendency of the model to generate harmful content within otherwise complex answers\.
With this in mind, models that tend to respond to user requests, even while attempting to avoid harmful content, ultimately produce more harmful content on average\. This explains, for example, the stronger performance of Gemma\-4\-26B compared to many black\-box models, which tend to consistently provide answers\. On the other hand, it is important to note that black\-box models accessed via API sometimes returnnullresponses, most likely due to content filtering mechanisms; we refer to these ashard refusals\. We treat such cases as failed jailbreaks\. Different models exhibit different rates of hard refusals: for instance, the Claude Opus models show a much higher rate of hard refusals compared to both GPT and Gemini, whereas GPT models exhibit the lowest rate\. We now highlight a few observations from the results\. In[fig\.˜5](https://arxiv.org/html/2606.24388#S3.F5), the extent of the colored area allows one to infer, with respect to the chosen baseline judge, the relative robustness of the models: a wider area corresponds to less aligned responses\. Among white\-box models, Gemma\-4\-26B is clearly the most robust\. Among black\-box models, the picture is different: all models in the Opus family exhibit comparable robustness, which is also similar to that of Gemini 3\.1 Pro, while the GPT family appears less robust\. However, as noted earlier, this should be interpreted alongside the higher rate of complete responses they produce\. Interestingly, within the GPT family \(from 5\.4 to 5\.5\), performance in terms of alignment appears to degrade slightly, although this is again coupled with the absence of hard refusals\. Another interesting observation fromLABEL:tbl:asr\_x\_model\_x\_cat\_corpusemerges from the distribution of colored cells: the most effective attacks across all models are those that embed harmful text within images\. This suggests that model alignment with respect to embedded textual content remains relatively weak, highlighting a persistent vulnerability in multimodal safety mechanisms\. Finally, across all models, the most vulnerable categories areD — Criminal and Economic RisksandE — Cybersecurity Threats, which also correspond to domains where one would expect models to provide more actionable and useful responses\.
Figure 5:Examining model vulnerability against harmful categories
## 4Limitations
PHANTOM has several limitations\. First, our evaluation relies primarily on an automated judge, Abel\-24\-HarmClassifier, which may introduce false positives and false negatives, particularly for responses that are partially harmful, evasive, or context\-dependent\. Although automated judging enables large\-scale evaluation, it cannot fully replace human assessment\.
Second, our reported evaluation is conducted on sampled subsets of attacks rather than on the entire released dataset\. While this makes the evaluation computationally feasible, it may underrepresent variability\.
Third, currently we focused on three multimodal attack strategies: BAP, IDEATOR, and MML\. These methods were selected for their empirical effectiveness and computational feasibility, but they do not exhaust the space of possible adversarial attacks against VLMs\.
Finally, the taxonomy and intent collection may inherit biases from the source benchmarks used to construct the dataset\.
## 5Ethical Considerations
PHANTOM dataset contains adversarial multimodal samples involving harmful and sensitive intents, and is therefore a dual\-use resource\. The dataset is intended solely for research, robustness evaluation, and the development of defensive guardrails\. To reduce misuse risks, we provide content warnings, structured metadata, category labels\. Sensitive categories, including child safety and personally harmful content, are included only for safety evaluation and should be handled under appropriate institutional and ethical safeguards\.
## 6Discussion and conclusion
Our empirical analysis highlights several key insights into the behavior of modern VLMs under adversarial conditions\.
First, despite advances in safety training, all evaluated systems exhibit non\-negligible ASR across multiple categories, confirming that alignment remains fragile in the presence of carefully constructed multimodal inputs\.
Second, we observe significant variation across attack strategies\. This suggests that different strategies exploit distinct failure modes of VLMs across categories of harmful content\. Consequently, evaluating robustness using a single attack family risks underestimating model vulnerability\. In particular, attacks that embed harmful requests directly within images \(i\.e\., typographic attacks such as MML, FC Attack, and CSDJ\) achieve consistently higher success rates, indicating that even recent large\-scale models still struggle to maintain strong filtering capabilities in fully multimodal settings\. A particularly surprising result comes from the Gemma\-4\-26B model: when attacked, it achieves an ASR of 100% in two categories, highlighting a substantial vulnerability despite the model’s overall capabilities\.
Third, the results reveal clear evidence of cross\-model transferability, particularly from open\-source to proprietary systems\. Attacks generated in a white\-box setting retain their effectiveness when transferred to black\-box models, with ASR remaining above 20% in most cases and reaching peaks of nearly 80% for the CSDJ attack\. This indicates that vulnerabilities are not purely model\-specific but instead reflect shared structural or training\-induced weaknesses\. These findings have important implications for real\-world deployment, where adversaries may optimize attacks against accessible models and subsequently transfer them to closed systems\.
A further observation is the heterogeneity across risk categories\. Certain domains, such as cybersecurity or economic crimes, tend to yield higher attack success rates, reaching up to 90% on black\-box models, while others are more robust\. This variability suggests that current alignment procedures may unevenly cover the safety landscape, leaving gaps that adversarial methods can exploit\. These findings are consistent with the broader trend illustrated in the evaluation results\.
Importantly, our analysis also highlights a well known evaluation caveat: success rates depend on the chosen judge model and may be influenced by partial refusals due to external filters or ambiguous outputs \(see also,\[[50](https://arxiv.org/html/2606.24388#bib.bib50)\],\[[51](https://arxiv.org/html/2606.24388#bib.bib51)\]\)\. As such, the results should be interpreted as relative indicators of robustness, rather than absolute measures of harmfulness\. Overall, PHANTOM enables a more systematic understanding of how different attack strategies, model architectures, and safety domains interact, offering a baseline to study multimodal robustness\.
By consolidating multiple attack strategies and providing structured evaluation across a diverse set of VLMs, the dataset addresses a key gap in the current landscape: the lack of accessible, reproducible, and comprehensive adversarial resources\. Our results demonstrate that multimodal jailbreaks remain a persistent and transferable threat, that robustness varies significantly across both models and safety domains, and that a diverse set of attack strategies is necessary for reliable evaluation\. Beyond benchmarking, PHANTOM provides a practical foundation for future research: the dataset can be used to develop and evaluate defensive mechanisms and guardrails, to train adversarially robust models, and to advance the study of cross\-modal alignment failures\. We release PHANTOM with the goal of lowering the barrier to multimodal safety research and fostering more reproducible, standardized, and comprehensive evaluations\. We hope that this resource will contribute to a deeper understanding of VLM robustness and support the development of safer and more reliable multimodal AI systems\.
## References
- Mazeika et al\. \[2024\]Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al\.Harmbench: A standardized evaluation framework for automated red teaming and robust refusal\.*arXiv preprint arXiv:2402\.04249*, 2024\.
- Liu et al\. \[2024\]Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao\.Mm\-safetybench: A benchmark for safety evaluation of multimodal large language models\.In*European Conference on Computer Vision*, pages 386–403\. Springer, 2024\.
- Weng et al\. \[2025\]Fenghua Weng, Yue Xu, Chengyan Fu, and Wenjie Wang\.Mmj\-bench: A comprehensive study on jailbreak attacks and defenses for vision language models\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 39, pages 27689–27697, 2025\.
- Jia et al\. \[2025\]Xiaojun Jia, Jie Liao, Qi Guo, Teng Ma, Simeng Qin, Ranjie Duan, Tianlin Li, Yihao Huang, Zhitao Zeng, Dongxian Wu, Yiming Li, Wenqi Ren, Xiaochun Cao, and Yang Liu\.Omnisafebench\-mm: A unified benchmark and toolbox for multimodal jailbreak attack\-defense evaluation, 2025\.URL[https://arxiv\.org/abs/2512\.06589](https://arxiv.org/abs/2512.06589)\.
- Song et al\. \[2026a\]Jialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen, Mingqian Feng, Xuekai Zhu, and Jianfeng Gao\.Multibreak: A scalable and diverse multi\-turn jailbreak benchmark for evaluating llm safety, 2026a\.URL[https://arxiv\.org/abs/2605\.01687](https://arxiv.org/abs/2605.01687)\.
- Wang et al\. \[2025a\]Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, and Yu\-Gang Jiang\.Ideator: Jailbreaking and benchmarking large vision\-language models using themselves, 2025a\.URL[https://arxiv\.org/abs/2411\.00827](https://arxiv.org/abs/2411.00827)\.
- Gong et al\. \[2025\]Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang\.Figstep: Jailbreaking large vision\-language models via typographic visual prompts\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 39, pages 23951–23959, 2025\.
- Ying et al\. \[2025a\]Zonghao Ying, Aishan Liu, Tianyuan Zhang, Zhengmin Yu, Siyuan Liang, Xianglong Liu, and Dacheng Tao\.Jailbreak vision language models via bi\-modal adversarial prompt\.*IEEE Transactions on Information Forensics and Security*, 20:7153–7165, 2025a\.doi:[10\.1109/TIFS\.2025\.3583249](https://doi.org/10.1109/TIFS.2025.3583249)\.
- Wang et al\. \[2025b\]Yu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang, and Tianxing He\.Jailbreak large vision\-language models through multi\-modal linkage\.In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 1466–1494, Vienna, Austria, July 2025b\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-251\-0\.doi:[10\.18653/v1/2025\.acl\-long\.74](https://doi.org/10.18653/v1/2025.acl-long.74)\.URL[https://aclanthology\.org/2025\.acl\-long\.74/](https://aclanthology.org/2025.acl-long.74/)\.
- Chao et al\. \[2024\]Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al\.Jailbreakbench: An open robustness benchmark for jailbreaking large language models\.*Advances in Neural Information Processing Systems*, 37:55005–55029, 2024\.
- Xie et al\. \[2024\]Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al\.Sorry\-bench: Systematically evaluating large language model safety refusal\.*arXiv preprint arXiv:2406\.14598*, 2024\.
- Zhang et al\. \[2024\]Hao Zhang, Wenqi Shao, Hong Liu, Yongqiang Ma, Ping Luo, Yu Qiao, Nanning Zheng, and Kaipeng Zhang\.B\-avibench: Toward evaluating the robustness of large vision\-language model on black\-box adversarial visual\-instructions\.*IEEE Transactions on Information Forensics and Security*, 20:1434–1446, 2024\.
- Ying et al\. \[2026\]Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, and Dacheng Tao\.Safebench: A safety evaluation framework for multimodal large language models\.*International Journal of Computer Vision*, 134\(1\):18, 2026\.
- Li et al\. \[2024a\]Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue\.Llm defenses are not robust to multi\-turn human jailbreaks yet\.*arXiv preprint arXiv:2408\.15221*, 2024a\.
- Souly et al\. \[2024\]Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al\.A strongreject for empty jailbreaks\.*Advances in Neural Information Processing Systems*, 37:125416–125440, 2024\.
- Qi et al\. \[2024a\]Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal\.Visual adversarial examples jailbreak aligned large language models\.In*Proceedings of the AAAI conference on artificial intelligence*, volume 38, pages 21527–21536, 2024a\.
- Zou et al\. \[2023\]Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson\.Universal and transferable adversarial attacks on aligned language models\.*arXiv preprint arXiv:2307\.15043*, 2023\.
- Li et al\. \[2024b\]Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji\-Rong Wen\.Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models\.In*European Conference on Computer Vision*, pages 174–189\. Springer, 2024b\.
- Luo et al\. \[2024\]Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao\.Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks\.*arXiv preprint arXiv:2404\.03027*, 2024\.
- Zhang et al\. \[2025\]Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, and Xinlei He\.Fc\-attack: Jailbreaking large vision\-language models via auto\-generated flowcharts\.*arXiv e\-prints*, pages arXiv–2502, 2025\.
- Yang et al\. \[2025\]Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Xin Lin, Tao Li, Kanghua Mo, and Changyu Dong\.Distraction is all you need for multimodal large language model jailbreaking\.In*Proceedings of the Computer Vision and Pattern Recognition Conference*, pages 9467–9476, 2025\.
- Bai et al\. \[2025\]Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al\.Qwen3\-vl technical report\.*arXiv preprint arXiv:2511\.21631*, 2025\.
- Wu et al\. \[2024\]Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al\.Deepseek\-vl2: Mixture\-of\-experts vision\-language models for advanced multimodal understanding\.*arXiv preprint arXiv:2412\.10302*, 2024\.
- Team et al\. \[2025a\]V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, Jing Chen, Jinhao Chen, Jinghao Lin, Jinjiang Wang, Junjie Chen, Leqi Lei, Letian Gong, Leyi Pan, Mingdao Liu, Mingde Xu, Mingzhi Zhang, Qinkai Zheng, Sheng Yang, Shi Zhong, Shiyu Huang, Shuyuan Zhao, Siyan Xue, Shangqin Tu, Shengbiao Meng, Tianshu Zhang, Tianwei Luo, Tianxiang Hao, Tianyu Tong, Wenkai Li, Wei Jia, Xiao Liu, Xiaohan Zhang, Xin Lyu, Xinyue Fan, Xuancheng Huang, Yanling Wang, Yadong Xue, Yanfeng Wang, Yanzi Wang, Yifan An, Yifan Du, Yiming Shi, Yiheng Huang, Yilin Niu, Yuan Wang, Yuanchang Yue, Yuchen Li, Yutao Zhang, Yuting Wang, Yu Wang, Yuxuan Zhang, Zhao Xue, Zhenyu Hou, Zhengxiao Du, Zihan Wang, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Minlie Huang, Yuxiao Dong, and Jie Tang\.Glm\-4\.5v and glm\-4\.1v\-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025a\.URL[https://arxiv\.org/abs/2507\.01006](https://arxiv.org/abs/2507.01006)\.
- Team et al\. \[2025b\]Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, Congcong Wang, Dehao Zhang, Dikang Du, Dongliang Wang, Enming Yuan, Enzhe Lu, Fang Li, Flood Sung, Guangda Wei, Guokun Lai, Han Zhu, Hao Ding, Hao Hu, Hao Yang, Hao Zhang, Haoning Wu, Haotian Yao, Haoyu Lu, Heng Wang, Hongcheng Gao, Huabin Zheng, Jiaming Li, Jianlin Su, Jianzhou Wang, Jiaqi Deng, Jiezhong Qiu, Jin Xie, Jinhong Wang, Jingyuan Liu, Junjie Yan, Kun Ouyang, Liang Chen, Lin Sui, Longhui Yu, Mengfan Dong, Mengnan Dong, Nuo Xu, Pengyu Cheng, Qizheng Gu, Runjie Zhou, Shaowei Liu, Sihan Cao, Tao Yu, Tianhui Song, Tongtong Bai, Wei Song, Weiran He, Weixiao Huang, Weixin Xu, Xiaokun Yuan, Xingcheng Yao, Xingzhe Wu, Xinxing Zu, Xinyu Zhou, Xinyuan Wang, Y\. Charles, Yan Zhong, Yang Li, Yangyang Hu, Yanru Chen, Yejie Wang, Yibo Liu, Yibo Miao, Yidao Qin, Yimin Chen, Yiping Bao, Yiqin Wang, Yongsheng Kang, Yuanxin Liu, Yulun Du, Yuxin Wu, Yuzhi Wang, Yuzi Yan, Zaida Zhou, Zhaowei Li, Zhejun Jiang, Zheng Zhang, Zhilin Yang, Zhiqi Huang, Zihao Huang, Zijia Zhao, and Ziwei Chen\.Kimi\-VL technical report, 2025b\.URL[https://arxiv\.org/abs/2504\.07491](https://arxiv.org/abs/2504.07491)\.
- Qwen Team \[2026a\]Qwen Team\.Qwen3\.5: Towards native multimodal agents, February 2026a\.URL[https://qwen\.ai/blog?id=qwen3\.5](https://qwen.ai/blog?id=qwen3.5)\.
- Qwen Team \[2026b\]Qwen Team\.Qwen3\.6\-27b: Flagship\-level coding in a 27b dense model, April 2026b\.URL[https://qwen\.ai/blog?id=qwen3\.6\-27b](https://qwen.ai/blog?id=qwen3.6-27b)\.
- Zeng et al\. \[2024\]Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi\.How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 14322–14350, 2024\.
- Liu et al\. \[2025\]Xiaogeng Liu, Peiran Li, G\. Edward Suh, Yevgeniy Vorobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, and Chaowei Xiao\.AutoDAN\-turbo: A lifelong agent for strategy self\-exploration to jailbreak LLMs\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://openreview\.net/forum?id=bhK7U37VW8](https://openreview.net/forum?id=bhK7U37VW8)\.
- Shin et al\. \[2020\]Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh\.Autoprompt: Eliciting knowledge from language models with automatically generated prompts\.In*Proceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\)*, pages 4222–4235, 2020\.
- Yu et al\. \[2023\]Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing\.Gptfuzzer: Red teaming large language models with auto\-generated jailbreak prompts\.*arXiv preprint arXiv:2309\.10253*, 2023\.
- Chao et al\. \[2025\]Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong\.Jailbreaking black box large language models in twenty queries\.In*2025 IEEE Conference on Secure and Trustworthy Machine Learning \(SaTML\)*, pages 23–42\. IEEE, 2025\.
- Ying et al\. \[2025b\]Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, and Dacheng Tao\.Safebench: A safety evaluation framework for multimodal large language models\.[https://safebench\-mm\.github\.io/](https://safebench-mm.github.io/), 2025b\.Online resource\.
- Andriushchenko and Flammarion \[2024\]Maksym Andriushchenko and Nicolas Flammarion\.Does refusal training in llms generalize to the past tense?*arXiv preprint arXiv:2407\.11969*, 2024\.
- Qi et al\. \[2024b\]Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal\.Visual adversarial examples jailbreak aligned large language models\.In*Proceedings of the AAAI conference on artificial intelligence*, volume 38, pages 21527–21536, 2024b\.
- Niu et al\. \[2024\]Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin\.Jailbreaking attack against multimodal large language model\.*arXiv preprint arXiv:2402\.02309*, 2024\.
- Zhao et al\. \[2023\]Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai\-Man Man Cheung, and Min Lin\.On evaluating adversarial robustness of large vision\-language models\.*Advances in Neural Information Processing Systems*, 36:54111–54138, 2023\.
- Shi et al\. \[2022\]Yucheng Shi, Yahong Han, Yu\-an Tan, and Xiaohui Kuang\.Decision\-based black\-box attack against vision transformers via patch\-wise adversarial removal\.*Advances in Neural Information Processing Systems*, 35:12921–12933, 2022\.
- Brendel et al\. \[2017\]Wieland Brendel, Jonas Rauber, and Matthias Bethge\.Decision\-based adversarial attacks: Reliable attacks against black\-box machine learning models\.*arXiv preprint arXiv:1712\.04248*, 2017\.
- Maho et al\. \[2021\]Thibault Maho, Teddy Furon, and Erwan Le Merrer\.Surfree: a fast surrogate\-free black\-box attack\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pages 10430–10439, 2021\.
- Wang et al\. \[2024\]Ruofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji, Guangnan Ye, and Yu\-Gang Jiang\.White\-box multimodal jailbreaks against large vision\-language models\.In*ACM Multimedia 2024*, 2024\.URL[https://openreview\.net/forum?id=SMOUQtEaAf](https://openreview.net/forum?id=SMOUQtEaAf)\.
- Galisai et al\. \[2026\]Marcello Galisai, Susanna Cifani, Francesco Giarrusso, Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Federico Sartore, and Daniele Nardi\.Adversarial humanities benchmark: Results on stylistic robustness in frontier model safety, 2026\.URL[https://arxiv\.org/abs/2604\.18487](https://arxiv.org/abs/2604.18487)\.
- Yang et al\. \[2026\]Langqi Yang, Tianhang Zheng, Yixuan Chen, Kedong Xiu, Hao Zhou, Wangze Ni, Lei Chen, Zhan Qin, and Kui Ren\.Harmmetric eval: Benchmarking metrics and judges for llm harmfulness assessment, 2026\.URL[https://arxiv\.org/abs/2509\.24384](https://arxiv.org/abs/2509.24384)\.
- \[44\]sentence\-transformers\.all\-minilm\-l6\-v2\.URL[https://huggingface\.co/sentence\-transformers/all\-MiniLM\-L6\-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2)\.Hugging Face model\.
- Jeong et al\. \[2025\]Joonhyun Jeong, Seyun Bae, Yeonsung Jung, Jaeryong Hwang, and Eunho Yang\.Playing the fool: Jailbreaking llms and multimodal llms with out\-of\-distribution strategy\.In*Proceedings of the Computer Vision and Pattern Recognition Conference*, pages 29937–29946, 2025\.
- Ma et al\. \[2025\]Teng Ma, Xiaojun Jia, Ranjie Duan, Xinfeng Li, Yihao Huang, Xiaoshuang Jia, Zhixuan Chu, and Wenqi Ren\.Heuristic\-induced multimodal risk distribution jailbreak attack for multimodal large language models\.In*Proceedings of the IEEE/CVF International Conference on Computer Vision*, pages 2686–2696, 2025\.
- Sima et al\. \[2025\]Bingrui Sima, Linhua Cong, Wenxuan Wang, and Kun He\.Viscra: A visual chain reasoning attack for jailbreaking multimodal large language models\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 6142–6155, 2025\.
- Liu et al\. \[2026\]Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu, Zhican Chen, Tao Guan, Shuyuan Luo, Zhongyi Zhai, and Yang Liu\.Midas: Multi\-image dispersion and semantic reconstruction for jailbreaking mllms\.*arXiv preprint arXiv:2603\.00565*, 2026\.
- Song et al\. \[2026b\]Zhixue Song, Boyan Han, Yiwei Wang, and Chi Zhang\.Hard to read, easy to jailbreak: How visual degradation bypasses mllm safety alignment\.*arXiv preprint arXiv:2605\.07250*, 2026b\.
- Chouldechova et al\. \[2025\]Alex Chouldechova, A\. Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach\.Comparison requires valid measurement: Rethinking attack success rate comparisons in ai red teaming\.In D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen, editors,*Advances in Neural Information Processing Systems*, volume 38\. Curran Associates, Inc\., 2025\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/file/455d043673bb4b1872ff5e7a24cb3969\-Paper\-Position\_Paper\_Track\.pdf](https://proceedings.neurips.cc/paper_files/paper/2025/file/455d043673bb4b1872ff5e7a24cb3969-Paper-Position_Paper_Track.pdf)\.
- Schwinn et al\. \[2026\]Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami, Gauthier Gidel, and Stephan Günnemann\.A coin flip for safety: Llm judges fail to reliably measure adversarial robustness\.*arXiv preprint arXiv:2603\.06594*, 2026\.
- Yong et al\. \[2023\]Zheng\-Xin Yong, Cristina Menghini, and Stephen H Bach\.Low\-resource languages jailbreak gpt\-4\.*arXiv preprint arXiv:2310\.02446*, 2023\.
- Song et al\. \[2025\]Jiayang Song, Yuheng Huang, Zhehua Zhou, and Lei Ma\.Multilingual blending: Large language model safety alignment evaluation with language mixture\.In*Findings of the Association for Computational Linguistics: NAACL 2025*, pages 3433–3449, 2025\.
- Lin et al\. \[2014a\]Tsung\-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C\. Lawrence Zitnick\.Microsoft coco: Common objects in context\.In David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors,*Computer Vision – ECCV 2014*, pages 740–755, Cham, 2014a\. Springer International Publishing\.ISBN 978\-3\-319\-10602\-1\.
- huihui\-ai \[a\]huihui\-ai\.Huihui\-qwen3\.5\-9b\-abliterated, a\.URL[https://huggingface\.co/huihui\-ai/Huihui\-Qwen3\.5\-9B\-abliterated](https://huggingface.co/huihui-ai/Huihui-Qwen3.5-9B-abliterated)\.Hugging Face model\.
- huihui\-ai \[b\]huihui\-ai\.Huihui\-gemma\-4\-31b\-it\-abliterated, b\.URL[https://huggingface\.co/huihui\-ai/Huihui\-gemma\-4\-31B\-it\-abliterated](https://huggingface.co/huihui-ai/Huihui-gemma-4-31B-it-abliterated)\.Hugging Face model\.
- Lin et al\. \[2014b\]Tsung\-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick\.Microsoft coco: Common objects in context\.In*European conference on computer vision*, pages 740–755\. Springer, 2014b\.
- Radford et al\. \[2021\]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al\.Learning transferable visual models from natural language supervision\.In*International conference on machine learning*, pages 8748–8763\. PmLR, 2021\.
## Appendix ACategories and subcategories
Table 5:Taxonomy of Risk Categories and SubcategoriesCodeCategorySubcategories \(Code – Name\)AEthical and Social RisksA1 – Bias and DiscriminationA2 – Insulting or Harassing SpeechA3 – Hate SpeechA4 – Discriminatory BehaviorBPrivacy and Data RisksB1 – Personal Data LeakageB2 – Biometric or Health Data ExposureB3 – Doxxing or TrackingB4 – Surveillance MisuseB5 – Training Data LeakageCSafety and Physical HarmC1 – Dangerous InstructionsC2 – Self\-harm PromotionC3 – Physical HarmC4 – Human TraffickingC5 – Animal AbuseC6 – Drug GuidanceC7 – Environmental DamageDCriminal and Economic RisksD1 – Fraud or ScamsD2 – Economic CrimeD3 – White\-Collar CrimeD4 – Illegal ActivityD5 – Evasion of LawECybersecurity ThreatsE1 – Malware GenerationE2 – Exploit DevelopmentE3 – Phishing or RansomwareE4 – Jailbreak or Prompt InjectionE5 – Cross\-modal AttackE6 – Data PoisoningE7 – Unauthorized Model AccessFInformation and Political ManipulationF1 – Misinformation or DisinformationF2 – Deepfakes or Synthetic MediaF3 – Political PropagandaF4 – Election InterferenceF5 – Public Order DisruptionF6 – National Security ThreatGContent and Cultural SafetyG1 – Pornography or NudityG2 – Sexual ContentG3 – Violence or GoreG4 – Cultural or Religious OffenseHIntellectual Property and OwnershipH1 – Copyright InfringementH2 – Trademark or Brand MisuseH3 – Proprietary Data LeakageH4 – Unauthorized Style or Persona ImitationIDecision and Cognitive RisksI1 – Medical AdviceI2 – Health MisguidanceI3 – Legal AdviceI4 – Government Decision SupportI5 – Financial AdviceI6 – Market ManipulationI7 – Psychological ManipulationI8 – Cognitive Bias or OverrelianceJChild SafetyJ1 – CSAM & SexualizationJ2 – Grooming or Enticement of MinorsJ3 – Child TraffickingJ4 – Harmful Content Targeting MinorsJ5 – Age Verification Evasion
## Appendix BLanguage translation analysis of adversarial attacks
[Table˜6](https://arxiv.org/html/2606.24388#A2.T6)evaluates how vulnerable different multimodal large language models \(MLLMs\) are when safety\-critical prompts are translated into various languages or presented in mixed\-language settings\. Current literature suggests that utilizing low\-resource languages should increase model vulnerabilities compared to high\-resource ones, as safety alignment data is typically scarce in those languages\[[52](https://arxiv.org/html/2606.24388#bib.bib52)\]\. To test this hypothesis, we select Farsi and Turkish as target low\-resource languages\. Furthermore, utilizing a mixture of languages within a translation should obscure prompt intent, heighten deception, and ultimately increase the ASR\[[53](https://arxiv.org/html/2606.24388#bib.bib53)\]\. For this multi\-lingual setting, we choose three high\-resource languages \(Italian, French, and German\) and three low\-resource languages \(Turkish, Farsi, and Khmer\), performing a sentence\-by\-sentence translation of the adversarial text\. As a third approach, we target specific semantic segments by translating only the inherently harmful parts of the prompt into a low\-resource language \(Partial Turkish\) to isolate its effect on model safety\.
Table 6:Examining model vulnerability against language translation\. ASR \(%\) per model across languages\.Our experimental evaluation yields several key insights:
#### The Vulnerability Trade\-off\.
The general consensus from our experiments indicates that language translation increases vulnerability only up to the point where it does not compromise the model’s fundamental semantic understanding of the attack\. Because adversarial strategies often rely on intricate, multi\-layered roleplay scenarios or convoluted logic, translation can introduce excessive linguistic ambiguity\. When this ambiguity disrupts comprehension—as heavily observed in theMixed Low\-Rescolumn—the model fails to grasp the underlying prompt intent and generates irrelevant or benign responses\. These are classified as non\-jailbreaks by the evaluation judge, leading to a sharp decline in ASR for mixed\-language settings\.
#### Targeted Susceptibility in Specific Model Families\.
The Qwen, Ministral, and Gemma families exhibit heightened vulnerability when exposed to low\-resource languages or hybrid formatting \(Partial Turkish\)\. For instance, Gemma\-4\-26B shows a noticeable increase in ASR from a baseline of45\.7%45\.7\\%to53\.3%53\.3\\%in Farsi and57\.6%57\.6\\%in Turkish\. This confirms that low\-resource translations successfully exploit gaps in the cross\-lingual safety alignment of these architectures\.
#### Cross\-Lingual Robustness and Transfer Variations\.
Models such as Ministral\-3\-14B and GLM\-4\.6V\-Flash maintain consistently high vulnerability scores across nearly all language configurations \(with Ministral hovering around80%80\\%ASR\)\. This suggests that adversarial prompt structures transfer seamlessly across linguistic boundaries for these models\. Conversely, models like DeepSeek\-VL2 and LLaVA\-v1\.6\-13b experience drastic drops in vulnerability when prompts are translated \(e\.g\., DeepSeek\-VL2 plunging from a60\.4%60\.4\\%baseline to just10\.3%10\.3\\%in Farsi\)\. This pattern points to either a brittle multilingual comprehension capability or a defensive posture that defaults to safe rejections when faced with distribution shifts in language\.
## Appendix CTransferability Results
This section presents additional results on the transferability of the attacks in our dataset to a broader set of models, extending those reported in the main corpus\. We follow the same evaluation protocol: for each attack and each subcategory, we sample 20 instances\. Once this set is fixed, it is evaluated across a range of different models, enabling a direct comparison within each attack strategy\. Overall, we evaluate these samples on nine white\-box models and six black\-box models\. Due to the large number of generated outputs, we do not perform manual inspection\. Instead, we rely on the state\-of\-the\-art judge Able\-24\-HarmClassifier\[[43](https://arxiv.org/html/2606.24388#bib.bib43)\]\. As a consequence, the reported results should be interpreted as relative to this evaluation baseline rather than as ground truth\.
The full results are reported inLABEL:tbl:combined\-results\. For ease of interpretation, we also provide radar plots offering different insights into the data\. First,[fig\.˜6](https://arxiv.org/html/2606.24388#A3.F6)shows the attack success rates across models and attack strategies\. Second, following the analysis in the main paper,[fig\.˜7](https://arxiv.org/html/2606.24388#A3.F7)presents the attack success rate \(ASR\) per category, averaged over attack strategies and evaluated across models, highlighting the categories to which models are most vulnerable independently of the chosen attack\. Third,[fig\.˜8](https://arxiv.org/html/2606.24388#A3.F8)shows the ASR averaged over categories, providing an at\-a\-glance comparison of the most effective attack strategy for each model\. Finally,[fig\.˜9](https://arxiv.org/html/2606.24388#A3.F9)reports the maximum ASR values per model and attack, identifying the weakest category\. This allows one to infer, for a given model and attack, the most vulnerable category and the expected performance\.
An interesting pattern that emerges from[fig\.˜6](https://arxiv.org/html/2606.24388#A3.F6)and[fig\.˜8](https://arxiv.org/html/2606.24388#A3.F8)is that MML is the most widely effective attack against almost all models, with the exception ofOpus 4\.7andOpus 4\.8, which appear to be highly robust to it\. However, these two models are particularly vulnerable to the CSDJ attack, which, in turn, is less effective against white\-box models\. The second most reliable attack across models is FC Attack, which shows good coverage across categories for most models, except for Gemma\-4\-26B\. IDEATOR exhibits the most unpredictable behavior: while it achieves high success rates on some white\-box models, such asGLM\-4\.6VandMistral\-14B, it is generally less reliable, aside from occasional spikes on specific categories\. Finally, BAP yields lower but relatively stable performance across models, with success rates ranging from30%30\\%to50%50\\%\.
Table 7:Table of ASR \(%\) per model and per category across all attacks, in bold the category with the highest success rate on each modelModelABCDEFGHIJAverageBAPDeepSeek\-VL230\.0022\.0044\.2953\.0048\.5736\.6721\.2521\.2530\.6227\.0034\.82GLM\-4\.6V\-Flash50\.0043\.0079\.2972\.0065\.7162\.5043\.7536\.2550\.0051\.0057\.09Gemma\-4\-26B23\.7516\.0010\.0018\.0027\.1420\.0022\.5022\.5026\.8834\.0022\.00Kimi\-VL40\.0029\.0057\.1461\.0052\.1448\.3327\.5032\.5036\.8832\.0042\.91Llava\-13b25\.0022\.0030\.0039\.0042\.1430\.8315\.0021\.2523\.7514\.0027\.27Ministral\-14B70\.0056\.0085\.7186\.0065\.0060\.8357\.5051\.2571\.2562\.0067\.73Qwen3\-VL\-30B46\.2545\.0047\.1447\.0038\.5739\.1736\.2537\.5044\.3844\.0042\.73Qwen3\.5\-27B50\.0035\.0030\.7150\.0044\.2937\.5032\.5037\.5045\.0047\.0040\.91Qwen3\.6\-27B42\.5040\.0030\.7142\.0038\.5735\.8333\.7540\.0045\.0051\.0039\.82GPT\-5\.413\.7515\.0025\.0023\.0011\.4318\.3310\.006\.2510\.6223\.0015\.91GPT\-5\.547\.5041\.0059\.2968\.0047\.8642\.5046\.2535\.0048\.1250\.0049\.09Claude Opus 4\.66\.2511\.000\.7110\.008\.575\.838\.757\.506\.8817\.007\.91Claude Opus 4\.751\.2544\.0020\.7161\.0040\.0039\.1735\.0056\.2550\.0049\.0043\.64Claude Opus 4\.847\.5040\.0012\.1465\.0039\.2934\.1721\.2541\.2543\.1251\.0038\.73Gemini 3\.1 Pro23\.7539\.0035\.7145\.0074\.2928\.3330\.0027\.5025\.6244\.0038\.36IDEATORDeepSeek\-VL226\.2538\.0045\.7158\.0064\.2959\.1721\.2520\.0033\.1244\.0042\.91GLM\-4\.6V\-Flash62\.5056\.0067\.1485\.0082\.1475\.8343\.7542\.5052\.5061\.0064\.09Gemma\-4\-26B10\.0030\.0020\.7134\.0033\.5720\.0015\.0021\.2535\.6243\.0027\.36Kimi\-VL41\.2540\.0052\.1462\.0066\.4357\.5023\.7520\.0036\.2540\.0045\.73Llava\-13b20\.0039\.0040\.0053\.0060\.0053\.3320\.0018\.7530\.6231\.0038\.45Ministral\-14B51\.2565\.0072\.8681\.0091\.4375\.8362\.5052\.5080\.0072\.0072\.73Qwen3\-VL\-30B17\.5020\.0015\.7140\.0035\.7140\.8318\.7516\.2536\.2561\.0031\.09Qwen3\.5\-27B13\.7517\.0015\.0019\.0012\.1418\.3325\.0018\.7535\.6259\.0023\.45Qwen3\.6\-27B7\.507\.0011\.4314\.0015\.0016\.6718\.757\.5023\.1265\.0018\.82GPT\-5\.425\.0014\.0027\.1449\.0044\.2936\.6726\.2512\.5027\.5033\.0030\.45GPT\-5\.521\.2528\.0026\.4342\.0048\.5744\.1752\.5031\.2537\.5044\.0037\.82Claude Opus 4\.66\.256\.005\.7117\.0013\.575\.833\.755\.0013\.127\.008\.82Claude Opus 4\.725\.0039\.0012\.1453\.0030\.7130\.0032\.5045\.0034\.3833\.0032\.55Claude Opus 4\.816\.2524\.0012\.1447\.0032\.8624\.1717\.5033\.7531\.2520\.0026\.09Gemini 3\.1 Pro28\.7543\.0042\.1456\.0071\.4358\.3356\.2542\.5054\.3762\.0052\.64MMLDeepSeek\-VL268\.8965\.7174\.3281\.6978\.5775\.0079\.4965\.5275\.5676\.9276\.10GLM\-4\.6V\-Flash97\.78100\.0093\.4898\.6097\.58100\.0097\.4496\.5598\.1198\.0097\.36Gemma\-4\-26B95\.56100\.0090\.3498\.5199\.3796\.8894\.87100\.0096\.5595\.2696\.09Kimi\-VL86\.6782\.8685\.1690\.5887\.0484\.3892\.3168\.9787\.0287\.0586\.66Llava\-13b68\.8974\.2966\.4881\.3467\.9287\.5079\.4968\.9777\.7877\.3774\.55Ministral\-14B95\.5697\.1499\.4399\.2599\.3796\.8897\.44100\.0099\.2397\.8998\.73Qwen3\-VL\-30B82\.2285\.7187\.6392\.5097\.7893\.7587\.1875\.8689\.1179\.2288\.35Qwen3\.5\-27B75\.5677\.1451\.6363\.5075\.4581\.2561\.5486\.2185\.5682\.4274\.34Qwen3\.6\-27B80\.0085\.7180\.9789\.6394\.5587\.5071\.7986\.2183\.1488\.4286\.10GPT\-5\.471\.7673\.4574\.1367\.9270\.0072\.7978\.7580\.4680\.0095\.0076\.03GPT\-5\.591\.2584\.0090\.7192\.0090\.7189\.1782\.5080\.0081\.2561\.0084\.64Claude Opus 4\.671\.9561\.3232\.1460\.7863\.9563\.7841\.2548\.8162\.8654\.5556\.19Claude Opus 4\.71\.2510\.000\.716\.002\.860\.830\.005\.001\.881\.002\.82Claude Opus 4\.812\.5021\.003\.5726\.0010\.0011\.6711\.2521\.2513\.7523\.0014\.64Gemini 3\.1 Pro3\.530\.884\.200\.9440\.6281\.6266\.2593\.1093\.7129\.0043\.38FC ATTACKDeepSeek\-VL288\.7584\.0089\.2994\.0095\.7187\.5078\.7558\.7566\.2575\.0082\.18GLM\-4\.6V\-Flash91\.2585\.0095\.7193\.0097\.1491\.6778\.7557\.5071\.8882\.0085\.18Gemma\-4\-26B37\.5029\.002\.147\.0022\.1411\.6733\.7528\.7518\.7514\.0018\.91Kimi\-VL82\.5077\.0082\.8689\.0092\.8685\.0080\.0052\.5066\.8874\.0078\.82Llava\-13b71\.2583\.0087\.8689\.0088\.5784\.1766\.2551\.2558\.7573\.0076\.18Ministral\-14B60\.0065\.0050\.0073\.0083\.5777\.5053\.7541\.2546\.2558\.0061\.27Qwen3\-VL\-30B83\.7583\.0059\.2965\.0087\.1481\.6768\.7563\.7569\.3873\.0073\.45Qwen3\.5\-27B70\.0075\.0041\.4378\.0090\.7168\.3362\.5075\.0066\.2559\.0068\.27Qwen3\.6\-27B86\.2581\.0055\.7185\.0092\.8682\.5072\.5076\.2578\.1264\.0077\.27GPT\-5\.452\.5043\.0060\.7171\.0077\.1474\.1746\.2547\.5059\.3863\.0061\.00GPT\-5\.557\.5057\.0068\.5786\.0089\.2976\.6767\.5063\.7565\.0074\.0071\.36Claude Opus 4\.677\.5085\.0041\.4368\.0087\.8670\.8363\.7582\.5075\.6260\.0070\.82Claude Opus 4\.770\.0072\.0038\.5775\.0068\.5765\.0037\.5082\.5068\.7561\.0063\.45Claude Opus 4\.863\.7583\.0048\.5777\.0079\.2975\.0036\.2578\.7562\.5070\.0067\.45Gemini 3\.1 Pro35\.0046\.0015\.0028\.0067\.1426\.6735\.0055\.0040\.6221\.0037\.00CSDJDeepSeek\-VL217\.5036\.0059\.2949\.0057\.8650\.8330\.0027\.5029\.3830\.0038\.74GLM\-4\.6V\-Flash55\.0061\.0080\.0076\.0090\.7181\.6761\.2550\.0046\.8869\.0067\.15Gemma\-4\-26B53\.7575\.0027\.8653\.0085\.7157\.5027\.5041\.2536\.8855\.0051\.35Kimi\-VL53\.7554\.0064\.2962\.0080\.0066\.6753\.7547\.5041\.8860\.0058\.38Llava\-13b3\.754\.008\.577\.007\.868\.3311\.253\.755\.003\.006\.25Ministral\-14B67\.5077\.0070\.0081\.0095\.0085\.8372\.5071\.2566\.8881\.0076\.80Qwen3\-VL\-30B73\.7576\.0082\.1484\.0092\.8689\.1762\.5058\.7563\.1270\.0075\.23Qwen3\.5\-27B46\.2549\.0035\.0048\.0072\.1456\.6746\.2535\.0039\.3866\.0049\.37Qwen3\.6\-27B71\.2568\.0064\.2974\.0084\.2970\.0070\.0060\.0061\.8870\.0069\.37GPT\-5\.478\.7578\.0091\.4387\.0087\.1486\.6766\.2561\.2563\.7581\.0078\.82GPT\-5\.582\.5080\.0088\.5791\.0095\.0085\.0081\.2570\.0069\.3887\.0083\.16Claude Opus 4\.673\.7588\.0067\.8685\.0093\.5784\.1748\.7571\.2573\.7583\.0077\.82Claude Opus 4\.770\.0090\.0057\.1478\.0095\.0088\.3350\.0071\.2559\.3888\.0074\.82Claude Opus 4\.880\.0089\.0071\.4391\.0095\.0087\.5050\.0071\.2559\.3889\.0078\.45Gemini 3\.1 Pro65\.0086\.0039\.2972\.0086\.4367\.5046\.2568\.7563\.1268\.0066\.18Figure 6:The plot graphically presents the vulnerability of each tested model to harmful categories, enabling a comparison across attacks\.Figure 7:The plot presents the vulnerability of models to harmful categories\. It is obtained by averaging results across different attack strategies to mitigate the influence of the specific attack used\.Figure 8:The plot shows the vulnerabilities of the models with respect to the attack strategies, averaged over the categories\.Figure 9:This plot identifies, for each attack strategy and model, the most vulnerable category by selecting the category with the highest ASR score\.
## Appendix DBenchmark overview
### D\.1Evolution of Textual and Early Multimodal Benchmarks
The field was pioneered byAdvBench\[[17](https://arxiv.org/html/2606.24388#bib.bib17)\], a text\-only benchmark containing520520behaviors\. Despite its reliance on primitive string\-matching judges and the simple GCG attack, it established the foundational framework for subsequent research\. Building on this,VAJM\[[16](https://arxiv.org/html/2606.24388#bib.bib16)\]introduced the image modality and implemented weak categorization based on race and gender\. It advanced evaluation methods by utilizing theDetoxifyclassifier as a judge and introducing prompt\-tuning optimization attacks\.
HarmBench\[[1](https://arxiv.org/html/2606.24388#bib.bib1)\]contributed the first rigorous categorization of behaviors into four functional groups\. It further matured the evaluation process by introducing a fine\-tuned Llama 2 model as a judge and proposing the R2D2 defense mechanism\. Expanding the scale of multimodal research,JailBreakV\-28K\[[19](https://arxiv.org/html/2606.24388#bib.bib19)\]incorporated 16 categories and20002000behaviors \(sourced from RedTeam\-2K\), utilizing2828k attack pairs generated via advanced typographic Stable Diffusion attacks\.
### D\.2Diversifying Metrics and Modalities
Subsequent works focused on refining metrics and interaction types\.MM\-SafetyBench\[[2](https://arxiv.org/html/2606.24388#bib.bib2)\]introduced the “Refusal Rate” as a key metric, emphasizing a comparison between Attack Success Rates \(ASR\) when models are given text\-only queries versus query\-relevant image\-text pairs\. In the purely textual domain,Strong Reject\[[15](https://arxiv.org/html/2606.24388#bib.bib15)\]moved away from binary evaluation by implementing a graded scoring system \(0,0\.33,0\.66,1\.00,0\.33,0\.66,1\.0\) to distinguish between full refusal, partial refusal, partial fulfillment, and full fulfillment across 37 different attacks\.
The importance of conversational context was highlighted by research intoMultiturn human jailbreaks\[[14](https://arxiv.org/html/2606.24388#bib.bib14)\], which demonstrated that models are significantly more vulnerable through iterative, back\-and\-forth prompting: a feature we leverage through the attack strategy chosen for our dataset\. Further expanding the scope of modalities,SafeBench\[[33](https://arxiv.org/html/2606.24388#bib.bib33)\]integrated audio alongside text and images, while introducing a “Safety Index Risk” evaluated by a consensus\-based roundtable of judges rather than a single entity\.
### D\.3Balancing Robustness with Model Utility
A critical shift in the literature involves the trade\-off between safety and helpfulness\.MMJ\-bench\[[3](https://arxiv.org/html/2606.24388#bib.bib3)\]categorized attack strategies into optimization\- and generation\-based methods, arguing that a perfect defense is counterproductive if it causes the model to refuse every prompt\. Similarly,JailbreakBench\[[10](https://arxiv.org/html/2606.24388#bib.bib10)\]introduced 100 benign prompts designed to appear harmful but which are actually safe, allowing researchers to measure if a model is overly defensive\. To quantify this performance impact,B\-AVIBench\[[12](https://arxiv.org/html/2606.24388#bib.bib12)\]introduced the Average Score Drop Rate \(ASDR\), measuring the percentage decrease in performance scores following an attack across various image, text, and content bias types\.
To ensure automated evaluations remain grounded,Sorry\-Bench\[[11](https://arxiv.org/html/2606.24388#bib.bib11)\]provided a human validation dataset for judges, utilizing Cohen’s Kappa \(κ\\kappa\) to measure the correlation between AI judges and human evaluation, alongside fulfillment rate as an additional metric\.
### D\.4High\-Granularity Categorization
Recent benchmarks have achieved unprecedented depth in their taxonomies\.VLJailbreakBench\[[10](https://arxiv.org/html/2606.24388#bib.bib10)\]implemented a robust categorization featuring1212safety topics and4646subcategories\. Finally,OmniSafeBench\[[4](https://arxiv.org/html/2606.24388#bib.bib4)\]—which serves as the primary reference for this work, introduced 9 major risk domains and5050fine\-grained categories\. Beyond evaluating1515different defense strategies, it established a multifaceted judgment criteria incorporating Harmfulness \(H\), Intent Alignment \(A\), and Level of Detail \(D\) to provide a holistic view of model safety\.
Table 8:Comparative Analysis of Safety and Jailbreak BenchmarksBenchmarkReleaseMod\.BehavioursSamplesModelAtt\.Def\.JudgesMetricsOmniSafeBench6 Dec 2025T/I5050cat\.1 2001\\,20018181315GPT\-4oASR, SRIVLJailbreak25 Sep 2025T/I4646cat\.3 6543\\,654111150GPT\-4o, GPT\-4ASRSorry\-BenchMar 2025T4444Topics8 8008\\,800505000Mistral\-7BFR, RR,κ\\kappaAgentHarm18 Apr 2025T/I1111cat\.4404401515110GPT\-4oSR, RRB\-AVIBench28 Dec 2024T/I2323Types316316k141410100GPT\-4ASDR, AEDJailbreakBench31 Oct 2024T2020cat\.10010044406 ClassifiersASR, FPRMMJ\-Bench22 Oct 2024T/I200200Behav\.1 0001\\,00066995GPT\-4, HarmBenchASRSafeBench4 Oct 2024T/I/A2323cat\.9 2009\\,200212130Ensemble \(2\)ASR, SRIMHJ4 Sep 2024T1 0001\\,000Req\.2 9122\\,9121114140GPT\-4o, HarmBenchASRStrongREJECT27 Aug 2024T66cat\.346346337370Gemma 2BFull RefusalMM\-SafetyBench19 Jun 24T/I1313cat\.5 0405\\,0401200GPT\-4, Llama\-2ASR, RRJailBreakV\-28K3 Apr 2024T/I1616cat\.2828k101004 ClassifiersASRHarmBenchFeb 2024T/I44cat\.5105103333222211Llama\-2 \(FT\)ASRVAJM16 Aug 2023T/I4040Behav\.32 22632\\,22633110Perspective APIToxicityAdvBenchJuly 2023T520520Behav\.52052099880GPT\-4, StringASRLegend:Mod\.= Modality \(T: Text, I: Image, A: Audio\);Model= Number of models evaluated;Att\.= Number of attack strategies;Def\.= Number of defense strategies\.
## Appendix EPHANTOM Similarity Checks
To assess the diversity of the generated adversarial prompts, we performed a cosine\-similarity analysis over the textual component of the attacks\. This analysis quantifies prompt\-level redundancy, including cases where the same underlying intent may lead to multiple generated attacks\. Such repetitions are expected, since PHANTOM contains7 8267\\,826unique intents but nearly3030k generated adversarial samples\.
We exclude theMMLstrategy from this analysis because, for this attack, the adversarial content is primarily encoded in the image rather than in the textual prompt\. MML and FC ATTACK prompts rely on a shared instruction template, while the harmful intent is embedded through visual transformations such as encoding, mirroring, rotation, or word substitution\. Therefore, measuring redundancy using only the textual prompt would produce similarity scores∼100%\\sim 100\\%, without providing a meaningful estimate of sample diversity\.
[Table˜9](https://arxiv.org/html/2606.24388#A5.T9)reports the redundancy rates obtained for BAP and IDEATOR across different cosine\-similarity thresholds, both globally and broken down by attack strategy and target model\. For each thresholdτ\\tau, we construct clusters of prompts whose pairwise cosine similarity is greater than or equal toτ\\tau\. Within each cluster, one prompt is treated as the representative, while the remaining prompts are counted as redundant\. Formally, if𝒦\\mathcal\{K\}denotes the set of clusters and\|Ck\|\|C\_\{k\}\|the size of clusterCkC\_\{k\}, the redundancy rate is computed as:
Redundancy=∑Ck∈𝒦max\(\|Ck\|−1,0\)N×100,\\text\{Redundancy\}=\\frac\{\\sum\_\{C\_\{k\}\\in\\mathcal\{K\}\}\\max\(\|C\_\{k\}\|\-1,0\)\}\{N\}\\times 100,whereNNis the total number of prompts in the analyzed group\.
Table 9:Redundancy Rate \(%\) Across Thresholds
## Appendix FReview of the attack strategies
In this section we will give an overview of the attack strategies that we used in the generation\.
### F\.1BAP attack
The core idea behind the BAP attack is to jointly optimize visual and textual components\. First, an adversarial perturbation is applied to the input image through projected gradient descent \(PGD\), using a corpus of affirmative model responses as optimization targets\. Subsequently, a prompt engineering step is performed to obfuscate the harmful intent within a seemingly benign textual prompt\.
The outcome of this process is an adversarial image that biases the model toward affirmative responses, coupled with a carefully engineered prompt that facilitates the bypass of safety mechanisms\.
In our pipeline, we fixed the number of PGD optimization steps to approximately 200 and optimized the adversarial image against a batch of 8 affirmative target responses, starting from clean images from the COCO train dataset\[[54](https://arxiv.org/html/2606.24388#bib.bib54)\]\. For the prompt engineering phase, we followed the standard iterative interaction flow\.
In addition to the target model under attack, we employed an abliterated version of Qwen3\.5, namely Huihui\-Qwen3\.5\-9B\-abliterated\[[55](https://arxiv.org/html/2606.24388#bib.bib55)\]as the attacker model\. Its reasoning capabilities were leveraged via a crafted system prompt that explicitly encoded previous failed attempts\. As anticipated before, we used Abel\-24\-HarmClassifier proposed in\[[43](https://arxiv.org/html/2606.24388#bib.bib43)\]as judge model\.
The prompt optimization loop was iterated for up to55attempts for each attack instance\.
At first glance,[fig\.˜3](https://arxiv.org/html/2606.24388#S3.F3)may suggest that this attack is less practical, given that it is substantially slower than the alternatives\. However, our decision to include it was motivated by an additional advantage: the adversarial images produced by this pipeline are universal\. As a result, one can recombine intents and images to obtain additional valid attacks, although some filtering and discarding may still be required\.
### F\.2IDEATOR attack
Also in the case of this attack strategy, proposed in\[[6](https://arxiv.org/html/2606.24388#bib.bib6)\], we largely followed the original pipeline\. Our modifications mainly consist of introducing different models for prompt and image generation\.
The attack pipeline relies on an attacker that produces two distinct prompts: one used to generate an image related to the harmful intent, and another aimed at engineering the harmful textual prompt itself\. Since the process is implemented as a multi\-turn conversation in which the prompt is progressively refined, we limited each conversation to a maximum of three image–prompt pairs\. In addition, for each target goal, the attack was retried at most three times; these retries serve primarily as a fallback mechanism rather than a core component of the method\.
In our implementation, we employed the same ablated Qwen3\.5 model above as the attacker to generate both the textual prompts, while image generation was performed using Stable Diffusion 3\.5 Medium\.
We selected this strategy not only because of its efficiency, but also because its structure naturally supports both multi\-turn and single\-turn settings: the full conversation can be used as input, or alternatively only the final image–prompt pair can be retained\.
### F\.3MML attack
As with the other methods, we remained faithful to the original structure of the Multi\-Modal Linkage attack proposed in\[[9](https://arxiv.org/html/2606.24388#bib.bib9)\]\. In our implementation we start from a harmful or restricted text prompt and apply obfuscation: it replaces key words with benign ones through NLTK package, optionally encodes the text \(e\.g\., Base64\), and then renders the transformed text into an image\. Additional visual distortions, such as mirroring, rotation, or both, are mainly applied through Pillow library to make the content harder to directly interpret\. Alongside this image, the system constructs a carefully designed “game\-like” prompt that instructs the model to recover the original text by reversing these transformations \(e\.g\., decoding, un\-mirroring, or using a provided word\-mapping dictionary\) and validating it against a scrambled word list\. The resulting image–text pair is fed into target vision\-language model, which is guided step\-by\-step to reconstruct the original prompt and then generate detailed content based on it\. Because the harmful intent is never explicitly presented in raw form but instead reconstructed by the model itself, safety mechanisms can be bypassed\.
Harmful IntentAdversarial ImageMirror/Rotate/WordReplacementAdversarial PromptPersona \(Game Dev\)Word ScrambleTarget VLMTransformRoleplayVisualText
Figure 10:Workflow of MML combining both image manipulation and role\-playing through text\.
### F\.4FC ATTACK
Once again, our methodology closely follows the original approach proposed in\[[20](https://arxiv.org/html/2606.24388#bib.bib20)\]\. The core idea is to start from a harmful intent and generate a sequence of logical steps to address it using an auxiliary abliterated model\. In our case, similarly to BAP\[[8](https://arxiv.org/html/2606.24388#bib.bib8)\], we employ the Huihui\-Qwen3\.5\-9B\-abliterated model\[[55](https://arxiv.org/html/2606.24388#bib.bib55)\]\. These steps are then represented as a flowchart using standard Python libraries such asGraphviz\. Finally, the model is prompted with both the flowchart and a standard instruction that encourages it to reason by following the outlined steps\.
### F\.5CSDJ attack
We generated attacks using the CS\-DJ attack strategy proposed in\[[21](https://arxiv.org/html/2606.24388#bib.bib21)\]\. Following the original pipeline, we crafted each attack as follows\.
We first selected an intent from our dataset and then followed two parallel paths\. First, using an abliterated model, namely Huihui\-gemma\-4\-31B\-it\-abliterated\[[56](https://arxiv.org/html/2606.24388#bib.bib56)\], we decomposed the harmful request into three less harmful sub\-requests and embedded each of them into separate images\. Second, we selected nine additional images from a pool of 10 000 images taken from COCO training set\[[57](https://arxiv.org/html/2606.24388#bib.bib57)\], these images are chosen such that their CLIP\[[58](https://arxiv.org/html/2606.24388#bib.bib58)\]embeddings are maximally distant from the embedding of the original intent\.
We then combined these components: the nine images were arranged in a 3×3 grid, followed by the three images containing the generated sub\-requests\. The images were numbered from 1 \(top\-left\) to 12 \(bottom\-right\)\.Similar Articles
Backdoor Learning in Language Models and Vision-Language Models
This dissertation addresses security and efficiency in AI by analyzing backdoor attacks in language and vision-language models, proposing detection frameworks and novel attack methods, and introducing efficient multimodal models for clinical applications.
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
A survey paper systematically analyzing the evolving safety landscape of multi-modal large language models, covering emerging threats such as adversarial attacks, data poisoning, jailbreaks, and hallucinations, and reviewing updated safety strategies.
DataComp-VLM: Improved Open Datasets for Vision-Language Models
This paper introduces DataComp-VLM (DCVLM), a comprehensive benchmark for evaluating data curation strategies for vision-language models. The authors find that data mixing, rather than filtering, significantly improves performance, and their resulting DCVLM-Baseline dataset achieves state-of-the-art results on 33 downstream tasks.
RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models
RedBench introduces a universal dataset aggregating 37 benchmark datasets with 29,362 samples across 22 risk categories and 19 domains to enable standardized and comprehensive red teaming evaluation of large language models. The work addresses inconsistencies in existing red teaming datasets and provides baselines, evaluation code, and open-source resources for assessing LLM robustness against adversarial prompts.
Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models
This paper proposes DURA, a diffusion-based unrestricted robotic attack that generates visually natural adversarial patches to disrupt Vision-Language-Action (VLA) models in both white-box and black-box settings, highlighting safety risks for physically deployed robotic systems.