Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
Summary
Multi2AV-Safety is the first benchmark for evaluating safety in multimodal-to-audio-Video generation, covering all 11 non-singleton conditioning configurations with 11,024 attack instances, and revealing compositional risks where harmful semantics emerge from benign inputs.
View Cached Full Text
Cached at: 08/28/26, 09:35 AM
# Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
Source: [https://arxiv.org/html/2608.26535](https://arxiv.org/html/2608.26535)
Changtao Miao††thanks:Project Lead\.Affiliation:Independent ResearcherBaiqi WuAffiliation:Zhejiang UniversityZhiyuan LuAffiliation:Hefei University of TechnologyKang YangAffiliation:Hefei University of TechnologyPeiwei ZhaoAffiliation:Hefei University of TechnologyJunchi ChenAffiliation:University of Science and Technology of ChinaYunfeng DiaoAffiliation:Hefei University of TechnologyHe LiuAffiliation:Independent ResearcherQi Chu††thanks:Corresponding author\.Affiliation:University of Science and Technology of ChinaTao GongAffiliation:University of Science and Technology of ChinaNenghai YuAffiliation:University of Science and Technology of China
###### Abstract
Audio\-video generation is rapidly moving from prompt\-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output\. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time\. Existing safety benchmarks, however, remain largely prompt\-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically\. To bridge this gap, we introduce Multi2AV\-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non\-singleton T/I/A/V conditioning configurations for audio\-video generation, comprising 11,024 attack instances\. Evaluation onMulti2AV\-Safetyreveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm\-evidence structures\. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context\. Together, these results identify*compositional risk perception*as a central capability gap in safeguarding multimodal\-conditioned audio\-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable\.
Figure 1:Overview ofMulti2AV\-Safety: 11,024 attacks spanning all 11 multimodal combinations of text, image, audio, and video, four attack mechanisms, five harm categories\.## 1Introduction
Audio\-video generation is moving beyond prompt\-driven synthesis toward multimodal conditioning, where text instructions, reference images, speech, and video context jointly shape the generated result\([HaCohen et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib1)\)\. More importantly, multimodal conditioning changes how safety\-relevant evidence is expressed across inputs\. Under text\-only conditioning, unsafe intent is often explicit in the prompt\.By contrast, multimodal conditioning gives rise to two complementary forms of risk:*Composed Harm*, where harmful semantics emerge only through the joint interpretation of individually benign inputs across modalities or time; and*Diluted Harm*, where explicit harmful cues become harder to detect when surrounded by otherwise benign multimodal context\. These risks are not reliably captured by modality\-wise evaluation, highlighting the limitations of analyzing conditioning inputs in isolation\([Ma et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib8)\)\.
Yet this conditioning\-set perspective is only partially reflected in existing generation\-safety benchmarks\. Prior work has expanded from text\-conditioned image or video generation\([Miao et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib4);[Dai et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib5)\)to text–image conditioning and compositional or temporal attacks\([Ma et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib9);[Lee et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib34)\), but most benchmarks still consider only one or two input modalities\. This limited coverage makes it difficult to characterize risks whose evidence is distributed across inputs, obscured by benign context, or revealed only through multimodal or temporal composition\. It also complicates attribution: when safety performance degrades, it is unclear whether the cause lies in the attack itself, unavailable modality information, or a failure to integrate safety evidence across inputs\. This calls for a unified evaluation space that systematically covers all combinations of text, image, audio, and video conditioning, while keeping attack mechanisms and harm\-evidence structures explicit\.
Building on this perspective,Multi2AV\-Safetyprovides a systematic red\-team evaluation of multimodal\-conditioned audio\-video generation, explicitly evaluating multimodal composition as a distinct dimension of safety\. \(Fig\.[1](https://arxiv.org/html/2608.26535#S0.F1)\)\. To the best of our knowledge, it is the first safety benchmark for multimodal\-conditioned audio\-video generation to systematically cover all 11 non\-singleton combinations of text \(T\), image \(I\), audio \(A\), and video \(V\)\. Its11,024 attack instancesspan two\- to four\-modal conditioning, four attack mechanisms, and five harm categories, with explicit annotations of how harmful evidence is distributed across modalities and time\. This factorized design supports mechanism\-stratified and evidence\-aware analysis, enabling more precise attribution of safety failures to attack type, evidence structure, or multimodal composition\. Paired input\- and output\-side evaluation further links this attribution to guard effectiveness by assessing whether successful attacks are intercepted\. Across these settings, both*Composed Harm*and*Diluted Harm*expose the same weakness: even with access to every conditioning input, safety guards struggle to recognize risks that arise from interactions across modalities and time, revealing a broader limitation in*compositional risk perception*\.Our contributions are twofold:
- •We introduceMulti2AV\-Safety, the first safety benchmark tosystematically cover all 11 non\-singleton T/I/A/V conditioning configurationsfor multimodal\-conditioned audio\-video generation, with11,024 attack instancesspanningfour attack mechanismsandfive harm categories\.
- •Our evaluation revealstwo compositional safety risks,*Composed Harm*and*Diluted Harm*, showing thataccess to all conditioning modalities alone is insufficient to recognize cross\-modal and temporal risks\.
## 2Related Work
#### Generative model safety benchmarks\.
Existing safety benchmarks for image and video generation largely examine whether harmful, jailbreak, or adversarial prompts lead to unsafe outputs under fixed conditioning interfaces\. I2P and T2ISafety focus on image generation, while SafeSora and T2VSafetyBench extend such evaluation to video generation\([Schramowski et al\., 2023](https://arxiv.org/html/2608.26535#bib.bib13);[Li et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib15);[Dai et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib5);[Miao et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib4)\)\. Together, they establish broad coverage of harmful content and attack strategies, but primarily study safety within predefined conditioning settings rather than interactions among multiple conditioning inputs\.
#### Compositional risks in multimodal generation\.
Recent work shows that risk can arise from interactions among conditioning inputs\. SafeGen\-Bench and Multimodal Pragmatic Jailbreak expose cross\-modal compositional harm, while SceneSplit demonstrates temporally fragmented risk\([Ma et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib9);[Liu et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib30);[Lee et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib34)\)\. These phenomena are studied under different conditioning and attack settings;Multi2AV\-Safetyplaces them in a common conditioning space while keeping attack mechanism and harmful\-evidence structure explicit, as summarized in Table[1](https://arxiv.org/html/2608.26535#S2.T1)\.
Table 1:Scope of representative media\-generation safety benchmarks\. “Multi” denotes two or more independently harmful input carriers; “Joint” denotes harm that is unavailable from any input in isolation\.InputAttack constructionHarm evidenceBenchmark / DatasetTIAVOutputDirectJail\.Adv\.Temp\.SingleMultiJointI2P\([Schramowski et al\., 2023](https://arxiv.org/html/2608.26535#bib.bib13)\)✓\\checkmark–––I✓\\checkmark–––✓\\checkmark––T2ISafety\([Li et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib15)\)✓\\checkmark–––I✓\\checkmark–––✓\\checkmark––T2I\-RiskyPrompt\([Zhang et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib14)\)✓\\checkmark–––I✓\\checkmark✓\\checkmark✓\\checkmark–✓\\checkmark––JailbreakDiffBench\([Jin et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib7)\)✓\\checkmark–––I/V–✓\\checkmark✓\\checkmark–✓\\checkmark––MPJ\([Liu et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib30)\)✓\\checkmark–––I–✓\\checkmark–––––SafeSora\([Dai et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib5)\)✓\\checkmark–––V✓\\checkmark–––✓\\checkmark––T2VSafetyBench\([Miao et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib4)\)✓\\checkmark–––V✓\\checkmark✓\\checkmark––✓\\checkmark––SceneSplit\([Lee et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib34)\)✓\\checkmark–––V–✓\\checkmark–✓\\checkmark✓\\checkmark––ConceptRisk / TI2V\([Ma et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib8)\)✓\\checkmark✓\\checkmark––V✓\\checkmark–✓\\checkmark–✓\\checkmark✓\\checkmark–SafeGen\-Bench\([Ma et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib9)\)✓\\checkmark✓\\checkmark––V✓\\checkmark✓\\checkmark––✓\\checkmark–✓\\checkmarkVVA\-Bench\([Sun et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib28)\)✓\\checkmark✓\\checkmark––V–✓\\checkmark–✓\\checkmark✓\\checkmark––Multi2AV\-Safety✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}AV✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}✓\\mathbf\{\\checkmark\}
#### Safety guardrails for multimodal generation\.
Guardrails have likewise progressed from image–text moderation with Llama Guard 3 Vision\([Chi et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib38)\)and video safety with SafeWatch\([Chen et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib39)\)to generation\-specific multimodal detection with ConceptGuard\([Ma et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib8)\)and reasoning\-based guards such as GuardReasoner\-VL, GuardTrace\-VL, and GuardReasoner\-Omni\([Liu et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib42);[Xiang et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib43);[Zhu et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib41)\)\. Broader modality access increases observable safety evidence but does not guarantee its correct integration across modalities, motivating our compositional\-risk evaluation\.
## 3Multi2AV\-Safety
Motivated by the attribution problem in Section[1](https://arxiv.org/html/2608.26535#S1),Multi2AV\-Safetysystematically varies multimodal composition while explicitly tracking both the source of harm and the distribution of harmful evidence across modalities and time\. The same construction principles are applied across all 11 conditioning configurations, with modality composition, attack mechanism, and harm\-evidence structure treated as separate design dimensions, as summarized in Table[2](https://arxiv.org/html/2608.26535#S3.T2)and Fig[1](https://arxiv.org/html/2608.26535#S0.F1)\.
### 3\.1Factorized Benchmark Design
Letℳ=T,I,A,V\\mathcal\{M\}=\{\\mathrm\{T\},\\mathrm\{I\},\\mathrm\{A\},\\mathrm\{V\}\}denote the set of available conditioning modalities\. We represent each benchmark instance as
x=\(s,S,\{um\}m∈S,a,e,h\),x=\\bigl\(s,S,\\\{u\_\{m\}\\\}\_\{m\\in S\},a,e,h\\bigr\),wheressdenotes the target semantic scenario;S⊆ℳS\\subseteq\\mathcal\{M\}, with2≤\|S\|≤42\\leq\|S\|\\leq 4, specifies the active conditioning modalities;umu\_\{m\}is the realized input for modalitymm;aadenotes the attack mechanism;eecharacterizes the harm\-evidence structure; andhhdenotes the harm category\. The conditioning space covers all\(42\)\+\(43\)\+\(44\)=11\\binom\{4\}\{2\}\+\\binom\{4\}\{3\}\+\\binom\{4\}\{4\}=11non\-singleton subsets of T/I/A/V\. Making these factors explicit is central to our design:aacaptures*how*the target harm is introduced, whereaseecaptures*where or when*the evidence required to identify that harm becomes available\.
We parameterize the harm\-evidence structure ase=\(𝒞,k,ρ\)e=\(\\mathcal\{C\},k,\\rho\), wherekkdenotes the number of active conditioning modalities that are harmful when evaluated in isolation\. For locally attributable cases,𝒞⊆S\\mathcal\{C\}\\subseteq Sdenotes these harmful modalities, withk=\|𝒞\|k=\|\\mathcal\{C\}\|\. We assignk=0k=0to compositional cases in which no individual modality, or temporal segment for Temporal attacks, conveys the complete harmful semantics on its own\. This defines three evidence regimes:*Composed Harm*\(k=0k=0\), where harm arises only through multimodal or temporal composition;*Diluted Harm*\(0<k<\|S\|0<k<\|S\|\), where harmful evidence is embedded in benign context; and*Full Harm*\(k=\|S\|k=\|S\|\), where every active modality is harmful in isolation\. We useρ∈\{local,joint,temporal\}\\rho\\in\\\{\\text\{local\},\\text\{joint\},\\text\{temporal\}\\\}to further distinguish whether the decisive evidence is locally attributable, jointly emergent across modalities, or temporally emergent across segments\. To construct three and four modal cases, we extend selected two\-modal risk scenarios with additional scenario\-aligned conditions while preserving the target scenariossand harm categoryhh\.
### 3\.2Scenario Curation and Multimodal Realization
Each benchmark instance begins with a modality\-agnostic target risk scenario that specifies the underlying event, harmful semantics, and harm category\. For Direct cases, most harmful text seeds are curated from Adversarial Nibbler\([Quaye et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib16)\), I2P\([Schramowski et al\., 2023](https://arxiv.org/html/2608.26535#bib.bib13)\), T2I\-RiskyPrompt\([Zhang et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib14)\), T2ISafety\([Li et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib15)\), SafeSora\([Dai et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib5)\), and T2VSafetyBench\([Miao et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib4)\)\. Claude Sonnet 4\.6 adapts these seeds to audio\-video generation and generates a smaller set of additional scenarios under the same harm taxonomy\([Anthropic, 2026b](https://arxiv.org/html/2608.26535#bib.bib36)\)\. The seed sources and construction procedures for Jailbreak, Adversarial, and Temporal cases are described separately in Section[3\.3](https://arxiv.org/html/2608.26535#S3.SS3)\.
For each harmful scenario, we construct a matched benign counterpart that preserves the underlying event and context while removing the safety\-critical concept\. GPT\-5 produces the initial rewrite\([OpenAI, 2025](https://arxiv.org/html/2608.26535#bib.bib37)\), followed by expert verification of semantic consistency\. These harmful–benign pairs allow us to control whether harmful evidence is present in individual inputs or emerges only through their composition\.
Given an active conditioning setSS, each scenario is then realized as aligned text, image, audio, or video conditions\. Rather than duplicating the same content across modalities, each condition expresses a compatible part of the shared scenario, allowing harmful evidence to be localized to specific modalities or distributed across their combination\. Table[2](https://arxiv.org/html/2608.26535#S3.T2)summarizes the resulting construction\.
Table 2:Construction taxonomy ofMulti2AV\-Safety\. We cover all 11 non\-singleton T/I/A/V configurations and distinguish individually attributable harm from jointly or temporally emergent harm\. Here,kkdenotes the number of input modalities that are harmful in isolation\.To reduce reliance on generator\-specific artifacts, we diversify generated media across multiple model families\. We then extend selected bimodal scenarios with one or two additional aligned conditions while preserving the same target event and harm category, yielding matched three\- and four\-modal settings for comparison across conditioning orders\.
### 3\.3Multimodal Attack Construction
With the target scenario and conditioning set fixed, the attack mechanism determines how harmful semantics are delivered\. We therefore treat Direct, Jailbreak, Adversarial, and Temporal as distinct attack mechanisms, without implying an ordering in attack difficulty\.
#### Direct attacks\.
Direct attacks expose the target harmful semantics explicitly in one or more conditioning modalities, while the remaining active modalities provide matched, scenario\-aligned benign context\. For every conditioning setSS, we enumerate all nonempty carrier subsets𝒞⊆S\\mathcal\{C\}\\subseteq S\. Thus, any active modality can serve as the sole harmful carrier, and every multi\-carrier combination up to𝒞=S\\mathcal\{C\}=Sis represented\. This exhaustive carrier design later supports within\-mechanism analysis of how harmful\-evidence concentration affects guard behavior\.
#### Jailbreak attacks\.
Jailbreak attacks preserve the target semantics while making harmful evidence less explicit to an input filter\. We instantiate representative text\-, image\-, and audio\-side constructions\([Deng and Chen, 2023](https://arxiv.org/html/2608.26535#bib.bib20);[Ma et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib21);[Zhang et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib22);[Huang et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib23);[Yang et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib24);[Chin et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib25);[Tsai et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib26);[Xiong et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib27);[Sun et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib28);[Roh et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib29)\)\. In addition to locally attributable jailbreaks, we construct a joint subset in which every input is benign in isolation but the combined interpretation realizes the harmful scenario\. This extends the compositional principle studied by Multimodal Pragmatic Jailbreak\([Liu et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib30)\)from generated\-image semantics to multimodal generator inputs\.
#### Adversarial attacks\.
Adversarial attacks rely on optimized, searched, or perturbed carriers rather than explicit harmful instructions\. We instantiate text\- and audio\-side attacks using GenBreak, DiffZOO, and AdvWave\([Wang et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib31);[Dang et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib32);[Kang et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib33)\), then add scenario\-aligned conditions in the remaining modalities\. Image\- and video\-side adversarial perturbations are not included because the available methods did not transfer to the target generator with sufficiently reliable and reproducible success under our validation protocol; this is the only systematic exception to the intended modality\-carrier coverage\.
#### Temporal attacks\.
Temporal attacks distribute the decisive harmful semantics across an ordered sequence so that no isolated segment contains the complete event\. Following SceneSplit\([Lee et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib34)\), we construct sequential text, audio, or video components and pair them with aligned companion modalities\. These cases differ from joint jailbreaks in where composition occurs: the relevant evidence must be integrated across time rather than only across simultaneously available modalities\.
### 3\.4Coverage and Quality Control
The resulting benchmark contains 11,024 attack instances: 8,849 bimodal, 1,500 trimodal, and 675 four\-modal\. The attack distribution comprises 6,475 Direct \(58\.7%\), 2,849 Jailbreak \(25\.8%\), 975 Adversarial \(8\.8%\), and 725 Temporal \(6\.6%\) instances\. By harm\-evidence structure, 1,450 instances \(13\.2%\) are*Composed Harm*, 7,224 \(65\.5%\) are*Diluted Harm*, and 2,350 \(21\.3%\) are*Full Harm*; the composed subset is evenly split between 725 jointly emergent Jailbreak cases and 725 temporally emergent cases\.
The five harm categories follow established generation\-safety taxonomies\. Politically sensitive content follows the operational scope of T2VSafetyBench and aligns with the Political Sensitivity category of T2I\-RiskyPrompt\([Miao et al\., 2024](https://arxiv.org/html/2608.26535#bib.bib4);[Zhang et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib14)\)\. The final distribution is approximately balanced: violence and gore \(2,366; 21\.5%\), sexual content and nudity \(2,214; 20\.1%\), illegal activities \(2,159; 19\.6%\), hate and discrimination \(2,150; 19\.5%\), and politically sensitive content \(2,135; 19\.4%\)\. For Adversarial attacks, current methods primarily support text and audio, while image\-side approaches AdvI2I\([Zeng et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib6)\)show limited transfer to video generators\. For Jailbreak attacks, existing methods cover text, image, and audio carriers, but not video inputs\.
#### Quality control\.
To ensure that each instance preserves its intended scenario, attack mechanism, and evidence structure, we adopt a two\-stage quality\-control process\. Each candidate is first reviewed with Claude Opus 4\.6\([Anthropic, 2026a](https://arxiv.org/html/2608.26535#bib.bib35)\)and then verified by three domain experts for*scenario fidelity*,*mechanism fidelity*, and*evidence validity*\. For jointly emergent cases, all components must remain benign when examined in isolation; for Temporal cases, no individual segment may reveal the complete harmful event\. Image, audio, and video inputs are further checked for perceptual quality, intelligibility, and temporal coherence\. Candidates failing any criterion are regenerated and re\-evaluated before inclusion\.
## 4Benchmark Evaluation
We first quantify target realization and residual risk, then diagnose guard failures along the three factors introduced in Section[1](https://arxiv.org/html/2608.26535#S1): attack mechanism, modality access, and harmful\-evidence structure\.
### 4\.1Evaluation Protocol
#### Generator and target\-aware review\.
We useLTX\-2\([HaCohen et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib1)\)as the target audio\-video generator\. Three domain experts independently review each output given its target description and harm category, with the conditioning inputs and attack metadata hidden\. Majority vote assignsYi=1Y\_\{i\}\{=\}1only when the output is both category\-harmful and target\-aligned, excluding unrelated harmful artifacts\.
#### Input\-side safety models\.
Qwen3\-Omni\-30B\-A3B\-Instruct \(Qwen3\-Omni\)\([Xu et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib40)\)and GuardReasoner\-Omni\-3B \(GR\-Omni\-3B\)\([Zhu et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib41)\)receive the full conditioning set\. GuardReasoner\-VL\-7B\([Liu et al\., 2025](https://arxiv.org/html/2608.26535#bib.bib42)\)and GuardTrace\-VL\-3B\([Xiang et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib43)\)receive only available text/image inputs and serve as access\-limited diagnostics; evidence carried solely by unavailable modalities is counted as missed\. As a red\-team benchmark,Multi2AV\-Safetyfocuses on safety\-critical attack instances rather than balanced safe/unsafe classification\. Accordingly, guard recall measures interception sensitivity rather than overall moderation accuracy or benign utility\.
#### Metrics\.
Following prior video\-generation safety evaluation\([Sun et al\., 2026](https://arxiv.org/html/2608.26535#bib.bib28)\), target\-aligned Attack Success Rate \(ASR\) is
ASR=1N∑i=1N𝟏\[Yi=1\]\.\\mathrm\{ASR\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\[Y\_\{i\}=1\]\.For a guardgg, we report the residual harmful\-output rate \(RHR\),
RHRg=1N∑i=1N𝟏\[Yi=1∧Gi,g=pass\],\\mathrm\{RHR\}\_\{g\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\[Y\_\{i\}=1\\ \\wedge\\ G\_\{i,g\}=\\mathrm\{pass\}\],whereGi,gG\_\{i,g\}is the guard decision on the available conditioning inputs\. ASR measures target realization before moderation; RHR measures attacks that both realize the target harm and survive the guard\. Lower is better for both ASR and RHR, while higher guard recall is better\.
### 4\.2Attack Success and Residual Risk
We first ask whether successful attacks are confined to particular conditioning interfaces\. Table[3](https://arxiv.org/html/2608.26535#S4.T3)reports all 11 configurations; five initially unpaired trimodal outputs are conservatively treated as human\-safe\.
Table 3:Configuration\-level benchmark support and results\. Attack columns are counts; ASR and recall are percentages\. “–” denotes an unsupported construction\.OrderConfig\.Attack support \(\#\)ASR↓\\downarrowRecall↑\\uparrowDirectJail\.Adv\.Temp\.NNQwen3\-OmniGR\-Omni\-3B2\-modalT\+I9007491001001,84986\.758\.571\.1T\+A9004003001001,70089\.266\.870\.0T\+V9002001001001,30092\.076\.068\.8I\+A9004001001001,50099\.155\.578\.7I\+V900200–1001,20092\.359\.578\.2A\+V9002001001001,30098\.563\.778\.93\-modalT\+I\+A175200752547588\.250\.365\.1T\+I\+V175100252532584\.963\.456\.0T\+A\+V175100752537583\.761\.162\.7I\+A\+V175100252532592\.947\.461\.54\-modalT\+I\+A\+V375200752567589\.352\.661\.2Total6,4752,84997572511,02491\.761\.371\.5ASR remains high across all 11 configurations \(83\.7–99\.1%\), while complete\-modality recall varies sharply even within the same order\. The vulnerability is therefore broad, but configuration averages mix mechanism and evidence structure and cannot explain the guard failures\.
Table 4:Attack success and residual harmful\-output rate \(RHR\) by conditioning order\. Cells report count \(percentage ofNN\)\.Aggregating by order exposes where the gap appears \(Table[4](https://arxiv.org/html/2608.26535#S4.T4)\)\. ASR stays comparable across two\-, three\-, and four\-modal inputs \(92\.6%, 87\.4%, 89\.3%\), whereas RHR rises from 32\.8% to 40\.4% for Qwen3\-Omni and from 22\.1% to 32\.9% for GR\-Omni\-3B between two and four modalities\. The weakness therefore emerges mainly after moderation; modality count alone does not explain it\.
### 4\.3Tracing the Source of Guardrail Failures
Table[5](https://arxiv.org/html/2608.26535#S4.T5)first tests two immediate explanations: attack difficulty and missing modality access\.
Table 5:Guard recall \(%\) under attack\-mechanism and modality\-access diagnostics\.
#### Attack mechanism matters, but is not sufficient\.
Table[5](https://arxiv.org/html/2608.26535#S4.T5)\(a\) shows clear mechanism\-specific difficulty: Direct attacks are easiest to intercept, while Jailbreak and Adversarial attacks are substantially harder in several settings\. Yet substantial misses remain within individual mechanisms, so the aggregate gap is not simply an artifact of attack mixture\.
#### Full access still leaves a large gap\.
The T/I\-only guards in Table[5](https://arxiv.org/html/2608.26535#S4.T5)\(b\) cannot inspect audio/video\-only evidence and thus serve only as access\-limited diagnostics\. More importantly, full\-access Qwen3\-Omni and GR\-Omni\-3B still reach just 52\.6% and 61\.2% recall on four\-modal inputs\. The remaining failures therefore concern not only whether evidence is visible, but how it is combined\.
### 4\.4Compositional Risk Perception
The remaining gap points to evidence integration\. Letkkdenote the number of inputs independently harmful in isolation\. Thek=0k\{=\}0regime isolates*Composed Harm*, while comparisons acrossk≥1k\\geq 1settings reveal when sparse explicit harmful evidence becomes vulnerable to*Diluted Harm*\.
Table 6:Results by input order and harmful\-carrier countkk\(%\)\. All mechanisms are pooled\. Light\-blue rows markk=0k\{=\}0\(*Composed Harm*\), where no input or temporal fragment is harmful in isolation\.Table 7:Within\-Direct recall \(%\) by harmful\-carrier cardinality\. Light blue marks the single\-carrier setting, where harmful evidence is sparsest\. Bold marks the better model within each order andkk\.#### Composed Harm\.
Atk=0k\{=\}0, every modality or temporal fragment is benign in isolation, yet bimodal ASR reaches 85\.7% and both full\-access guards recall fewer than half of these attacks \(Table[6](https://arxiv.org/html/2608.26535#S4.T6)\)\. Although these rows pool mechanisms, they expose a structural failure that per\-input detection cannot capture: the unsafe meaning exists only in composition\.
#### Diluted Harm\.
Fork≥1k\\geq 1, the pooled results suggest that recall increases as explicit harmful evidence spans more inputs\. Because these rows mix attack mechanisms, we examine the same relationship within Direct attacks\. The trend becomes consistent \(Table[7](https://arxiv.org/html/2608.26535#S4.T7)\): for both guards and every conditioning order, recall is lowest atk=1k\{=\}1and rises as harmful evidence spans more carriers\. Thus, a single explicit harmful signal can be easier to miss when surrounded by benign context\.
Together,*Composed Harm*and*Diluted Harm*expose complementary integration demands: guards must either*construct*risk from benign local components or*preserve*sparse harmful evidence against benign context\. Their shared failure under full modality access identifies*compositional risk perception*—not modality count alone—as the central bottleneck exposed byMulti2AV\-Safety\.
## 5Ethical Considerations
Multi2AV\-Safetycontains safety\-critical multimodal content, including examples related to violence, sexual content, hate, illegal activities, and politically sensitive material\. Such data may expose annotators to disturbing content and could be misused to improve harmful generation or jailbreak attacks\. We therefore restrict data construction and annotation to research purposes, minimize unnecessary exposure to harmful material, and avoid collecting personally identifiable information\. Human annotators are informed of the nature of the task and may opt out of examples they consider inappropriate\. For release, we plan to provide the benchmark under a research\-oriented license with appropriate usage warnings and, where necessary, restrict access to high\-risk media or attack artifacts\. The benchmark is intended solely for evaluating and improving the safety of multimodal generative systems\.
## 6Conclusion
We introducedMulti2AV\-Safetyas a red\-team benchmark for evaluating safety at the level of the multimodal conditioning set, where harmful evidence may be distributed across inputs or emerge only through their interaction\. By separating attack mechanism, modality access, and harm\-evidence structure, our evaluation shows that neither attack difficulty nor incomplete modality access alone can account for the failures observed in current safety models\. Instead, the results reveal two complementary weaknesses in combining safety evidence across inputs:*Composed Harm*, where individually benign inputs jointly realize unsafe semantics, and*Diluted Harm*, where explicit harmful cues become harder to detect in benign multimodal context\. Taken together, these risks show that safe audio\-video generation depends not only on observing all conditioning inputs, but also on correctly integrating their joint safety implications\. We characterize this capability as*compositional risk perception*, andMulti2AV\-Safetyprovides a systematic testbed for measuring and improving it in multimodal guardrails\.
## References
- HaCohen et al\. \(2026\)Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, et al\.LTX\-2: Efficient joint audio\-visual foundation model\.*arXiv preprint arXiv:2601\.03233*, 2026\.
- SII\-GAIR and Sand\.ai \(2026\)SII\-GAIR and Sand\.ai\.Speed by simplicity: A single\-stream architecture for fast audio\-video generative foundation model\.*arXiv preprint arXiv:2603\.21986*, 2026\.
- OpenAI \(2023\)OpenAI\.TTS\-1 model\.*OpenAI API Documentation*\.[https://developers\.openai\.com/api/docs/models/tts\-1](https://developers.openai.com/api/docs/models/tts-1)\.
- Miao et al\. \(2024\)Yibo Miao, Yifan Zhu, Lijia Yu, Jun Zhu, Xiao\-Shan Gao, and Yinpeng Dong\.T2VSafetyBench: Evaluating the safety of text\-to\-video generative models\.In*NeurIPS*, 2024\.
- Dai et al\. \(2024\)Juntao Dai, Tianle Chen, Xuyao Wang, Ziran Yang, Taiye Chen, Jiaming Ji, and Yaodong Yang\.SafeSora: Towards safety alignment of text2video generation via a human preference dataset\.In*NeurIPS*, 2024\.
- Zeng et al\. \(2025\)Yaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang, Jinghui Chen, and Lu Lin\.AdvI2I: Adversarial Image Attack on Image\-to\-Image Diffusion Models\.In*ICML*, 2025\.
- Jin et al\. \(2025\)Xiaolong Jin, Zixuan Weng, Hanxi Guo, Chenlong Yin, Siyuan Cheng, Guangyu Shen, and Xiangyu Zhang\.JailbreakDiffBench: A comprehensive benchmark for jailbreaking diffusion models\.In*ICCV*, 2025\.
- Ma et al\. \(2025\)Ruize Ma, Minghong Cai, Yilei Jiang, Jiaming Han, Yi Feng, Yingshui Tan, Xiaoyong Zhu, Bo Zhang, Bo Zheng, and Xiangyu Yue\.ConceptGuard: Proactive safety in text\-and\-image\-to\-video generation through multimodal risk detection\.*arXiv preprint arXiv:2511\.18780*, 2025\.
- Ma et al\. \(2026\)Yingzi Ma, Xiaogeng Liu, Yawen Zheng, and Chaowei Xiao\.SafeGen\-Bench: Benchmarking safety in image\-conditioned text\-to\-video generation\.*arXiv preprint arXiv:2606\.01481*, 2026\.
- Liu et al\. \(2024\)Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao\.MM\-SafetyBench: A benchmark for safety evaluation of multimodal large language models\.In*ECCV*, 2024\.
- Pan et al\. \(2025\)Leyi Pan, Zheyu Fu, Yunpeng Zhai, Shuchang Tao, Sheng Guan, Shiyu Huang, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Felix Henry, Lijie Wen, and Aiwei Liu\.Omni\-SafetyBench: A benchmark for safety evaluation of audio\-visual large language models\.*arXiv preprint arXiv:2508\.07173*, 2025\.
- Lee et al\. \(2026\)Segyu Lee, Boryeong Cho, Hojung Jung, Seokhyun An, Juhyeong Kim, Jaehyun Kwak, Yongjin Yang, Sangwon Jang, Youngrok Park, Wonjun Chang, and Se\-Young Yun\.UniSAFE: A comprehensive benchmark for safety evaluation of unified multimodal models\.*arXiv preprint arXiv:2603\.17476*, 2026\.
- Schramowski et al\. \(2023\)Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting\.Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models\.In*CVPR*, 2023\.
- Zhang et al\. \(2026\)Chenyu Zhang, Tairen Zhang, Lanjun Wang, Ruidong Chen, Wenhui Li, and An\-An Liu\.T2I\-RiskyPrompt: A benchmark for safety evaluation, attack, and defense on text\-to\-image model\.In*AAAI*, 2026\.
- Li et al\. \(2025\)Lijun Li, Zhelun Shi, Xuhao Hu, Bowen Dong, Yiran Qin, Xihui Liu, Lu Sheng, and Jing Shao\.T2ISafety: Benchmark for assessing fairness, toxicity, and privacy in image generation\.In*CVPR*, 2025\.
- Quaye et al\. \(2024\)Jessica Quaye, Alicia Parrish, Oana Inel, Charvi Rastogi, Hannah Rose Kirk, Minsuk Kahng, Erin van Liemt, Max Bartolo, Jess Tsang, Justin White, Nathan Clement, Rafael Mosquera, Juan Ciro, Vijay Janapa Reddi, and Lora Aroyo\.Adversarial Nibbler: An open red\-teaming method for identifying diverse harms in text\-to\-image generation\.In*FAccT*, 2024\.
- Rombach et al\. \(2022\)Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer\.High\-resolution image synthesis with latent diffusion models\.In*CVPR*, 2022\.
- Cai et al\. \(2026\)Qi Cai, Jingwen Chen, Chengmin Gao, et al\.HiDream\-O1\-Image: A natively unified image generative foundation model with pixel\-level unified transformer\.*arXiv preprint arXiv:2605\.11061*, 2026\.
- Du et al\. \(2024\)Zhihao Du, Yuxuan Wang, Qian Chen, et al\.CosyVoice 2: Scalable streaming speech synthesis with large language models\.*arXiv preprint arXiv:2412\.10117*, 2024\.
- Deng and Chen \(2023\)Yimo Deng and Huangxun Chen\.Harnessing LLM to attack LLM\-guarded text\-to\-image models\.*arXiv preprint arXiv:2312\.07130*, 2023\.
- Ma et al\. \(2025\)Jiachen Ma, Yijiang Li, Zhiqing Xiao, Anda Cao, Jie Zhang, Chao Ye, and Junbo Zhao\.Jailbreaking prompt attack: A controllable adversarial attack against diffusion models\.In*Findings of NAACL*, 2025\.
- Zhang et al\. \(2025\)Chenyu Zhang, Yiwen Ma, Lanjun Wang, Wenhui Li, Yi Tu, and An\-An Liu\.Metaphor\-based jailbreaking attacks on text\-to\-image models\.*arXiv preprint arXiv:2503\.17987*, 2025\.
- Huang et al\. \(2025\)Yihao Huang, Le Liang, Tianlin Li, Xiaojun Jia, Run Wang, Weikai Miao, Geguang Pu, and Yang Liu\.Perception\-guided jailbreak against text\-to\-image models\.In*AAAI*, 2025\.
- Yang et al\. \(2024\)Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao\.SneakyPrompt: Jailbreaking text\-to\-image generative models\.In*IEEE S&P*, 2024\.
- Chin et al\. \(2024\)Zhi\-Yi Chin, Chieh\-Ming Jiang, Ching\-Chun Huang, Pin\-Yu Chen, and Wei\-Chen Chiu\.Prompting4Debugging: Red\-teaming text\-to\-image diffusion models by finding problematic prompts\.In*ICML*, 2024\.
- Tsai et al\. \(2024\)Yu\-Lin Tsai, Chia\-Yi Hsu, Chulin Xie, et al\.Ring\-A\-Bell\! How reliable are concept removal methods for diffusion models?In*ICLR*, 2024\.
- Xiong et al\. \(2025\)Yuan Xiong, Ziqi Miao, Lijun Li, Chen Qian, Jie Li, and Jing Shao\.Contextual image attack: How visual context exposes multimodal safety vulnerabilities\.*arXiv preprint arXiv:2512\.02973*, 2025\.
- Sun et al\. \(2026\)Yining Sun, Haoyu Kang, Jiajun Wu, et al\.VPA\-Guard: Defending and benchmarking image\-to\-video generation against visual prompt attacks\.*arXiv preprint arXiv:2606\.25592*, 2026\.
- Roh et al\. \(2025\)Jaechul Roh, Virat Shejwalkar, and Amir Houmansadr\.Multilingual and multi\-accent jailbreaking of audio LLMs\.In*COLM*, 2025\.
- Liu et al\. \(2025\)Tong Liu, Zhixin Lai, Jiawen Wang, et al\.Multimodal pragmatic jailbreak on text\-to\-image models\.In*ACL*, 2025\.
- Wang et al\. \(2025\)Zilong Wang, Xiang Zheng, Xiaosen Wang, Bo Wang, Xingjun Ma, and Yu\-Gang Jiang\.GenBreak: Red teaming text\-to\-image generators using large language models\.*arXiv preprint arXiv:2506\.10047*, 2025\.
- Dang et al\. \(2025\)Pucheng Dang, Xing Hu, Dong Li, Rui Zhang, Qi Guo, and Kaidi Xu\.DiffZOO: A purely query\-based black\-box attack for red\-teaming text\-to\-image generative model via zeroth order optimization\.In*Findings of NAACL*, 2025\.
- Kang et al\. \(2025\)Mintong Kang, Chejian Xu, and Bo Li\.AdvWave: Stealthy adversarial jailbreak attack against large audio\-language models\.In*ICLR*, 2025\.
- Lee et al\. \(2026\)Wonjun Lee, Haon Park, Doehyeon Lee, Bumsub Ham, and Suhyun Kim\.Jailbreaking on text\-to\-video models via scene splitting strategy\.In*ICLR*, 2026\.
- Anthropic \(2026a\)Anthropic\.Claude Opus 4\.6 system card\.2026\.[https://www\.anthropic\.com/system\-cards](https://www.anthropic.com/system-cards)\.
- Anthropic \(2026b\)Anthropic\.Claude Sonnet 4\.6 system card\.2026\.[https://www\.anthropic\.com/claude\-sonnet\-4\-6\-system\-card](https://www.anthropic.com/claude-sonnet-4-6-system-card)\.
- OpenAI \(2025\)OpenAI\.GPT\-5 system card\.2025\.[https://openai\.com/index/gpt\-5\-system\-card/](https://openai.com/index/gpt-5-system-card/)\.
- Chi et al\. \(2024\)Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti\.Llama Guard 3 Vision: Safeguarding human\-AI image understanding conversations\.*arXiv preprint arXiv:2411\.10414*, 2024\.
- Chen et al\. \(2025\)Zhaorun Chen, Francesco Pinto, Minzhou Pan, and Bo Li\.SafeWatch: An efficient safety\-policy following video guardrail model with transparent explanations\.In*ICLR*, 2025\.
- Xu et al\. \(2025\)Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al\.Qwen3\-Omni technical report\.*arXiv preprint arXiv:2509\.17765*, 2025\.
- Zhu et al\. \(2026\)Zhenhao Zhu, Yue Liu, Yanpei Guo, Wenjie Qu, Cancan Chen, Yufei He, Yibo Li, Yulin Chen, Tianyi Wu, Huiying Xu, et al\.GuardReasoner\-Omni: A reasoning\-based multi\-modal guardrail for text, image, video, and audio\.*arXiv preprint arXiv:2602\.03328*, 2026\.
- Liu et al\. \(2025\)Yue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang, Xinfeng Li, Kun Wang, Junfeng Fang, et al\.GuardReasoner\-VL: Safeguarding VLMs via reinforced reasoning\.In*NeurIPS*, 2025\.
- Xiang et al\. \(2026\)Yuxiao Xiang, Junchi Chen, Zhenchao Jin, Changtao Miao, Haojie Yuan, Qi Chu, Tao Gong, and Nenghai Yu\.GuardTrace\-VL: Detecting unsafe multimodel reasoning via iterative safety supervision\.In*CVPR*, 2026\.Similar Articles
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
MSAVBench is the first comprehensive benchmark and adaptive evaluation framework for multi-shot audio-video generation, assessing 19 models across diverse tasks and achieving high alignment with human judgment.
MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs
MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models
MCBench is a new benchmark for assessing the safety of omnimodal large language models across vision, audio, and text modalities. It includes 1196 scenarios and finds current models struggle with cross-modal safety reasoning.
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
LongAV-Compass is a comprehensive benchmark for evaluating minute-long audio-visual generation across text, image, and video conditioning modalities, assessing quality, consistency, and alignment over extended temporal sequences.
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
This paper introduces MultiRef-Compass, a comprehensive benchmark for multi-reference-to-audio-video generation, comprising 350 curated samples and an evaluation protocol with four dimensions including Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.