@LeeLeepenkman: Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models https://pap…

X AI KOLs Timeline Papers

Summary

The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.

Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models https://t.co/NljjVIHrnS https://t.co/NGrf1icnF1
Original Article
View Cached Full Text

Cached at: 08/26/26, 07:24 AM

Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

https://t.co/NljjVIHrnS https://t.co/NGrf1icnF1


Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

Source: https://papers.app.nz/view/paper?id=9107387 Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable perturbations in a text-anchored semantic space. TA-SPA integrates Text-Anchored Semantic Factorization (TASF), which encourages the separation of cross-modal semantic factors from modality-specific residuals, with Semantic-Preserving Augmentation (SPA), which diversifies harmful target anchors while preserving semantic consistency. Experiments show strong attack effectiveness and transfer to commercial MLLMs, with competitive performance under representative defenses. Additional controls and probing support the intended factorization without implying perfect disentanglement, motivating representation-level safety alignment beyond input-level filtering.

Similar Articles