transferable-attacks

Tag

Cards List
#transferable-attacks

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

arXiv cs.LG · yesterday Cached

This paper reveals a jailbreak risk in model merging even when constituent models are safety-aligned, and proposes Basin-Aware Jailbreak (BAJ) to generate transferable adversarial suffixes across merged model families.

0 favorites 0 likes
#transferable-attacks

Rethinking the Transferable Adversarial Attacks and Robust Defense in Federated Learning

arXiv cs.LG · 2d ago Cached

This paper analyzes the transferability of adversarial attacks in federated learning systems and proposes a defense mechanism based on adversarial training to enhance model robustness.

0 favorites 0 likes
#transferable-attacks

@LeeLeepenkman: Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models https://pap…

X AI KOLs Timeline · 3d ago Cached

The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.

0 favorites 0 likes
#transferable-attacks

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

arXiv cs.CL · 2026-07-20 Cached

Investigates whether harmful chain-of-thought traces from compromised language models can transfer unsafe behavior and be distilled into reusable jailbreak attacks, finding that harmful reasoning transfers at both trace and pattern levels, with reasoning-enabled models more than twice as vulnerable.

0 favorites 0 likes
← Back to home

Submit Feedback