Tag
This paper reveals a jailbreak risk in model merging even when constituent models are safety-aligned, and proposes Basin-Aware Jailbreak (BAJ) to generate transferable adversarial suffixes across merged model families.
This paper analyzes the transferability of adversarial attacks in federated learning systems and proposes a defense mechanism based on adversarial training to enhance model robustness.
The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.
Investigates whether harmful chain-of-thought traces from compromised language models can transfer unsafe behavior and be distilled into reusable jailbreak attacks, finding that harmful reasoning transfers at both trace and pattern levels, with reasoning-enabled models more than twice as vulnerable.