Tag
This paper introduces an energy landscape framework to analyze why jailbreak attacks succeed in diffusion-based language models, deriving three complementary, training-free detection signals for safety alignment.
The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.
This paper introduces Multi-Turn Certified Robustness (MTCR), a framework for certifying the safety of large language models against multi-turn jailbreak attacks by providing tighter bounds through compositional methods and safety persistence.
A tool that automates over 10,000 jailbreaks and adversarial attacks to test AI agents before users do, ensuring security for chat, code, and voice agents.
STAR-Teaming introduces a multiplex-network-driven multi-agent framework that automates LLM red-teaming, achieving higher attack success rates with lower compute by organizing attack strategies into interpretable semantic communities.