jailbreak-attacks

Tag

Cards List
#jailbreak-attacks

Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

arXiv cs.AI ↗ · 6d ago Cached

This paper introduces an energy landscape framework to analyze why jailbreak attacks succeed in diffusion-based language models, deriving three complementary, training-free detection signals for safety alignment.

0 favorites 0 likes
#jailbreak-attacks

@LeeLeepenkman: Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models https://pap…

X AI KOLs Timeline ↗ · 2026-08-26 Cached

The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.

0 favorites 0 likes
#jailbreak-attacks

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

arXiv cs.AI ↗ · 2026-08-24 Cached

This paper introduces Multi-Turn Certified Robustness (MTCR), a framework for certifying the safety of large language models against multi-turn jailbreak attacks by providing tighter bounds through compositional methods and safety persistence.

0 favorites 0 likes
#jailbreak-attacks

@svpino: This will let you break your agent before your users do. This works with any agent, including chat, code, and voice age…

X AI KOLs Timeline ↗ · 2026-08-19 Cached

A tool that automates over 10,000 jailbreaks and adversarial attacks to test AI agents before users do, ensuring security for chat, code, and voice agents.

0 favorites 0 likes
#jailbreak-attacks

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

arXiv cs.CL ↗ · 2026-04-22 Cached

STAR-Teaming introduces a multiplex-network-driven multi-agent framework that automates LLM red-teaming, achieving higher attack success rates with lower compute by organizing attack strategies into interpretable semantic communities.

0 favorites 0 likes
← Back to home

Submit Feedback