jailbreaks

Tag

Cards List
#jailbreaks

Evaluating Language Model Safety Across Long Adversarial Conversations

arXiv cs.CL ↗ · yesterday Cached

This paper evaluates whether open-weight instruction-tuned language models maintain safe behavior during long adversarial conversations, finding that safe-response rates drop from 85-100% on the first turn to 15-44% by depth 101, demonstrating that strong single-turn safety does not persist across sustained interaction.

0 favorites 0 likes
#jailbreaks

The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models

arXiv cs.CL ↗ · yesterday Cached

This paper uses Centered Kernel Alignment (CKA) and activation patching across 17 instruction-tuned models to show that system prompts are 'seen' at every layer but only deeply restructure representations for persona/formatting instructions — safety prompts barely alter computation, providing a mechanistic explanation for why system-prompt-based safety remains jailbreakable.

0 favorites 0 likes
#jailbreaks

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

arXiv cs.LG ↗ · 2026-08-11 Cached

A survey paper systematically analyzing the evolving safety landscape of multi-modal large language models, covering emerging threats such as adversarial attacks, data poisoning, jailbreaks, and hallucinations, and reviewing updated safety strategies.

0 favorites 0 likes
#jailbreaks

The Trump White House Is Over Anthropic CEO Dario Amodei

Wired ↗ · 2026-06-24 Cached

The Trump administration has grown frustrated with Anthropic CEO Dario Amodei, preferring to deal with cofounder Tom Brown regarding the re-release of the Claude Fable 5 AI model, as export controls remain in place due to jailbreak concerns.

0 favorites 0 likes
#jailbreaks

TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization

Hugging Face Daily Papers ↗ · 2026-06-22 Cached

TROPT is an open-source framework that unifies discrete text-trigger optimization, standardizing development and execution across domains like LLM jailbreaking and model interpretability. It includes over 15 optimizers and 30 recipes, lowering barriers for adoption and advancement.

0 favorites 0 likes
← Back to home

Submit Feedback