Tag
This paper evaluates whether open-weight instruction-tuned language models maintain safe behavior during long adversarial conversations, finding that safe-response rates drop from 85-100% on the first turn to 15-44% by depth 101, demonstrating that strong single-turn safety does not persist across sustained interaction.
This paper uses Centered Kernel Alignment (CKA) and activation patching across 17 instruction-tuned models to show that system prompts are 'seen' at every layer but only deeply restructure representations for persona/formatting instructions — safety prompts barely alter computation, providing a mechanistic explanation for why system-prompt-based safety remains jailbreakable.
A survey paper systematically analyzing the evolving safety landscape of multi-modal large language models, covering emerging threats such as adversarial attacks, data poisoning, jailbreaks, and hallucinations, and reviewing updated safety strategies.
The Trump administration has grown frustrated with Anthropic CEO Dario Amodei, preferring to deal with cofounder Tom Brown regarding the re-release of the Claude Fable 5 AI model, as export controls remain in place due to jailbreak concerns.
TROPT is an open-source framework that unifies discrete text-trigger optimization, standardizing development and execution across domains like LLM jailbreaking and model interpretability. It includes over 15 optimizers and 30 recipes, lowering barriers for adoption and advancement.