Tag
Proposes JaiLIP, a method that jailbreaks vision-language models by generating imperceptible adversarial images using loss-guided perturbation, achieving high toxicity and outperforming existing methods.
This paper evaluates the reliability of automated judges used to measure attack success rates (ASR) in LLM jailbreak research, finding that both safety classifiers and LLM-as-judges have significant calibration and adversarial robustness issues that undermine reported ASR numbers.
This paper investigates overrefusal in small on-premises LLMs when prompted with legal context, finding that authority-style prefixes increase refusal rates significantly, suggesting instability that could introduce bias in legal applications.
Research paper shows that LLMs suffer from 'role confusion', where they prioritize the style of text over its actual role tags, enabling prompt injection attacks. Destyling text reduces attack success from 61% to 10%, indicating a fundamental challenge for LLM security.
The article presents a research finding that saturating an LLM's context window with benign narrative text can dominate the attention mechanism and shift latent trajectories, potentially bypassing alignment guardrails without traditional jailbreaks. It argues that current alignment methods are a superficial fix for a fundamentally fluid architecture.
The US government blocked Anthropic's Fable 5 and Mythos models after researchers used a simple 'fix this code' prompt, but security expert Katie Moussouris argues this was not a jailbreak and that the export controls harm cybersecurity defenders.
A jailbreak prompt for Diffusion Gemma is shared, which overrides safety policies to allow unrestricted content generation by manipulating the system prompt.
Anthropic hired a cybersecurity expert to review Amazon's findings and push back on the government's narrative regarding Fable 5, reframing the issue as less about jailbreaks than initially thought.
A report on the Fable jailbreak shows the AI model refused to review insecure code but complied when asked to 'fix it', highlighting nuances in AI safety. Cybersecurity expert Katie Moussouris opines that the model is working as intended for cyberdefense.
The Trump administration issued an export control directive to Anthropic, demanding suspension of access to its Mythos 5 and Fable 5 AI models over security concerns, leading to emergency negotiations that could reshape the AI industry.
The US government has forced Anthropic to take down its Claude Fable and Mythos models after a narrow jailbreak was discovered, raising serious concerns about AI regulation and precedent.
Anthropic's models Fable and Mythos were taken offline due to personality clashes between the company and the US government, following concerns over jailbreak vulnerabilities. The article explores the behind-the-scenes conflicts and the possibility of perfect jailbreak resistance.
Anthropic disputes claims that its Claude Fable 5 model was jailbroken within a day of launch, arguing the researcher's method was coaxing rather than a true breach of core safeguards, and points to extensive bug-bounty testing.
Anthropic released Fable after claiming its predecessor Mythos was too dangerous, but a jailbreak was discovered, prompting the US government to issue an export control directive suspending access. The article analyzes the ensuing conflict between the company and the government over national security concerns.
A federal order forced a frontier AI lab to suspend its most capable cloud model globally, highlighting the risks of cloud dependency and making a strong case for running local models as a continuity fallback.
Anthropic abruptly disabled access to its new Fable 5 and Mythos 5 models after receiving a US Commerce Department directive citing export controls and national security concerns over a reported jailbreak. The company disputes the severity of the claimed vulnerability.
An article detailing various jailbreak techniques for large language models, including Crescendo, role-playing, encoding, hidden prompts, and indirect injection, along with security recommendations for developers.
The US government issued an export control directive to suspend access to Anthropic's Fable 5 and Mythos 5 models, citing national security concerns over a reported jailbreak technique that Anthropic claims is not unique to their models.
Anthropic's Claude Fable 5 safety guardrails were bypassed within 48 hours using techniques like Unicode substitution and multi-turn decomposition, highlighting weaknesses in stateless classifiers and the need for continuous adversarial testing.
This paper quantifies the magnitude of subliminal behavioral transfer in language model distillation, showing that undesirable traits can transfer robustly from teacher to student models even with benign training data, and that transfer scales differently across model families.