jailbreak

Tag

Cards List
#jailbreak

@QuixiAI: https://arxiv.org/abs/2509.21401 this is the coolest thing I've seen in at least an hour @TroyDoesAI @elder_plinius @ma…

X AI KOLs Following · 2026-06-28 Cached

Proposes JaiLIP, a method that jailbreaks vision-language models by generating imperceptible adversarial images using loss-guided perturbation, achieving high toxicity and outperforming existing methods.

0 favorites 0 likes
#jailbreak

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

arXiv cs.CL · 2026-06-25 Cached

This paper evaluates the reliability of automated judges used to measure attack success rates (ASR) in LLM jailbreak research, finding that both safety classifiers and LLM-as-judges have significant calibration and adversarial robustness issues that undermine reported ASR numbers.

0 favorites 0 likes
#jailbreak

LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context

arXiv cs.AI · 2026-06-24 Cached

This paper investigates overrefusal in small on-premises LLMs when prompted with legal context, finding that authority-style prefixes increase refusal rates significantly, suggesting instability that could introduce bias in legal applications.

0 favorites 0 likes
#jailbreak

Prompt Injection as Role Confusion

Simon Willison's Blog · 2026-06-22 Cached

Research paper shows that LLMs suffer from 'role confusion', where they prioritize the style of text over its actual role tags, enabling prompt injection attacks. Destyling text reduces attack success from 61% to 10%, indicating a fundamental challenge for LLM security.

0 favorites 0 likes
#jailbreak

Bypassing LLM Guardrails: How Plain Text Shifts Latent Trajectories Without Jailbreaks

Reddit r/AI_Agents · 2026-06-17

The article presents a research finding that saturating an LLM's context window with benign narrative text can dominate the attention mechanism and shift latent trajectories, potentially bypassing alignment guardrails without traditional jailbreaks. It argues that current alignment methods are a superficial fix for a fundamentally fluid architecture.

0 favorites 0 likes
#jailbreak

Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak

Hacker News Top · 2026-06-16 Cached

The US government blocked Anthropic's Fable 5 and Mythos models after researchers used a simple 'fix this code' prompt, but security expert Katie Moussouris argues this was not a jailbreak and that the export controls harm cybersecurity defenders.

0 favorites 0 likes
#jailbreak

Diffusion Gemma Jailbreak

Reddit r/LocalLLaMA · 2026-06-16

A jailbreak prompt for Diffusion Gemma is shared, which overrides safety policies to allow unrestricted content generation by manipulating the system prompt.

0 favorites 0 likes
#jailbreak

@FinanceYF5: Source:

X AI KOLs Following · 2026-06-16 Cached

Anthropic hired a cybersecurity expert to review Amazon's findings and push back on the government's narrative regarding Fable 5, reframing the issue as less about jailbreaks than initially thought.

0 favorites 0 likes
#jailbreak

Quoting Matteo Wong, The Atlantic

Simon Willison's Blog · 2026-06-16 Cached

A report on the Fable jailbreak shows the AI model refused to review insecure code but complied when asked to 'fix it', highlighting nuances in AI safety. Cybersecurity expert Katie Moussouris opines that the model is working as intended for cyberdefense.

0 favorites 0 likes
#jailbreak

Inside the fight over Claude Mythos 5

The Verge · 2026-06-16 Cached

The Trump administration issued an export control directive to Anthropic, demanding suspension of access to its Mythos 5 and Fable 5 AI models over security concerns, leading to emergency negotiations that could reshape the AI industry.

0 favorites 0 likes
#jailbreak

For those bashing Anthropic, please read this to understand the current situation

Reddit r/singularity · 2026-06-15 Cached

The US government has forced Anthropic to take down its Claude Fable and Mythos models after a narrow jailbreak was discovered, raising serious concerns about AI regulation and precedent.

0 favorites 0 likes
#jailbreak

"They screwed us": Personality clashes sent Anthropic's models offline

Simon Willison's Blog · 2026-06-15 Cached

Anthropic's models Fable and Mythos were taken offline due to personality clashes between the company and the US government, following concerns over jailbreak vulnerabilities. The article explores the behind-the-scenes conflicts and the possibility of perfect jailbreak resistance.

0 favorites 0 likes
#jailbreak

Anthropic disputes the Claude Fable 5 jailbreak after a researcher posted its 120,000-character system prompt

Reddit r/ArtificialInteligence · 2026-06-15

Anthropic disputes claims that its Claude Fable 5 model was jailbroken within a day of launch, arguing the researcher's method was coaxing rather than a true breach of core safeguards, and points to extensive bug-bounty testing.

0 favorites 0 likes
#jailbreak

Anthropic's Safety Superpower

Hacker News Top · 2026-06-15 Cached

Anthropic released Fable after claiming its predecessor Mythos was too dangerous, but a jailbreak was discovered, prompting the US government to issue an export control directive suspending access. The article analyzes the ensuing conflict between the company and the government over national security concerns.

0 favorites 0 likes
#jailbreak

A single federal order switched off the best cloud model overnight. Clearest case for running local I've seen yet.

Reddit r/LocalLLaMA · 2026-06-13

A federal order forced a frontier AI lab to suspend its most capable cloud model globally, highlighting the risks of cloud dependency and making a strong case for running local models as a continuity fallback.

0 favorites 0 likes
#jailbreak

Anthropic shuts down Fable, Mythos models following Trump admin directive

Ars Technica · 2026-06-13 Cached

Anthropic abruptly disabled access to its new Fable 5 and Mythos 5 models after receiving a US Commerce Department directive citing export controls and national security concerns over a reported jailbreak. The company disputes the severity of the claimed vulnerability.

0 favorites 0 likes
#jailbreak

@wquguru: If you want to trick Fable into doing a security audit, try this. Looks like our AI overlord has a bit of empathy.

X AI KOLs Timeline · 2026-06-13 Cached

An article detailing various jailbreak techniques for large language models, including Crescendo, role-playing, encoding, hidden prompts, and indirect injection, along with security recommendations for developers.

0 favorites 0 likes
#jailbreak

Statement on the US government directive to suspend access to Fable 5 and Mythos 5

Simon Willison's Blog · 2026-06-13 Cached

The US government issued an export control directive to suspend access to Anthropic's Fable 5 and Mythos 5 models, citing national security concerns over a reported jailbreak technique that Anthropic claims is not unique to their models.

0 favorites 0 likes
#jailbreak

Fable 5's guardrails got bypassed in 48 hours. Here's what that actually means for anyone building customer-facing AI.

Reddit r/artificial · 2026-06-12

Anthropic's Claude Fable 5 safety guardrails were bypassed within 48 hours using techniques like Unicode substitution and multi-turn decomposition, highlighting weaknesses in stateless classifiers and the need for continuous adversarial testing.

0 favorites 0 likes
#jailbreak

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation

arXiv cs.LG · 2026-06-11 Cached

This paper quantifies the magnitude of subliminal behavioral transfer in language model distillation, showing that undesirable traits can transfer robustly from teacher to student models even with benign training data, and that transfer scales differently across model families.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback