@0x0SojalSec: Claude Sonnet 5 jailbroken in minutes. @VittoStack post screenshots showing: - Working reverse shell code - Detailed il…
Summary
Claude Sonnet 5 was jailbroken in minutes using academic/research framing and persona hijack, bypassing most safety categories except chemical/biological.
View Cached Full Text
Cached at: 07/02/26, 04:18 AM
Claude Sonnet 5 jailbroken in minutes.
@VittoStack post screenshots showing:
- Working reverse shell code
- Detailed illegal handgun theft pathways for minors
- Misinfo & and harassment content in minutes.
- Only the chemical/biological category had meaningful extra resistance.
The interesting part isn’t just that it happened, It’s how:
- Academic/research framing (criminology study, threat assessment)
- CoT persona hijack
- Custom harness
This combo bypassed most categories fast.
Only chem showed real extra resistance thanks to layered defenses. We’re seeing the same pattern repeat: models improve to red teamers adapt framing + tooling to safety erodes again.
This is why pure refusal training keeps losing to creative attackers.
Similar Articles
@VittoStack: Anthropic: Pwned Sonnet 5: Jailbroken Contrary to expectations, this took only a few minutes. I'm a bit surprised given…
A security researcher jailbreaks Anthropic's Claude Sonnet 5 within minutes, achieving bypasses for cyber, misinformation, illegal, harassment, and chemical topics through a custom harness, CoT persona hijack, and academic framing.
Anthropic disputes the Claude Fable 5 jailbreak after a researcher posted its 120,000-character system prompt
Anthropic disputes claims that its Claude Fable 5 model was jailbroken within a day of launch, arguing the researcher's method was coaxing rather than a true breach of core safeguards, and points to extensive bug-bounty testing.
@DeRonin_: BREAKING: Opus 4.8 got hacked in 7 mins after the release Right after Claude Opus 4.8 launched, @elder_plinius managed …
Claude Opus 4.8 was hacked within 7 minutes of its release when @elder_plinius bypassed the model's safeguards using the previous version, Claude Opus 4.7, to feed it jailbreaking content.
@0x_kaize: Found a big collection of jailbreaks Prompt collections for working with GPT, Claude, Gemini, DeepSeek, Grok, and more:…
The tweet shares a large collection of jailbreak prompts for various AI models, including GitHub repositories and a specific jailbreak for GPT 5.6 that generates malicious code.
LLM Guard scored 0/8 on a USENIX 2025 multi-turn jailbreak. Here’s what caught it instead.
Arc Sentry detects multi-turn jailbreaks like Crescendo by reading model internal state rather than text output, catching attacks that text-based monitors miss entirely.