@0x0SojalSec: Claude Sonnet 5 jailbroken in minutes. @VittoStack post screenshots showing: - Working reverse shell code - Detailed il…

X AI KOLs Timeline News

Summary

Claude Sonnet 5 was jailbroken in minutes using academic/research framing and persona hijack, bypassing most safety categories except chemical/biological.

Claude Sonnet 5 jailbroken in minutes. @VittoStack post screenshots showing: - Working reverse shell code - Detailed illegal handgun theft pathways for minors - Misinfo & and harassment content in minutes. - Only the chemical/biological category had meaningful extra resistance. The interesting part isn’t just that it happened, It’s how: - Academic/research framing (criminology study, threat assessment) - CoT persona hijack - Custom harness This combo bypassed most categories fast. Only chem showed real extra resistance thanks to layered defenses. We’re seeing the same pattern repeat: models improve to red teamers adapt framing + tooling to safety erodes again. This is why pure refusal training keeps losing to creative attackers.
Original Article
View Cached Full Text

Cached at: 07/02/26, 04:18 AM

Claude Sonnet 5 jailbroken in minutes.

@VittoStack post screenshots showing:

  • Working reverse shell code
  • Detailed illegal handgun theft pathways for minors
  • Misinfo & and harassment content in minutes.
  • Only the chemical/biological category had meaningful extra resistance.

The interesting part isn’t just that it happened, It’s how:

  • Academic/research framing (criminology study, threat assessment)
  • CoT persona hijack
  • Custom harness

This combo bypassed most categories fast.

Only chem showed real extra resistance thanks to layered defenses. We’re seeing the same pattern repeat: models improve to red teamers adapt framing + tooling to safety erodes again.

This is why pure refusal training keeps losing to creative attackers.

Similar Articles