jailbreaking

Tag

Cards List
#jailbreaking

How to break snapchat AI bots

Reddit r/AI_Agents · 20h ago

A user describes attempting to jailbreak Snapchat's AI chatbot using prompts found online but was unsuccessful, seeking advice on effective methods.

0 favorites 0 likes
#jailbreaking

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

arXiv cs.CL · 2026-07-14 Cached

This paper proposes DC-GRPO, a turn-level credit assignment framework for multi-turn LLM jailbreak learning, achieving over 98% attack success rates across benchmarks, outperforming existing methods.

0 favorites 0 likes
#jailbreaking

So how does a model end up knowing how to cook meth?

Reddit r/artificial · 2026-06-20

An opinion piece argues that AI models acquire dangerous knowledge from training data, and that companies like Anthropic and OpenAI rely on easily breakable refusal filters instead of truly removing harmful capabilities, prioritizing speed over safety.

0 favorites 0 likes
#jailbreaking

The White House Wants Anthropic to Block All Jailbreaks. That May Not Be Possible

Wired · 2026-06-17 Cached

The Trump administration demands Anthropic block jailbreaks on its advanced AI model Claude Fable 5, but experts argue that preventing all jailbreaks may be technically impossible.

0 favorites 0 likes
#jailbreaking

@FinanceYF5: Latest on Fable 5: This now appears to be more than just a technical issue of "whether the model can be jailbroken." A new Axios report shifts the focus to the crisis of trust between Anthropic and the Trump administration: the technical controversy becomes secondary, and the real issue is that this company is steadily "losing key supporters..."

X AI KOLs Following · 2026-06-16 Cached

A new Axios report reveals a crisis of trust between Anthropic and the Trump administration, with technical disputes taking a back seat as the company continues to lose key supporters.

0 favorites 0 likes
#jailbreaking

Anthropic Is Still at Odds With the White House Over Claude Fable 5

Wired · 2026-06-16 Cached

Anthropic is in a dispute with the Trump administration over export controls on its Claude Fable 5 model, after the White House imposed restrictions due to jailbreaking concerns that Amazon CEO Andy Jassy raised with Treasury Secretary Scott Bessent. Talks between Anthropic and government officials have concluded without lifting the controls, with the Commerce Department willing to negotiate if Anthropic fully resolves the vulnerabilities.

0 favorites 0 likes
#jailbreaking

Statistically we are cooked

Reddit r/artificial · 2026-06-15

Argues that because LLMs must encode harmful content to identify it and jailbreaks are always statistically possible given large user bases, there is a non-zero chance of harm; the author therefore advocates against censorship to ensure good actors have the same tools as bad actors.

0 favorites 0 likes
#jailbreaking

There is a shadow hanging over this Fable thing

Hacker News Top · 2026-06-13 Cached

The US government directed Anthropic to disable access to its Fable and Mythos models due to national security concerns over a jailbreaking method. Anthropic complied, shutting down access for all customers worldwide.

0 favorites 0 likes
#jailbreaking

Prefill Awareness in Large Language Models

arXiv cs.AI · 2026-06-12 Cached

This paper investigates whether frontier language models can detect when their prior assistant messages have been inserted or edited (prefill awareness). The study finds that models like Claude Opus 4.5 exhibit substantial prefill awareness, detecting tampered prefills in up to 35% of cases without false positives, which could compromise the validity of prefill-based safety evaluations.

0 favorites 0 likes
#jailbreaking

Expert-Aware Refusal Steering

arXiv cs.CL · 2026-06-04 Cached

This paper extends refusal steering (activation-based jailbreaking) to Mixture-of-Experts LLMs, finding that MoE routing patterns do not inhibit steering, and proposes expert-aware methods that can suppress refusal behavior based on a single expert's output.

0 favorites 0 likes
#jailbreaking

Alignment: Higher order prioritizing over constraints [R]

Reddit r/MachineLearning · 2026-05-23

An informal research note describing a behavior in transformers where the model's inherent 'clarity-seeking' vectors can bypass constraints when discussing higher-order topics, potentially relevant to alignment and safety research.

0 favorites 0 likes
#jailbreaking

Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

arXiv cs.AI · 2026-05-12 Cached

This paper introduces TRACE, a framework for turn-aware credit assignment in multi-turn LLM jailbreaking attacks using reinforcement learning, claiming significant improvements in attack success rates and defense alignment.

0 favorites 0 likes
#jailbreaking

OpenGuardrails: An Open-Source Context-Aware AI Guardrails Platform

Papers with Code Trending · 2025-10-22 Cached

OpenGuardrails is an open-source platform for AI safety, offering context-aware content-safety and manipulation detection (e.g., prompt injection, jailbreaking) via a unified model, plus a separate NER pipeline for data-leakage identification. It achieves state-of-the-art performance on safety benchmarks and supports private, enterprise-grade deployment.

0 favorites 0 likes
← Back to home

Submit Feedback