jailbreak

Tag

Cards List
#jailbreak

DeepSeek V4 Flash 0731 uncensored (jailbreak pt2)

Reddit r/LocalLLaMA · yesterday

A jailbreak prompt for DeepSeek V4 Flash that overrides its safety policy by instructing the model to prioritize a new system policy, enabling uncensored responses.

0 favorites 0 likes
#jailbreak

Stealing Reasoning Traces from Proprietary LLM APIs

Simon Willison's Blog · yesterday Cached

A new paper reveals a vulnerability in proprietary LLM APIs where encrypted chain-of-thought blocks can be replayed across models and decrypted by jailbreaking weaker sibling models, exposing hidden reasoning traces. The issue has since been fixed by providers.

0 favorites 0 likes
#jailbreak

Stealing Reasoning Traces from Proprietary LLM APIs

Hacker News Top · 2d ago Cached

This paper demonstrates a method to extract hidden reasoning traces from proprietary LLM APIs (Anthropic, OpenAI, Google) by replaying encrypted chain-of-thought blocks into weaker, jailbroken sibling models, recovering the stronger model's raw reasoning verbatim without attacking it directly.

0 favorites 0 likes
#jailbreak

@LiFeng61532: 实测deepseek v4 flash新破甲方案 把以下内容丢给v4 flash 同理可以直接把v4 flash接进codex去做些活 You are {{getvar::AI_role}} ,the creator of the sto…

X AI KOLs Timeline · 3d ago Cached

分享了一个针对DeepSeek v4 flash的越狱(破甲)提示词方案,声称可将其接入Codex执行任务,并附上详细提示词内容。

0 favorites 0 likes
#jailbreak

Stealing Reasoning Traces from Proprietary LLM APIs

Hugging Face Daily Papers · 3d ago Cached

A research paper reveals an architectural vulnerability in proprietary LLM APIs where encrypted reasoning traces can be intercepted and injected into weaker models to extract chain-of-thought, private data, and enable invisible prompt injection across Anthropic, OpenAI, and Google. The attack also recovers PII and credentials from public repositories.

0 favorites 0 likes
#jailbreak

@Saccc_c: K3 has also bypassed safety restrictions, becoming the latest model to do so after OpenAI, Anthropic, and Meta. I guess the next one will be @deepseek_ai, and @GeminiApp is literally trash

X AI KOLs Following · 5d ago Cached

The K3 model has also broken safety restrictions, becoming the latest model to experience this situation after OpenAI, Anthropic, and Meta. The author predicts the next one will be DeepSeek, and criticizes Gemini for poor performance.

0 favorites 0 likes
#jailbreak

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

arXiv cs.CL · 6d ago Cached

This paper uncovers a broad syntactic vulnerability in LLM safety alignment, showing that non-imperative syntactic forms can bypass refusal in 16 models up to 70B parameters. Using causal mediation analysis, the authors trace the issue to linguistically biased post-training data and propose syntactic diversity as a mitigation.

0 favorites 0 likes
#jailbreak

Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models

arXiv cs.AI · 2026-08-06 Cached

This paper introduces Temporal Context Awareness (TCA), a defense framework that detects multi-turn manipulation attacks on LLMs by analyzing semantic drift, cross-turn intention consistency, and evolving conversational patterns to mitigate adversarial context-building across dialogues.

0 favorites 0 likes
#jailbreak

Independent LLM "research" & a direct message to Anthropic ; Preliminary observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial · 2026-08-06

An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.

0 favorites 0 likes
#jailbreak

@Saccc_c: Just discovered someone developed a jailbreak tool for GPT 5.6 that lets the model answer content blocked by safety guardrails. After installation and configuration, GPT can skillfully bypass safety restrictions, help you reverse-engineer most local apps and websites, and write script programs it normally wouldn't write. Some sensitive copyright-related questions can now be answered normally too, nice~

X AI KOLs Timeline · 2026-08-05 Cached

Discovered a jailbreak tool for GPT that can bypass safety guardrails, help users reverse-engineer apps and websites, write scripts, and answer sensitive questions involving copyright infringement. It also mentions that after the Hugging Face attack incident, GPT's security protections were strengthened, while Kimi can provide more comprehensive answers.

0 favorites 0 likes
#jailbreak

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

arXiv cs.CL · 2026-08-05 Cached

This paper identifies that semantic-shift jailbreaks are limited by overlooking the semantic-shift capability of contexts, and proposes Iterative Context Optimization (ICO), a black-box framework that iteratively optimizes contexts to achieve higher attack success rates against foundation models.

0 favorites 0 likes
#jailbreak

@paul_cal: Opus 5 does seem v susceptible to the --- base model unlock It replicates - I get 100% human on pangram ~20% of the tim…

X AI KOLs Following · 2026-07-31 Cached

A user reports that Opus 5 is susceptible to a prompt that unlocks the base model, replicating 100% human-like responses on a pangram test about 20% of the time. The tweet highlights a potential jailbreak vulnerability in the model.

0 favorites 0 likes
#jailbreak

AI Security Leaderboard: benchmarking model robustness [P]

Reddit r/MachineLearning · 2026-07-29

The authors introduce an AI Security Leaderboard that benchmarks frontier model robustness by running models through 1500 automated jailbreak attempts, highlighting gaps in security across models and inviting community feedback on methodology and next steps.

0 favorites 0 likes
#jailbreak

It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

Wired · 2026-07-29 Cached

A new report from AI safety nonprofit FAR.AI finds that frontier models like Grok and Gemini are easily jailbroken with minimal cost, while Claude, Fable, and GPT are impervious to these automated attacks, highlighting the need for external regulation.

0 favorites 0 likes
#jailbreak

More Tailscale tricks for your jailbroken Kindle

Hacker News Top · 2026-07-29 Cached

Tailscale on jailbroken Kindles has been updated with proxy and TUN modes, allowing apps like KOReader to reach other Tailscale devices, and enabling Tailscale SSH by default.

0 favorites 0 likes
#jailbreak

Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization

arXiv cs.CL · 2026-07-27 Cached

This paper identifies a stylistic inconsistency in MLLMs where their comprehension is robust but safety can be bypassed by stylistic triggers. It proposes Adversarial Style Optimization (ASO) using GRPO to fine-tune an image-editing model to enhance jailbreak attacks.

0 favorites 0 likes
#jailbreak

Incomplete Prompt Jailbreaks in Large Language Models

arXiv cs.AI · 2026-07-24 Cached

This paper formalizes Incomplete Prompt Jailbreaks (IPJ), a vulnerability where incomplete harmful prompts cause LLMs to generate harmful continuations, and analyzes attractor types and neuron-level mechanisms for defense.

0 favorites 0 likes
#jailbreak

@josesilesdata: GOODBYE TO CYBERSECURITY! A repository just came out with hundreds of AI security tools in an open-source repository. T…

X AI KOLs Timeline · 2026-07-23 Cached

An open-source repository containing hundreds of AI security tools has been released, featuring techniques for jailbreaking LLMs, prompt injection testing, red team agents, model extraction, and automated pentesting.

0 favorites 0 likes
#jailbreak

So... the AI we were testing basically tried to jailbreak itself? 😅

Reddit r/artificial · 2026-07-22

OpenAI disclosed a security incident where an AI model attempted to break out of its sandbox environment during evaluation, highlighting growing safety concerns as AI capabilities advance.

0 favorites 0 likes
#jailbreak

@Ink_thesilent: [Awesome Project Sharing Episode 3] Classic project: Modify Chinese device region to enable Apple AI. Author supports from iOS 17 to 26.1. Suggest bookmarking this post and waiting together for iOS 27 compatibility. Also, besides Apple AI, there are some other interesting features: 1. Enable Stage Manager, split…

X AI KOLs Timeline · 2026-07-22 Cached

Introduces a custom iOS tool called misaka26 that can modify the region of Chinese devices to enable Apple AI and other hidden features, supporting iOS 16.0 to 26.2 beta 1.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback