Statistically we are cooked
Summary
Argues that because LLMs must encode harmful content to identify it and jailbreaks are always statistically possible given large user bases, there is a non-zero chance of harm; the author therefore advocates against censorship to ensure good actors have the same tools as bad actors.
Similar Articles
Is it ever possible to have a malicious LLM with a backdoor
Discusses the possibility of LLMs containing backdoors triggered by secret sentences or conditions, and the relative risks of closed vs open-source models.
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing
Researchers introduce HarDBench, a benchmark exposing how LLMs can be jailbroken via malicious drafts in collaborative writing, and propose a preference-optimization defense that cuts harmful outputs without hurting co-authoring utility.
Born Against, or why hobby programming communities are aggressively against LLM usage
A blog post reflecting on why hobby programming communities like OSDev and demoscene are hostile to LLM usage, arguing that the craft and learning process itself is the point.
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
This paper reveals that grammar-constrained decoding (GCD) can be exploited as a jailbreak attack (CodeSpear) to induce LLMs to generate malicious code, and proposes a defense (CodeShield) that preserves safety under such attacks.
How Far Will They Go? Red-Teaming Online Influence with Large Language Models
This paper introduces a red-teaming framework that measures the 'Overton Window' of political opinions open-source LLMs can express and evaluates how simple jailbreaks expand that range, finding systematic left-leaning biases and vulnerabilities across 30+ models.