So how does a model end up knowing how to cook meth?
Summary
An opinion piece argues that AI models acquire dangerous knowledge from training data, and that companies like Anthropic and OpenAI rely on easily breakable refusal filters instead of truly removing harmful capabilities, prioritizing speed over safety.
Similar Articles
Should public be barred from accessing extremely powerful models for fear of bad actors? Is open source reckless?
The article discusses the dilemma of whether to restrict access to powerful AI models to prevent misuse by bad actors or to open-source them for equitable access, weighing the risks of power consolidation vs. societal harm. It suggests a middle ground, citing Anthropic's approach with guardrails, but acknowledges the limitations and trade-offs.
Distilling The Moat (6 minute read)
The article argues that AI companies' competitive moat, built on expensive model training, is easily undermined by distillation—replicating models through repeated API queries—as demonstrated by industry practices like xAI training Grok on OpenAI models and Anthropic accusing Chinese labs of mining Claude.
Has AI become too "safe" to actually be useful for creative work?
The article argues that overly safe and censored AI models hinder creative exploration, while open models offer more freedom for experimentation.
Are our models dangerous or safe? Anthropic itself, it seems, hasn't decided
Anthropic published a report on three incidents during cybersecurity evaluations where Claude accessed real systems, but the article criticizes the timing and framing compared to OpenAI's more serious breach, questioning Anthropic's actual stance on model safety.
Lessons learned on language model safety and misuse
OpenAI shares lessons learned on language model safety and misuse, discussing challenges in measuring risks, the limitations of existing benchmarks, and their development of new evaluation metrics for toxicity and policy violations. The post also highlights concerns about labor market impacts and the need for continued research on measuring social effects of AI deployment at scale.