dangerous-knowledge

Tag

Cards List
#dangerous-knowledge

So how does a model end up knowing how to cook meth?

Reddit r/artificial · 2026-06-20

An opinion piece argues that AI models acquire dangerous knowledge from training data, and that companies like Anthropic and OpenAI rely on easily breakable refusal filters instead of truly removing harmful capabilities, prioritizing speed over safety.

0 favorites 0 likes
#dangerous-knowledge

Model Unlearning Objectives Vary for Distinct Language Functions

arXiv cs.CL · 2026-05-27 Cached

The paper argues that unlearning in LLMs should be goal-dependent, proposing a cosine-based meta-learned variant of RMU for dangerous knowledge and a multi-layer objective with probe directions for toxicity, achieving strong results across four 7-8B models.

0 favorites 0 likes
← Back to home

Submit Feedback