Tag
Anthropic accused Alibaba of creating tens of thousands of fake Claude accounts for distillation attacks, leading to stricter guardrails that are affecting legitimate users.
This article warns that current and upcoming AI models significantly lower the barrier to creating bioweapons, citing distillation attacks on open-weight models and the inability to prevent safety ablation. It calls for public funding of broad-spectrum countermeasures as a necessary response.
This paper studies distillation attacks where model outputs can enable imitation, proposing a minimax game framework and a forward-pass-only defense called Product-of-Experts, showing that adaptive students recover more capability than passive evaluation suggests.