Tag
This paper introduces a certification framework for concept unlearning in text-to-image diffusion models, providing high-confidence guarantees on residual leakage and demonstrating that standard attack-based evaluations often underestimate safety risks.