Tag
This paper defines multi-image implicit toxicity (MIIT), where individually benign images become toxic when combined, and proposes MiShield, a model trained with progressively distilled reasoning supervision to detect MIIT. Experiments show MiShield-8B outperforms existing moderation services.
ToxiREX is a new multilingual dataset of Reddit comments annotated for implicit toxicity using a toxic reasoning schema, covering six languages and multiple events.
The paper introduces CITA, a framework for generating implicit toxicity attacks in Chinese to evaluate and improve LLM toxicity detectors, finding high attack success rates across tested models.