implicit-toxicity

Tag

Cards List
#implicit-toxicity

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

arXiv cs.CL · 2026-07-02 Cached

This paper defines multi-image implicit toxicity (MIIT), where individually benign images become toxic when combined, and proposes MiShield, a model trained with progressively distilled reasoning supervision to detect MIIT. Experiments show MiShield-8B outperforms existing moderation services.

0 favorites 0 likes
#implicit-toxicity

ToxiREX: A Dataset on Toxic REasoning in ConteXt

arXiv cs.CL · 2026-06-29 Cached

ToxiREX is a new multilingual dataset of Reddit comments annotated for implicit toxicity using a toxic reasoning schema, covering six languages and multiple events.

0 favorites 0 likes
#implicit-toxicity

Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting

arXiv cs.CL · 2026-05-22 Cached

The paper introduces CITA, a framework for generating implicit toxicity attacks in Chinese to evaluate and improve LLM toxicity detectors, finding high attack success rates across tested models.

0 favorites 0 likes
← Back to home

Submit Feedback