llm-moderation

Tag

Cards List
#llm-moderation

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

arXiv cs.CL · 2d ago Cached

This paper examines the susceptibility of LLM-based hate speech moderation to annotator-style rebuttals, showing that such attacks degrade performance and reveal directional asymmetries between whitewashing and smearing manipulations.

0 favorites 0 likes
← Back to home

Submit Feedback