Tag
This paper examines the susceptibility of LLM-based hate speech moderation to annotator-style rebuttals, showing that such attacks degrade performance and reveal directional asymmetries between whitewashing and smearing manipulations.