Tag
A multi-stage agentic framework for generating, refining, and evaluating counter-narratives to combat hate speech and misinformation, with human validation showing improved persuasiveness and safety over expert-written counterspeech.
This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.