Tag
This paper evaluates end-to-end trade-offs in moderation for conversational AI, comparing filter placement (input, response, both) and actions (blocking vs rewriting) using customer-outcome metrics like Usefulness and Harmful Exposure instead of component accuracy.