Tag
This paper introduces influence-guided response rewriting to intervene on influential training examples, showing that rewriting responses creates stronger and more persistent behavioral shifts in language models than conventional reweighting.
This paper evaluates end-to-end trade-offs in moderation for conversational AI, comparing filter placement (input, response, both) and actions (blocking vs rewriting) using customer-outcome metrics like Usefulness and Harmful Exposure instead of component accuracy.
PILA reformulates LLM-native advertising as a conditional response rewriting problem, decoupling ad insertion from upstream generation via a lightweight, model-agnostic sidecar module that preserves response quality while enabling controllable ad exposure.