Tag
This paper presents C-Guard, a constitution-grid instrument for data-efficient RL alignment of safety guards, using C-LIM learnability scores to prune, densify, amend, or expand training data. It reduces over-refusal from 22.4% to 12.8% and improves data efficiency.
A developer details how their AI agent's silence was caused by safety guards failing closed, timeouts, and nested JSON issues, emphasizing that silent failures are worse than wrong answers in customer-facing chatbots.