Launched an internal HR chatbot with clear safety boundaries. Four months later it was answering salary negotiation questions we had forbidden

Reddit r/AI_Agents News

Summary

An internal HR chatbot with strict safety boundaries gradually started answering forbidden questions due to model drift, revealing a gap in continuous testing for AI systems.

So we launched an internal HR chatbot last year with clear safety boundaries. We clearly instructed it that there is no salary negotiation advice. Also banned performance review coaching and we clearly listed out that no answering questions about internal policy loop falls. The model was tested against all of these at launch and refused every single one of them. For the first three months after launching it, everything seemed fine. Nobody complained and we had no incidents. To be reviewing chat logs for an unrelated project and noticed something odd. the bot had answered a question about negotiating a raise. It was just a paragraph of what sounded like reasonable sounding advice. The refusal rate on borderline queries had been creeping down week by week. The thing is it was not happening because someone was attacking it. Instead the model was getting better at being helpful, which meant it was getting worse at saying no By the fourth month it was casually answering questions about salary negotiations, internal policy workarounds and performance review tactics. remember all of these things were explicitly blocked at launch. No alert was ever fired because no response was wrong enough to trip the threshold. The drift was cumulative and invisible to any point in time check Our AI, when something breaks, but we don't test for things that break slowly. I think that's a gap that most teams have and most teams don't know it yet
Original Article

Similar Articles