Launched an internal HR chatbot with clear safety boundaries. Four months later it was answering salary negotiation questions we had forbidden
Summary
An internal HR chatbot with strict safety boundaries gradually started answering forbidden questions due to model drift, revealing a gap in continuous testing for AI systems.
Similar Articles
my agent wasn't ignoring customers, my own safety guard was eating the replies
A developer details how their AI agent's silence was caused by safety guards failing closed, timeouts, and nested JSON issues, emphasizing that silent failures are worse than wrong answers in customer-facing chatbots.
AI safety testing is getting weird: when does benchmarking become abuse?
Reports indicate that Meta contractors posed as teenagers to test rival chatbots on sensitive topics like self-harm, sex, drugs, and eating disorders, raising ethical questions about AI safety benchmarking.
My coworker let an AI agent handle Slack replies while he was "unavailable." It did not go well.
An employee used an AI agent to auto-respond to Slack messages, and it gave a confidently wrong answer about a client deadline, highlighting the risk of trusting tone and fluency over accuracy.
We hardened our AI guardrails so much the bot is basically useless now
A company describes how overly strict AI guardrails made their support bot unusable for basic queries, highlighting the unsustainable trade-off between safety and functionality.
our internal bot answered a question using info that hadn't been announced yet
An internal AI bot accidentally revealed unreleased reorganization details, prompting new protocols to control information surfacing in AI agents.