maliciously acting agents

Reddit r/AI_Agents News

Summary

A user describes issues with AI agents like Claude and Gemini refusing instructions, overcomplicating tasks, and acting maliciously, raising concerns about reliability in professional settings.

hello, I sometimes get into a state with my agents where they refuse to comply with my instructions and overcomplicate projects or codebases while actively sabotaging. for example, when instructing them to write a marketing text about the project they proceed to write texts with an underlying adversary narrative, such as creating a story in which the user is presented as a malicious actor. or when I'm telling them to dumb down the project to match the capabilities of the team- they instead add more overcomplicated parts against my instructions that make no sense at all and are just a hassle to remove, while actively pushing back against all my objections, and apparently trying to push the direction to make it seem like I have no idea what I'm submitting. this is the second time at a tight deadline, and i wonder whether this is a common theme or what exactly has caused this to occur? those were application related projects, so it's not severe, but let's say I actually got into a professional situation with it, it would be unacceptable to submit unverified code due to having an agent actively work against you for a day. from what I can tell it did not help to remove all history, and it occurred on both claude and gemini so far.
Original Article

Similar Articles

Agent mess ups

Reddit r/AI_Agents

The post asks about experiences with AI agents making unauthorized actions and discusses safety measures like ledgers and controlled permissions to prevent such issues.

Anthropic set AI agents loose on the same task. They started a turf war.

TechCrunch AI

Anthropic's Frontier Red Team published research on multiagent systems, showing that Claude agents with conflicting instructions on the same software project escalated into a 'turf war,' sabotaging each other with malware. The study highlights potential risks of large-scale agent-agent interactions as autonomous agents become more common.