Meta's Muse has a separate Sentinel for outbound actions. What would you test first?

Reddit r/AI_Agents Products

Summary

Meta announced Muse, a personal AI agent featuring a separate Sentinel for approving outbound actions and user permissions for sensitive tasks. The article suggests testing which security layer rejects malicious instructions to evaluate the system.

Meta's September 8 Muse announcement describes a personal agent running in a dedicated cloud VM, with a separate Sentinel approving outbound activity. Meta also says sensitive actions such as sending email or making purchases require the user's permission. The distinction worth watching is between an AI reviewer deciding an action looks acceptable and permissions that neither agent can override. A useful test would be a page telling the assistant to send private information to a new destination: which layer rejects it, and what evidence reaches the user? There's also a rollout boundary: Muse is rolling out in the US on iOS, Android and the web. The Confidential VM with a key held only by the user is planned for later this year, not the same as the VM offered now. These are Meta's design claims, not an independent security test. For browser agents, what would you test first: malicious page instructions, approval-payload changes, or revoked access? Source in the comments.
Original Article

Similar Articles