Follow-up: 22 replies to my "checks on the call" post taught me more than writing it did. Six lessons, credited.

Reddit r/AI_Agents Tools

Summary

A follow-up post collecting 22 community replies on hardening tool-call checks in agents (rule-naming refusals, choke points over name-blocking, resolved-path validation, idempotency keys), plus the public launch of Paveo, a source-available in-process guardrail library for agent tool calls.

Three days ago I posted four patterns for putting a hard check in front of every tool call an agent makes (link in the comments). The replies were better than the post, so here they are in one place, with credit. Name the rule, never the number. (u/QuanTradin) A refusal that says "amount must be under 500" invites a 499. Say which rule it hit and what to do next ("ask the user"), and keep anything about changing the policy in the log for the human. Gating by tool name leaks if the agent has a shell. (u/mastafied) Block send_mail and a determined model writes ten lines of Python that hit the same API. The real fix is a choke point: credentials live only behind it, so the shell has nothing to reach. And for force pushes, branch protection on the remote beats any pattern match. Replace dangerous arguments instead of validating them. (u/AdministrativeBad752) A URL is close to impossible to allowlist: redirects, userinfo like [email protected], hex IPs, DNS that answers differently after the check. Pass an opaque handle that maps to a reviewed upstream, and treat whatever argument is left as hostile. Check arguments after they're resolved, not as typed. (u/ianreboot) "./build/../.env" or a symlink named notes.txt passes a literal check while the write lands somewhere else. Check the source of truth, not the agent's memory. (u/federicodonatone) Their outbound-email check reads the sent folder before every send, because drafts, scheduled sends and undo windows all look like "sent" from the agent's side. A reservation needs an idempotency key. (u/Excellent-Nebula-313) If the hold succeeds but the reply is lost, the retry should recover the same hold, not reserve twice. And a hold nobody ever settles gets charged its worst case: "we don't know if it ran" isn't "it didn't run". The thread that kept coming back: a pattern check is a guard rail, not a sandbox. It catches the common case cheaply, and it should say so honestly. Disclosure: the original four patterns are the core of a library I've been building, Paveo, and it went public today. Measured against this thread, honestly: Does: refusals name the rule and tell the model to stop and ask, without the policy's numbers or how to change it (#1). A hold that's never settled is charged its worst case (#6, second half). Rate and repeat limits per tool. Doesn't yet: resolve paths before comparing (#4). That's the next thing I'm building, because two people asked. It runs in your process, so it's a check, not a choke point that holds credentials (#2): use it alongside one, not instead. Watch out: its rate limit counts the calls it has seen, which is the agent-side view #5 warns about. Where there's a real source of truth, check that too. Python 3.11+, runs in your process, makes no network calls, free for up to two agents, source-available. Link in the comments. What would you add to the list?
Original Article

Similar Articles

Instructions didn't stop my agents. Checks on the call did. Four patterns that held up

Reddit r/AI_Agents

A practitioner shares four hard guardrail patterns that enforce checks on AI agent tool calls themselves—rather than relying on prompt instructions—validated against two months of Claude Code history, showing that runtime argument validation, ordering constraints, worst-case budget reservation, and explicit refusal messages prevent agents from bypassing instructions.