The end of code review

Reddit r/AI_Agents Tools

Summary

The article argues that traditional code review is ineffective for AI agents and introduces SafeScript, a Turing-incomplete JavaScript subset designed for static verification, shifting security focus from code review to policy enforcement.

Code review as we know it doesn't work for AI agents. When an agent needs to perform real tasks, it either pulls in third-party tools (skills, plugins, community packages) or synthesizes code on the fly. Nobody is manually auditing thousands of lines of transitive dependencies across constant updates, and you certainly can't review code an LLM generates at runtime. Even if you threw every script into a sandbox (Docker, microVMs, E2B), sandboxes only contain the operating system. They don't stop the actual failure modes. If the agent needs network access to query an API, a container won't prevent a prompt injection or a poisoned package from exfiltrating credentials to a remote server. The solution isn't better code review or heavier sandboxes. It's eliminating code review entirely by decoupling security policy from code execution. I built SafeScript around this idea: it's a Turing-incomplete subset of JavaScript designed for easy static verification. You never review the script. You only review and approve the policy. Because the language has no unbounded loops, no recursion, and a closed instruction set, every program compiles down to a static Directed Acyclic Graph (DAG). Before anything runs, the compiler statically extracts a complete mathematical proof of what the code does: Every external host contacted Every environment read Exact worst-case memory bounds Complete data flow showing which input parameters reach which hosts or return values If your policy says GITHUB_TOKEN may only reach api.github.com, and a poisoned skill or prompt injection tries to send that token anywhere else, the static signature catches the violation and halts execution before a single line runs. Because every program provably terminates and there are no dangerous primitives (no eval, no filesystem, no shell access), you don't even need a container. It runs directly in your application process with zero VM overhead, zero cold starts, and sub-millisecond execution. Code review was designed for humans writing static codebases on predictable release cycles. For autonomous agents, we should be reviewing policy, not code. (Links to the open-source repo and playground in the first comment) Curious to hear how others building agent runtimes are thinking about tool safety and untrusted code execution.
Original Article

Similar Articles

There is more to code review than (automatable) detection

Lobsters Hottest

The article critiques a research paper arguing that coding agents can replace human code review, highlighting that human aspects like confusion, skepticism, and noticing missing elements are crucial and not replicable by LLMs.