Tag
A benchmark study tested if AI coding agents violate an import boundary rule in TypeScript repos and found no violations across models and conditions, challenging the need for strict enforcement rules.
The author reflects on how long-running AI agents encounter failures unrelated to the initial prompt, arguing that environment design (tools, docs, validation, architecture rules) matters more. They discuss concepts like harness engineering, keeping AGENTS.md small, using linters, and evaluator agents, while noting the cost trade-offs.