Long-running agents: is the bottleneck the model or the scaffolding around it?
Summary
A practitioner discussion exploring whether long-running AI agent failures stem from model capabilities or from the scaffolding around them, highlighting error compounding, context pollution, and weak self-correction as key failure modes.
Similar Articles
Your agent isn't failing because of the model, it's failing because nobody built a stop button
The article argues that the primary failure point for AI agents in production is not the model itself, but the lack of infrastructure such as stop buttons, billing oversight, and traceability for tool calls.
Are we focusing too much on models and not enough on agent infrastructure?
An opinion piece questioning whether the AI community is overemphasizing model capabilities at the expense of building robust agent infrastructure.
Running agents all day, I keep noticing the bottleneck is me defining "good", not the model
The author reflects that the primary bottleneck in running AI agents is not the model's capability but the human's ability to precisely define what 'good' or 'done' means, drawing parallels to managing people.
AI agents fail in ways nobody writes about. Here's what I've actually seen.
The article highlights practical system-level failures in AI agent workflows, such as context bleed and hallucinated details, arguing that these are often infrastructure issues rather than model defects.
Long-running AI agents may have a bigger continuity problem than memory
Reflects on the continuity problem for long-running AI agents, arguing that a deterministic control layer is needed to manage authoritative state, and questions whether existing infrastructure like IAM, transactions, and provenance is sufficient.