Long-running agents: is the bottleneck the model or the scaffolding around it?

Reddit r/artificial News

Summary

A practitioner discussion exploring whether long-running AI agent failures stem from model capabilities or from the scaffolding around them, highlighting error compounding, context pollution, and weak self-correction as key failure modes.

Something I keep noticing with agent setups: every individual step is easy for the model, but the full chain still falls apart on long tasks. I think it comes down to three things: Error compounding. At 95% accuracy per step, a 20-step chain only succeeds about a third of the time (0.95^20 is roughly 0.36). Small errors stack up fast. Context pollution. After enough tool calls, the context is full of old outputs, dead ends and failed attempts, and the model slowly loses track of the original goal. Weak self-correction. When a step goes wrong, the model usually keeps building on top of it instead of backing up and fixing it. What I'm unsure about is where the real fix comes from. One camp says it's mostly better models. The other says a well-designed loop (checkpoints, verifier steps, state stored outside the context window) matters more than the model itself. Curious what people here have seen in practice: Do you keep the full history in context, or summarize/prune as you go? Has a separate critic or verifier model actually helped, or does it just add latency and cost? At what point in a task do your agents usually start breaking, and what actually fixed it for you? Would love to hear what's worked and what hasn't.
Original Article

Similar Articles