Tag
A survey shows that most teams keep agents on a short leash, and data indicates narrow-scope agents succeed 65% of the time vs 16% for broad scope, suggesting the reliability issue may be more about scope than model capability.