Tag
The article discusses stress-tests on autonomous research agents, revealing that failures stem from metacognition issues rather than capability and introduces ARFT, a failure taxonomy with 45 patterns.
Multi-agent systems can cost 15-50x more than a single agent, yet most failures stem from specification ambiguity and coordination breakdowns, not model capability. Treating handoffs as API contracts and adding explicit verification is recommended.