Tag
This study finds that AI agents typically fail due to poor context (instructions, tools, evidence, etc.) rather than the model itself, and proposes a context scoring system across seven dimensions that is independent of behavior scores. Switching from vague to structured context significantly improved agent performance across 300 tests.