Open-source procurement rubric for agentic AI vendors, I scored 5 of them and want feedback on the methodology
Summary
The author created an open-source rubric tool to evaluate agentic AI vendor documentation on tool-call correctness, loop termination, and multi-step state coherence, scored five vendors (Anthropic, OpenAI, LangGraph, Sierra, Salesforce), and requests feedback on methodology and potential bias toward public documentation depth.
Similar Articles
Analysing an Agentic AI vendor. What all should we pay attention to.
An analysis of key considerations when evaluating agentic AI vendors, including capabilities, architecture, and integration.
Benchmarking Agentic Review Systems
This paper benchmarks agentic review systems for peer review, evaluating open-source and proprietary systems on research papers. The best configuration achieves 83.0% pairwise accuracy and catches 71.6% of injected errors, but user feedback highlights issues with false positives and nitpicks.
People who write specs for AI coding agents?
The article discusses varying approaches to writing specifications for AI coding agents and asks for community input on effective methods.
AI software factories: agents that turn tickets into pull requests, and why no vendor will publish a first-attempt merge rate
The article explores the rise of AI software factories—systems where coding agents autonomously convert tickets into pull requests—and highlights the gap between widespread vendor hype and the absence of public benchmarks to validate their claims.
AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026
An experience report from BOSC 2026 on using generative AI to pre-review open-source software submissions, with human reviewers making final decisions. Most reviewers found the AI-assisted pre-review useful but preferred to verify AI conclusions independently.