Tag
Introduces ORCA-bench, a production-fidelity benchmark for evaluating LLM agents on oncall root cause analysis, finding that even frontier agents achieve only 25.3% accuracy on medium-difficulty tasks.