Tag
Introduces TelemetrySuffBench, a benchmark for evaluating whether agent telemetry is sufficient for failure-origin diagnosis. Finds that full telemetry enables high origin-step accuracy for some models, but coarse/OpenTelemetry-compatible views create a strong detection-localization gap, and several models struggle with safe abstention on ambiguous traces.