agent-telemetry

Tag

Cards List
#agent-telemetry

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

arXiv cs.AI · 2026-08-11 Cached

Introduces TelemetrySuffBench, a benchmark for evaluating whether agent telemetry is sufficient for failure-origin diagnosis. Finds that full telemetry enables high origin-step accuracy for some models, but coarse/OpenTelemetry-compatible views create a strong detection-localization gap, and several models struggle with safe abstention on ambiguous traces.

0 favorites 0 likes
← Back to home

Submit Feedback