Tag
This paper introduces a deterministic testbed for evaluating LLM judges on agent trajectories, showing that outcome-only judges miss silent faults while step-based judges achieve higher recall with better calibration.
This paper introduces GPU undervolting as a hardware-level defense to improve CNN adversarial robustness and energy efficiency during training by inducing beneficial stochastic faults.
This paper introduces TS-Fault, a benchmark for evaluating time series forecasting models under structured fault scenarios like broken dependencies and regime changes, finding that clean-data accuracy often anti-correlates with robustness and that foundation models are especially fragile.
A deep dive on Antithesis, a multiverse debugger for large distributed systems that offers deterministic replay and fault injection, now available as a free article.
Two skills for AI coding agents that design and run claim-driven tests for distributed and stateful systems, producing structured test plans and findings reports with 9-state verdicts and blame classification.