Tag
The paper presents AgentChaosBench, a benchmark for detecting and localizing runtime faults in LLM-based agentic systems, and evaluates it using zero-shot LLM baselines, revealing significant challenges in fault diagnosis.