Tag
This study examines the action-level reliability of clinical LLM agents by rerunning tasks with identical inputs and comparing orders, finding significant divergence that benchmarks may miss and proposing enhanced evaluation methods.