Tag
This paper studies the evaluation of agentic AI systems that repair decision policies when per-state expert action labels are unavailable, using a hotel-pricing simulator with region-level diagnostic feedback. It finds that aggregate alignment can be misleading and proposes evaluating policy repair by closed-loop outcome rather than behavioral distance.