Tag
The paper presents EvalCEGAR, a method for automatically evolving evaluation metrics using a pool of Python operators that flag specific defects in AI outputs, improving accuracy over hand-written operators and LLM judges.