Tag
A blog post from TensorZero argues that even very noisy LLM evaluators can be useful for offline agent selection and improvement, as noise averages out over many samples to reliably rank agents.