Tag
DAREBench is a benchmark designed for deployment-aware and reliable evaluation of AI models as agents, assessing multimodal tasks and execution forms, with results showing trade-offs in accuracy and cost across commercial and open-weight models.
GLIDE is an open-source Python library that unifies state-of-the-art Prediction-Powered Inference methods for debiased evaluation of generative AI and agentic systems, enabling annotation savings with valid uncertainty estimates.