reliable-evaluation

Tag

Cards List
#reliable-evaluation

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

arXiv cs.AI · 2026-09-10 Cached

DAREBench is a benchmark designed for deployment-aware and reliable evaluation of AI models as agents, assessing multimodal tasks and execution forms, with results showing trade-offs in accuracy and cost across commercial and open-weight models.

0 favorites 0 likes
#reliable-evaluation

Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

arXiv cs.AI · 2026-06-01 Cached

GLIDE is an open-source Python library that unifies state-of-the-art Prediction-Powered Inference methods for debiased evaluation of generative AI and agentic systems, enabling annotation savings with valid uncertainty estimates.

0 favorites 0 likes
← Back to home

Submit Feedback