Tag
WhatWorkedBench is a benchmark designed to measure the experimental understanding of AI agents by evaluating their accuracy in predicting outcomes after budgeted experimentation across various tasks and configurations.