Tag
This paper introduces a ground-truth-as-code framework for evaluating data-science agents on continuously updated data, achieving improved agreement with human evaluators and reduced token consumption.