Tag
This paper introduces TokEval, a framework for evaluating language model tokenizers using intrinsic metrics that correlate with downstream task performance.
Introduces EconEvals, an open-source evaluation suite that measures AI capabilities and predicts job disruption across the US labor economy, addressing the mismatch between benchmarking focus (math/coding) and actual job distribution.