evaluation-suite

Tag

Cards List
#evaluation-suite

TokEval: A Tokenizer Evaluation Suite

arXiv cs.CL · 2026-08-19 Cached

This paper introduces TokEval, a framework for evaluating language model tokenizers using intrinsic metrics that correlate with downstream task performance.

0 favorites 0 likes
#evaluation-suite

@alexwan55: 40% of benchmarking effort targets math/coding, but the related occupations are only 3.5% of US jobs. We introduce Econ…

X AI KOLs Following · 2026-06-24 Cached

Introduces EconEvals, an open-source evaluation suite that measures AI capabilities and predicts job disruption across the US labor economy, addressing the mismatch between benchmarking focus (math/coding) and actual job distribution.

0 favorites 0 likes
← Back to home

Submit Feedback