Tag
The article examines the issues with MMLU benchmark scores, showing that different model builds can yield incomparable accuracies due to open evaluation variables, and proposes a content-addressed framework for better traceability.