How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Summary
UK AISI and EvalEval are collaborating to openly share AI evaluation results using a standardized schema and platform, enhancing reproducibility and transparency in benchmarking for AI models.
View Cached Full Text
Cached at: 09/22/26, 09:01 PM
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Source: https://huggingface.co/blog/evaleval-aisi TheEvalEval Coalitionis thrilled to share that theUK AI Security Institute (AISI)is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.
AISI and EvalEval have previously collaborated on research that began at ajoint workshop alongside NeurIPS 2025, and feedback from the Institute has helped shape theEvery Eval Ever (EEE) schema. This next phase of the collaboration puts that shared infrastructure into practice.
https://huggingface.co/blog/evaleval-aisi#why-reproducible-evaluation-reporting-mattersWhy reproducible evaluation reporting matters
As AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance. Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them. Running the evaluations again may itself be prohibitively expensive.
EvalEval’s mission is to improve this ecosystem through a shared reporting schema,Every Eval Ever, and an open platform,Evaluation Cards, that brings evaluation results and the information needed to interpret them into a common structure.
This builds naturally on AISI’s work to make evaluation more efficient throughOptStop, more statistically rigorous throughHiBayES, and more standardised in areas including transcript analysis and capability elicitation. Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.
https://huggingface.co/blog/evaleval-aisi#what-aisi-is-sharingWhat AISI is sharing
Transcript-level transparency matters not only for reproducibility, but also for analysis and diagnosis. In this new phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate. The release includes verified results, context, and configuration information for the five benchmarks in the paper’s main experiment:
- HealthBench
- FrontierMath
- Humanity’s Last Exam
- SWE-Bench Pro
- Terminal-Bench 2.0
These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models. The data accompany AISI’s paper,How Inference Compute Shapes Frontier LLM Evaluation, which studies how benchmark performance depends on inference-time compute and evaluation protocol.
Performance on Humanity’s Last Exam changes with evaluation protocol and inference compute. Each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task. When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased.
When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem. Where other reports lack these details, releases like AISI’s provide verified reference points for interpreting evaluations in context—for example, by helping researchers understand how setup choices may influence reported performance. As more evaluators adopt EEE, open comparisons like these can support broader and more reliable meta-research.
AISI’s Terminal-Bench 2.0 results alongside other reported evaluations for the same models, under different evaluation setups.
We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.
https://huggingface.co/blog/evaleval-aisi#contribute-to-the-shared-missionContribute to the shared mission
- Model developers:Report verified evaluation results.
- **Evaluation developers:**Report benchmarks and run data using theEvery Eval Ever schema.
- Evaluation, governance, and policy researchers:Explore Evaluation Cardsby benchmark or model, or use it to examine the state of evaluation reporting as a whole.
https://huggingface.co/blog/evaleval-aisi#about-the-evaleval-coalitionAbout the EvalEval Coalition
The EvalEval Coalition is a research community developing scientifically grounded research and robust deployment infrastructure for the evaluation ecosystem. Its goal is to improve evaluation science, address the lack of consensus around documenting evaluation applicability and utility, and broaden coverage of the impacts that matter for scientific research and policy analysis.
The coalition’s flagship projects includeEvery Eval Ever, a shared schema and repository for evaluation results, andEvaluation Cards, which combines benchmark metadata, evaluation-run data, and model metadata into interpretable records. Together, they make it easier to understand when apparently similar scores were produced under meaningfully different conditions.
https://huggingface.co/blog/evaleval-aisi#about-the-uk-ai-security-instituteAbout the UK AI Security Institute
TheUK AI Security Instituteis a research organisation within the UK government’s Department for Science, Innovation and Technology. Its mission is to equip governments with a scientific understanding of the risks posed by advanced AI. AISI conducts research and builds infrastructure to understand advanced AI capabilities and impacts, develop and test mitigations, and inform policy.
https://huggingface.co/blog/evaleval-aisi#further-readingFurther reading
Similar Articles
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Introduces Every Eval Ever, a shared schema and community-crowdsourced repository for standardizing AI evaluation results, with automatic converters and a hosted database spanning over 22k models and 2.2k benchmarks.
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
This paper introduces EvalCards, an operational framework that standardizes AI evaluation reporting by composing benchmark metadata, evaluation run data, and model metadata into a unified record with interpretive signals for reproducibility, completeness, provenance, risk, and score comparability. The authors deploy a monitoring tool across thousands of models and benchmarks, revealing systematic gaps in current reporting practices.
@OpenAI: As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchmarks help the field u…
OpenAI emphasizes the need for more rigorous and trustworthy evaluations for coding AI models to better measure real progress.
Featuring Every Eval Ever Results on Hugging Face Model Pages
Every Eval Ever (EEE) and Hugging Face Community Evals are now intercompatible, allowing standardized cross-posting of AI evaluation results to model pages, improving trust and comparability.
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
This paper introduces Benchmarking-Cultures-25, a dataset analyzing how AI model builders selectively highlight benchmarks in press releases. It finds a fragmented evaluation landscape with limited cross-model comparability, arguing that benchmarks are used as narrative devices for market positioning rather than standardized scientific measurement.