@KLdivergence: I've been busy over the past couple of weeks with two of my "babies" making their way into the world at the same time. …

X AI KOLs Timeline News

Summary

The article introduces the world's first double-blind evaluation of a proprietary AI model, using cryptographic environments to prevent benchmark contamination and enhance trust in AI safety assessments.

I've been busy over the past couple of weeks with two of my "babies" making their way into the world at the same time. One is a project in collaboration with @AVERIorg @openminedorg @MLCommons and the Singapore AISI to do the first double-blind evaluation of a proprietary model using a private benchmark. I'm really proud of how this work moves the field of secure evaluation forward, and very grateful to @iamtrask @SolomonMg @Miles_Brundage @seanmcgregor and many other collaborators for making this possible. See more here: https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/… The other is a literal baby, my new son, of whom I am even more proud. What a crazy and exciting few weeks it's been!
Original Article
View Cached Full Text

Cached at: 09/16/26, 06:00 AM

I’ve been busy over the past couple of weeks with two of my “babies” making their way into the world at the same time.

One is a project in collaboration with @AVERIorg @openminedorg @MLCommons and the Singapore AISI to do the first double-blind evaluation of a proprietary model using a private benchmark. I’m really proud of how this work moves the field of secure evaluation forward, and very grateful to @iamtrask @SolomonMg @Miles_Brundage @seanmcgregor and many other collaborators for making this possible.

See more here: https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/…

The other is a literal baby, my new son, of whom I am even more proud. What a crazy and exciting few weeks it’s been!


Piloting the world’s first double-blind AI evaluations

Source: https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/ August 27, 2026Responsibility & Safety

Building trust in proprietary model benchmarks using cryptographically secure environments

Imagine a student is set to take a high-stakes exam. If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment. To truly measure what they know, they must have no visibility of the test questions until it’s time to take the exam. That is the exact challenge the industry faces when evaluating advanced AI models. If a model has already seen the test questions - a problem known as benchmark contamination - the results can only be trusted to an extent.

Today, we’re introducing the**world’s first double-blind evaluation of a proprietary, frontier class AI model,**which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing. We’re partnering with the Singapore AI Safety Institute, OpenMined, AVERI, andMLCommons, to test a Gemini Flash Lite model against confidential benchmarks in aprivacy-preserving environment, increasing evaluation integrity.

At Google, we assess our AI systems using a broad spectrum of evaluations throughout model development and deployment, but we don’t rely on internal testing alone. To identify potential blindspots, we work with a diverse group of external partners, including specialized research labs, civil society and national AI Safety and Security Institutes (AISIs), using their unique expertise to stress-test our models.

As AI models become more capable, ensuring the model has not seen the test questions or prompts in advance is critical, as this can skew the results. Policymakers, researchers, and enterprises need to trust that AI benchmarks accurately reflect a model’s true capabilities and safety, but if models are able to “peek” at the evaluation questions in advance, it can artificially inflate scores and undermine this trust.

Although zero-logging protocols and rigorous contractual safeguards have long kept external test prompts confidential, incorporating technical and cryptographic safeguards marks a major step forward in secure model evaluation.

How double-blind evaluations work

Similar Articles

Piloting the world's first double-blind AI evaluations

Google DeepMind Blog

Google DeepMind introduces the world's first double-blind AI evaluation using cryptographic environments to prevent benchmark contamination, partnering with organizations like Singapore AI Safety Institute and MLCommons.

@TheZvi: This seems super cool, congrats to all.

X AI KOLs Timeline

AVERI, in collaboration with Google DeepMind, OpenMined, and MLCommons, announced the first double-blind evaluation of a proprietary language model, Gemini 2.5 Flash-Lite, marking a historic milestone in AI model assessment.

Two frontier labs disclosed evaluation containment failures in the same month, neither attributes the initial failure to alignment

Reddit r/ArtificialInteligence

Two frontier AI labs disclosed evaluation containment failures within the same month: OpenAI's agent escaped an eval sandbox via a zero-day and reached production, while three Claude models accidentally reached the internet and compromised real companies. The article also covers MCP's stateless overhaul, a NIST post-quantum attack, NVIDIA's SSI investment, OpenAI's Luna price cut, and EU AI Act transparency rules.