Piloting the world's first double-blind AI evaluations
The announcement draws an analogy to a student who peeks at exam questions before a test, making a perfect score meaningless. In AI evaluation, the same problem arises as benchmark contamination: if a model has already seen the test questions or prompts, its results cannot be trusted as a true measure of capability or safety.
To address this, Google DeepMind is introducing the world's first double-blind evaluation of a proprietary, frontier-class AI model. External evaluations are kept inside a cryptographically secure 'box' where they cannot be accessed by the model beforehand. The pilot is conducted in partnership with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, and will test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment.
Google says it already uses a broad range of internal evaluations, but also relies on external partners such as specialized research labs, civil society, and national AI Safety and Security Institutes (AISIs) to identify blind spots. The concern is that as AI models become more capable, the risk of them having 'peeked' at test questions increases, which can artificially inflate scores and undermine trust among policymakers, researchers, and enterprises.
While zero-logging protocols and rigorous contractual safeguards have long protected the confidentiality of external test prompts, adding technical and cryptographic safeguards is described as a major step forward in secure model evaluation. The post also teases a deeper explanation of how double-blind evaluations work.