Google DeepMind pilots the first 'double-blind' evaluation of a proprietary frontier model with Singapore's AISI and MLCommons
On Aug 27, 2026 Google DeepMind described what it calls the world's first double-blind evaluation of a proprietary frontier-class model. A Gemini Flash-Lite model was tested inside a cryptographically attested confidential-computing environment on Google Cloud, so the evaluators (Singapore's AI Safety Institute, OpenMined, AVERI and MLCommons) never saw the model weights and Google never saw the test prompts. The aim is to prevent benchmark contamination without forcing labs to hand over weights.
Key facts
- Partners: Singapore AI Safety Institute, OpenMined, AVERI and MLCommons (DeepMind)
- Uses Google Cloud Confidential Space to cryptographically verify that the evaluation data and the proprietary model stay private to their owners; The New Stack reports the run used an NVIDIA H100 confidential GPU instance
- Model tested: 'a Gemini Flash Lite model' (DeepMind); The New Stack identifies it as Gemini 2.5 Flash Lite, run against reserve prompts from MLCommons' AILuminate safety benchmark and a separate private prompt set for the Singapore context
- Removes the old trade-off in which evaluators either hand over test prompts (contamination risk) or labs hand over weights (IP risk); DeepMind says it matters most for sensitive cyber and government evaluations
What happened
External safety testing usually forces a trade-off. Either the evaluator gives the lab its test prompts, which can then leak into training and inflate scores, or the lab gives the evaluator its model weights, which it does not want to do. Google DeepMind and four partners ran an evaluation where neither happened. Inside a hardware-isolated, attested Google Cloud environment, the evaluators' confidential benchmarks ran against a Gemini Flash-Lite model: "The evaluator cannot see the Gemini model weights, and Google cannot see the evaluator's test prompts." DeepMind published its methodology and findings with the announcement.
Why it matters
As AI safety institutes and outside auditors gain a formal role (e.g. California SB 813's independent assessors, OpenAI's Sept 22 principles for third-party assessments), a trusted way to test closed models on secret benchmarks becomes basic infrastructure. The pilot used a small model; it is not yet shown to scale to frontier models.
Changelog
- 2026-09-30: created (official-blog audit)
Related events
Sources (3)
- officialGoogle DeepMind: Piloting the world's first double-blind AI evaluations
- pressThe New Stack: Google found a way to test Gemini without seeing the questions
- pressMLQ: Google DeepMind pilots sealed external tests for Gemini model
id: 2026-08-27-deepmind-double-blind-ai-evaluations · updated 2026-09-30 · open in the interactive timeline