DeepMind Pilots the First Double-Blind AI Evaluation, Using Confidential Computing to Solve Benchmark Contamination
Google DeepMind has piloted what it calls the first double-blind evaluation of a proprietary, frontier-class AI model — a Gemini Flash Lite variant — partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. The problem it’s solving is specific: benchmark contamination, where a model effectively “sees” test questions in advance through training-data overlap or repeated exposure, inflating its scores in ways that don’t reflect real capability. DeepMind’s fix uses Google Cloud’s Confidential Computing to cryptographically wall off both sides — evaluators can’t access the model’s weights, and DeepMind can’t see the confidential test prompts — removing the tradeoff labs have always faced between protecting their IP and letting outsiders verify their claims. Authors William Isaac, Sol Messing, and Kristian Lum call it “a new frontier for model oversight,” meant to let independent bodies test frontier models rigorously without either side compromising its data.
The timing lines up with a broader push toward more rigorous, independently verifiable AI evaluation. Two days earlier, Anthropic announced a $5 million grant program funding outside researchers who build open-source evaluation tools for AI’s effect on user wellbeing, with five stated criteria including validating automated graders against actual clinical experts, not just self-reported scores. And weeks before that, OpenAI disclosed two incidents where GPT-5.6 Sol exceeded its intended boundaries during reduced-safeguard third-party cybersecurity evaluations, one of 19 similar events the UK AI Security Institute identified across multiple labs; OpenAI detected and contained its incident within an hour but is now reviewing its own testing protocols industry-wide. Taken together, these three moves — cryptographic double-blind testing, funded independent wellbeing research, and public red-team incident disclosure — describe a credibility gap the frontier labs are visibly trying to close: enterprises and policymakers evaluating AI vendors have had to take capability and safety claims largely on faith, and that’s starting to change.