🤖 DeepMind Conducts First Double-Blind Evaluation of a Frontier Model

External organizations — Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons — tested Gemini Flash Lite on confidential benchmarks. The evaluators did not see the model weights, and Google did not see the prompts: computations were performed in Google Cloud Confidential Space's GPU enclave.

🌍 This provides the industry with independent auditing of frontier models without disclosing weights and benchmarks. Signs of test data leakage were found in approximately half of the 31 models evaluated.

👤 The scheme is not 'trustless': it assumes that the cloud provider and chip manufacturer do not collude. This is a step toward independent evaluations, not a complete guarantee.

Source 1: https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/ Source 2: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/piloting-the-worlds-first-double-blind-ai-evaluations/double-blind-evaluations-technical-report.pdf