Pre-deployment evaluation is the structured assessment of an AI model’s capabilities, limitations, and risks conducted before it is released or deployed into production use. It typically combines capability benchmarking, red-teaming, and safety testing to detect dangerous capabilities, misuse potential, or unexpected behaviour ahead of exposure to real users. Pre-deployment evaluation is a central mechanism of frontier AI governance frameworks, including commitments made at the Bletchley Declaration.

Provenance