A structured evaluation and oversight framework produced by the UK AI Safety Institute (AISI) to assess the catastrophic risks posed by frontier AI models prior to and following their public release. The framework specifies pre-deployment testing protocols, thresholds for dangerous capability uplift, and post-deployment monitoring obligations that developers of frontier models are expected to satisfy. It represents the UK government’s primary technical instrument for operationalising AI safety commitments made at the Bletchley Park AI Safety Summit of November 2023.
Content
- The AISI was established in September 2023, ahead of the November 2023 Bletchley Park AI Safety Summit, where leading AI developers and twenty-eight governments signed the Bletchley Declaration committing to cooperative frontier AI risk assessment. AISI’s framework emerged from that mandate, drawing on earlier work by Anthropic, DeepMind, and OpenAI on responsible scaling policies and their own internal capability thresholds.
- The framework’s technical architecture centres on a tiered evaluation process. Initial automated capability evaluations use standardised benchmarks to detect dangerous capability uplift; positive signals trigger structured human red-team exercises designed to elicit harmful outputs. Evaluation domains prioritise catastrophic risk categories—CBRN uplift, autonomous replication, and severe cyber-offence capabilities—before broader societal harm assessments. AISI publishes methodology notes and, selectively, capability evaluation findings.
- Operationally, AISI has established bilateral memoranda of understanding with major frontier AI developers, gaining pre-release model access for evaluation. It has also collaborated with the US AI Safety Institute (AISI’s American counterpart established by the Biden administration’s October 2023 Executive Order) and the OECD on evaluation standardisation. The framework is designed to evolve as capabilities advance and evaluation techniques improve.
- By 2024–2025, the framework has conducted evaluations of several major model releases and has published findings indicating that, while no model tested has exceeded dangerous capability thresholds, margins are narrowing in certain cyber-offence domains. The UK’s decision in 2024 to rename AISI as the “AI Security Institute” reflected a broadening remit, and the framework continues to be updated to address agentic and multi-model system configurations that were not originally in scope.