Capability evaluation is the systematic measurement of what an AI model can do, especially the elicitation and assessment of potentially dangerous capabilities such as autonomous replication, cyber-offence, or assistance with weapons. It combines benchmarks, structured tasks, and adversarial elicitation (including red-teaming) to establish upper bounds on model behaviour under best-effort prompting and tooling. Results feed safety cases and trigger the thresholds defined in responsible scaling and preparedness frameworks.

Overview

  • Capability evaluation asks not whether a model is safe by default but what it could be made to do given strong prompting, fine-tuning, scaffolding, and tools.
  • Because under-elicitation can hide latent capability, evaluators invest in adversarial elicitation, agentic harnesses, and expert-designed task suites.
  • Evaluations are calibrated to thresholds: crossing a capability level triggers heightened safeguards under preparedness and responsible-scaling commitments.

Key aspects

  • Dangerous-capability focus: cyber, biological, autonomy, and persuasion domains receive targeted assessment.
  • Elicitation rigour: best-effort prompting, tool access, and fine-tuning probes establish robust upper bounds.
  • Agentic tasks: multi-step autonomous tasks test planning, tool use, and self-improvement.
  • Threshold mapping: results are tied to defined capability levels that gate deployment decisions.
  • Reproducibility: standardised task suites and scoring support comparison across models and labs.

Applications

  • Pre-deployment risk assessment for frontier models.
  • Triggering safeguards and deployment gates in responsible scaling and preparedness frameworks.
  • Informing governance, disclosure, and third-party auditing.
  • Tracking capability trends across model generations.

Provenance