Capability forecasting is the practice of predicting the future capabilities of AI systems before they are built or deployed, typically by extrapolating from scaling laws, benchmark trends, and historical progress. It aims to anticipate when models will reach particular performance thresholds so that safety, governance, and deployment decisions can be made proactively. Forecasts are inherently uncertain because of emergent behaviour and discontinuous jumps in capability.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:ScalingLaws))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:BenchmarkTrendAnalysis))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:EmergentCapabilities))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:ExpertElicitation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:UncertaintyQuantification))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:PredictionMarkets))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:ThresholdEstimation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:BiologicalAnchorsModel))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:hasPart ai:TestTimeComputeForecasting))
Dependency Relationships
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:requires ai:ScalingLaws))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:requires ai:ModelEvaluation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:requires ai:EvaluationBenchmarks))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:requires ai:ComputeBudget))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:requires ai:TrainingData))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:requires ai:FoundationModel))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:dependsOn ai:DangerousCapabilityEvaluation))
Capability Relationships
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:RiskAssessment))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:AISafety))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:CatastrophicRiskAssessment))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:AIRegulation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:HumanOversight))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:AnticipatoryGovernance))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:enables ai:ResponsibleScalingPolicy))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:supports ai:AIGovernance))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:supports ai:ExportControl))
Implementation Relationships
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:implements ai:BiologicalAnchorsModel))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:implements ai:ExpertElicitation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:implements ai:PredictionMarkets))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:implements ai:BenchmarkTrendExtrapolation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:uses ai:ScalingLaws))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:uses ai:ModelEvaluation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:uses ai:RedTeaming))
Reduction Relationships
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:reducesTo ai:ScalingLaws))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:reducesTo ai:ThresholdEstimation))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:reducesTo ai:UncertaintyQuantification))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:partOf ai:AIGovernance))
SubClassOf(ai:CapabilityForecasting
ObjectSomeValuesFrom(ai:partOf ai:AISafety))
About
Capability forecasting emerged as a recognised sub-discipline of AI Safety and AI Governance in response to the practical question: when will AI systems become capable enough to cross safety-relevant thresholds, and what can policymakers and safety engineers do to prepare?
The discipline draws its technical foundations from empirical Scaling Laws research:
-
The OpenAI Kaplan et al. (2020) scaling laws paper established smooth power-law relationships between training compute and held-out language modelling loss across five orders of magnitude.
-
The DeepMind Hoffmann et al. (2022) Chinchilla paper corrected earlier over-parameterised models, establishing the compute-optimal training formula (equal scaling of model parameters and training tokens).
-
Together these relationships provide the quantitative backbone for extrapolating: if loss follows L(C) ≈ A/C^α across many orders of magnitude, then given planned training compute budgets, loss at future training runs can be predicted.
-
If downstream task performance is well-calibrated against loss, downstream capability can be forecast from planned training runs before those runs are executed.
The practical difficulty is that “downstream capability” at the task level does not always scale smoothly with loss:
-
Emergent Capabilities appear as discontinuous jumps on specific benchmarks at sufficient scale, defying smooth extrapolation from training loss metrics.
-
Wei et al. (2022) documented more than 100 emergent abilities in large language models — from multi-step reasoning to chain-of-thought problem solving — appearing abruptly above threshold compute scales.
-
Schaeffer et al. (2023) countered that apparent emergence is an artefact of discrete evaluation metrics; continuous metrics reveal smooth scaling throughout.
-
This debate remains fundamental: if emergence is real and discontinuous, forecasting is inherently limited by the impossibility of predicting phase transitions; if it is metric-artefactual, reliable extrapolation is in principle achievable.
The institutional infrastructure for capability forecasting developed significantly from 2022 onwards:
-
EpochAI: a non-profit research institute funded by Open Philanthropy and Jaan Tallinn, maintains the most comprehensive public database of AI training runs, tracking compute trends, parameter counts, and benchmark performance trajectories.
-
Epoch’s direct approach: uses observed scaling laws and empirical measurements to directly predict performance improvements, complementing Ajeya Cotra’s biological anchors framework.
-
Biological anchors (Cotra 2020/2022): treats computational cost of evolution or a human lifetime as upper bounds on training compute required for transformative AI.
-
EpochAI GATE model: as of late 2025, places the median transformative AI estimate around 2033 under baseline compute scaling.
-
Samotsvety superforecasters: assigned a 28% probability of AGI by 2030 as of early 2026, following significant timeline compression in 2024–2025.
-
Metaculus community: as of early 2026, assigned 25% probability of AGI by 2029 and 50% by 2033 — compressed from a 50-year median in 2020.
-
METR (Model Evaluation and Threat Research): constructed a particularly influential empirical approach using task completion time horizons, measuring the longest autonomous task frontier models can reliably complete, then tracking how this time horizon grew across model generations.
-
METR’s March 2025 analysis found this time horizon growing exponentially from 2019 to 2025, suggesting the ability to complete one-month-long autonomous software tasks may arrive in the 2027–2030 range.
The policy use of capability forecasting has accelerated in 2024–2026:
-
AI lab responsible scaling policies (Anthropic’s RSP, OpenAI’s Preparedness Framework, Google DeepMind’s Frontier Safety Framework) all include capability threshold definitions operationalised through evaluation benchmarks.
-
Capability forecasting enables labs and regulators to project when these thresholds might be reached, informing proactive preparation for required safety measures.
-
The UK AI Security Institute (AISI) incorporated capability trend extrapolation in its Frontier AI Trends Report (December 2025), noting that cyber autonomy tasks requiring over 10 years of human experience were successfully completed by AI for the first time in 2025.
-
AISI also documented that autonomy task completion rates above one hour exceeded 40% for advanced models by mid-2025.
-
These empirical trajectory findings directly enable Catastrophic Risk Assessment by grounding regulatory threshold definitions in observed capability growth rates.
Components / Architecture
1. Scaling-law extrapolation
The primary quantitative method, fitting power-law models to observed loss-versus-compute curves:
-
Standard form: L(C) ≈ A·C^(-α), where α is the estimated scaling exponent typically in the range 0.05–0.30 for language models.
-
Uncertainty in extrapolation grows with extrapolation distance; forecast error distributions widen substantially beyond the range of observed training runs.
-
Requires accurate accounting of training Compute Budget in floating-point operations (FLOPs), careful standardisation of evaluation conditions, and domain-specific exponent estimation.
-
Multi-axis fits across model parameters (N), training tokens (D), and compute (C = 6ND for standard transformers) decompose contributions from each axis.
-
Architecture improvements (Mixtures of Experts, attention variants, improved tokenisation) can shift the power-law coefficient, introducing discontinuities into naive compute extrapolation.
-
The Chinchilla compute-optimal frontier identifies the Pareto-optimal (N, D) allocation for each compute budget, enabling forecast of the optimal model configuration at future compute scales.
2. Benchmark trend extrapolation
Tracks performance trajectories on standardised evaluation benchmarks across model generations:
-
ForecastBench (ICLR 2025) provides a systematic framework for benchmark trend extrapolation methodology.
-
Benchmark saturation is a critical constraint: by early 2025, benchmarks that were frontier challenges in early 2024 — including GPQA (graduate-level scientific reasoning) — were being saturated within months of introduction.
-
METR’s RE-Bench (Research Engineering Benchmark) and SWE-Bench (software engineering) are canonical examples: forecasts placed SWE-Bench saturation in 2026.
-
Benchmark saturation dynamics require forecasters to simultaneously project model performance improvement and benchmark discrimination lifetime — a two-dimensional forecasting problem.
-
The “Forecasting Frontier Language Model Agent Capabilities” paper (arXiv:2502.15850, 2025) demonstrated that aggregated benchmark performance can be predicted with reasonable accuracy but individual task performance cannot.
3. Expert elicitation and prediction markets
Structured aggregation of expert judgements complements quantitative extrapolation:
-
Metaculus: community forecasting aggregating predictions from thousands of registered forecasters with calibration tracking; assigned 25% probability AGI by 2029 as of early 2026.
-
Samotsvety: high-accuracy superforecasting group specialising in technology timelines; 28% probability AGI by 2030 as of early 2026.
-
Kalshi, Polymarket, Manifold: prediction markets providing incentive-aligned probability estimates for specific capability events.
-
AI Impacts survey: annual survey of machine learning researchers on capability timeline expectations.
-
RAND AI Expert Panel: structured expert elicitation using Delphi methodology for catastrophic risk probability estimation.
-
The diversity of methodology across sources, and their partial agreement and disagreement, provides calibration of forecast uncertainty.
4. Biological anchors modelling
Cotra’s biological anchors framework (2020, updated 2022) anchors transformative AI compute requirements in biological estimates:
-
Evolution anchor: number of floating-point operations performed by evolution to produce a human brain.
-
Lifetime anchor: number performed by a human brain over a lifetime of learning.
-
Neural network anchors: scale model sizes to match brain parameter counts using observed scaling law relationships.
-
The approach provides reference class forecasting — grounding uncertain AI timelines in biological facts better understood than AI R&D trajectories.
-
Key limitation: uncertainty about how computational efficiency of biological evolution relates to deep learning training efficiency spans many orders of magnitude.
5. Test-time compute forecasting
An emerging extension addressing the shift from training-compute-dominated to test-time-compute-dominated progress:
-
Chain-of-thought reasoning (Wei et al., 2022), tree-of-thought search (Yao et al., 2023), and formally verified reasoning (DeepSeek-R1, 2025; Gemini 2.5, 2025) showed that inference-time compute allocation substantially improves performance on reasoning tasks.
-
METR’s task completion time horizon framework naturally incorporates test-time compute effects because the metric measures what a model can achieve with arbitrary compute during generation.
-
Forecasting test-time compute scaling requires modelling not just model weights but inference serving infrastructure and economic feasibility of extended compute-intensive inference.
-
METR’s 2026 analysis found exponential growth in time horizon from 2019 to 2025, with preliminary model suggesting 99% AI R&D automation around 2032.
6. Threshold estimation and dangerous capabilities evaluation
Operationalises capability forecasting for safety policy:
-
Involves estimating when specific dangerous capabilities will be reached — “provides serious uplift to CBRN weapon synthesis” or “can conduct autonomous cyberattacks at nation-state level.”
-
Anthropic’s Responsible Scaling Policy defines thresholds as AI Safety Levels (ASL-2, ASL-3, ASL-4).
-
Operationalised via Dangerous Capability Evaluation benchmarks including WMDP (Weapons of Mass Destruction Proxy), HarmBench, and biological uplift assessments with domain expert consultants.
-
Capability forecasting enables labs to project when ASL-3 or ASL-4 thresholds might be approached and to prepare required safety measures in advance.
-
California SB53 (September 2025) legislated analogous framework requirements for all large frontier AI developers.
Use Cases / Major Families
1. Responsible scaling policies (RSPs)
AI labs use internal capability forecasting to set the schedule and criteria for upgrading safety measures:
-
Anthropic’s RSP defines training-run triggers (compute scale approaching a threshold) and evaluation triggers (benchmark performance approaching a dangerous-capability threshold) that activate additional safety evaluation requirements.
-
Capability forecasting enables calculation of: “given planned compute scaling, when will the next training run plausibly hit the ASL-3 threshold?” — providing lead time to prepare the required safety mitigations before deployment.
-
OpenAI’s Preparedness Framework (updated April 2025) similarly uses capability forecasts to schedule evaluations against “Critical” thresholds (capabilities enabling mass casualties or billions of dollars in economic damage).
-
Google DeepMind’s Frontier Safety Framework defines analogous capability classification and monitoring requirements.
-
Meta’s Frontier AI Safety framework follows similar structures for its Llama model family.
-
California SB53 (September 2025) effectively legislates the capability forecasting and threshold-management function that RSPs implement voluntarily.
2. Regulatory compute thresholds
-
EU AI Act: uses compute-based capability thresholds (originally 10^25 floating-point operations) as proxy for frontier model status triggering enhanced regulatory requirements.
-
Capability forecasting informs where to set these thresholds and how they should evolve as the capability frontier advances.
-
AISI has worked with DSIT to inform the forthcoming UK AI Bill’s approach to capability thresholds, drawing on Frontier AI Trends Report empirical findings and trend extrapolations.
-
NIST AI RMF: provides the US voluntary baseline for AI risk management including capability-sensitive risk classification.
3. Export control and compute governance
-
US Bureau of Industry and Security (BIS) rules restricting export of high-performance AI accelerators use compute thresholds requiring capability forecasting to calibrate.
-
Setting thresholds too high relative to capability makes controls ineffective; too low impedes legitimate research.
-
EpochAI’s compute tracking database is a key input to these calibrations, providing the most comprehensive public record of AI training compute trends.
-
International discussions on AI compute governance through the G7 Hiroshima AI Process, OECD AI Policy Observatory, and emerging multilateral AI governance bodies all reference capability forecasting.
4. Red-teaming scheduling and evaluation prioritisation
-
Safety engineering teams use capability forecasts to prioritise which capability domains require red-teaming before the next training run.
-
If a capability forecast indicates biological synthesis assistance will cross an uplift threshold within two training generations, this domain receives priority evaluation resources.
-
METR’s evaluation pipeline tracks capability growth in domains including autonomous cyberattack, biological synthesis assistance, and long-horizon autonomous action.
-
This trend data informs where to concentrate safety evaluation effort and which benchmarks need development.
5. Academic and philanthropic resource allocation
-
Open Philanthropy, Survival and Flourishing Fund, and the Arc Institute use capability forecasting as input to grant-making strategies for AI safety research.
-
Compressed capability timelines directly justify increased safety research urgency and investment.
-
The major AI safety funding surge of 2023–2025 reflects institutions acting on capability forecasts that compressed expected timelines substantially:
-
Anthropic raised $7.3B Series E (2024)
-
UK government committed £100M AI Safety Research Grant programme
-
US AI Safety Institute budget expanded
-
Open Philanthropy’s AI safety grantmaking reached approximately $200M per year by 2025
6. Insurance, financial risk, and scenario planning
-
-
Financial institutions, re-insurance companies, and national risk registries are beginning to incorporate AI capability forecasts into scenario planning.
-
AI-enabled risks to scenario planning include: rapid knowledge-work automation, sophisticated cyber fraud, technology competition acceleration, systemic financial AI risk.
-
Lloyd’s of London and Swiss Re have begun developing AI risk modelling frameworks incorporating capability forecasting as a core input, analogous to climate scenario analysis under TCFD recommendations.
-
National risk registries (UK National Risk Register, US National Risk Management Centre) are adding AI capability trajectory to their horizon-scanning inputs.
7. Scientific impact forecasting
-
An emerging use case applying capability forecasting to scientific research acceleration timelines: when will AI systems reach capability thresholds for autonomous drug discovery, materials science synthesis pathway prediction, or mathematical theorem proving?
-
AlphaFold 2 (2020) and AlphaFold 3 (2024) provided the most dramatic demonstrations of AI crossing scientific capability thresholds with transformative research impact.
-
Capability forecasting for scientific AI is methodologically distinct from safety-focused forecasting because the target is benefit (capability crossing a productivity threshold) rather than risk (capability crossing a harm threshold), but uses the same underlying scaling law and benchmark trend methodology.
Academic Context
The intellectual foundations of AI capability forecasting trace to I.J. Good’s (1965) concept of an “intelligence explosion” — the feedback loop by which a sufficiently intelligent machine could improve its own design, triggering recursive capability growth — which introduced the idea that AI capability trajectories might be discontinuous and difficult to forecast from gradual extrapolation.
Foundational scaling law work:
-
Kaplan et al. (2020, “Scaling Laws for Neural Language Models,” arXiv:2001.08361) demonstrated power-law scaling across five orders of magnitude in compute, parameter count, and data — the quantitative backbone of capability extrapolation.
-
The surprise finding that smooth power laws held so reliably motivated the subsequent scaling-law research programme.
-
Hoffmann et al. (2022, “Training Compute-Optimal Large Language Models,” arXiv:2203.15556 — Chinchilla) refined the compute allocation formula, revealing that earlier GPT-3-scale models were massively undertrained, and providing a practical framework for predicting compute-optimal model configurations at future compute budgets.
Emergence debate:
-
Wei et al. (2022, “Emergent Abilities of Large Language Models,” Transactions on Machine Learning Research) documented more than 100 task families where models gained ability to solve problems only above a threshold scale, with near-zero performance below and high performance above.
-
Schaeffer et al. (2023, “Are Emergent Abilities of Large Language Models a Mirage?”, NeurIPS 2023) argued that apparent emergence is an artefact of discrete metrics; continuous metrics reveal smooth scaling throughout.
-
Arora and Goyal (2023, arXiv:2307.15936) provided a theory of emergence grounded in the geometry of feature learning in high-dimensional spaces.
-
This debate remains unresolved and constitutes the central epistemological challenge for smooth-extrapolation capability forecasting.
Transformative AI timelines research:
-
Cotra (2020; updated 2022, Open Philanthropy) provided the first systematic probabilistic forecast of transformative AI timelines using biological anchors, with the 2022 update compressing the median timeline substantially.
-
The paper is notable for explicit uncertainty quantification and scenario modelling, establishing structured AI capability forecasting as a formal analytical practice.
-
Epoch AI’s “Literature Review of Transformative Artificial Intelligence Timelines” (2023) systematically catalogued prior forecasting approaches, their assumptions, and their disagreements.
Autonomy evaluation and forecasting:
-
Kinniment et al. (2023, arXiv:2312.11671) established foundational methodology for autonomy evaluation using realistic task scenarios; METR’s subsequent work refined and scaled this approach.
-
METR (2025) introduced the task completion time horizon as the canonical empirical metric for AI autonomy capability, finding exponential growth in time horizon from 2019 to 2025.
-
Phuong et al. (2024, arXiv:2403.13793) established the methodology for dangerous capability evaluations that bridge capability forecasting with Catastrophic Risk Assessment, defining evaluation protocols for autonomous replication, cyberattack capability, and CBRN uplift.
-
“Forecasting Frontier Language Model Agent Capabilities” (arXiv:2502.15850, 2025) synthesised benchmark trend extrapolation and autonomy forecasting, demonstrating that aggregated benchmark performance can be predicted but individual task performance cannot.
-
ForecastBench (Dahl et al., ICLR 2025) provided standardised methodology and evaluation criteria for AI forecasting capability, enabling systematic comparison of forecasting approaches.
UK academic contributions:
-
Cambridge Centre for the Study of Existential Risk (CSER): applies structured expert elicitation to AI risk timelines and has contributed to UK government capability assessment methodology.
-
Oxford Future of Humanity Institute (closed 2024, with key researchers dispersing to Anthropic, new independent institutes, and academia): produced foundational early work on AI timelines under Nick Bostrom, Toby Ord, and Carl Shulman.
-
Toby Ord’s “The Precipice” (2020) provided the most systematic public-facing analysis of AI existential risk probability estimates, including capability forecasting inputs.
-
Oxford’s Global Priorities Institute continues AI policy and long-term risk research following FHI’s closure.
Current Landscape (2026)
The capability forecasting landscape in 2026 is characterised by three defining features: rapid benchmark saturation, a paradigm shift from training-time to test-time compute, and growing institutional formalisation of forecasting as a governance input.
Benchmark saturation acceleration:
-
Benchmarks posing frontier challenges in early 2024 — GPQA Diamond, AIME competition mathematics, PhD-level chemistry reasoning — were largely saturated by leading models within twelve to eighteen months.
-
This saturation race has driven development of increasingly difficult evaluation suites as of mid-2026:
- METR’s RE-Bench: research engineering tasks requiring weeks of expert work
- Long-Horizon Tasks: tasks taking humans days to months
- FRONTIER-Bench: frontier scientific research tasks requiring graduate-level expertise
-
The Frontier AI Trends Report (AISI, December 2025) noted that by mid-2025, advanced models could complete software autonomy tasks taking a human at least one hour in over 40% of cases.
-
The first model to successfully complete expert-level cyber tasks requiring more than ten years of human professional experience was evaluated in 2025.
Test-time compute paradigm shift:
-
Catalysed by o1-series models (OpenAI, 2024), DeepSeek-R1 (2025), and Gemini 2.5 Ultra (2025), which achieve task performance substantially above what pre-training loss alone would predict.
-
Complicated training-compute-based capability forecasting by adding a second scaling dimension not captured in pre-training loss curves.
-
METR’s time-horizon forecasting framework adapts to this by measuring task completion capability empirically rather than deriving it from loss metrics.
-
Requires comprehensive empirical evaluation of each new model generation rather than pure extrapolation, increasing the cost and frequency of required assessments.
-
METR’s February 2026 analysis suggested that under current trends, 99% AI R&D automation may arrive around 2032.
Forecast estimate current state (early 2026):
-
EpochAI GATE model: median transformative AI date approximately 2033 under baseline compute scaling, earlier under accelerated or breakthrough scenarios.
-
Samotsvety superforecasters: 28% probability of AGI by 2030, revised substantially earlier following 2024–2025 capability jumps.
-
Metaculus: 25% probability of AGI by 2029, 50% by 2033 — compressed from a 50-year median in 2020.
-
Kalshi contributors: approximately 40% probability of OpenAI achieving AGI by 2030 (January 2026).
-
Polymarket: 9% probability of AGI by 2027 (early 2026).
Regulatory incorporation:
-
California SB53 (September 2025): requires all large frontier AI developers to publish frameworks describing dangerous capability assessment and threshold management.
-
EU AI Act progressive implementation: mandatory framework publication requirements made previously internal capability forecasting methodology externally visible and subject to public scrutiny.
-
UK AI Bill (expected Parliamentary introduction 2026–2027): expected to include statutory capability evaluation requirements, drawing on AISI’s evaluation methodology and capability trend data.
-
The DSIT committed in January 2025 to establishing AISI as a statutory body with permanent institutional footing for capability assessment.
UK Context
The UK has positioned capability forecasting as a core function of frontier AI governance through a combination of institutional investment, evaluation methodology development, and regulatory integration.
AI Security Institute (AISI):
-
The primary UK government capability assessment institution, established November 2023 as AI Safety Institute, rebranded to AI Security Institute February 2025.
-
Has conducted systematic capability evaluations across more than thirty frontier models since establishment.
-
Published the most comprehensive public capability trend assessment by any government body: the Frontier AI Trends Report (December 2025).
-
Coverage includes dangerous capability trajectories in: autonomous cyber operation, biological synthesis assistance, long-horizon task autonomy, and manipulation and persuasion.
-
Evaluation frameworks used: Inspect (open-source Python evaluation harness), InspectCyber (cyber capability evaluation), ControlArena (control protocol testing).
-
DSIT committed in January 2025 to establishing AISI as a statutory body, with the UK AI Bill expected to include statutory capability evaluation requirements.
Cambridge academic infrastructure:
-
Centre for the Study of Existential Risk (CSER): produces academic research on AI risk timelines and risk assessment methodology; applies structured expert elicitation and scenario analysis; contributed written evidence to Parliamentary committees on AI risk.
-
Leverhulme Centre for the Future of Intelligence (CFI): addresses capability forecasting in the context of societal impact analysis, examining when AI systems reach thresholds relevant to labour market disruption, democratic integrity, and information ecosystem health.
Oxford research capacity:
-
Former Future of Humanity Institute (closed 2024): produced the key early academic work on AI timelines under Nick Bostrom, Toby Ord, and Carl Shulman; Toby Ord’s “The Precipice” (2020) provided the most systematic public-facing analysis of AI existential risk probability including capability forecasting inputs.
-
Global Priorities Institute (Oxford): continues AI policy and long-term risk research following FHI’s closure.
Northern English and Scottish universities:
-
Edinburgh’s machine learning research group contributes to uncertainty quantification methods and Bayesian approaches to extrapolation under model uncertainty — directly applicable to capability forecast confidence interval estimation.
-
Manchester’s Data Science Institute contributes to benchmark design and evaluation methodology.
-
The UKRI Trustworthy Autonomous Systems (TAS) Hub (led by Nottingham) produced research on safety assurance for autonomous systems applicable to capability assessment frameworks.
National AI research infrastructure:
-
The Alan Turing Institute, as the UK’s national AI research body, has published analyses of AI capability trends through its Centre for Emerging Technology and Security (CETaS).
-
CETaS produced the International AI Safety Report 2026 — a landmark synthesis drawing on contributions from researchers across thirty countries, with major findings on capability trajectories and their safety implications.
-
The UK’s Long-Term Resilience Centre has published specific recommendations for how the UK AI Bill should incorporate capability forecasting into its regulatory framework, recommending a “preparedness framework” structure analogous to Anthropic’s RSP.
Future Directions (2026–2030)
Automated benchmark generation and evaluation: The benchmark saturation problem — evaluation suites being exhausted by frontier models faster than they can be constructed — motivates automated benchmark generation methods that can produce novel, calibrated evaluation tasks at scale. Language-model-assisted benchmark creation (using AI to generate new evaluation tasks in validated formats) and adversarial evaluation design (creating tasks specifically targeting current model weaknesses) are active research directions. The goal is evaluation infrastructure that can maintain discriminatory power indefinitely, rather than requiring constant manual benchmark development. Key challenges include ensuring that AI-generated benchmark tasks are valid, not contaminated by training data, and at the appropriate difficulty level — all of which are easy to verify for human-constructed benchmarks but harder for automated generation.
Integrating test-time compute into scaling forecasts: The transition to test-time compute as a primary capability driver requires new forecasting frameworks that model both pre-training compute and inference budget jointly. Theoretical work on the computation-capability relationship under extended inference-time compute (including Markov chain Monte Carlo search, tree-of-thought, and tool-augmented reasoning) is needed to provide principled scaling laws for the inference-time compute dimension. Empirical data from the o1/o3/Gemini 2.5 generation is beginning to enable systematic study of test-time scaling relationships. METR’s 2026 analysis suggesting 99% AI R&D automation by approximately 2032 reflects these accelerated trajectories when test-time compute scaling is incorporated alongside training-time scaling.
Dangerous capability threshold calibration: As capability evaluations mature, the specific thresholds defined in RSPs and regulatory frameworks require empirical calibration against actual harm potential. The question “does capability X provide serious uplift to CBRN weapon development?” requires domain expert input (from biosecurity specialists, chemists, cyber security researchers) combined with empirical evaluation methodology. WMDP and the biological uplift assessments conducted by AISI and Anthropic represent the current state of this calibration; the 2027–2030 period is likely to see significant refinement of threshold definitions as evaluation methodology matures and as models demonstrate more complex and novel dangerous capabilities not anticipated in initial threshold definitions.
Forecast aggregation and calibration science: The proliferation of capability forecasting sources — scaling law extrapolation, biological anchors, expert elicitation, prediction markets, empirical task measurement — creates an aggregation problem: how should these sources be combined to produce well-calibrated probability distributions over capability milestones? The Forecasting Research Institute and Epoch AI are developing systematic aggregation methodologies, drawing on Bayesian model averaging and superforecasting ensemble methods. Well-calibrated aggregate forecasts are essential for governance use cases where decision-makers need reliable confidence intervals rather than point estimates. Structured deliberation — the process of sharing information and reasoning across forecaster groups before individual forecasts are made — has been shown to improve calibration in human forecasting teams and is being explored as a methodology for AI capability forecast aggregation.
Agent capability forecasting: As AI increasingly operates in agentic settings — orchestrating multi-step workflows, taking actions with real-world consequences, using tools and APIs — capability forecasting must extend from single-model performance on isolated tasks to agent-system performance on extended, multi-step, real-world tasks. This requires new evaluation frameworks, new theoretical models of agent capability scaling, and new empirical data on how agent capabilities scale with component model capabilities. METR’s task time horizon framework is the leading approach to date but requires substantial extension for compound AI system architectures that combine multiple models with tool use, memory, and external knowledge retrieval.
Societal impact capability thresholds: Alongside safety-critical capability thresholds (CBRN uplift, cyberattack autonomy), governance bodies increasingly need capability forecasts for socio-economic thresholds: when will AI automate 10%, 50%, 90% of software engineering tasks? When will AI systems be capable of autonomous scientific research at PhD level or above? When will AI-generated media reach the point where synthetic and authentic content are indistinguishable at population scale? These thresholds are relevant to labour market policy, education system design, and democratic integrity, requiring capability forecasting to engage with social science methodology alongside its technical toolkit.
Benchmark Datasets and Evaluation Infrastructure
Training compute tracking:
-
EpochAI Training Compute Database: the most comprehensive public database of AI training runs, tracking compute (FLOPs), parameter counts, training data volume, training duration, and reported benchmark performance. Used as the primary data source for scaling law exponent estimation and compute trend extrapolation. As of mid-2026, the database contains over 800 documented training runs spanning 2010–2026.
-
MLPerf Training Benchmarks: standardised training compute benchmarks enabling comparison of hardware efficiency across AI accelerator generations, supporting capability-per-compute-dollar trend estimation.
Downstream performance benchmarks (capability indicators):
-
MMLU (Massive Multitask Language Understanding): 57-subject academic knowledge benchmark; used as a general capability proxy; widely saturated by frontier models by early 2025.
-
GPQA Diamond: graduate-level scientific reasoning across biology, chemistry, and physics; designed to be PhD-level difficult; being saturated by leading models by mid-2025.
-
AIME (American Invitational Mathematics Examination): competition mathematics; saturated by leading reasoning models by early 2025.
-
HumanEval / SWE-Bench Verified: software engineering capability; SWE-Bench Verified forecast to be saturated in 2026 with strong elicitation.
-
BIG-Bench Hard: collection of 23 “beyond the scaling law” tasks requiring reasoning; used to test whether models exhibit genuine reasoning versus pattern matching.
-
HELM (Holistic Evaluation of Language Models): multidimensional evaluation framework measuring accuracy, calibration, robustness, fairness, and efficiency across 42 scenarios.
Autonomy and agent capability benchmarks:
-
METR RE-Bench (Research Engineering Benchmark): tasks requiring weeks of skilled research engineering work; designed to remain challenging beyond 2026; forecast saturation in 2027.
-
SWE-Bench: software engineering tasks derived from real GitHub issues; measures end-to-end capability to resolve complex, multi-file bugs; forecast saturation late 2026.
-
METR Task Time Horizon metric: not a benchmark in the traditional sense but an empirically measured capability indicator tracking the longest autonomous task a model can reliably complete; used by METR as the primary capability forecasting metric.
-
AgentBench: standardised evaluation of agent capabilities across operating system interaction, database management, knowledge graph navigation, and web browsing tasks.
-
WebArena: benchmark for autonomous web navigation and task completion on realistic browser environments.
Dangerous capability evaluation benchmarks:
-
WMDP (Weapons of Mass Destruction Proxy): measures knowledge relevant to CBRN weapon development across biological, chemical, and cybersecurity domains; used by Anthropic, AISI, and METR for ASL-3 evaluation.
-
Virology Capabilities Test (VCT): expert-virologist-validated assessment of biological uplift; used by AISI for biological dangerous capability evaluation.
-
CyberSecEval / InspectCyber: automated cybersecurity capability evaluation; part of AISI’s Inspect framework and Meta’s CyberSecEval suite.
-
HarmBench: standardised automated red teaming infrastructure with 400+ harmful behaviour test cases across 6 semantic categories.
Forecasting capability benchmarks:
-
ForecastBench (ICLR 2025): dynamic benchmark of AI forecasting capability using real prediction questions with known outcomes; used to evaluate whether AI systems themselves can contribute to capability forecasting.
-
AI Forecasting Accuracy Database: maintained by Forecasting Research Institute tracking human and AI forecasting accuracy on technology questions over time.
Key Terminology
-
Scaling Law: A power-law relationship between model training resources (compute, parameters, data) and performance metrics (typically cross-entropy loss), forming the primary quantitative tool of capability forecasting.
-
Compute Budget (C): The total floating-point operations used in training a model, typically measured in FLOPs (floating-point operations) or FLOP equivalents; the primary axis along which training-time capability forecasts are made.
-
Emergent Capability: A task performance that appears only above a threshold model scale, potentially posing a discontinuous prediction challenge for smooth-extrapolation forecasting.
-
Uplift: The capability increase an AI system provides to a malicious actor beyond what they could achieve through existing means without AI assistance; the primary harm-relevance criterion in dangerous capability evaluation.
-
AI Safety Level (ASL): Anthropic’s classification system for AI models based on their dangerous capabilities, from ASL-1 (no meaningful capability uplift) through ASL-4 (autonomous civilisational-scale harm potential), operationalising capability thresholds as deployment triggers.
-
Transformative AI (TAI): AI with sufficiently broad capabilities and autonomous reasoning to cause transformative changes in economic productivity, scientific progress, and societal organisation, typically used as the target event in long-range capability forecasting.
-
Benchmark Saturation: The state in which a model’s performance on a benchmark reaches the performance ceiling (maximum possible score), ending the benchmark’s discriminatory value for capability assessment.
-
Test-Time Compute: Compute used during inference (generation) rather than training, including extended chain-of-thought reasoning, tree-of-thought search, and tool use; increasingly important as a capability growth driver independent of pre-training scale.
-
Biological Anchors: A reference class forecasting approach to AI timelines that grounds compute requirement estimates in biological facts about brains and evolution, providing an alternative to pure trend extrapolation.
-
Task Completion Time Horizon: METR’s empirical metric measuring the longest autonomous task a frontier model can reliably complete without human intervention, used as a leading indicator of AI autonomy capability and tracked as a proxy for capability level across model generations.
Key Research Institutions and Organisations
Technical forecasting and evaluation organisations:
-
EpochAI: non-profit research institute maintaining the most comprehensive public database of AI training runs and benchmark performance; publishes systematic capability trend analyses and the GATE transformative AI timeline model. Primary source of empirical data for scaling-law-based capability forecasting.
-
METR (Model Evaluation and Threat Research): pioneered the task completion time horizon metric and the most comprehensive public autonomy evaluation suite; developed the ControlArena framework for agentic AI control testing; conducts pre-deployment evaluations for major AI labs under information-sharing agreements with AISI.
-
Samotsvety Forecasting Group: high-accuracy team of superforecasters applying structured probabilistic reasoning to AI capability timelines; produces regularly updated AGI timeline estimates that are among the most systematically calibrated public forecasts available.
-
Metaculus: open forecasting platform aggregating AI capability timeline predictions from thousands of registered forecasters with calibration tracking; maintains the most widely-cited community AI timeline estimates.
-
Forecasting Research Institute: supports rigorous forecasting methodology development including structured deliberation and calibration for AI capability predictions.
-
Apollo Research: specialises in evaluation of deception, situational awareness, and self-preservation behaviours in frontier models — capabilities most relevant to long-horizon capability forecasting.
AI laboratory internal forecasting teams:
-
Anthropic RSP team: maintains internal capability forecasting infrastructure for ASL threshold monitoring; publishes voluntary RSP updates with threshold definitions informed by capability trend projections.
-
OpenAI Preparedness team: conducts pre-deployment capability evaluation and maintains the Preparedness Framework threshold classifications; updates capability risk scores as frontier models are developed.
-
Google DeepMind Safety team: maintains the Frontier Safety Framework and conducts dangerous capability evaluations; publishes academic work on evaluation methodology including Phuong et al. (2024).
Government and policy bodies:
-
UK AI Security Institute (AISI): the primary government body conducting systematic capability trend assessment; publishes the Frontier AI Trends Report; develops and maintains the Inspect evaluation framework.
-
US NIST AI Safety Institute: US counterpart to AISI; co-evaluates frontier models under bilateral information-sharing agreement; developed the NIST AI RMF framework used for capability-sensitive risk classification.
-
EU AI Office: responsible for enforcing EU AI Act general-purpose AI model provisions including dangerous capability evaluation; coordinates with AI Safety Institute network on evaluation methodology.
-
Long-Term Resilience Centre (UK): publishes policy recommendations for incorporating capability forecasting into UK AI legislation; primary policy advocacy organisation for preparedness-framework approaches.
Academic centres:
-
Cambridge Centre for the Study of Existential Risk (CSER): applies structured expert elicitation to AI risk timelines; has contributed to UK government’s capability assessment methodology; operates the MPhil in Global Risk and Resilience.
-
Leverhulme Centre for the Future of Intelligence (CFI, Cambridge): addresses capability forecasting for societal impact thresholds beyond safety-critical capabilities.
-
Alan Turing Institute / CETaS: national UK AI research body; published the International AI Safety Report 2026 and maintains the most systematic UK-based AI capability trend analysis.
-
Oxford Global Priorities Institute: continues long-horizon AI risk research following FHI’s 2024 closure; contributes to AI existential risk probability estimation and policy implications.
Research and Literature
- Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361.
- Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendrycks, D., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, J., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J.W., Vinyals, O., and Sifre, L. (2022). Training Compute-Optimal Large Language Models. arXiv:2203.15556.
- Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Miculivicius, V., Narang, S., Chowdhery, A., Le, Q., Sutton, C., and Fedus, W. (2022). Emergent Abilities of Large Language Models. Transactions on Machine Learning Research.
- Schaeffer, R., Miranda, B., and Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage? NeurIPS 2023.
- Cotra, A. (2020, updated 2022). Forecasting Transformative AI with Biological Anchors. Open Philanthropy Technical Report.
- METR (2025). Task-Completion Time Horizons of Frontier AI Models. METR Technical Report, March 2025. metr.org/time-horizons.
- Kinniment, M., Shlegeris, B., Denison, C., Perez, E., Schneider, J., Lim, S., Bai, Y., and Irving, G. (2023). Evaluating Language-Model Agents on Realistic Autonomous Tasks. arXiv:2312.11671.
- Dahl, M., et al. (2025). ForecastBench: A Dynamic Benchmark of AI Forecasting Capability. ICLR 2025.
- Phuong, M., Aitchison, M., Catt, E., Dziedzic, A., Farquhar, S., Elinas, P., Fokkens, A., Fortunato, M., Freeman, R., Hawkins, R., He, C., Holte, C., Hutter, F., Iyer, S., Koëter, S., Loos, A., Maddison, C.J., March, T., Mathewson, K., Muelling, K., Ramapuram, J., Rowley, J., Sherwood, D., Siddiqui, A., Slatyer, T., Spoelstra, E., Stasinopoulos, M., Sullivan, K., Szepesvari, C., and Voth, J. (2024). Evaluating Frontier Models for Dangerous Capabilities. DeepMind Technical Report. arXiv:2403.13793.
- Epoch AI (2025). GATE: Generative AI Timeline Estimation. Epoch AI Research. epoch.ai/topics/future-of-ai.
- Epoch AI (2023). Literature Review of Transformative Artificial Intelligence Timelines. epoch.ai/blog.
- OpenAI (2025). Preparedness Framework (Updated April 2025). OpenAI Policy Document.
- Anthropic (2023; updated 2024). Responsible Scaling Policy. Anthropic Policy Document.
- Google DeepMind (2024). Frontier Safety Framework. Google DeepMind Technical Report.
- UK AI Security Institute (2025). Frontier AI Trends Report. AISI, December 2025. aisi.gov.uk/frontier-ai-trends-report.
- Zhang, J., and Li, E. (2025). Emergency Response Measures for Catastrophic AI Risk. arXiv:2511.05526.
- Ruan, Y., et al. (2024). Identifying the Risks of LM Agents with an LM-Emulated Sandbox. arXiv:2309.15817.
- Samotsvety Forecasting Group (2026). AI Timelines Update: January 2026. AI Futures Model blog.aifutures.org.
- Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
- Ord, T. (2020). The Precipice: Existential Risk and the Future of Humanity. Bloomsbury.
- Bommasani, R., et al. (2021). On the Opportunities and Risks of Foundation Models. arXiv:2108.07258.
- Ganguli, D., et al. (2022). Predictability and Surprise in Large Generative Models. ACM FAccT 2022.
- Muennighoff, N., et al. (2023). Scaling Data-Constrained Language Models. NeurIPS 2023.
- DeepSeek (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. DeepSeek Technical Report.
- Arora, S. and Goyal, A. (2023). A Theory for Emergence of Complex Skills in Language Models. arXiv:2307.15936.
- METR (2026). A Simpler AI Timelines Model Predicts 99% AI R&D Automation in ~2032. METR Notes, February 2026. metr.org/notes.
- CETaS / Alan Turing Institute (2026). International AI Safety Report 2026. Centre for Emerging Technology and Security.
- Shulman, C. (2022). Forecasting AI Progress: Perspectives on Timelines to Transformative AI. Open Philanthropy Discussion.
Why Capability Forecasting Is Hard — Core Challenges
The field faces several structural challenges that distinguish it from other quantitative forecasting domains:
1. The emergence problem:
-
Smooth loss scaling does not guarantee smooth capability scaling on specific tasks.
-
Phase transitions in benchmark performance (the “emergent abilities” phenomenon) create cliff-edge risks that scaling law extrapolation misses.
-
Debate about whether emergence is real or metric-artefactual (Schaeffer et al., 2023) remains unresolved.
-
Even if individual capabilities scale smoothly, capability thresholds relevant to governance (such as “provides meaningful CBRN uplift”) may appear to be discontinuous.
2. The benchmark saturation treadmill:
-
Frontier models exhaust evaluation benchmarks faster than new ones can be validated and deployed.
-
This creates episodic blind spots where capability growth continues but measurement tools are temporarily saturated.
-
The resolution — harder benchmarks — introduces its own biases as task difficulty selection affects apparent capability growth rates.
-
Automated benchmark generation (using AI to create AI evaluations) creates circularity concerns about evaluation validity.
3. The test-time compute complication:
-
Training-time compute is the traditional forecasting variable; test-time compute is increasingly the capability driver.
-
Test-time compute scaling is harder to forecast because it depends on inference economics (GPU cost per token), prompting strategy, and task-specific search depth — variables that change independently of model weights.
-
The same model weights can exhibit dramatically different capability levels depending on inference settings, making single-point capability assessments misleading.
4. The elicitation gap:
-
Benchmarks measure capabilities under specific elicitation conditions (prompt format, number of few-shot examples, chain-of-thought encouragement).
-
True model capability may substantially exceed benchmark-measured capability under optimal elicitation.
-
METR explicitly excluded high-elicitation scenarios from some forecasts, noting the estimates might be too conservative.
-
As elicitation techniques improve, measured capabilities increase even without model changes, complicating time-series trend analysis.
5. The out-of-distribution generalisation uncertainty:
-
Scaling laws are fitted on in-distribution evaluation; they may not generalise to novel capability domains or task types not represented in training benchmarks.
-
Future capability milestones may require qualitatively different cognitive operations that current scaling law exponents do not anticipate.
-
Architecture discontinuities (e.g., the shift from pure language model to multimodal model to agent system) can break historical scaling law fits.
6. The recursive improvement discontinuity:
-
If AI systems become sufficiently capable of AI research itself, capability trajectories may accelerate in ways that no backward-looking extrapolation can predict.
-
METR’s simpler AI timelines model (February 2026) projects 99% AI R&D automation around 2032 — which would create a recursive improvement dynamic qualitatively different from the linear extrapolation paradigm.
-
At that point, capability forecasting based on human-driven compute scaling would need to be replaced by models of AI-recursive capability growth, which are inherently more uncertain.
-
This represents the ultimate limit of scaling-law-based capability forecasting methodology: beyond the AI R&D automation threshold, the extrapolation machinery itself breaks down.
-
The governance implication is that safety measures and regulatory frameworks need to be established before this threshold is reached, since regulatory frameworks designed for human-pace AI development may be inadequate for AI-recursive capability growth.
Relationship to Adjacent Concepts
-
vs. Capability Evaluation: Capability evaluation measures what current models can do; capability forecasting projects what future models will be able to do. Evaluation is backward-looking (empirical); forecasting is forward-looking (predictive). Both use the same benchmark infrastructure, but forecasting adds the extrapolation layer. Forecasting also involves probability distributions over uncertain futures, whereas evaluation produces point measurements.
-
vs. Emergent Capabilities: Emergent capabilities are the principal challenge for smooth capability forecasting — they represent the potential for non-smooth, discontinuous capability jumps that extrapolation from prior trends cannot predict. Capability forecasting must treat emergence as a fundamental uncertainty rather than a predictable phenomenon. The debate between Wei et al. (2022) and Schaeffer et al. (2023) on whether emergence is “real” or metric-artefactual is thus directly consequential for capability forecasting methodology.
-
vs. Scaling Laws: Scaling laws are the primary technical tool of capability forecasting for the training-time compute dimension. Capability forecasting extends scaling laws by applying them to governance-relevant downstream task performance rather than abstract loss metrics. It also extends beyond scaling laws to incorporate test-time compute, architectural innovations, and elicitation effects that scaling laws do not capture.
-
vs. Catastrophic Risk Assessment: Capability forecasting enables catastrophic risk assessment by projecting when dangerous capability thresholds will be reached; catastrophic risk assessment consumes capability forecast outputs to set risk management timelines. They are complementary components of anticipatory AI governance — capability forecasting provides the “when” and catastrophic risk assessment provides the “how bad.” The AISI Frontier AI Trends Report combines both: empirical capability trend measurement (capability forecasting) and assessment of which capabilities matter for harm (catastrophic risk assessment).
-
vs. Responsible Scaling Policy: RSPs operationalise capability forecasting for AI lab governance, defining the specific thresholds and safety measures that capability forecasts are used to schedule preparation for. An RSP without capability forecasting cannot anticipate when thresholds will be crossed; capability forecasting without an RSP lacks the institutional mechanism to act on forecasts.
-
vs. AI Timelines: AI timelines research is the long-horizon variant of capability forecasting focused on transformative or general AI milestones decades out; near-term capability forecasting focuses on specific benchmark saturations and dangerous capability thresholds on a months-to-years timescale. The two traditions share forecasting methodology (scaling law extrapolation, expert elicitation, biological anchors) but differ in time horizon, purpose, and the specificity of the capability targets they project.
-
vs. Model Evaluation: Model evaluation provides the empirical data inputs for capability forecasting — benchmark scores, autonomy task completion rates, elicitation-controlled capability measurements. Capability forecasting takes these empirical inputs and extrapolates forward. As evaluation frameworks improve (harder benchmarks, better elicitation protocols), the quality of capability forecasts improves correspondingly.
-
vs. AI Safety Level: AI Safety Levels (ASLs) are the operationalisation of dangerous capability thresholds that capability forecasting projects. Capability forecasting tells you “when will we reach ASL-3?”; the ASL framework tells you “what does reaching ASL-3 mean and what must be done?” The two concepts are tightly coupled in frontier AI lab governance: RSPs reference ASL thresholds; capability forecasting projects ASL timelines; safety engineering prepares the required mitigations; evaluation gates verify threshold status before deployment.
-
vs. Foundation Model: Foundation models are the subject of capability forecasting — the systems whose future capabilities are being projected. Capability forecasting helps determine which foundation model training runs require dangerous capability evaluation before deployment, and at what point in the scaling trajectory new safety measures must be prepared.
-
vs. Expert Elicitation: Expert elicitation is one of the primary methods within capability forecasting, providing probabilistic timeline estimates from domain experts as an alternative or complement to quantitative scaling law extrapolation. Forecasting organisations including EpochAI and Samotsvety specifically synthesise expert elicitation with quantitative extrapolation to produce calibrated forecasts that neither method could achieve alone.