AGI timelines are structured probabilistic forecasts estimating when artificial general intelligence — AI matching or exceeding human cognitive ability across most economically valuable tasks — might be achieved. Such forecasts aggregate expert surveys (AI Impacts ESPAI 2022/2023), compute-scaling extrapolations (biological anchors model, Cotra 2020/2022), benchmark-progress curves (Metaculus, Manifold Markets), and Bayesian models of research milestones to produce probability distributions over future dates. As of mid-2026, median estimates cluster in the 2030–2047 range depending on methodology. Timelines are inherently uncertain and contested, and they materially influence AI safety research prioritisation, alignment investment, compute governance thresholds, and public policy.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:ExpertSurveyComponent))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:BiologicalAnchorsModel))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:BenchmarkProgressCurve))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:PredictionMarketAggregate))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:ProbabilityDistributionOverDates))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:AGIDefinitionCriteria))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:UncertaintyQuantification))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:hasPart ai:MilestoneDecomposition))

Dependency Relationships

SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:requires ai:ArtificialGeneralIntelligenceDefinition))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:requires ai:ScalingLaws))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:requires ai:CapabilityBenchmarks))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:requires ai:ComputeScalingData))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:requires ai:ExpertElicitation))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:requires ai:BayesianAggregationMethod))

Capability Relationships

SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:AIGovernanceDecisions))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:SafetyInvestmentPrioritisation))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:ComputeThresholdPolicy))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:AlignmentResearchPlanning))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:PublicPolicyFormulation))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:LabourMarketPreparation))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:ExistentialRiskAssessment))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:enables ai:InternationalTreatyNegotiation))

Implementation Relationships

SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:implements ai:ProbabilisticForecasting))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:implements ai:ExpertElicitationProtocol))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:implements ai:BenchmarkExtrapolation))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:implements ai:ComputeScalingModel))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:implements ai:PredictionMarketAggregation))

Reduction Relationships

SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:ProbabilityDistribution))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:CapabilityForecast))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:ResearchMilestoneAssessment))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:ExpertConsensusEstimate))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:MilestoneDecompositionSchedule))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:BayesianBeliefDistribution))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:PolicyPrioritySignal))
SubClassOf(ai:AGITimelines
  ObjectSomeValuesFrom(ai:reducesTo ai:ScenarioConditionalAnalysis))

About

  • AGI timelines constitute one of the most consequential and methodologically contested subdisciplines within AI research. The central question — when will AI systems achieve general cognitive capability matching or exceeding that of humans across economically valuable tasks? — sits at the intersection of Deep Learning research, forecasting methodology, Cognitive Science, political economy, and existential risk analysis. Unlike most scientific forecasting problems, AGI timelines forecasting faces the unusual challenge that the phenomenon being forecast is deeply uncertain in its definition, is produced by rapidly accelerating research that defies stable trend extrapolation, and has consequences so large that the methodology of forecasting is itself politically contested by actors with strong incentives to produce particular outcomes.

  • The core methodological challenge is that any timeline forecast must first grapple with definition: “AGI” has no universally accepted operationalisation, and the same underlying reality can appear near or far depending on which definition the forecaster adopts. A forecaster who defines AGI as “AI that can autonomously complete any economically valuable software engineering task” might estimate AGI in 2027–2028, given observed progress on coding benchmarks. A forecaster who defines AGI as “AI that can autonomously conduct novel scientific research across all disciplines” might estimate 2040–2060. A forecaster who defines AGI as “AI that can perform every human economic task better and more cheaply” might estimate 2050–2080. And a forecaster defining AGI as “AI that poses an existential threat to human civilisation if misaligned” might not give a specific date at all, arguing that the risk threshold precedes the full economic capability threshold. This definitional variance is not merely academic: it produces survey results where participants answering “the same question” with different underlying definitions produce bimodal distributions that look, when averaged, like a single moderate estimate, when in reality there may be two distinct population subgroups with very different underlying beliefs.

  • The four conceptual targets most actively used in the forecasting community are: HLMI (High-Level Machine Intelligence, AI Impacts definition: “unaided machines can accomplish every task better and more cheaply than human workers”) — a definition that requires full economic substitutability across all sectors and consequently produces the latest estimates; TAI (Transformative AI, Open Philanthropy definition: a level of automation producing economic growth comparable to the Industrial Revolution) — a definition focused on economic impact and consequently somewhat earlier in most models than HLMI; AGI per Anthropic/OpenAI (AI systems broadly better than humans at almost everything in terms of Reasoning and task performance) — an inside-view framing that leading AI laboratories use, which typically produces estimates in the 2026–2030 range for the most aggressive insiders; and benchmark-operationalised AGI (specific milestones such as ARC-AGI grand prize at 85%+, FrontierMath >80%, or SWE-Bench >90%) — a prediction market framing that produces the earliest estimates because it requires only narrow performance thresholds rather than general economic substitutability.

  • The four principal forecasting methodologies each have distinct assumptions, characteristic outputs, and known failure modes. Direct Expert Elicitation asks practising ML researchers to give probability distributions over years until HLMI, producing the AI Impacts ESPAI series; the 2022 survey (n=738) gave 50% by 2059, the 2023 re-survey (n=2,778, conducted seven months after GPT-4’s release) compressed this to 2047 — a 12-year shift in a single year and the largest single-year update in the survey’s history; known failure modes include anchoring on recent salient events (GPT-4’s capabilities), overconfidence in the near-term extrapolation of current scaling trends, and the Dunning-Kruger problem where the most capable ML researchers (who know the most about remaining challenges) give later estimates while less experienced respondents give earlier ones. The biological anchors model (Ajeya Cotra for Open Philanthropy, 2020 and 2022) derives timeline estimates from the compute required to match human brain information processing, combined with hardware cost curve projections; it estimated a 50% chance of TAI by 2050 in 2020, updated to 2040 in 2022 after incorporating two years of observed scaling progress; its failure modes include the circularity of using brain compute as the reference threshold without justifying why that specific threshold corresponds to AGI capability, and high sensitivity to algorithmic efficiency gain assumptions. Community Prediction Markets (Metaculus, Manifold, Kalshi) aggregate thousands of individual forecasters via market mechanisms; Metaculus’ “first AGI” question median compressed from 2052 in 2020 to approximately 2033 by early 2026, and Manifold Markets estimated 47% probability of AGI before 2028 as of early 2026; failure modes include resolution criteria ambiguity (different platforms use different definitions), information cascade effects where early high-status forecasters anchor community beliefs, and the absence of meaningful financial skin-in-the-game on most platforms. Superforecasting groups (Samotsvety, Good Judgment Project) produce calibrated ensemble estimates from track-record experts; Samotsvety’s January 2026 update (8 forecasters) placed AGI probability by 2030 at approximately 28%; failure modes include over-reliance on outside-view base rates that may not account for the fundamental novelty of the AI capability trajectory, and the difficulty of applying historical forecasting track records to a problem class (unprecedented technological transition) that has no historical precedent for calibration.

    Definitional Architecture and Measurement

  • AGI timelines forecasts are composed of several interacting methodological layers, each introducing assumptions that propagate through to the final probability distribution, and understanding these layers is essential for interpreting any specific timeline estimate or comparing estimates across methodologies and research groups.

  • The Definition Layer is the single largest source of variance in timeline estimates — models otherwise identical in methodology can diverge by 15–25 years depending solely on which target definition of AGI the forecaster adopts. The four definitions in active use each reflect different theoretical frameworks and produce characteristically different outputs. HLMI (High-Level Machine Intelligence, the AI Impacts definition requiring that “unaided machines can accomplish every task better and more cheaply than human workers”) demands full economic substitutability across all sectors and all task types, from physical labour to creative work to scientific research, producing the latest median estimates in the surveyed literature (AI Impacts 2023: median 2047 among ~2,778 researchers). TAI (Transformative AI, the Open Philanthropy definition specifying a level of AI automation producing economic growth comparable to the Industrial Revolution or to the agricultural revolution) is more precise about economic impact and less demanding about task-by-task substitutability, producing moderately earlier estimates in biological anchor models (50% by 2040 in the 2022 Cotra update). Inside-view capability thresholds as used by Anthropic and OpenAI (“AI systems broadly better than humans at almost everything in terms of Reasoning and task performance”) are intentionally vague but tend to be interpreted as requiring only broad superiority rather than full economic substitutability, producing earlier estimates from inside-view informants. Benchmark-operationalised definitions — specific performance thresholds on existing evaluation suites such as ARC-AGI at 85%+, FrontierMath at 80%+, or SWE-Bench at 90%+ — produce the earliest, most specific estimates by requiring only narrow demonstrated performance rather than general economic substitutability across all domains, and they have the advantage of providing prediction market resolution criteria with unambiguous verification procedures. Resolution criteria for prediction market questions using benchmark definitions are publicly available and binding, reducing but not eliminating definitional ambiguity — though market operators regularly debate whether specific AI systems “count” as satisfying stated criteria, revealing remaining conceptual ambiguity even in seemingly precise benchmark-based definitions.

  • The Evidence Aggregation Layer synthesises heterogeneous inputs with different update frequencies, reliability profiles, and susceptibility to systematic bias. Expert surveys (AI Impacts ESPAI 2016 with n=80, 2022 with n=738, 2023 with n=2,778; RAND Corporation 2024 survey; LEAP Wave 8 conducted April-May 2026 with approximately 100 domain experts and 200 superforecasters) produce the most authoritative estimates in terms of domain expertise coverage but are conducted only annually or less frequently, making them systematically lagged relative to capability developments — the 2023 survey captured researcher reactions to GPT-4 but was conducted seven months after its release, already partially outdated given Claude 3, Llama 2, Gemini Pro releases. Compute scaling models (biological anchors: Cotra 2020, 2022; Chinchilla optimal scaling laws: Hoffmann et al. 2022; algorithmic efficiency tracking: Epoch AI database) provide mechanistic grounding in physics and economics rather than expert opinion, enabling more principled extrapolation than pure expert elicitation, but require strong and potentially invalid assumptions about the relationship between compute and capability. Benchmark trajectory extrapolation uses observed performance curves on evaluation suites — MMLU progression from 25% (GPT-3 base) to 90%+ (GPT-4 and successors) over 2020–2024, HumanEval coding from near-zero to 90%+ over the same period, ARC-AGI from near-zero in 2019 to ~50% by late 2025 — as basis for extrapolating when further performance thresholds will be reached. Prediction market prices from Metaculus, Manifold Markets, and Kalshi aggregate thousands of individual forecasters in continuous real-time updating, incorporating new information with minimal latency, though subject to cascade effects from influential early forecasters and to the problem that most market participants have no meaningful financial incentive to be calibrated rather than strategically positioned. Superforecasting group estimates from Samotsvety (8 core forecasters, January 2026 update: 28% probability of AGI by 2030) and Good Judgment Project panels apply rigorous calibration methodology and forecaster track-record weighting, providing the most epistemically disciplined but smallest-N input to timeline aggregation. Inside-view lab estimates from public statements by OpenAI leadership, Anthropic’s core views documents, DeepMind published strategy, and Meta AI research forecasts are among the most informationally rich inputs (these organisations have the most detailed knowledge of current capability trajectories) but are among the least reliable due to their enormous incentive to publish estimates supporting favourable regulatory, investment, and competitive positioning.

  • The Uncertainty Model distinguishes three qualitatively different categories of forecasting uncertainty, each requiring different treatment in a rigorous timeline model. Fundamental (deep) uncertainty refers to unknown unknowns — potential algorithmic breakthroughs (e.g., a new Learning paradigm beyond transformer-based architectures and Scaling Laws-based training) or fundamental barriers (computational complexity limits on certain cognitive tasks, generalisation barriers that do not yield to additional compute) that cannot be estimated from current evidence because their existence and nature are unknown. This is the type of uncertainty that caused virtually all pre-2017 AGI timeline models to systematically underestimate the trajectory of Large Language Models — the transformer architecture and the scaling paradigm were not legible as breakthrough directions until after the fact. Epistemic uncertainty — structured uncertainty from insufficient data — includes the question of whether current Scaling Laws trajectories will continue (the Chinchilla laws describe current compute-efficient training regimes but do not specify whether the same relationships hold at 100× or 1000× current compute budgets), the question of what the relevant compute threshold for AGI capability actually is (the biological anchors estimate is internally consistent but its reference class of human brain computation is not validated as the correct threshold), and the question of how to extrapolate benchmark performance curves to general capability (models that saturate one benchmark may not generalise to other benchmarks at all). Aleatory uncertainty arises from the stochastic nature of research progress — which groups solve which problems, funding decisions, regulatory actions, key individual contributions, and serendipitous breakthroughs — which are in principle unpredictable even with perfect knowledge of current state.

  • The Output Format of a rigorous AGI timeline forecast is a full probability distribution over years — not a point estimate — enabling policymakers and safety researchers to read off specific percentile thresholds relevant to their planning horizon. Metaculus early 2026 community aggregate, for example, puts 10% probability before 2027, 25% before 2030, 50% before 2033, 75% before 2038, and 90% before 2045. Summary statistics P10/P25/P50/P75/P90 are tracked across successive survey rounds as a quantitative record of how estimates shift in response to new evidence, with the P50 (median) receiving most attention in public reporting but the P10 (optimistic tail) being most policy-relevant for near-term AI Safety preparation. The Update Mechanism for AGI timelines forecasts is the subject of explicit methodological research: the Forecasting Research Institute’s LEAP series tracks update patterns at 6-month intervals across expert and superforecaster populations, measuring the direction and magnitude of updates in response to specific capability developments, with LEAP Wave 8 (April-May 2026) showing that expert and superforecaster estimates are converging around the 2028–2033 window for near-AGI milestones. The emerging Milestone Decomposition methodology breaks holistic “AGI” forecasting into sequenced intermediate milestones — “AI achieves 90% on MMLU across all language groups,” “AI solves ARC-AGI novel tasks at 85%+,” “AI autonomously completes a 1-week software engineering project without human guidance,” “AI publishes a peer-reviewed paper without co-authorship” — each independently forecastable with shorter horizons and unambiguous resolution criteria, reducing the definition-layer noise that inflates disagreement in surveys asking directly about “AGI.”

    Capability Milestones 2022–2026

  • The period 2022–2026 has seen dramatic compression of median AGI timeline estimates across all methodologies, driven by qualitatively new capability demonstrations that were not anticipated by most pre-2022 timeline models and that repeatedly crossed thresholds previously considered markers of human cognitive distinctiveness.

  • The ChatGPT release (November 2022) demonstrated instruction-following Large Language Models capability at mass-market scale for the first time, producing the first widespread public encounter with frontier model capability and prompting rapid upward revision of near-term AI impact estimates among commentators, investors, and policymakers who had not previously engaged with frontier model research. Within days of release, it was clear that instruction-following models trained via RLHF on large-scale corpora could handle a far wider range of tasks than previous chatbots and code assistants, including complex multi-step Reasoning, creative writing, code generation, and multi-domain question answering. The subsequent six months saw major timeline compression among commentators — not yet in formal surveys, which require months to conduct and analyse — as the magnitude of the capability step from previous generation models became apparent.

  • The GPT-4 release (March 2023) produced the most concentrated set of AGI-relevant capability demonstrations to that date: performance at the 90th percentile among human test-takers on the US bar exam (a complex multi-step Reasoning, legal knowledge, and essay task); >90th-percentile performance on the LSAT, SAT, GRE verbal reasoning, AP Biology, AP Chemistry, and AP Environmental Science exams; medical licensing exam performance at or above passing threshold; and dramatic improvements on MMLU, HumanEval coding, and mathematical Reasoning benchmarks. GPT-4 was the direct catalyst for the AI Impacts 2023 expert survey’s 13-year median compression (from 2059 to 2047 vs 2022 baseline), with respondents explicitly citing GPT-4’s professional-examination performance as updating their beliefs about how far current-generation models had progressed beyond their 2022 expectations. It also catalysed the first mainstream policy responses, including the EU AI Act’s compute thresholds and the US Executive Order 14110.

  • The AlphaProof and AlphaGeometry 2 milestone (July 2024) — gold-medal equivalent performance at the 2024 International Mathematical Olympiad — demonstrated formal mathematical Reasoning at elite human level in a domain that had been specifically cited by sceptics as a stronghold of human cognitive capability resistant to language model approaches. The result updated many expert timelines meaningfully because formal mathematics (proof verification, novel theorem construction, combinatorial problem-solving under time constraints) requires compositional multi-step Reasoning over abstract structures that does not obviously benefit from the statistical pattern matching underlying Large Language Models, yet the Google DeepMind systems achieved it. This milestone is regularly cited in 2025–2026 prediction market discussions as evidence that the set of cognitive tasks “immune” to AI progress is shrinking faster than previous outside-view historical extrapolations would suggest.

  • The o1 and o3 reasoning model series from OpenAI (September 2024 through early 2025) introduced extended chain-of-thought Reasoning within the inference process itself — models that “think through” problems across hundreds of internal reasoning steps before producing output, rather than producing single-pass completions. This architectural shift produced substantial improvements on ARC-AGI benchmark performance (previously regarded as highly resistant to Large Language Models because it requires novel task adaptation rather than memorisation), with o3 reportedly achieving performance in the 75-85% range on ARC-AGI evaluation sets. The demonstration that extended Reasoning compute at inference time could substitute for certain kinds of generalisation challenged the view that ARC-AGI required fundamentally different capabilities from language modelling.

  • The LEAP Wave 8 survey results (April-May 2026) representing the most recent systematic cross-population forecasting data, revealed significant heterogeneity across respondent types: the expert-domain-researcher median for “AI achieving 80% success rate on randomly sampled 8-hour software engineering tasks” was 2030; the superforecaster median was 2028; and the general-public sample gave a much later estimate of 2037 — a 9-year gap between the professionally calibrated forecast and the public estimate, illustrating the significant public communication gap around AGI timeline evidence. The 8-hour software engineering task benchmark was specifically chosen by LEAP as a milestone requiring sustained autonomous planning, debugging, and self-correction across a realistic time horizon — a capability threshold substantially beyond current frontier model autonomous agency performance as of mid-2026.

  • Counter-evidence and timeline extension signals from 2025–2026: The dominant narrative of accelerating timeline compression was partially qualified in early 2026 by several convergent observations. Progress on fully autonomous research tasks — AI completing open-ended scientific projects with no human guidance over multi-week horizons — proved substantially slower than optimistic extrapolations from 2023–2024 coding task performance suggested. Novel ARC-AGI task families (constructed specifically to prevent solution by memorisation or near-retrieval) continued to present significant difficulty to o3-class models, suggesting that genuine novel-task generalisation remained a bottleneck even for extended-inference systems. Evidence of diminishing capability returns per unit of additional training compute — relative to the trajectory observed from GPT-3 to GPT-4 — without new architectural innovations led several prominent forecasters including Metaculus community aggregate trackers to slightly extend their median estimates in early 2026, partially reversing the extreme compression of 2023–2025.

    Use Cases / Major Families

  • AGI timelines serve as foundational inputs to several distinct decision-making contexts, each operationalising probability distributions differently.

  • AI Safety Research Prioritisation: The primary practical application of AGI timelines research is informing how urgently AI Safety Research and Alignment Research need to advance relative to capability development. Short timelines — AGI within 5–10 years — constrain the available window for longitudinal alignment experiments that require years of iterative development and empirical validation, for interpretability methods that need to scale to frontier model sizes (currently requiring specialised hardware and months of compute per experiment), and for building the institutional oversight capacity that would be required to deploy near-AGI systems safely, including regulatory infrastructure, international coordination mechanisms, and public deliberation processes. Long timelines — AGI in 30+ years — afford significant latitude for foundational safety research to mature without urgency pressure, enabling longitudinal studies of model alignment properties, iterative improvement of oversight techniques, and democratic participation in shaping AI governance frameworks. Open Philanthropy’s annual giving process — one of the largest philanthropic funders of AI Safety Research — explicitly uses biological anchor and survey-based timeline estimates to calibrate research area prioritisation across a multi-hundred-million-dollar annual budget. The consequentialist logic underlying this application is straightforward: the expected value of safety research investments scales with the probability that the research will be usable before misaligned AGI is deployed, which scales inversely with timeline length (shorter timelines mean safety research must produce usable results faster, reducing the value of approaches requiring long development cycles) — making precise timeline estimation a direct input to research portfolio optimisation at organisations like Anthropic, DeepMind safety teams, MIRI, ARC, and government bodies like AISI.

  • Compute Governance Thresholds: AGI timelines research provides the theoretical foundation for compute-based AI regulation, an area of significant active policy development globally. The EU AI Act’s 10^25 FLOP training compute threshold for high-risk General Purpose AI (GPAI) system designation directly embeds compute-to-capability assumptions derived from biological anchors-style models — the threshold was set at a level expected to encompass only the most powerful frontier systems based on 2023 compute data, but its adequacy depends critically on whether compute-to-capability relationships remain stable as algorithmic efficiency improves. The US Executive Order 14110 (October 2023) compute reporting threshold for AI models uses the same biological anchors-derived framework, requiring developers of frontier models above the threshold to report to the Department of Commerce on safety evaluations and dual-use potential. AGI timelines research informs regulators on where these thresholds should be set, how frequently they need to be reviewed as hardware costs fall (lowering the compute required to train a given capability level and thus “graduating” more systems into the regulated tier), and whether capability metrics from benchmark evaluations provide more policy-relevant triggers than raw compute proxies. The Oxford AIGI White Paper on AI Thresholds (August 2025, n=47 expert respondents) found that current regulatory compute thresholds have already been exceeded by deployed systems not exhibiting the transformative capabilities the thresholds were intended to capture — suggesting that compute-only thresholds are insufficiently discriminating — and recommends capability-based rather than compute-based triggers, with specific benchmark performance thresholds on safety-relevant evaluations as the preferred approach.

  • Investment and Venture Capital:

    • Timeline estimates calibrate investment time horizons: 3–5 year product roadmaps appropriate for long-timeline worlds; 1–3 year windows for short-timeline worlds
    • Key decision: whether portfolio companies in automation, robotics, and software have 3 years or 15 years before general-purpose AI disrupts their markets
    • Andreessen Horowitz, Sequoia AI investment theses incorporate timeline assumptions explicitly
    • Labour Market Disruption analysis for insurance, pension funds, and sovereign wealth funds uses timeline estimates to calibrate pace of white-collar automation
  • Benchmark Milestone Tracking: The benchmark milestone tracking application treats AGI timelines not as abstract probability distributions over years but as concrete forecasting targets around specific, unambiguously resolvable capability demonstrations — a methodological shift that makes timeline forecasting substantially more empirically tractable and useful for near-term planning. The LEAP Wave 8 survey (April-May 2026) specifically asked participants to forecast AI achieving 80% success on randomly sampled 8-hour software engineering tasks, receiving expert median predictions of 2030, superforecaster median predictions of 2028, and a general public estimate of 2037. Samotsvety’s January 2026 AGI aggregate included 49% probability that the ARC-AGI grand prize (1 trillion by January 2027 — demonstrating that serious forecasters now treat near-term milestone timing, not just distant AGI arrival, as the active frontier of timelines research. The shift to milestone decomposition also enables systematic measurement of forecaster track records at shorter horizons, providing calibration data for improving long-horizon AGI forecasting quality.

  • Public Policy, Regulation, and International Governance:

    • UK AI Safety Institute explicitly monitors Frontier AI capability trajectories; AISI Frontier AI Trends Report (2025) presents empirical measurements as inputs to timeline estimates and threshold calibration
    • EU AI Act mandates biennial review of GPAI thresholds, creating institutionalised demand for systematic timeline input every two years
    • International treaty negotiations anticipated to accelerate after Bletchley-Seoul-Paris summit sequence require agreed capability assessment methodologies
    • US-UK AI Safety Partnership (announced at Bletchley Park, November 2023) established collaborative evaluation frameworks incorporating timeline-informed threshold setting
  • Workforce and Economic Transition Planning: AGI timeline uncertainty creates a fundamental planning challenge for governments, institutions, and individuals making long-horizon decisions about education, retraining investment, pension fund allocation, and infrastructure development. The IMF World Economic Outlook 2024 report on AI and the labour market addressed this by conditioning projections on two explicit timeline scenarios — “near AGI” (AGI achieved in the 2030–2035 window) and “far AGI” (AGI delayed to 2050 or later) — producing scenario-conditional labour market displacement estimates that can be read as probability-weighted. The near-AGI scenario implies acute Labour Market Disruption affecting 40–60% of white-collar occupations within a decade of AGI deployment, primarily through automation of information-processing and Reasoning-intensive tasks currently performed by knowledge workers, creating a structural adjustment challenge faster than historical transitions (the industrial revolution unfolded over decades; near-AGI scenario compressed transition in 10 years). The far-AGI scenario allows gradual structural adjustment through natural workforce turnover and retraining capacity. Central banks and treasury departments in the G7 are increasingly developing scenario-conditioned analyses of fiscal and monetary policy for the AI transition period, with the UK Treasury’s 2025 AI economy workstream explicitly incorporating timeline-conditional projections for the UK’s financial and professional services sectors — both disproportionately exposed to Large Language Models-mediated automation in the near-timeline scenario. This application makes AGI timelines research directly consequential for the £80+ billion UK public spending on social safety net programmes that would need to respond to automation-driven unemployment.

  • Investment and Venture Capital:

    • Timeline estimates calibrate investment time horizons: 3–5 year product roadmaps appropriate for long-timeline worlds; 1–3 year windows for short-timeline worlds

    • Key decision: whether portfolio companies in automation, robotics, and software have 3 years or 15 years before general-purpose AI disrupts their markets

    • Andreessen Horowitz, Sequoia AI investment theses incorporate timeline assumptions explicitly

    • Labour Market Disruption analysis for insurance, pension funds, and sovereign wealth funds uses timeline estimates to calibrate pace of white-collar automation

      Academic Context

  • The systematic study of AGI timelines has a surprisingly long history, with informal forecasts predating the modern machine learning era and the field evolving from speculation to rigorous probabilistic methodology over six decades.

  • The historical trajectory of AGI forecasting begins with the first AI overestimation at the 1956 Dartmouth Conference, where Minsky, McCarthy, and colleagues predicted that “every aspect of learning or any other feature of intelligence can be so precisely described that a machine can be made to simulate it” within a single summer’s research programme — a prediction that proved wrong by at least fifty years. The subsequent AI winters of the 1970s and 1980s taught researchers the danger of near-term optimism, producing a generation of AI researchers who systematically underestimated progress rates during the deep learning era. Ray Kurzweil’s “The Singularity Is Near” (2005) revived compute-extrapolation arguments for AGI by 2029, popular but methodologically informal — Kurzweil’s extrapolation assumed continuous exponential hardware scaling and algorithmic efficiency improvement without accounting for potential capability discontinuities or data limitations. The AI Impacts ESPAI 2016 survey (n=80) established the first rigorous survey methodology, producing a median HLMI estimate of approximately 2060. ESPAI 2022 (n=738, pre-GPT-4) updated this to a median of 2059. ESPAI 2023 (n=2,778, conducted seven months after GPT-4’s release in March 2023) compressed this dramatically to 2047 — the largest single-year shift in the survey’s history, driven by the demonstration that Large Language Models could pass professional-level examinations in law, medicine, and finance.

  • The biological anchors framework developed by Ajeya Cotra for Open Philanthropy (2020, substantially updated 2022) is the most influential methodological innovation in the AGI timelines field. The framework asks a precisely specified question: how much training compute would be required to produce a model with the same information-processing capacity as a human brain, measured in floating-point operations across a learning lifetime? Using neuroscientific estimates of the human brain’s computational capacity (approximately 10^15 FLOPs per second sustained processing, roughly 10^24 FLOPs of synaptic updates across a human learning lifetime from birth to adult expertise), combined with projections of GPU hardware cost curves (following historical trends showing roughly 2× cost reduction per year for equivalent compute) and training efficiency improvement trends, the model produces probability distributions over the year in which sufficient training compute becomes economically accessible for frontier AI research organisations. The 2020 version estimated a 50% chance of transformative AI by 2050 and a 15% chance by 2030. The 2022 update, incorporating two years of observed algorithmic efficiency gains and scaling law discoveries (including the Kaplan et al. and Chinchilla laws), revised these to a 50% chance by 2040 and a 30% chance by 2030 — a decade of compression from a single model update, illustrating the model’s high sensitivity to algorithmic efficiency gain assumptions. Tom Davidson’s subsequent “take-off speeds” model (Open Philanthropy, 2023) extends Cotra’s framework to estimate not just when AGI might be achieved but how rapidly the economic transformation after AGI achievement would unfold — whether “fast take-off” (transformative economic impact within months of AGI deployment) or “slow take-off” (gradual impact over years and decades) depending on AGI deployment constraints, physical capital complementarity requirements, and regulatory responses.

  • The field draws on several disciplinary traditions that each contribute distinct conceptual tools. Superforecasting methodology, developed by Philip Tetlock over decades of forecasting tournament research and summarised in “Superforecasting: The Art and Science of Prediction” (Tetlock and Gardner, 2015), provides the calibration framework applied to AI by Good Judgment Project panels and Samotsvety. The key Superforecasting insight most relevant to AGI timelines is that calibration — giving 70% probability only to events that occur 70% of the time — matters more than resolution (making bold confident predictions), and that most domain experts are systematically overconfident. Bayesian Inference provides the philosophical framework for treating prior beliefs as probability distributions and updating them on evidence via likelihood functions, enabling formal integration of new capability milestones into existing timeline models. Prediction Markets theory (Hanson 2003, 2007) provides the mechanism-design foundation for aggregation platforms like Metaculus and Manifold Markets — properly incentivised market prices should converge to well-calibrated probabilities, making market data a useful complement to survey-based elicitation. Cognitive Science and the philosophy of mind contribute the conceptual apparatus for defining the target: what counts as “intelligence,” “understanding,” “general capability,” and “reasoning” at the level required for AGI? François Chollet’s “On the Measure of Intelligence” (2019), which introduced the ARC-AGI benchmark, made the most precise operationalisation of this question available: general intelligence is efficient novel task adaptation rather than accumulated task performance, and any system that achieves high performance purely through memorisation of training patterns has not demonstrated general intelligence even if it performs well on standard benchmarks.

  • The Epoch AI research organisation (founded 2021) has built the most systematic empirical infrastructure supporting AGI timelines forecasting. Their compute database tracking training runs across published models from 2010 to 2026 documents a roughly 4× annual increase in training compute for frontier models over the decade — though with some deceleration since 2022 as the largest training runs (Gemini Ultra, GPT-4, Claude 3, Llama 3 405B) become extremely expensive and further scaling faces both economic and hardware constraints. Their algorithmic efficiency tracking finds approximately 2× efficiency improvement per year across various ML tasks (meaning the same performance can be achieved with half the compute each year), implying that “effective compute” (actual compute × efficiency) doubles roughly every 6–9 months. This effective compute doubling rate, significantly faster than hardware-alone Moore’s Law, underpins the optimistic end of AGI timeline estimates. Epoch AI’s “Literature Review of Transformative AI Timelines” (2023) synthesised all credible existing timeline models and concluded that TAI by 2030–2060 is the modal range, with significant probability mass outside that window in both directions — an honest acknowledgement that the range of reasonable estimates spans three full decades.

  • Key empirical phenomena that have substantially updated timeline models in the 2020–2026 period include: the Scaling Laws discovery (Kaplan et al., 2020; Hoffmann et al. Chinchilla paper, 2022) demonstrating predictable cross-entropy loss curves as a function of compute and data, enabling quantitative extrapolation of training compute requirements for future capability thresholds; Emergent Behaviour research (Wei et al., 2022) documenting qualitative capability jumps at specific compute scales — in-context learning, chain-of-thought Reasoning, arithmetic, and other capabilities appearing discontinuously at certain model sizes — creating uncertainty about whether AGI-level capabilities might emerge discontinuously from a single training run rather than gradually from years of incremental improvement; chain-of-thought prompting (Wei et al., 2022) demonstrating that Large Language Models could reason across multiple steps without architectural changes by simply formatting examples with intermediate Reasoning steps, substantially expanding apparent model capability with no new training; and the IMO 2024 gold-medal equivalent performance of AlphaProof and AlphaGeometry 2 (Google DeepMind, July 2024), which demonstrated elite formal mathematical Reasoning — a domain previously considered a gold standard for human cognitive capability — suggesting that at least some forms of high-level human intellectual performance are now within reach of frontier AI systems.

    Current Landscape (2026)

  • As of June 2026, the AGI timelines landscape is characterised by unusually wide disagreement, rapid update frequency, and intense policy salience.

  • Current estimate landscape:

    • Metaculus community aggregate (early 2026): 25% probability by 2029, 50% by 2033
    • Samotsvety Forecasting (January 2026): approximately 28% probability of AGI by 2030
    • AI Impacts ESPAI 2023 (October 2023, n=2,778): median HLMI by 2047 (13 years earlier than 2022 survey)
    • LEAP survey (April–May 2026): expert median for 80% success on 8-hour software tasks: 2030; superforecasters: 2028; public: 2037
    • Anthropic policy documents (March 2025): expects “powerful AI systems” by late 2026 or early 2027
    • Manifold Markets (early 2026): 47% probability of AGI before 2028
  • 2025–2026 partial reversal trend: timeline compression from 2022–2025 partially reversed in early 2026 among some groups:

    • Metaculus, Dario Amodei, and analyst Peter Wildeford all pushed estimates slightly later compared to 2025 positions
    • Attributed to: slower-than-expected progress on autonomous research tasks; continued difficulty of novel ARC-AGI task families; evidence of diminishing capability returns per unit compute without new architectural innovations
  • UK AI Safety Institute empirical monitoring:

    • AISI renamed from AI Safety Institute to AI Security Institute (2024), operating as DSIT directorate
    • AISI measurements (2025): AI models completing expert-level cyber tasks at >50% success rates, up from 10% in early 2024 — rapid trajectory illustrating why timelines compress
    • AISI Frontier AI Trends Report (2025) provides empirical capability data directly feeding into timeline estimates and UK policy threshold calibration
  • EU AI Act compute threshold reassessment:

    • 10^25 FLOP training compute threshold (set 2024) already exceeded by systems without producing qualitatively transformative capabilities
    • Oxford AIGI White Paper (August 2025, n=47 experts) recommends capability-based rather than compute-based regulatory triggers
    • EU AI Act mandatory biennial review mechanism creates the institutional pathway for such revisions — precisely the governance recalibration AGI timelines research is designed to enable
  • Prediction market evolution: platforms now track AGI-adjacent milestones rather than just “AGI achieved”: FrontierMath >80%, ARC-AGI grand prize, 8-hour autonomous software tasks, AI-authored peer-reviewed paper — enabling more granular, resolvable tracking than the binary “AGI question”

    UK Context

  • The UK occupies a distinctive position in AGI timelines research and policy, with the world’s first dedicated government frontier AI safety evaluation body.

  • AI Safety Institute (AISI):

    • Established by the Conservative Government in 2023; operating as a directorate of DSIT
    • World’s first government body for Frontier AI safety evaluation with explicit capability assessment mandate
    • AISI Frontier AI Trends Report tracks empirical model capabilities across cyber, autonomous agency, and Reasoning tasks, providing data-driven complement to survey-based timeline forecasts
  • University of Oxford:

    • Stuart Armstrong (Future of Humanity Institute, Oxford, now at ARC): co-authored influential work on machine intelligence research and the ethics of AGI timelines in “Intelligence Explosion: Evidence and Import”
    • Future of Humanity Institute (FHI, founded by Nick Bostrom at Oxford): primary academic home for AGI timelines and Existential Risk research before closure in 2024; intellectual legacy continues through Centre for the Governance of AI (GovAI, Oxford) and the Legal Priorities Project
    • GovAI researchers including Allan Dafoe and Markus Anderljung have published on governance implications of short AGI timelines, informing UK AI policy documents and international coordination
    • Oxford AIGI White Paper on AI Thresholds (August 2025): 47 expert survey on appropriate compute and capability thresholds; finding that current regulatory thresholds already exceeded without triggering intended transformative capability concerns
  • University of Cambridge:

    • Leverhulme Centre for the Future of Intelligence (LCFI): philosophers, computer scientists, and social scientists assessing societal implications of AGI including epistemic challenges of timeline forecasting under deep uncertainty
  • Alan Turing Institute: AI safety and governance research relevant to timelines; hosted fellowships on frontier capability assessment methodology

  • ARIA (Advanced Research and Invention Agency):

    • “Safeguarded AI” programme under programme manager David Dalrymple: targeting interpretability and verifiable alignment research
    • ARIA’s investment posture embeds an implicit short-timeline assumption by prioritising foundational alignment over longer-horizon capability research
  • Northern England:

    • University of Manchester: Advanced AI Research Group includes work on Reasoning and generalisation bearing on capability assessment; Responsible AI group produces policy-relevant analysis for UK Parliament Science and Technology Committee
    • University of Sheffield: Responsible AI group produces AI governance analysis including AGI timeline implications for Northern Powerhouse industries; connection to Labour Market Disruption research for manufacturing regions
    • University of Leeds: Economic Impact of AI research for UK cities affected by near-term automation, particularly white-collar job displacement in financial and professional services sectors
  • UK policy process: UK’s Bletchley Park AI Safety Summit (November 2023), Bletchley Declaration signed by 28 countries, established UK as international hub for frontier AI safety governance — directly incorporating AGI timeline uncertainty as a motivating concern for international coordination

    Future Directions (2026–2030)

  • Five structural changes are likely to reshape the AGI timelines field between 2026 and 2030, making the discipline more empirical, more institutionalised, and more consequentially linked to governance.

  • Operationalised milestone forecasting:

    • Shift from abstract HLMI/AGI definitions to concrete, unambiguously evaluable milestones
    • Examples: “AI system completes 1-year research project autonomously producing novel published results”; “AI achieves ARC-AGI grand prize at 85%+”; “AI designs and implements significant architectural self-improvement”
    • Reduces definitional noise inflating disagreement in current surveys — different forecasters can now have evidence-based disagreements about specific milestone timing rather than disguised disagreements about target definition
    • Platforms like Metaculus building structured decomposition tools breaking “AGI” into sequenced measurable milestones with explicit resolution criteria
  • Continuous real-time tracking:

    • Epoch AI’s compute tracking database combined with automated academic literature parsing for algorithmic efficiency signals enables continuous (rather than annual-survey-based) timeline updating
    • As new Frontier AI model releases provide capability benchmarks, they can be automatically incorporated into rolling timeline estimates, reducing latency between capability developments and forecast updates from months to days
    • Blurs the distinction between AGI timeline forecasting and frontier AI capability monitoring — the two activities converging into a single continuous assessment framework
  • Institutionalisation in regulatory processes:

    • EU AI Act mandatory biennial review of GPAI thresholds creates formal governance mechanism requiring systematic timeline input every two years
    • UK AISI frontier capability assessment mandate creates similar demand in UK context
    • International AI treaty negotiations anticipated to require agreed methodologies for capability assessment feeding into treaty compliance monitoring
    • AGI timelines research becoming an institutionalised, government-adjacent discipline with formal review processes and accountability — analogous to climate modelling in international climate policy
  • AI-assisted and self-referential forecasting:

    • Models will increasingly participate in timeline assessment by analysing research literature for capability signals, simulating research trajectories, or proposing structured Expert Elicitation questions
    • Self-referential element — AI systems forecasting AGI arrival — requires careful attention to strategic bias (labs have incentives to publish estimates supporting their funding or regulatory positioning) and to systematic underestimation (systems cannot simulate capabilities they do not yet have)
    • Frameworks for AI-assisted forecasting with Human-in-the-Loop oversight will be a necessary methodological development
  • Integration with Alignment Research roadmaps:

Provenance