An AI Model is a computational artefact — comprising a parameterised mathematical function, its learned weights, and associated configuration — that encodes patterns extracted from training data and can be applied to new inputs to generate predictions, classifications, embeddings, or generative outputs. AI models range from simple linear regressors to billion-parameter deep neural networks and constitute the core intellectual and commercial asset of modern AI systems.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:ModelWeights))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:ModelCheckpoint))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:AiModelArchitecture))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:LossFunction))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:TrainingData))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:Tokenizer))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:HyperparameterConfiguration))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:hasPart ai:EvaluationMetric))Dependency Relationships
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:requires ai:ComputeInfrastructure))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:requires ai:ModelTraining))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:requires ai:TrainingData))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:dependsOn ai:DeepLearning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:dependsOn ai:GradientDescent))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:dependsOn ai:NeuralNetwork))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:dependsOn ai:SelfSupervisedLearning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:dependsOn ai:AiChips))Capability Relationships
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:AiInference))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:GenerativeAi))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:NaturalLanguageProcessing))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:ComputerVision))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:AiAgent))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:EmergentCapabilities))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:InContextLearning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:enables ai:AgenticWorkflow))Implementation Relationships
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:implements ai:LearningAlgorithm))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:implements ai:SelfSupervisedLearning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:implements ai:ReinforcementLearning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:implements ai:TransferLearning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:supports ai:FineTuning))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:supports ai:ModelQuantization))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:supports ai:ModelDistillation))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:supports ai:RetrievalAugmentedGeneration))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:supports ai:PromptEngineering))Reduction Relationships
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:reducesTo ai:ParameterisedFunction))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:reducesTo ai:LearnedRepresentation))
SubClassOf(ai:AiModel
ObjectSomeValuesFrom(ai:reducesTo ai:ProbabilityDistribution))About
An AI model is a discrete, portable computational artefact — a parameterised mathematical function — that encodes statistical regularities extracted from large corpora of Training Data and can be applied to previously unseen inputs at AI Inference time. The concept of an “AI model” as a separable, redistributable file is historically recent: it crystallised with the emergence of the deep learning framework ecosystem (TensorFlow SavedModel format, 2016; PyTorch checkpoint conventions, 2017; ONNX cross-framework exchange, 2017) and was catalysed by the rise of the Hugging Face Hub (launched 2019), which by 2025 hosted over 1.5 million public model repositories and had become the de facto distribution channel for the field. Earlier AI paradigms — symbolic AI, Expert System, logic programming — did not produce “model artefacts” in this sense; their knowledge was encoded as rules or structured knowledge bases that were not straightforwardly transferable. The shift to learned, data-driven models constituted a fundamental reconceptualisation of what it means to build an AI system: rather than explicitly encoding human expertise as if/then rules, practitioners instead curate a dataset of examples, specify a parameterised function class and an objective, and run an optimisation procedure to extract a model that generalises from those examples. This inversion — from specification to learning, from rules to weights — is the defining epistemic move of the modern AI model paradigm.
The mathematical substrate of a modern AI model is almost always a Neural Network: a composition of parameterised linear transformations interleaved with non-linear activation functions, whose parameters are optimised via Gradient Descent over a scalar Loss Function measuring prediction error on training examples. The dominant architectural family is the Transformer Architecture (Vaswani et al., 2017), which uses self-attention mechanisms to capture long-range dependencies in sequential data and underlies virtually all frontier Large Language Models as well as most leading Computer Vision and Multimodal AI systems. The attention mechanism computes pairwise affinities between all positions in a sequence, allowing the model to selectively weight information from any position when computing its output at any other position; this capability — combined with massively parallel training on GPUs — makes transformers far more efficient to train at scale than the recurrent architectures (LSTMs, GRUs) that previously dominated sequential modelling. Alternative architectures in active research and deployment include Convolutional Neural Network for image processing tasks (where spatial locality and weight sharing yield strong inductive biases for visual data), Graph Neural Network for structured relational data (where edges encode semantic relationships between entities), state space models (Mamba, RWKV) offering O(n) rather than O(n^2) attention complexity for long sequences, and Diffusion Model architectures for generative image and video synthesis via iterative denoising from Gaussian noise. Since 2022, Mixture of Experts (MoE) architectures have become prominent at the frontier: by activating only a subset of expert sub-networks per inference call (typically 2 out of 8, 16, or more experts), MoE models achieve the representational capacity of dense models at lower per-token compute cost. Meta’s Llama 4 Behemoth (2026) uses MoE with over 2 trillion total parameters and approximately 288 billion active parameters per forward pass; Mistral Large 3 (December 2025) implements a 675B total / 41B active MoE architecture under Apache 2.0 licence.
The lifecycle of an AI model spans multiple phases that collectively constitute the AI Model Development discipline: (1) problem framing and data strategy — defining the task, collecting and curating a Training Data corpus, applying data-cleaning and augmentation pipelines, deduplicating web-crawled text, filtering for quality and safety; (2) architecture selection and initialisation — choosing or adapting an AI Model Architecture to the problem domain, initialising Model Weights via random initialisation or weight inheritance from a prior model; (3) pre-training or training — iteratively optimising Model Weights via Gradient Descent over the Loss Function, typically using stochastic mini-batch variants such as Adam or AdamW, with gradient clipping, learning-rate warmup, and cosine annealing schedules, across large-scale Compute Infrastructure comprising thousands of AI Chips (GPUs, TPUs, or custom accelerators); (4) post-training alignment — for instruction-following language models, applying supervised fine-tuning (SFT) on curated instruction-response pairs, followed by RLHF or direct preference optimisation (DPO) to align model behaviour with human preferences for helpfulness, harmlessness, and honesty; (5) Model Evaluation — benchmarking on held-out test sets and standardised benchmarks (MMLU for general knowledge, HumanEval for coding, MATH for mathematical reasoning, BIG-Bench for difficult tasks, SWE-bench for autonomous software engineering) to characterise capabilities and failure modes, and conducting Red Teaming to discover unsafe outputs before release; (6) post-processing and compression — Model Quantization (reducing weight precision from float32 to INT8, INT4, or lower via techniques including GPTQ, AWQ, and GGUF format conversion), Model Distillation (training a smaller student model to mimic the output distribution of a larger teacher model, often achieving 90%+ of teacher performance at a fraction of the size), and structured pruning (removing entire attention heads or feed-forward neurons with low activation magnitude); (7) Model Deployment via serving infrastructure including inference servers (TensorRT-LLM, vLLM, Ollama for local deployment), API gateways, and auto-scaling cloud infrastructure; (8) production monitoring and retraining as data distribution shifts occur, new tasks emerge, or model performance degrades. MLOps platforms (MLflow, Weights and Biases, Kubeflow, Vertex AI Pipelines) provide tooling spanning the lifecycle from experiment tracking through deployment automation.
A critical structural feature of the AI model paradigm is the distinction between the training compute budget and the inference compute budget. Training a frontier AI model is a one-time (or infrequent) expense measured in millions of GPU-hours and hundreds of millions of dollars; inference is an ongoing cost incurred for every input processed. Historically, training dominated discourse; by 2026, approximately 63% of frontier model lifecycle energy is inference, compared to 37% for training, reflecting the massive deployment scale relative to training frequency. This asymmetry has driven the inference optimisation industry: Model Quantization, Model Distillation, speculative decoding, continuous batching, and tensor parallelism are all techniques for reducing the marginal cost of serving a model at inference time. The reasoning model paradigm introduced by o1 (December 2024) complicates this balance further: extended chain-of-thought at inference time increases per-query compute (and cost) while improving accuracy on complex tasks, reintroducing inference compute as a first-class concern in the design of capable AI systems.
The commercialisation of AI models has created a stratified market with distinct competitive dynamics. At the frontier, a small number of well-capitalised organisations — OpenAI, Anthropic, Google DeepMind, Meta AI, xAI, Mistral, Cohere — compete on raw capability, deploying the largest models with the most extensive alignment work. Below the frontier, a much larger ecosystem of fine-tuning practitioners, vertical-domain model builders, and deployment specialists adapt open-weight base models (Llama 4, Qwen 2.5, Mistral Large 3, Gemma 3) to specific industries (healthcare, legal, finance, manufacturing) using parameter-efficient Fine Tuning methods (LoRA, QLoRA, prefix tuning, adapter layers). This ecosystem enables businesses without frontier-scale training budgets to deploy state-of-the-art capabilities by standing on the shoulders of base models produced by the frontier labs. The Hugging Face ecosystem — Hub, Transformers library, PEFT library, Diffusers library — provides the open-source infrastructure that makes this layered model economy possible. Generative AI reached 53% population adoption within three years of widespread frontier model availability, a faster consumer adoption rate than the personal computer or the internet, driven by the accessibility of chatbot interfaces and the zero-shot utility of instruction-tuned models.
Components / Architecture
The constitutive elements of an AI model as an artefact include:
-
Model Weights: The parameter tensors — billions of floating-point numbers — that encode the statistical knowledge extracted from Training Data. In a Transformer Architecture-based model, these include embedding matrices (mapping discrete tokens to continuous vector representations), attention projection matrices (Q, K, V, O — query, key, value, and output projections for each attention head), feed-forward layer weights (two linear transformations with a non-linear activation such as GeLU or SwiGLU between them), and layer normalisation parameters (gain and bias). Weights are stored in checkpoint files (PyTorch
.pt, SafeTensors for memory-safe loading, ONNX for cross-framework portability, GGUF for efficient CPU inference via llama.cpp) that can be loaded for AI Inference without retraining. The weight count of frontier models ranges from approximately 7 billion for efficient small models to over 2 trillion for Mixture of Experts frontier models; quantisation reduces the storage and memory footprint without necessarily retraining the model. -
AI Model Architecture: The structural blueprint specifying how computational operations are organised: the number and type of layers (attention blocks, feed-forward blocks, convolutional layers, or SSM blocks), the dimensionality of internal representations (hidden dimension, intermediate dimension, number of attention heads, head dimension), skip/residual connections (enabling gradient flow through deep networks without vanishing), normalisation placement (pre-norm vs. post-norm, RMSNorm vs. LayerNorm), positional encoding scheme (absolute sinusoidal, rotary (RoPE), ALiBi, or no positional encoding for SSMs), and the input-output interface (vocabulary size, context window length, output head structure). Architecture determines the inductive biases the model brings to learning, defining what patterns it can efficiently represent and what it must approximate poorly.
-
Loss Function: The scalar objective used during training to measure prediction error or generation quality; the form chosen determines what regularities the model learns to encode. Cross-entropy over next-token predictions is the standard loss for autoregressive language models; masked cross-entropy for masked language models (BERT); reconstruction loss plus KL divergence for variational autoencoders; adversarial minimax loss for GANs; denoising score matching for diffusion models; contrastive InfoNCE loss for representation learning (CLIP, DINO). The choice of loss function is deeply consequential for emergent model behaviour: next-token prediction over diverse internet text produces Emergent Capabilities in reasoning and instruction-following that were not explicitly specified in the loss.
-
Model Checkpoint: A snapshot of Model Weights and optimiser state (momentum vectors, second-moment estimates for Adam) at a particular point in training, enabling resumption if training is interrupted, comparison of capabilities across training stages, selection of the best checkpoint by held-out evaluation metric, and rollback if subsequent training (especially post-training alignment) degrades performance. Checkpoint management is a non-trivial infrastructure concern for frontier models; SafeTensors format was introduced to eliminate security vulnerabilities in pickle-based PyTorch checkpoints.
-
Tokenisation / pre-processing configuration: The pipeline that maps raw inputs (text, pixels, audio waveforms) to the discrete or continuous tokens the model processes. For language models, the tokeniser (BPE for GPT-series, SentencePiece/Unigram for multilingual models, tiktoken for OpenAI models) governs vocabulary (typically 32K–200K tokens), subword segmentation, and effective sequence length; its configuration must accompany weights to enable correct inference. For vision models, image patch tokenisers divide images into fixed-size patches and embed them linearly; multimodal models require tokeniser alignment across modalities.
-
Hyperparameter configuration: Architectural hyperparameters (layer count, attention head count, hidden dimension, context window) and training hyperparameters (peak learning rate, warmup steps, decay schedule, batch size, weight decay coefficient, gradient clipping threshold, data mixture proportions) are documented in AI Model Card records and configuration files (typically JSON or YAML) to ensure reproducibility. The Chinchilla paper revealed that hyperparameter choices — particularly the ratio of training tokens to model parameters — have first-order effects on final model capability, rivalling architectural choices in importance.
-
Inference configuration: Sampling hyperparameters (temperature, top-p / nucleus sampling, top-k, repetition penalty, beam search parameters) used at AI Inference time to control the stochasticity and diversity of generated outputs. These are not part of the trained model per se but are part of its operational specification and significantly affect output quality for generative models.
-
System prompt and alignment specification: For instruction-tuned models, the system prompt encodes the model’s behavioural persona, safety constraints, and task framing. While not stored in the weights, the combination of a base model and a system prompt constitutes the effective AI model as experienced by end users. The EU AI Act requires transparency obligations to extend to the “general-purpose AI model” system as a whole, including post-training modifications.
Use Cases / Major Families
AI models are categorised along several orthogonal dimensions; the most analytically useful are modality, scale, training regime, and deployment architecture.
By primary modality:
-
Language models (Large Language Models): trained on text corpora via autoregressive next-token prediction; capabilities include reasoning, coding, summarisation, dialogue, question answering, and tool use. The GPT family (OpenAI), Claude family (Anthropic), Gemini family (Google DeepMind), and Llama family (Meta) are the leading model lineages. Frontier instances in 2026 include Claude Opus 4.5 (Anthropic), GPT-5 (OpenAI, launched August 2025 as a native multimodal system handling text, image, audio, and code), and Llama 4 Maverick (Meta, open-weight Mixture of Experts achieving GPT-4o-class performance). Claude Sonnet 4.5 achieved 77.2% on SWE-bench Verified, representing the current frontier for autonomous software engineering by AI models. Llama 3.3 70B demonstrates that improved training data curation and recipe optimisation can yield a 70B model with performance comparable to earlier 405B models, illustrating that parameter count is not the sole determinant of capability.
-
Vision models (Computer Vision): Convolutional Neural Network architectures (ResNet, EfficientNet, ConvNeXt) and vision Transformer Architecture instances (ViT, DeiT, Swin Transformer) for image classification, object detection (YOLO variants, DETR), instance segmentation, depth estimation, and dense prediction tasks. The segmentation-anything model (SAM, Meta, 2023) demonstrated Emergent Capabilities in zero-shot segmentation from a single prompt. Medical imaging AI models (radiology report generation, pathology slide classification, retinal disease screening) are among the highest-value vertical applications, subject to medical device regulation.
-
Multimodal AI models: Foundation models trained on paired or interleaved text-image (and increasingly text-image-audio-video) data, processing multiple modalities through unified representations. Gemini 2.0 Ultra and GPT-5 handle text, image, audio, and video natively in a single forward pass. Flamingo (DeepMind, 2022) established the paradigm of cross-attention-based vision-language grounding; CLIP (OpenAI, 2021) established contrastive vision-language alignment. Multimodal models power document intelligence (parsing complex layouts in PDFs, tables, forms), visual question answering, image captioning, and Agentic Workflow systems that perceive their environment through screenshots or camera feeds.
-
Diffusion Model families: Generative AI models trained via denoising score matching that produce high-fidelity image, video, and audio outputs by learning to reverse a Gaussian diffusion process. Image generation: Stable Diffusion 3.5 (Stability AI), DALL-E 3/4 (OpenAI), Imagen (Google), Flux (Black Forest Labs). Video synthesis: Sora (OpenAI), Kling (Kuaishou), Runway Gen-3, CogVideo (Tsinghua). Audio generation: AudioLDM, MusicGen (Meta), Stable Audio. Diffusion Model architectures have largely supplanted GANs for image and video generation due to higher output quality and more stable training dynamics.
-
Code models: Specialised on source code corpora (GitHub, documentation, Stack Overflow) for code generation, completion, debugging, refactoring, and test synthesis. Codex (OpenAI, 2021) established the paradigm; GitHub Copilot (built on Codex successors) reportedly reduced developer task completion times by 55% in controlled studies. Current leading code models include Claude Sonnet 4.5 (77.2% SWE-bench Verified), DeepSeek-Coder, and Llama-based code fine-tunes.
-
Scientific Foundation Model instances: Domain-adapted models for biology (protein language models ESMFold, AlphaFold 2/3 — the latter achieving Gaussian experimental accuracy across diverse biomolecular structure prediction tasks), chemistry (Uni-Mol, ChemBERTa for molecular property prediction), climate modelling (FourCastNet, Pangu-Weather, GraphCast for weather prediction faster than numerical weather prediction systems), and materials science (Cambridge MACE, a leading materials foundation model supported by UK sovereign AI compute).
By scale:
-
Small models (1–8B parameters): Efficient, deployable on consumer hardware (a single GPU with 8–24GB VRAM, or modern smartphones with NPU acceleration). Examples include Llama 4 Scout (17B total / 4B active MoE, running at small-model cost), Mistral 7B, Gemma 3 (2B/9B/27B), Phi-3 Mini. Small models enable privacy-preserving local deployment in healthcare, legal, and enterprise settings where data cannot be sent to third-party cloud APIs. Model Quantization (GGUF format, INT4/INT8) reduces memory requirements further, enabling 7B models to run on devices with 4GB of VRAM.
-
Mid-range models (10–70B parameters): Balance capability and deployment cost; suitable for single or dual GPU inference without extensive infrastructure. Llama 3.3 70B (Meta, 2024) offers performance comparable to earlier 405B models due to improved training data curation and recipe optimisation — demonstrating that scaling laws must account for data quality, not only quantity. Mistral Large 3 (675B total / 41B active, December 2025) achieves frontier-class performance with mid-range per-token inference cost via sparse Mixture of Experts activation.
-
Frontier models (100B–2T+ parameters): State-of-the-art capability on complex reasoning, multimodal understanding, and autonomous task completion; require large-scale multi-GPU Compute Infrastructure for inference and are typically accessed via cloud API rather than local deployment. These models are the primary targets of EU AI Act GPAI model obligations.
By training and adaptation regime:
-
Pre-trained base models: Trained from scratch on large diverse corpora using self-supervised objectives; form the substrate for downstream task-specific adaptation. Useful for research and specialised fine-tuning but not directly deployed as user-facing systems (outputs are unaligned with human preferences).
-
Instruction-tuned / chat models: Base models subjected to supervised fine-tuning on curated instruction-response datasets followed by RLHF or direct preference optimisation (DPO); produce the assistant behaviour familiar from consumer chatbot products.
-
Reasoning models: Base or instruction-tuned models further trained to produce extended internal chain-of-thought traces (extended thinking) before emitting a final answer; trade latency and inference compute for substantially higher accuracy on mathematics, coding, and complex multi-step reasoning tasks. OpenAI o1/o3, Claude’s extended thinking mode, and Gemini Flash Thinking are the leading instances in 2026.
-
Open-weight models: Weights publicly released, enabling community fine-tuning, Model Quantization, Model Distillation, and local deployment without API dependency; Llama 4, Mistral Large 3, Qwen 2.5, and Gemma 3 are leading examples. Open-weight models enable vertical-domain specialisation (legal AI, medical AI, financial AI) through parameter-efficient Fine Tuning (LoRA, QLoRA, prefix tuning) without full retraining.
By application domain:
-
Natural Language Processing: text generation, summarisation, machine translation, named entity recognition, sentiment analysis, question answering, and information extraction.
-
Computer Vision: image classification, object detection, semantic segmentation, video understanding, and visual grounding.
-
Generative AI: creative content generation (text, image, video, music, code, synthetic data), personalisation, and content augmentation.
-
AI Agent and Agentic Workflow orchestration: autonomous task completion, tool use, web browsing, code execution, multi-agent coordination, and long-horizon planning in software engineering, scientific research, and business process automation.
-
Retrieval-Augmented Generation: enterprise knowledge management, document question answering, and RAG-based chatbots that ground language model outputs in retrieved evidence from proprietary corpora.
-
Multimodal perception for robotics and embodied systems: Multimodal AI models that process camera feeds and produce action sequences for robotic manipulation and autonomous navigation.
Academic Context
The modern AI model paradigm has theoretical roots in several converging research traditions spanning seven decades. The multilayer perceptron and the concept of a learnable function approximator emerged from Rosenblatt’s perceptron (1957) and Minsky and Papert’s influential critique (1969), which temporarily suppressed interest in neural models. The critical algorithmic breakthrough came with Rumelhart, Hinton, and Williams’s backpropagation paper (1986), which provided an efficient method for computing gradients through multi-layer networks, enabling the first demonstrations of useful representation learning. LeCun et al.’s convolutional network for handwritten digit recognition (1989, formalised in the 1998 LeNet paper) established that spatial inductive biases in architecture design could dramatically improve learning efficiency for visual data. Hochreiter and Schmidhuber’s LSTM (1997) addressed the vanishing gradient problem for sequences, enabling the first generation of capable language and speech models. Probabilistic graphical models (Dempster, Laird, and Rubin’s EM algorithm, 1977; Baum-Welch for HMMs; Restricted Boltzmann Machines and Deep Belief Networks, Hinton and Salakhutdinov 2006) provided an alternative generative modelling tradition.
The deep learning resurgence was catalysed by Krizhevsky, Sutskever, and Hinton’s AlexNet (2012), which demonstrated GPU-accelerated deep convolutional networks achieving a decisive 10.9 percentage-point improvement on ImageNet top-5 error over the second-place competitor — a margin that established deep learning as categorically superior to prior feature-engineering-based computer vision approaches. The availability of NVIDIA CUDA (2007) and large labelled datasets (ImageNet, 2009) were equally critical enabling conditions. The generative modelling revolution began with the Generative Adversarial Network (Goodfellow et al., 2014) and the Variational Autoencoder (Kingma and Welling, 2013), establishing adversarial training and variational inference as two distinct paths to learning generative AI models. The transformer breakthrough (Vaswani et al., “Attention Is All You Need”, 2017, NeurIPS) introduced scaled dot-product self-attention as the dominant architectural primitive, enabling parallelisable training at scales that RNN-based models could not support. BERT (Devlin et al., 2018) demonstrated bidirectional contextual representations from masked language modelling pre-training, establishing the transfer learning paradigm for NLP; GPT-2 (Radford et al., 2019) demonstrated generative pre-training and few-shot text generation; GPT-3 (Brown et al., 2020) demonstrated in-context learning as a qualitatively new capability of sufficiently large AI models.
Scaling laws formalised by Kaplan et al. (2020, “Scaling Laws for Neural Language Models”) and revised by Hoffmann et al. (2022, “Training Compute-Optimal Large Language Models”, the “Chinchilla” paper) established that model performance follows predictable power-law relationships with compute, data, and parameters, and that prior frontier models were significantly undertrained relative to the compute-optimal point — a finding that prompted a generation of models trained on far more tokens per parameter. The concept of the Foundation Model as a distinct category was introduced by Bommasani et al. (2021, “On the Opportunities and Risks of Foundation Models”, Stanford HAI), which named and characterised the homogenising role large pre-trained models play across AI research and application domains, and also introduced the concept of risk from homogenisation — that dependence on a small number of foundation models concentrates societal vulnerability. Wei et al. (2022, “Emergent Abilities of Large Language Models”) formally characterised the phenomenon of Emergent Capabilities: task-specific abilities that appear abruptly above certain scale thresholds, not present in smaller models, raising questions about predictability and safety evaluation.
Key research institutions contributing to AI model science include: Google DeepMind (AlphaFold 1/2/3, Gemini multimodal models, transformer and attention variants, Gato multi-task models); OpenAI (GPT series, RLHF methodology, DALL-E image generation, Codex code models, reasoning models o1/o3); Anthropic (Constitutional AI, mechanistic interpretability, Claude series, Responsible Scaling Policy); Meta AI (Llama series, open model ecosystem, FAIR foundational research); Stanford HAI and CRFM (foundation model characterisation, policy analysis, Foundation Model Transparency Index); CMU LTI (NLP models, multilingual AI); New York University (mechanistic interpretability, DALL-E); Edinburgh School of Informatics (NLP, neural architectures — ranked #1 UK for NLP research, foundational work on neural machine translation by Cho et al. and the Edinburgh NLP group); UCL Centre for Artificial Intelligence (leading a generative AI hub spanning Imperial College London, Cambridge, Oxford, Manchester, Edinburgh, Cardiff, and Surrey, with industry partners including DeepMind and IBM, and a Google DeepMind Academic Fellow appointment in March 2026); Cambridge (MACE materials Foundation Model, Leverhulme Centre for the Future of Intelligence, Bayesian machine learning tradition); Imperial College London (Nightingale AI health foundation model, London AI Technology Centre in partnership with Lenovo, 2026); and Microsoft Research (RLHF and alignment, LLM safety, multimodal AI).
Current Landscape (2026)
As of June 2026, the AI model landscape is characterised by several defining dynamics that collectively represent a shift from the “AI model as research artefact” paradigm of 2020 to “AI model as industrial infrastructure” — a transformation with profound implications for both the technical character of AI development and its governance.
Frontier capability acceleration: Industry produced over 90% of notable frontier models in 2025, with academic institutions largely unable to compete at the training scale required for frontier capability. Multiple frontier models now meet or exceed average human performance on PhD-level science questions (GPQA Diamond benchmark: GPT-5 and Claude Opus 4.5 both exceed 75%), competition mathematics (AIME 2025: o3 scored above 96th percentile of human competitors), and multimodal reasoning (MMMU benchmark). U.S. private AI investment reached 12.4 billion — with corporate spending on AI infrastructure (GPUs, data centres, cloud AI services) exceeding $130 billion. This investment concentration in the US and, to a lesser extent, China and the UK is reshaping the global distribution of AI capability.
Multimodal as the floor: Every major frontier model in 2025–2026 handles text, image, and document input as a minimum capability. Multimodal has shifted from differentiator to baseline expectation: the question for frontier labs is no longer whether a model is multimodal but which additional modalities (audio, video, structured data, code) are natively supported. GPT-5 (OpenAI, August 2025) handles text, image, audio, and code in a single unified model. Gemini 2.0 Ultra handles text, image, audio, and video natively and can process hour-long videos with a 2-million-token context window. Claude Opus 4.5 processes text, image, and document inputs with strong performance on complex visual reasoning tasks.
Reasoning models as a distinct paradigm: The release of OpenAI o1 (December 2024) established extended chain-of-thought at inference time as a distinct AI model class, qualitatively different from single-pass generation: the model generates internal reasoning traces (sometimes thousands of tokens long) before producing its final answer, trading latency and inference compute for substantially higher accuracy on mathematics, coding, science, and complex multi-step reasoning tasks. By mid-2026, most frontier labs have released reasoning model variants: o3 and o4 (OpenAI), Claude’s extended thinking mode, Gemini 2.0 Flash Thinking, DeepSeek-R1 (and its open-source derivative DeepSeek-R1-Distill). The reasoning model paradigm has reopened the inference compute scaling debate: how much quality improvement can be obtained by giving a model more time (and compute) at inference to think?
Open-weight competition closing the proprietary gap: Meta’s Llama 4 series (early 2026) with Llama 4 Maverick (17B active / 400B total Mixture of Experts) achieving GPT-4o-class performance on open weights, and Mistral Large 3 (December 2025, Apache 2.0 licence, 675B total / 41B active MoE, 256K context window, 200+ language support) have significantly narrowed the proprietary-open capability gap. DeepSeek-V3 (December 2024, released open-source) demonstrated that Chinese AI labs could achieve frontier performance at a fraction of US training costs, raising questions about the effectiveness of export controls as a Compute Governance tool. Qwen 2.5 (Alibaba) and Gemma 3 (Google) add further competition at the open-weight tier. The proliferation of capable open-weight models is driving enterprise AI adoption at lower cost and enabling privacy-preserving local deployment.
Inference cost compression: Inference costs have fallen approximately 10x per year for equivalent capability, driven by architectural improvements (Mixture of Experts reducing active parameters per token), Model Quantization (INT4/INT8, GPTQ, AWQ, GGUF formats reducing memory footprint by 4–8x with minimal quality loss), Model Distillation (smaller student models matching larger teacher performance), speculative decoding (using a small draft model to propose tokens verified by a large model at high throughput), and hardware improvements (NVIDIA H200, Blackwell B100, and AMD MI300X GPUs offering higher HBM bandwidth). By 2026, approximately 63% of frontier model lifecycle energy is consumed at inference, compared to 37% at training — an inversion from the previous pattern, as deployment scale vastly outpaces training frequency.
Agentic Workflow deployment becoming mainstream: AI models are increasingly deployed not as single-shot question-answering systems but as orchestrating AI Agent systems that plan, use tools, browse the web, write and execute code, manage files, coordinate sub-agents, and maintain state across long multi-step tasks. Claude Sonnet 4.5 achieved 77.2% on SWE-bench Verified — a benchmark requiring autonomous resolution of real GitHub software engineering issues — reflecting this paradigm shift. Enterprise adoption of agentic AI is growing rapidly: according to Stanford HAI’s 2026 AI Index, over 60% of Fortune 500 companies have at least one agentic AI workflow in production as of 2025, up from approximately 15% in 2023.
Regulatory pressure materialising: The EU AI Act’s General-Purpose AI (Foundation Model) model obligations — including technical documentation, copyright compliance, capability evaluations, and transparency disclosures — applied from August 2025, requiring providers including OpenAI, Anthropic, Google, Meta, Mistral, and others to file conformity documentation with the European AI Office. An Omnibus amendment (political agreement May 2026) extended some compliance deadlines, simplified conformity assessment requirements for lower-risk high-risk systems, and added specific provisions for AI-generated intimate images. The EU AI Act’s GPAI systemic-risk designation (triggered by training compute exceeding 10^26 FLOPs) imposes heightened obligations including adversarial Red Teaming, incident reporting to the AI Office, and cooperation with model evaluations by competent authorities.
Sustainability and environmental policy: AI model training and inference energy consumption is attracting mandatory disclosure requirements. The International Energy Agency estimated in 2024 that data centres consumed approximately 460 TWh globally; AI workloads are the fastest-growing component. GPT-4’s training run was estimated to have consumed approximately 50 GWh; frontier 2025–2026 training runs are estimated to require 2–5x this. Water consumption for data centre cooling is also attracting policy attention. The EU’s Corporate Sustainability Reporting Directive (CSRD) and the UK’s SECR (Streamlined Energy and Carbon Reporting) framework now require AI-intensive companies to disclose energy consumption, bringing environmental accountability into the AI model governance landscape.
UK Context
The UK is a significant node in the global AI model research and development ecosystem:
-
Edinburgh School of Informatics is ranked #1 in the UK for Natural Language Processing research and is globally prominent in AI model architecture research. The school has produced foundational work in neural machine translation, probabilistic NLP models, and efficient transformers.
-
UCL Centre for Artificial Intelligence leads a generative AI hub spanning Imperial College London, Cambridge, Oxford, Manchester, Edinburgh, Cardiff, and Surrey, with industry partners including IBM, BT, Google DeepMind, and Cisco. A Google DeepMind Academic Fellow was appointed to the UCL AI Centre in March 2026.
-
Imperial College London partnered with Lenovo in 2026 to establish the London AI Technology Centre at its White City Deep Tech Campus, focusing on Foundation Model deployment, Agentic Workflow, and intelligent systems coordination. Imperial researchers are developing Nightingale AI — a foundation model for health — supported by sovereign AI compute resources.
-
Cambridge researchers are developing MACE (a leading materials Foundation Model) supported by UK sovereign AI compute. Cambridge hosts the Leverhulme Centre for the Future of Intelligence and has a strong tradition in Bayesian machine learning.
-
Manchester, Leeds, and Sheffield contribute through the N8 Research Partnership, with Manchester’s Alan Turing Institute collaboration supporting AI model research in industrial AI applications including manufacturing, logistics, and healthcare.
-
UK government AI compute: The AI Opportunities Action Plan commits to building national AI compute infrastructure, with the Isambard-AI supercomputer (Bristol) and Dawn (Cambridge) providing sovereign GPU capacity for academic model training.
-
The UK’s leading AI companies — DeepMind (acquired by Google), Wayve (autonomous driving), Stability AI, Graphcore (IPU hardware) — have historically been founded or based in the UK and have made significant contributions to AI model architecture and hardware innovation.
Future Directions (2026-2030)
The AI model field is undergoing several simultaneous architectural, algorithmic, and sociotechnical transitions whose trajectories will shape the capability and deployment landscape through 2030.
-
Post-transformer architectures: State space models (Mamba, Mamba-2), linear attention variants (RWKV-v6, RetNet), and hybrid transformer-SSM architectures (Jamba, Zamba) challenge the standard transformer’s dominance by offering O(n) rather than O(n^2) attention complexity with respect to sequence length, enabling processing of context windows measured in millions or hundreds of millions of tokens — orders of magnitude beyond current transformer practical limits — at proportionally lower memory and compute cost. Research groups at AI2, Carnegie Mellon, and several UK institutions are pursuing hybrid architectures that combine SSM long-range efficiency with attention-based precision for high-salience positions.
-
Efficient scaling and algorithmic improvements: Beyond architectural change, improvements in training data curation (deduplication, quality filtering, synthetic data generation), better optimisers (Muon, Shampoo), improved learning-rate schedules, and curriculum learning approaches continue to shift the Pareto frontier of capability-per-FLOP, enabling smaller models to match the performance of larger predecessors. Model Distillation at scale — training student models to match teacher distributions — is expected to produce a proliferation of highly capable compact models suitable for on-device deployment without cloud infrastructure dependency.
-
Multimodal unification and any-to-any generation: Single Foundation Model instances trained jointly on text, image, audio, video, code, and structured data (tabular, molecular, genome sequences) may replace specialised per-modality architectures over the 2026–2030 horizon, enabling richer cross-modal reasoning — for example, a model that simultaneously understands spoken language, reads from a displayed screen, and generates code to solve a task demonstrated visually. Native any-to-any generation (producing image, audio, or video outputs directly from the same model that processes text) is an active research frontier.
-
Reasoning model maturation and integration: The extended chain-of-thought paradigm (o1/o3/o4 class at OpenAI; Claude extended thinking; Gemini 2.0 Flash Thinking) will be integrated into multimodal and Agentic Workflow contexts, enabling systems that reason over images, code, tool outputs, and environmental feedback with human-PhD-level or superhuman reliability on specific task classes. The key open research question is whether extended thinking at inference time can be made sufficiently efficient in token use (and therefore cost) to deploy economically in high-volume production settings.
-
Open Source AI ecosystem growth and democratisation: The proliferation of open-weight models — driven by Meta’s Llama series, Mistral’s Apache-licensed releases, Alibaba’s Qwen series, and Google’s Gemma releases — combined with parameter-efficient Fine Tuning (LoRA, QLoRA, prefix tuning, adapter layers) and Model Quantization tooling (llama.cpp, Ollama, ExLlamaV2) will continue to democratise AI model deployment, enabling vertical-industry specialists to build highly capable domain-adapted models (medical, legal, financial, manufacturing) without frontier-scale training budgets or cloud API dependency.
-
Regulatory formalisation and AI Model Card convergence: AI model documentation (model cards, system cards, technical reports) will transition from voluntary best practice to legally mandated transparency artefacts under the EU AI Act (which requires GPAI model providers to publish training data summaries and capability evaluations), the UK AI Bill (anticipated 2026), and successor regulations globally. AI Model Card standards are expected to converge internationally through ISO/IEC JTC1 SC42 standardisation work, creating harmonised templates that simplify compliance for multinational model providers.
-
Compute Governance institutionalisation: Policies targeting frontier model training compute via FLOP thresholds, export controls on AI Chips (extending current GPU and HBM controls to new hardware generations), mandatory training-run registration with national authorities, and international sharing of model evaluation results are expected to mature from ad-hoc executive orders and voluntary commitments into institutionalised regulatory frameworks. The emerging network of national AI Safety Institutes (UK, US, Japan, Singapore, EU) is expected to develop standardised pre-deployment evaluation protocols that provide a global safety baseline for frontier AI models.
-
Embodied AI and world models: AI models trained not only on text and images but on video demonstrations of physical tasks and on simulated environment interactions will drive advances in robotics (manipulation, locomotion, open-vocabulary task following), autonomous vehicles (end-to-end learned driving policies), and scientific simulation (molecular dynamics, climate modelling, drug design). World models — AI models that maintain an internal representation of the physical state of an environment and can predict the consequences of actions — represent the frontier of AI model capability extension from pattern recognition toward physical-world planning and control.
-
AI model interpretability and Explainable AI: The mechanistic interpretability research programme — reverse-engineering the computational structure of trained AI models to understand what circuits, features, and representations they use — is expected to mature from single-circuit analyses of small models to scalable tools applicable to frontier systems, providing a scientific foundation for AI Safety evaluation and Bias Mitigation at the level of internal model representations rather than surface-level behaviour auditing.
-
AI Alignment and value learning advances: Techniques including Constitutional AI, scalable oversight (debate, weak-to-strong generalisation), and automated red-teaming are expected to mature into more robust and scalable alignment pipelines that can ensure frontier AI models remain helpful, honest, and harmless as capabilities scale further. The convergence of AI Alignment as a technical discipline with AI Policy as a governance framework will become increasingly important as models approach or exceed human performance on high-stakes cognitive tasks.
Research and Literature
- Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). “Attention Is All You Need.” NeurIPS 2017. The transformer architecture paper; foundational for virtually all frontier AI models.
- Kaplan, J., McCandlish, S., Henighan, T., et al. (2020). “Scaling Laws for Neural Language Models.” arXiv:2001.08361. Established power-law relationships between compute, data, parameters, and performance.
- Hoffmann, J., Borgeaud, S., Mensch, A., et al. (2022). “Training Compute-Optimal Large Language Models.” arXiv:2203.15556 (Chinchilla paper). Revised scaling law estimates, showing frontier models were significantly undertrained.
- Bommasani, R., Hudson, D. A., Aditi, E., et al. (2021). “On the Opportunities and Risks of Foundation Models.” Stanford HAI Center for Research on Foundation Models. Named and characterised the Foundation Model paradigm.
- Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). “ImageNet Classification with Deep Convolutional Neural Networks.” NeurIPS 2012. AlexNet; catalysed the deep learning resurgence and the shift to learned AI models.
- Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). “Learning representations by back-propagating errors.” Nature 323, 533–536. Backpropagation as the foundational training algorithm for neural networks.
- LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998). “Gradient-Based Learning Applied to Document Recognition.” Proceedings of the IEEE 86(11), 2278–2324. Convolutional Neural Network for AI model training.
- Hochreiter, S. and Schmidhuber, J. (1997). “Long Short-Term Memory.” Neural Computation 9(8), 1735–1780. LSTM architecture; underpinned the first wave of sequence model AI models.
- Brown, T. B., Mann, B., Ryder, N., et al. (2020). “Language Models are Few-Shot Learners.” NeurIPS 2020. GPT-3; demonstrated In-Context Learning as an emergent capability of large AI models.
- Ouyang, L., Wu, J., Jiang, X., et al. (2022). “Training Language Models to Follow Instructions with Human Feedback.” NeurIPS 2022. InstructGPT; established RLHF as the dominant post-training alignment method.
- Hu, E. J., Shen, Y., Wallis, P., et al. (2022). “LoRA: Low-Rank Adaptation of Large Language Models.” ICLR 2022. Parameter-efficient Fine Tuning enabling task adaptation without full retraining.
- Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). “QLoRA: Efficient Finetuning of Quantized LLMs.” NeurIPS 2023. Model Quantization combined with Fine Tuning for consumer-hardware deployment.
- Shazeer, N., Mirhoseini, A., Maziarz, K., et al. (2017). “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.” ICLR 2017. Foundational paper for Mixture of Experts AI model architectures.
- Wei, J., Tay, Y., Bommasani, R., et al. (2022). “Emergent Abilities of Large Language Models.” Transactions on Machine Learning Research. Characterised Emergent Capabilities arising from scale.
- Jiang, A. Q., Sablayrolles, A., Roux, A., et al. (2024). “Mixtral of Experts.” arXiv:2401.04088. Sparse MoE LLM; influenced subsequent open-weight frontier model development.
- Touvron, H., Martin, L., Stone, K., et al. (2023). “Llama 2: Open Foundation and Fine-Tuned Chat Models.” arXiv:2307.09288. Established open-weight model release norms; predecessor to Llama 4.
- Radford, A., Kim, J. W., Hallacy, C., et al. (2021). “Learning Transferable Visual Models From Natural Language Supervision.” ICML 2021. CLIP; foundational Multimodal AI model training paradigm.
- Ho, J., Jain, A., and Abbeel, P. (2020). “Denoising Diffusion Probabilistic Models.” NeurIPS 2020. Diffusion Model training; enables generative AI models for image and video synthesis.
- Jumper, J., Evans, R., Pritzel, A., et al. (2021). “Highly Accurate Protein Structure Prediction with AlphaFold.” Nature 596, 583–589. Scientific foundation model application; AlphaFold demonstrated AI model transfer to structural biology.
- Bai, Y., Jones, A., Ndousse, K., et al. (2022). “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.” arXiv:2204.05862. Anthropic Constitutional AI precursor; RLHF for Responsible AI model alignment.
- Google DeepMind. (2024). “Gemini: A Family of Highly Capable Multimodal Models.” arXiv:2312.11805v4. Frontier multimodal AI model; unified text, image, audio, and video processing.
- Anthropic. (2024). “The Claude 3 Model Family: Opus, Sonnet, Haiku.” Anthropic Technical Report. Characterises the capability, safety, and alignment properties of the Claude family of AI models.
- Meta AI. (2024). “The Llama 3 Herd of Models.” arXiv:2407.21783. Llama 3 design, training data curation, Benchmark Evaluation methodology, and open-weight release rationale.
- Mitchell, M., Wu, S., Zaldivar, A., et al. (2019). “Model Cards for Model Reporting.” FAccT 2019. AI Model Card framework; now embedded in EU AI Act transparency obligations.
- Bommasani, R., Shetty, S., Narayanan, A., et al. (2023). “Foundation Model Transparency Index.” Stanford CRFM. Systematic evaluation of transparency practices across major AI model providers.
- Stanford HAI. (2026). “2026 AI Index Report.” Stanford Human-Centered AI. Comprehensive statistics on AI model development, industry investment ($285.9B US private AI investment 2025), and deployment trends.
- Hugging Face. (2025). “The State of Open Models 2025.” Hugging Face Blog. Survey of 1.5M+ open-weight AI models, fine-tuning ecosystem, and quantisation adoption patterns.
- European Parliament and Council. (2024). “Regulation (EU) 2024/1689 — AI Act.” EUR-Lex. Regulatory framework establishing obligations for AI models including GPAI model transparency, systemic risk assessment, and technical documentation requirements.