Foundation models are large-scale neural networks trained on broad, diverse datasets via self-supervised learning that acquire general-purpose representations transferable to a wide range of downstream tasks. They are characterised by massive parameter counts, emergent capabilities not explicitly trained for, and the ability to be fine-tuned or prompted for specialised applications. Prominent examples include GPT-4, BERT, CLIP, and Stable Diffusion, spanning language, vision, and multimodal domains. Their scale and generality make them qualitatively distinct from narrow task-specific models.
Content
- The term “foundation model” was introduced by the Stanford CRFM in 2021 to capture the paradigm shift whereby a single large model trained on internet-scale data becomes the basis for a wide ecosystem of applications. Prior to this paradigm, practitioners trained bespoke models for each task; foundation models inverted this by making pre-training the expensive, centralised step and task adaptation cheap and distributed.
- Training dynamics of foundation models are governed by scaling laws relating compute, data, and parameter count to predictable gains in loss. The Chinchilla scaling laws (Hoffmann et al., 2022) demonstrated that most large models at the time were significantly under-trained relative to their parameter budgets, reshaping how organisations allocate training compute. Modern foundation models balance parameter count with dataset size and training duration.
- Foundation models exhibit emergent capabilities: behaviours that appear discontinuously as model scale crosses thresholds. Chain-of-thought reasoning, few-shot learning, and in-context learning were not explicitly trained objectives but arise from scale and data diversity. This unpredictability creates both opportunity—unexpected utility—and risk, motivating alignment research and capability evaluation.
- The architectural backbone of most language foundation models is the Transformer with self-attention, while vision models employ vision transformers (ViT) or hybrid CNN-Transformer designs. Multimodal models such as CLIP align image and text representations through contrastive pre-training, enabling zero-shot classification and cross-modal retrieval. Diffusion-based foundation models like Stable Diffusion use latent representations to generate high-fidelity images conditioned on text prompts.
- Governance and access to foundation models present significant societal considerations. Centralised training concentrates capability in a small number of actors, raising questions about equitable access, dual-use risks, and accountability. Open-weight models (LLaMA, Mistral) partially democratise access but shift risk profiles. The EU AI Act’s treatment of general-purpose AI models directly targets foundation models, imposing documentation, transparency, and risk-assessment obligations on providers of sufficiently large models.