The transformer is a neural network architecture introduced by Vaswani et al. in ‘Attention Is All You Need’ (2017). It replaces recurrence and convolution with multi-head self-attention and position-wise feed-forward layers, enabling fully parallel sequence processing, and underpins large language models and modern vision, speech, and protein-structure systems.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:MultiHeadAttention))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:SelfAttention))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:FeedForwardNetwork))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:PositionalEncoding))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:LayerNormalization))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:ResidualConnection))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:KVCache))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:hasPart ai:TokenEmbedding))

## Dependency Relationships
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:requires ai:Tokenization))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:requires ai:TrainingData))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:requires ai:GradientDescent))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:requires ai:GPUCompute))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:dependsOn ai:LinearAlgebra))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:dependsOn ai:InformationTheory))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:dependsOn ai:Softmax))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:dependsOn ai:Backpropagation))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:dependsOn ai:MixedPrecisionTraining))

## Capability Relationships
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:enables ai:LargeLanguageModel))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:enables ai:VisionTransformer))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:enables ai:SpeechRecognition))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:enables ai:CodeGeneration))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:enables ai:MachineTranslation))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:enables ai:ProteinStructurePrediction))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:supports ai:NaturalLanguageProcessing))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:supports ai:ComputerVision))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:supports ai:AudioProcessing))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:supports ai:TimeSeriesForecasting))

## Implementation Relationships
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:ScaledDotProductAttention))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:RotaryPositionEmbedding))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:SwiGLUActivation))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:RMSNorm))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:FlashAttention))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:GroupedQueryAttention))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:implements ai:SparseMixtureOfExperts))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:uses ai:AdamOptimizer))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:uses ai:DropoutRegularization))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:uses ai:BytePairEncoding))

## Reduction Relationships
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:reduces ai:SequentialComputationBottleneck))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:reduces ai:LongRangeDependencyPathLength))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:reduces ai:VanishingGradientProblem))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:reduces ai:RecurrenceInductiveBias))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:reduces ai:TrainingTimeVsRNN))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:contrasts ai:RecurrentNeuralNetwork))
SubClassOf(ai:Transformers
  ObjectSomeValuesFrom(ai:contrasts ai:Mamba))

DataPropertyAssertion(ai:hasIdentifier ai:Transformers "AI-0742"^^xsd:string)
DataPropertyAssertion(ai:authorityScore ai:Transformers "0.87"^^xsd:decimal)
DataPropertyAssertion(ai:attentionComplexity ai:Transformers "quadratic"^^xsd:string)
DataPropertyAssertion(ai:originalHeads ai:Transformers "8"^^xsd:integer)
DataPropertyAssertion(ai:originalLayers ai:Transformers "6"^^xsd:integer)
DataPropertyAssertion(ai:yearIntroduced ai:Transformers "2017"^^xsd:integer)

About Transformers

  • The Transformer is the foundational neural architecture of modern Artificial Intelligence, introduced in “Attention Is All You Need” (Vaswani et al., NeurIPS 2017). It discarded the sequential inductive bias of Recurrent Neural Networks and the local spatial bias of Convolutional Neural Networks, replacing both with a general-purpose mechanism: multi-head scaled dot-product attention. This architectural decision granted transformers two decisive advantages — full parallelism during training and O(1)-distance paths between any two positions in a sequence.
  • The original model achieved a then-state-of-the-art BLEU score of 28.4 on WMT 2014 English-to-German translation, training in 3.5 days on 8 P100 GPUs versus weeks for comparable RNN systems. The key insight was that attention over a full sequence — previously used only as an auxiliary mechanism bolted onto RNNs — could serve as the complete computational primitive, eliminating recurrence entirely.
  • Since 2017, every major AI breakthrough — BERT, GPT 3, AlphaFold 2, DALL-E, Whisper, LLaMA, Gemini Multimodal Language Model, Claude — has been built on transformer foundations. The architecture has proven to be a general-purpose learner that scales smoothly with compute and data, conforming to power-law scaling laws across 10 orders of magnitude in parameter count.

Components / Architecture

Overview of the Transformer Block

  • A single transformer block consists of: (1) a multi-head attention sub-layer, (2) a residual connection adding the attention output to the input, (3) a layer/RMS normalisation, (4) a position-wise feed-forward sub-layer, and (5) a second residual + normalisation. These two sub-layers are the only learned components; positional information is injected once before the first layer via a positional encoding scheme. Modern configurations apply normalisation before each sub-layer (pre-norm) rather than after (post-norm).
  • Stacking L such blocks creates the full encoder or decoder. Encoders use bidirectional attention (full n×n mask); decoders use causal masking (upper-triangular zeros). Encoder-decoder models include an additional cross-attention sub-layer in each decoder block allowing decoder positions to query encoder outputs.

Scaled Dot-Product Attention

  • The core computation projects a sequence of d_model-dimensional input vectors into query (Q), key (K), and value (V) matrices via learned linear projections W^Q, W^K, W^V ∈ ℝ^(d_model × d_k). Attention weights are computed and applied as:
  • Attention(Q, K, V) = softmax(QK^T / √d_k) · V
  • The √d_k scaling prevents dot products from growing large in magnitude, which would push softmax into regions of near-zero gradient. For causal (decoder) attention, a triangular mask zeroes out future positions before softmax, ensuring autoregressive generation cannot attend to tokens not yet generated. The attention matrix is of size n×n for sequence length n, giving O(n²) time and memory complexity — the fundamental bottleneck that subsequent work has sought to reduce.

Multi-Head Attention (MHA)

  • Rather than a single attention function, MHA runs h parallel attention operations on projected subspaces of dimension d_k = d_v = d_model/h, concatenating results and projecting back: MultiHead(Q,K,V) = Concat(head_1,…,head_h) W^O where head_i = Attention(QW_i^Q, KW_i^K, VW_i^V).
  • Vaswani et al. used h=8, d_model=512. Modern large models use h=32-128 heads with d_model=4096-16384. Each head can specialise: circuit-analysis work (Olsson et al. 2022) finds individual heads implementing induction (attending to tokens following a previous occurrence of the current token), syntactic agreement, and positional copying operations. The h subspace decomposition also allows the model to simultaneously attend to content from different representational perspectives — semantic similarity in one head, syntactic proximity in another, positional distance in a third.

Feed-Forward Network (FFN)

  • Each transformer layer contains a position-wise FFN applied identically to all positions: FFN(x) = max(0, xW_1 + b_1)W_2 + b_2. Typically d_ff = 4 × d_model (e.g. 512 → 2048 in the original, or 4096 → 16384 in GPT-3).
  • The SwiGLU variant (Shazeer 2020, adopted by LLaMA, PaLM, Gemini Multimodal Language Model) replaces ReLU with a gated linear unit: SwiGLU(x, W, V) = Swish(xW) ⊙ (xV), where Swish(x) = x·σ(x). SwiGLU is typically paired with a reduced expansion ratio of ~8/3 × d_model to maintain comparable parameter counts, consistently outperforming ReLU/GELU at matched compute budget. The FFN layers collectively store the model’s factual knowledge — ablation studies show FFN layers encode entity-attribute associations that attention heads retrieve via key lookup.

Positional Encodings: Sinusoidal, RoPE, ALiBi, YaRN

  • Because attention is permutation-equivariant, position information must be injected explicitly. Four main families exist in practice:
  • Sinusoidal (Vaswani 2017): PE(pos, 2i) = sin(pos/10000^(2i/d)), PE(pos, 2i+1) = cos(pos/10000^(2i/d)). Fixed, not learned. Extrapolates modestly beyond training length. Wavelengths range from 2π (highest frequency, dimension 0) to 10000·2π (lowest frequency, dimension d−1), creating a unique position fingerprint. Used by the original transformer, some T5 variants, and early BERT pre-training.
  • Learned absolute: Learned embedding table over positions (BERT, GPT-2). Simple and effective within training context length; fails catastrophically beyond it as out-of-distribution position IDs receive no meaningful encoding.
  • Rotary Position Embedding (RoPE, Su et al. 2021): Encodes relative position by rotating query/key vectors in 2D subspaces: the dot product Q_m^T K_n = f(m−n) depends only on relative offset, not absolute positions. Implemented efficiently as element-wise complex multiplication in the frequency domain. Adopted by LLaMA, GPT-NeoX, Mistral, Phi, Gemma, Qwen, Gemini Multimodal Language Model. RoPE enables context extension via NTK-aware interpolation and YaRN without retraining.
  • ALiBi (Press et al. 2021): Adds a linear distance penalty (−|i−j|·m_h) to attention logits with per-head slope m_h, removing position embeddings entirely. Strong length generalisation at training length and beyond; used by BLOOM, MPT, and long-context variants. ALiBi’s simplicity makes it robust: no positional embedding parameters to learn or extrapolate.
  • YaRN (Peng et al. 2023): Frequency-aware NTK interpolation of RoPE extending LLaMA-2’s 4K training context to 128K+ tokens with only ~0.1% fine-tuning compute. Identifies that high-frequency RoPE dimensions should not be interpolated (they carry fine-grained positional signals) while low-frequency dimensions can be scaled. YaRN + LLaMA 2 achieves competitive long-context perplexity with full retraining on long documents.

Normalisation: LayerNorm, RMSNorm, Pre-Norm vs Post-Norm

  • LayerNorm (Ba et al. 2016): Normalises across d_model dimensions within each position independently, learning scale γ and bias β: LN(x) = γ·(x−μ)/σ + β. Stabilises training by reducing covariate shift within each layer. Distinct from BatchNorm (normalises across batch dimension) — for variable-length sequences BatchNorm is impractical.
  • RMSNorm (Zhang & Sennrich 2019): Removes mean-centring, normalises by root-mean-square: RMS(x) = x / rms(x) · γ where rms(x) = √(mean(x²)). Approximately 10% faster computation and equally stable to LayerNorm; adopted by LLaMA, T5, Mistral, Phi, Gemma. The bias β is also dropped, reducing parameter count slightly.
  • Post-Norm (original Vaswani 2017): Normalisation after residual addition: x’ = LN(x + Sublayer(x)). Theoretically appealing but difficult to train without careful warm-up scheduling; deep models (>24 layers) exhibit instability during early training due to large residual activations.
  • Pre-Norm (modern standard): Normalisation before each sub-layer: x’ = x + Sublayer(Norm(x)). Used by GPT-3, all LLaMA variants, PaLM, Claude. Gradient flows cleanly through the residual stream: ∂L/∂x = ∂L/∂x’ · (1 + ∂Sublayer(Norm(x))/∂x), where the identity path ensures gradient magnitude ≥ 1 regardless of depth. Enables stable training of 100+ layer models.

Attention Variants: MHA, MQA, GQA, Sliding Window

  • Multi-Query Attention (MQA, Shazeer 2019): All query heads share a single K/V head. During inference, the KV cache is reduced from (h × n × d_k) to (1 × n × d_k), an h-fold memory reduction enabling larger batch sizes and higher throughput. Quality loss is minor for generation tasks. Used by Falcon, early PaLM 2, Mistral v0.1.
  • Grouped-Query Attention (GQA, Ainslie et al. 2023): Query heads divided into G groups, each group sharing one K/V head (G typically 4-8). KV cache reduced by factor h/G while quality degrades less than MQA. GQA is trained from scratch or converted from MHA checkpoints via mean-pooling of K/V heads. Adopted by LLaMA 2 70B (G=8), Mistral 7B (G=8), Gemma, Claude 3 series, Phi-3. Now the default for production-scale decoder models.
  • Sliding Window Attention: Each token attends only to the W nearest tokens. O(n·W) compute vs O(n²). Longformer (Beltagy et al. 2020) combines window attention with global tokens for CLS and task-specific positions. Mistral 7B uses 4096-width sliding window with interleaved full-attention layers to maintain global information flow. Effective for tasks where local context dominates (e.g. character-level, document chunking) but sacrifices retrieval of distant tokens.

Mixture of Experts (MoE)

  • MoE replaces dense FFN layers with N parallel expert networks, routing each token to top-k (typically k=2) based on a learned gating network: gate(x) = softmax(x·W_g), route to top-k by score. Only k/N of parameters are active per token, enabling trillion-scale capacity at proportional inference cost.
  • Switch Transformer (Fedus et al. 2021): 1.6T parameters, k=1 routing (top-1 simplifies load balancing), 2048 experts per layer, achieves 7× training speedup versus comparable dense T5-XXL. Load-balancing auxiliary loss L_aux = α·∑_i f_i·P_i penalises uneven token assignment where f_i is fraction of tokens assigned to expert i and P_i is routing probability.
  • Mixtral 8x7B (Jiang et al. 2024): 8 experts per layer, k=2 routing; 47B total parameters but only ~13B active per token. Outperforms LLaMA 2 70B (dense) on MMLU, HumanEval, and GSM8K while using ~4× less compute per token. Architecture: 32 transformer layers, each with 8 FFN experts of d_ff=14336, using GQA (8 heads, 2 KV heads), SwiGLU. Mixtral 8x22B (2024): 141B total/39B active.
  • DeepSeek-V2 (2024): 236B total/21B active per token, 64 experts with k=6, Multi-Head Latent Attention (MLA) compressing KV cache via low-rank decomposition — reduces KV cache 93% vs standard MHA while maintaining quality. Trained on 8.1T tokens, demonstrating frontier-class performance at MoE efficiency.
  • Key MoE challenges: expert collapse (token routing concentrates on few experts, others go unused); load imbalance under auxiliary loss weighting; deployment complexity (all expert shards must be loaded across device memory even though only k are active per token); communication overhead in distributed inference requiring expert routing across network.

Encoder-Only, Encoder-Decoder, Decoder-Only

  • Encoder-only models (BERT family) apply bidirectional attention — each token attends to all others. Pre-training via masked language modelling (MLM): 15% of tokens masked, model predicts originals from bidirectional context. This forces rich contextual representations. Suited to classification, NER, QA, and semantic similarity where full context is available at inference. Typical size: 110M-340M parameters (BERT-base/large). Encoder-only models dominated NLU benchmarks (GLUE, SuperGLUE) 2019-2022 before being displaced by few-shot LLMs.
  • Encoder-decoder (original Transformer, T5, BART, mT5): Encoder reads the source sequence with bidirectional attention; decoder generates the target autoregressively, attending to encoder outputs via cross-attention. Cross-attention allows each decoder position to directly query all encoder positions, providing rich source context for each generated token. Natural for seq2seq tasks: machine translation, text summarisation, speech recognition (Whisper), code generation from specifications.
  • Decoder-only (GPT series, LLaMA, Claude, Mistral, Gemini Multimodal Language Model, Phi): Causal (left-to-right) masking — each token attends only to preceding tokens. Pre-trained via next-token prediction on massive corpora. Dominates large-scale language modelling since GPT-3 demonstrated that sufficiently large causal models perform seq2seq tasks via in-context conditioning without cross-attention. Task specification, few-shot examples, and chain-of-thought prompting all fit naturally into the autoregressive prompt paradigm.

Use Cases / Major Families

Architecture at a Glance

  • Before diving into families, the common architectural template as of 2026: decoder-only, pre-norm (RMSNorm), causal masking, RoPE positional encoding, SwiGLU FFN (expansion ~8/3 × d_model), GQA (G=4-8 KV groups), flash attention for training efficiency, BF16 mixed precision, BPE tokeniser (32K-128K vocabulary). This template applies to LLaMA 3, Mistral, Gemma, Phi-3, Qwen, DeepSeek-V3, and the majority of 2024-2026 open-weight releases.

GPT Series (OpenAI, decoder-only)

  • GPT-1 (Radford et al. 2018, 117M parameters): First large-scale pre-training + fine-tuning demonstration for NLP. Showed that a 12-layer transformer pre-trained on BooksCorpus transferred to 11 downstream tasks with minimal modification. GPT-2 (2019, 1.5B, 48 layers): Zero-shot task generalisation — summarisation, translation, QA without task-specific fine-tuning. Infamously withheld from release due to misuse concerns, then released incrementally.
  • GPT-3 (Brown et al. 2020, 175B, 96 layers, 96 heads): Established in-context learning at scale. With a handful of examples in the prompt context, GPT-3 matched fine-tuned models on many NLP benchmarks. Trained on 300B tokens (Common Crawl, WebText2, Books, Wikipedia) in 3.14×10²³ FLOP. Launched the modern API-based LLM economy.
  • ChatGPT (2022): RLHF fine-tuning of GPT-3.5 producing a conversational assistant aligned with human preferences. Reached 100M users in 2 months — the fastest consumer product adoption in history. GPT-4 (2023): Multimodal (text + image input), 128K context (Turbo), dramatically improved reasoning and instruction following. GPT-4o (2024): Natively multimodal training across text, image, and audio with faster inference.

BERT Family (Google, encoder-only)

  • BERT (Devlin et al. 2019, 110M/340M parameters): Bidirectional transformer pre-trained with MLM and next-sentence prediction (NSP) on BooksCorpus + Wikipedia. Fine-tuning with a single linear head achieves state-of-the-art on GLUE (80.4), SQuAD 1.1 (93.2 F1), and NER benchmarks. Established the fine-tuning paradigm for NLU. RoBERTa (Liu et al. 2019): Removed NSP, used dynamic masking, larger batches (8K), and 160GB training data; gained 4+ GLUE points over BERT. DeBERTa (He et al. 2021): Disentangled position-content attention; enhanced mask decoder for MLM; surpassed human performance on SuperGLUE (90.3 vs 89.8 human).
  • ALBERT (Lan et al. 2020): Factorised embedding parameterisation and cross-layer parameter sharing, achieving BERT-large quality with 18× fewer parameters. DistilBERT (Sanh et al. 2020): Knowledge distillation producing a 66M-parameter model retaining 97% of BERT performance at 60% smaller size and 60% faster inference.

T5 and PaLM (Google, encoder-decoder and decoder-only)

  • T5 (Raffel et al. 2020): Unified all NLP tasks as text-to-text problems (“translate English to German: …”) using encoder-decoder architecture. Trained on C4 (Colossal Clean Crawled Corpus, 750GB). Conducted the largest hyperparameter and architecture sweep in NLP history (thousands of ablations) establishing that: (1) pre-norm outperforms post-norm, (2) relative position biases outperform sinusoidal, (3) multi-task pre-training helps. Sizes 60M-11B. T5-11B held the GLUE and SuperGLUE records briefly.
  • PaLM (Chowdhery et al. 2022, 540B): Decoder-only, trained on 780B tokens using Google’s Pathways system on 6144 TPU v4 chips in two pods. Introduced chain-of-thought few-shot prompting as a standard evaluation. Achieved breakthrough on BIG-Bench tasks that GPT-3/Chinchilla could not solve (e.g. grade-school math word problems). PaLM 2 (2023): Smaller but better trained on more multilingual and code data; powers Gemini 1.0 Nano.

LLaMA Family (Meta, decoder-only, open-weight)

  • LLaMA 1 (Touvron et al. Feb 2023, 7B/13B/33B/65B): Trained on 1-1.4T tokens from public sources (Common Crawl, C4, GitHub, Wikipedia, Books, ArXiv, StackExchange). Used RoPE, pre-norm, SwiGLU — establishing what became the standard modern decoder configuration. Despite smaller parameter counts, outperformed GPT-3 (175B) on most benchmarks due to compute-optimal training per Chinchilla. Leaked immediately; catalysed the open-source LLM movement.
  • LLaMA 2 (Touvron et al. July 2023, 7B/13B/34B/70B): 2T training tokens; GQA (70B); 4K context (extendable to 8K); commercial license. LLaMA 2 Chat variants add RLHF alignment. 70B outperforms GPT-3.5 on most academic benchmarks. LLaMA 3 (Meta, April 2024, 8B/70B/405B): 15T+ training tokens; Tiktoken BPE vocabulary (128K tokens); context 8K (extended to 128K in Llama 3.1); strong multilingual and code performance. LLaMA 3 405B approaches GPT-4 quality with MMLU 87.3%.

Claude (Anthropic, decoder-only)

  • Constitutional AI (Bai et al. 2022): Defines a set of principles, has the model critique and revise its own outputs according to these principles, then trains on revised outputs using RLHF with AI-generated (RLAIF) preference labels. Reduces dependence on costly human labelling of harmful content while maintaining alignment. Claude 1 (2022), Claude 2 (2023, 100K context), Claude 3 (March 2024, Opus/Sonnet/Haiku variants, 200K context, multimodal vision).
  • Claude 3.5 Sonnet (June 2024): Led HumanEval (92.0% pass@1), MATH, and GPQA Diamond benchmarks. Claude 3.7 Sonnet (February 2025): Introduced extended thinking mode — explicit chain-of-thought reasoning with configurable thinking token budgets, achieving frontier performance on advanced reasoning while allowing users to inspect reasoning traces.

Mistral / Mixtral (Mistral AI, decoder-only)

  • Mistral 7B (Jiang et al. 2023): Sliding window attention (4096 window), GQA (8 query groups/1 KV head), grouped RoPE — outperforms LLaMA 2 13B on all benchmarks at half the active parameter count. Demonstrates careful architecture design compensates for parameter size. Mixtral 8x7B (Jan 2024): 8 experts / top-2 routing / 47B total / 13B active. Outperforms LLaMA 2 70B while using 4× less compute per token. Mixtral 8x22B (April 2024): 141B/39B active. All Mistral models use 32K context window.

Vision Transformers (ViT, Swin, BEiT)

  • ViT (Dosovitskiy et al. 2021): Split 224×224 image into 14×14 or 16×16 patches (196 or 196 patches), linearly project each to d_model, prepend [CLS] token, add 2D positional embeddings, pass through standard transformer encoder. Classification from [CLS] representation. Requires large-scale pre-training (ImageNet-21K, JFT-300M) to match CNN accuracy. ViT-G/14 (1.8B parameters, trained on 4B images) achieves 90.45% top-1 on ImageNet.
  • Swin Transformer (Liu et al. 2021): Hierarchical feature maps with patch merging layers; shifted window attention capturing cross-window interactions. Provides dense prediction outputs for object detection (COCO mAP 58.7 with Swin-L) and semantic segmentation. DINO (Caron et al. 2021): Self-supervised ViT training via self-distillation with no labels, discovering semantically meaningful patch features. DINOv2 (Oquab et al. 2023): Large-scale self-supervised ViT with 1B parameters; strong frozen features for dense tasks without fine-tuning. SAM (Kirillov et al. 2023, Meta): 1B-parameter ViT-H for zero-shot image segmentation, trained on 11M images with 1.1B masks.

Audio Transformers

  • Whisper (Radford et al. 2022, OpenAI): 680M encoder-decoder. Audio encoded as log-mel spectrograms (80 mel bins, 25ms frames, 10ms hop → 1500 tokens per 30s window after 2D Conv). Decoder generates text autoregressively with language token and task token conditioning. Trained on 680K hours from the web in 99 languages via weak supervision (no manual transcription). Achieves 5.5% WER on LibriSpeech test-clean, human-level for English. Multilingual ASR and translation capabilities.
  • wav2vec 2.0 (Baevski et al. 2020): CNN encoder extracts local features from raw waveforms; transformer contextualises; contrastive learning distinguishes true quantised latents from distractors. Fine-tuning on 10 minutes of labelled speech achieves 5.2% WER on LibriSpeech test-clean. AudioLM (Borsos et al. 2023): Hierarchical transformer modelling semantic tokens (from HuBERT) then acoustic tokens (SoundStream), generating high-quality long-form audio preserving speaker identity. MusicLM (Agostinelli et al. 2023): Conditions AudioLM on text embeddings from MuLan for text-to-music synthesis.

Transformers in Science

  • AlphaFold 2 (Jumper et al. 2021, Nature, DeepMind): Evoformer module processes multiple sequence alignments (MSA) via row-wise gated self-attention (over residues sharing a sequence) and column-wise attention (over sequences at a given residue), with pair representation updated via triangle attention enforcing geometric consistency. Structure module applies Invariant Point Attention (IPA) directly on 3D backbone frames. Achieves median TM-score >0.9 (near-experimental accuracy) on CASP14. AlphaFold DB covers >200M structures (all of UniProt); freely available at alphafold.ebi.ac.uk in collaboration with EMBL-EBI (Hinxton, Cambridge, UK).
  • ESM-2 (Lin et al. 2023, Meta): 15B-parameter protein language model pre-trained on 250M UniRef50 sequences with masked language modelling. Residue representations capture evolutionary and structural information; ESMFold (paired regression head) achieves AlphaFold-class structure prediction 60× faster. ESM-3 (2024) integrates sequence, structure, and function as three parallel tracks in a unified generative model, enabling programmable protein design.
  • Scientific reasoning: Minerva (Lewkowycz et al. 2022): PaLM fine-tuned on mathematical documents achieving 50.3% on MATH competition problems via chain-of-thought. AlphaCode (Li et al. 2022, DeepMind): 41.4B encoder-decoder trained on GitHub code achieving competitive performance on Codeforces (top 50th percentile). AlphaCode 2 (2023) reaches top 15th percentile on competitive programming.

Academic Context

  • The transformer emerged from Google Brain and Google Research in 2017, with eight co-authors: Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin. The architecture built on earlier attention work by Bahdanau et al. (2015, additive attention for RNN-based MT) and Luong et al. (2015, multiplicative attention). The key innovation was “self-attention” — attending within a single sequence rather than across encoder-decoder pairs — enabling full encoder stacks without recurrence.
  • Theoretical understanding has advanced via multiple programmes. Tay et al. (2022, Efficient Transformers: A Survey) catalogued 40+ efficient variants. Elhage et al. (2021, “A Mathematical Framework for Transformer Circuits”, Anthropic) formalised the residual stream view: each transformer layer additively writes to a shared residual stream, enabling independent attribution. Olsson et al. (2022) identified induction heads as elementary in-context learning circuits appearing during a sharp phase transition at ~100M training tokens.
  • Scaling laws (Kaplan et al. 2020, OpenAI; Hoffmann et al. 2022 “Chinchilla”, DeepMind) provided compute-optimal training recipes. Kaplan et al. found loss L ∝ N^(-α) with α≈0.076 (parameters) and L ∝ D^(-β) with β≈0.095 (tokens). Chinchilla (Hoffmann et al.) corrected that most 2020-era models were undertrained: a 70B model trained on 1.4T tokens outperforms GPT-3 175B trained on 300B tokens. Post-Chinchilla, the field shifted toward smaller, heavily-trained models (LLaMA series, Mistral, Phi).
  • Mechanistic interpretability has matured into a systematic programme. Nanda et al. (2023) reverse-engineered the algorithm implemented by groking transformers on modular arithmetic, discovering a Fourier-based computation strategy. Marks et al. (2024, Anthropic) scaled sparse autoencoder analysis to Claude 3 Sonnet, decomposing residual stream superposition into thousands of interpretable monosemantic features. The linear representation hypothesis (Park et al. 2023) shows transformers systematically encode conceptual categories as linear directions enabling arithmetic operations on representations.

Current Landscape (2026)

  • As of May 2026, decoder-only transformers with RoPE positional encoding, GQA, SwiGLU FFNs, RMSNorm pre-norm, and FlashAttention constitute the near-universal standard for large language models. The architectural recipe has stabilised: organisations now compete on training data quality, scale, and post-training alignment rather than architecture innovation.
  • MoE adoption has accelerated substantially. Mixtral 8x7B (open), DeepSeek-V2 (236B/21B, open), Grok-1 (314B/86B, open), and Qwen1.5-MoE (14.3B/2.7B active, Apache 2.0) demonstrate that sparse routing enables frontier capability at proportional inference cost. The trend toward MoE is driven by inference economics: serving a 13B active-parameter MoE model costs the same per-token as a 13B dense model, but the MoE model has seen 50-100B parameter training capacity.
  • Context windows have extended from 2K (GPT-2 era) through 128K (Claude 3, GPT-4 Turbo) to 1M+ (Gemini 1.5 Pro via ring attention). Long-context quality remains heterogeneous: models often retrieve effectively from the first and last ~30% of a 128K context but lose information from the middle (“lost in the middle” phenomenon, Liu et al. 2023).
  • State-space model alternatives — particularly Mamba (Gu & Dao 2023) with selective state spaces — have shown competitive perplexity on language modelling benchmarks at ≤3B parameter scale with O(n) inference. Hybrid Mamba-attention models (Jamba: 52B interleaved SSM/attention; Zamba2: 7B MoE-SSM; Falcon Mamba: 7B pure SSM) represent the active research frontier for architectures beyond pure attention.
  • Inference efficiency has improved dramatically independent of architecture changes: PagedAttention (vLLM, 23× throughput improvement), continuous batching, speculative decoding (2-4× latency), and FP8 training/inference (H100 Transformer Engine, 2× compute utilisation) have collectively reduced cost-per-token by ~10× between 2023 and 2026.

UK Context (Imperial / Edinburgh / UCL / Cambridge / Manchester / Northern England)

  • DeepMind (London and Cambridge): AlphaFold 2 (Jumper et al., Nature 2021) was developed primarily at DeepMind London with EMBL-EBI Cambridge as database partner. The Evoformer transformer variant achieves sub-ångström protein structure accuracy and is freely accessible via alphafold.ebi.ac.uk. DeepMind also produced Gato (2022, 1.2B generalist transformer across 600+ tasks), Flamingo (80B vision-language, Alayrac et al. 2022), and the Gemini series (post-Google Brain merger 2023), with substantial UK-based research contributing to architecture and training methodology.
  • University of Edinburgh (School of Informatics): Rico Sennrich co-invented byte-pair encoding (BPE, Sennrich, Haddow & Birch, ACL 2016) — now the dominant subword tokenisation method in virtually every deployed transformer LLM. The Edinburgh NLP group produced influential neural machine translation work (Nematus system) pre-dating and post-dating the transformer, and continues active research in low-resource MT, structured prediction, and multilingual modelling. Adam Lopez’s group works on structured prediction and transformer interpretability.
  • University of Cambridge (Computer Laboratory and Language Technology Lab): Phil Woodland’s speech group interfaces with transformer-based ASR (Cambridge LVCSR system contributions to HTK/Kaldi). The statistical machine learning group (Zoubin Ghahramani, now VP at Google DeepMind but long-term Cambridge faculty) has contributed Bayesian perspectives on uncertainty in neural networks applied to transformers. EMBL-EBI (European Bioinformatics Institute, Wellcome Genome Campus, Hinxton) hosts AlphaFold DB and collaborates on ESM-based protein structure prediction.
  • Imperial College London (I-X Centre and Department of Computing): Björn Schuller’s AESOP group applies transformers to affective computing, speech emotion recognition, and audio-visual sentiment analysis. Murray Shanahan (cognitive robotics, now at Google DeepMind) published seminal thinking on role-playing as a transformer-based agent (2023) relevant to AI alignment. The I-X Centre for Intelligent Systems and their intersecting disciplines fund multimodal transformer research at the intersection of computing and medicine.
  • UCL (Centre for Artificial Intelligence and Department of Computer Science): Pontus Stenetorp works on NLP, few-shot learning, and cross-lingual transfer with transformer fine-tuning. Sebastian Riedel (formerly UCL, now VP Research at Meta AI) contributed transformer-based relation extraction and open-domain QA. UCL collaborates with the Alan Turing Institute on transformer interpretability and fairness.
  • University of Manchester (National Centre for Text Mining, NACTEM): Sophia Ananiadou’s group is internationally recognised for biomedical text mining using transformer-based NLP. Their PubMedBERT variants and clinical NLP systems are deployed in NHS data pipelines for information extraction from clinical notes, processing millions of documents across UK Biobank and CPRD.
  • Alan Turing Institute (London, with distributed UK nodes): Cross-institutional transformer research programme spanning healthcare AI (NHS Federated Data Platform interactions), climate modelling (transformer-based weather prediction analogues to FourCastNet), and fairness and interpretability (with Edinburgh, Manchester, UCL, Cambridge). The UK AI Safety Institute (AISI, Bletchley Park 2023) conducts pre-deployment evaluations of frontier transformer models for dangerous capability thresholds.
  • Northern England Industry and Research: Sheffield NLP group (Mark Stevenson, Rob Gaizauskas) applies transformers to legal text analysis, clinical IE, and low-resource languages; informs Sheffield-based legal-tech startups. Leeds Data Analytics and intelligence (LIDA) uses transformer models for health informatics. Newcastle Digital Institute works on transformer-based accessibility tools (automatic sign language recognition, augmentative communication). Manchester-based AI companies (e.g. Peak AI, Robiquity) deploy fine-tuned transformer models for industrial analytics and supply-chain optimisation.

Future Directions (2026-2030)

  • Hybrid SSM-Attention Architectures: Jamba, Mamba-2, and Zamba interleave SSM layers (O(n) inference complexity) with sparse attention layers (O(n²) training expressiveness), targeting the quadratic bottleneck without full quality degradation. The fundamental question — whether selective state-space recurrence is computationally equivalent to attention for all practically relevant tasks — remains open and is the subject of theoretical work in 2025-2026.
  • Mixture-of-Depths (MoD): Raposo et al. (2024) route tokens to bypass transformer layers entirely for easy tokens, reducing active compute proportionally. Complementary to MoE: MoE selects which FFN expert, MoD selects whether to apply the layer at all. Early experiments show 50% compute reduction with <1% quality loss on language modelling.
  • Extended Context to 10M+ Tokens: Ring attention and derivative techniques enable arbitrary context lengths limited only by total device memory. Applications: codebase-level reasoning (entire Linux kernel as context), novel-length document understanding, long-horizon planning. Key unsolved challenges: positional encoding extrapolation, attention sink phenomena, and attention heat diffusion degrading retrieval at extreme lengths.
  • Natively Multimodal Training: Moving beyond late-fusion (separate encoders + cross-attention at the language model boundary) to unified tokenisation of text, image patches, audio frames, video clips, and action sequences in a single autoregressive model. GPT-4o and Gemini 1.5 represent early instances; fully unified training at frontier scale is the next milestone.
  • Mechanistic Interpretability at Scale: Scaling sparse autoencoder and circuit-analysis methods to 100B+ parameter models. The UK groups — Anthropic London, Edinburgh, UCL — are central to this agenda. If successful, mechanistic interpretability will enable verification of alignment properties rather than empirical red-teaming alone.
  • Hardware Co-Design: Transformers drove custom silicon (Google TPU, Cerebras WSE, Groq LPU). Future architectures — particularly MoE and SSM hybrids — are co-designed with hardware from the start. Groq’s LPU achieves 300-500 tokens/sec for 70B models via deterministic SRAM execution; Cerebras WSE-3 eliminates HBM bottleneck with 44GB on-chip SRAM. Hardware-aware architecture search will increasingly influence what transformer variants are deployed at scale.

Training Transformers: Objectives, Tokenisation, and Infrastructure

Pre-Training Objectives

  • Masked Language Modelling (MLM): Used by BERT and encoder-only models. Replace 15% of input tokens with [MASK] (80%), a random token (10%), or leave unchanged (10%); predict the original tokens from bidirectional context. Provides rich contextual representations capturing both syntactic and semantic relationships. Cannot generate autoregressively without modification (BERT cannot extend a sequence). Cross-entropy loss computed only on masked positions reduces per-step gradient signal compared to CLM.
  • Causal Language Modelling (CLM): Standard for decoder-only models. Predict each token x_t from all preceding tokens x_{<t}: L = −∑_t log P(x_t | x_{<t}; θ). Scales naturally to generation tasks with no architectural modification. GPT-3 trained on 300B tokens; LLaMA 3 on 15T tokens. The autoregressive structure allows the model to condition on arbitrary prompt prefixes during inference, enabling in-context learning.
  • Span Corruption (T5): Mask contiguous spans of tokens of variable length, replace each span with a single sentinel token, predict the missing spans in the decoder. More information-dense than MLM: each masked sentinel in the target sequence requires predicting multiple tokens. Encoder-decoder models train naturally with this objective.
  • Contrastive Learning (CLIP, wav2vec 2.0): InfoNCE loss maximises agreement between matched pairs and minimises agreement for mismatched pairs using a temperature-scaled similarity. Enables cross-modal alignment without explicitly labelled datasets.

Tokenisation

  • Byte-Pair Encoding (BPE, Sennrich, Haddow & Birch, ACL 2016, Edinburgh NLP group): Starts from character-level vocabulary, iteratively merges the most frequent symbol pair until a target vocabulary size is reached. Vocabulary sizes: GPT-2 (50,257), LLaMA (32,000), LLaMA 3 / GPT-4 (128,256 via Tiktoken). Handles out-of-vocabulary words by decomposing to sub-word units; byte-level BPE (GPT-2) falls back to raw UTF-8 bytes, eliminating unknown tokens entirely.
  • SentencePiece (Kudo & Richardson 2018, Google): Unsupervised tokenisation treating input as a raw character stream, enabling language-agnostic tokenisation without pre-tokenisation whitespace splitting. Used by T5, LLaMA 1/2, multilingual models. Unigram language model variant selects the most probable tokenisation under a unigram language model trained with EM.
  • Tiktoken (OpenAI): BPE implementation optimised for speed with special token handling. Used by GPT-4, ChatGPT, LLaMA 3. Vocabulary of 100,277 tokens (GPT-4) encoding English at ~1.3 tokens/word average.
  • Token fertility (average tokens per word) varies by language and tokeniser training corpus. English: ~1.2-1.4 tokens/word. Code: ~2-4 tokens/word (delimiters and whitespace tokenised separately). Languages not well-represented in training corpus (e.g. Swahili, Tamil) can reach 5-10+ tokens/word, severely disadvantaging those languages in context-length-limited applications — a significant equity concern in multilingual deployment.

Distributed Training Infrastructure

  • Training a 70B transformer model requires ~1.4 million GPU-hours on A100s (roughly 1000 GPUs for ~58 days). Four parallelism strategies compose hierarchically: Data Parallelism (DP) replicates model across devices, splits batches, synchronises gradients via all-reduce after each step; Tensor Parallelism (TP, Megatron-LM) shards individual matrix multiplications along head or hidden dimensions across devices (typically 8-16 within a node); Pipeline Parallelism (PP) assigns consecutive layer groups to device stages with micro-batching to overlap forward/backward passes across stages; Expert Parallelism (EP) distributes MoE expert networks across device groups with all-to-all communication for routing.
  • ZeRO (Zero Redundancy Optimizer, Rajbhandari et al. 2020, DeepSpeed): Partitions optimiser states (ZeRO-1), gradients (ZeRO-2), and parameters (ZeRO-3) across data-parallel ranks, eliminating redundant storage. ZeRO-3 enables training models with 100× more parameters per GPU than naive DDP with only 1.5× communication overhead. GPT-3 was trained with Megatron + ZeRO combination.
  • Mixed precision (BF16/FP16 + FP32 master weights): Forward and backward passes in reduced precision reduce memory 2× and leverage NVIDIA Tensor Core throughput. BF16 preferred over FP16 for stability (larger dynamic range, fewer overflow/underflow issues). H100 Transformer Engine with FP8 further doubles effective throughput with hardware-managed scaling factors.
  • Gradient checkpointing (activation rematerialisation): Instead of storing all intermediate activations for the backward pass (O(L·n·d) memory for L layers), only store activations at checkpoint boundaries and recompute intermediate activations during backward. Reduces activation memory from O(L·n·d) to O(√L·n·d) at cost of ~33% additional compute. Essential for training deep models with long sequences under GPU memory constraints.

Fine-Tuning and Alignment

Supervised Fine-Tuning (SFT)

  • After pre-training, models are fine-tuned on high-quality human-written demonstrations of target behaviours. SFT datasets: OpenAssistant OASST (161K human-written conversation turns); FLAN collection (2,000+ NLP tasks formatted as instructions); Alpaca (52K GPT-4-generated instruction-response pairs); ShareGPT (100K+ real ChatGPT conversations from volunteer users). SFT typically requires 1-3 epochs on 10K-100K examples — representing <0.01% of pre-training compute — yet significantly improves instruction following and reduces harmful outputs.
  • The catastrophic forgetting problem (fine-tuning on narrow distributions degrades pre-trained general capabilities) is mitigated by: mixing SFT data with pre-training replay samples; using elastic weight consolidation (EWC) penalty; or applying PEFT methods that keep most parameters frozen.

RLHF (Reinforcement Learning from Human Feedback)

  • Standard pipeline (Ouyang et al. 2022): (1) SFT on demonstrations; (2) Preference data collection (50-100K prompts, annotators rank outputs); (3) Reward model training via Bradley-Terry loss; (4) PPO optimisation with KL penalty β·KL(π_θ||π_ref) to prevent reward hacking. Known failure modes: reward model misspecification, sycophancy (agreeing with preferred statements regardless of truth), length bias.

DPO (Direct Preference Optimisation)

  • Rafailov et al. (2023): Direct optimisation of RLHF objective from preference pairs without a separate reward model, using reparametrised implicit reward. Eliminates reward model training and PPO loop; reduces alignment to supervised learning on preference pairs. Widely adopted (Zephyr, Tulu, OpenHermes). Variants: IPO (regularised), SimPO (reference-free, length-normalised), KTO (binary per-sample feedback).

Parameter-Efficient Fine-Tuning (PEFT)

  • LoRA (Hu et al. 2022): Add low-rank matrices ΔW = BA ∈ ℝ^(d×k) with rank r ≪ min(d,k) to frozen weight matrices W, training only B ∈ ℝ^(d×r) and A ∈ ℝ^(r×k). Typical r=4-64 training 0.1-1% of total parameters while maintaining quality comparable to full fine-tuning. A initialised from random Gaussian, B from zeros ensuring ΔW=0 at initialisation. Applied to Q, K, V, O projection matrices in attention layers.
  • QLoRA (Dettmers et al. 2023): Combines 4-bit NormalFloat (NF4) quantisation of frozen base model weights with LoRA adapter training. Double quantisation further compresses quantisation constants. Enables fine-tuning of 65B-parameter LLaMA on 2×48GB GPUs, democratising alignment of frontier models. The quantisation error is absorbed into the adapter training signal.
  • Prefix Tuning (Li & Liang 2021): Prepend L trainable “virtual tokens” to key-value context at each layer, attending over these prefixes alongside the regular sequence. Only 0.1% of parameters updated. Less stable than LoRA for large models; superseded in most applications by LoRA variants.
  • IA³ (Infused Adapter by Inhibiting and Amplifying Inner Activations): Rescale keys, values, and FFN activations with learned vectors, training only 3 × L × d_model parameters (three vectors per layer). Extremely parameter-efficient; suited to many-task inference where different adapters are swapped per request.

Efficient Attention and Long-Context Techniques

Flash Attention and IO-Aware Computation

  • FlashAttention (Dao et al. NeurIPS 2022) reorders the standard attention computation to tile across the sequence length dimension, keeping each tile in GPU SRAM (on-chip, fast) rather than materialising the full n×n attention matrix in HBM (off-chip, 20× slower bandwidth). Each tile computes a numerically stable softmax via running max/sum tricks. Memory complexity reduces from O(n²) to O(n); runtime 2-4× faster on A100 vs PyTorch standard attention at 4K sequence length, 10× faster at 16K.
  • FlashAttention-2 (Dao 2023): Eliminates non-matrix multiplications in the inner loop; exploits sequence parallelism across thread blocks; achieves 73% of A100 peak FLOPS for attention forward pass. FlashAttention-3 (2024): Targets H100 hardware with async WGMMA (Warp Group Matrix Multiply-Accumulate) and ping-pong pipeline overlapping GEMM with softmax, reaching 75% H100 peak FLOPS. All variants produce bit-exact identical outputs to naive attention — not approximations.

Sparse and Linear Attention

  • Sparse Transformer (Child et al. 2019): Strided attention alternates between local and strided patterns, reducing O(n²) to O(n√n). Each head sees either the n nearest tokens (local) or every √n-th token (strided), combining to cover all positions in two heads. Linformer (Wang et al. 2020): Projects K and V to r-dimensional representations (r ≪ n) via learned projection matrices E, F ∈ ℝ^(n×r), reducing attention to O(n·r) with controlled approximation error growing as r increases.
  • Performer (Choromanski et al. 2021): Approximates softmax attention using random feature maps φ such that exp(q·k/√d) ≈ φ(q)^T φ(k). Enables exact O(n) attention via kernel trick. Flash Linear Attention (Yang et al. 2024): Parallelisable recurrent formulation of linear attention matching Flash Attention efficiency while extending to hardware-efficient implementation.

Ring Attention and Context Parallelism

  • Ring attention (Li et al. 2023): Distributes the sequence across a ring of T devices, with each device holding n/T tokens. During attention, each device’s queries are applied to its own KV block, then KV blocks rotate around the ring. One full rotation computes complete attention for all query positions. Enables arbitrary sequence lengths limited only by total device memory (T × VRAM). Gemini 1.5 Pro (1M-token context) uses ring attention across a ring of TPU pods.
  • Context Parallelism (NVIDIA Megatron-LM CP): Similar to ring attention but with flash attention kernel integration, enabling sequence parallelism within a node alongside tensor parallelism and pipeline parallelism for comprehensive 4D parallelism during training.

Benchmark Performance (2025-2026)

Language Understanding and Reasoning

  • MMLU (Massive Multitask Language Understanding, 57 academic subjects, 5-shot): GPT-4o (2024): 88.7%; Claude 3 Opus (2024): 86.8%; Gemini Ultra (2024): 90.0%; LLaMA 3 405B (2024): 87.3%; Human expert: 89.8%. Claude 3.7 Sonnet with extended thinking (2025): exceeds 90% on several reasoning-heavy MMLU subcategories.
  • MATH (mathematical competition, 4-shot chain-of-thought): GPT-4o: 76.6%; Claude 3.5 Sonnet: 71.1%; LLaMA 3 70B: 58.4%; o1-preview (OpenAI reasoning model): 94.2%. GPQA Diamond (expert-level science QA): Claude 3.7 Sonnet with thinking: 84.8%; GPT-4o: 53.6%; PhD-level human: 69.7%.

Code Generation

  • HumanEval (Python function synthesis, pass@1): Claude 3.5 Sonnet (June 2024): 92.0%; GPT-4o: 90.2%; LLaMA 3 70B: 81.7%; Mistral Large: 45.1%; GPT-3.5: 48.1%. SWE-bench (real-world GitHub issue resolution): Claude 3.7 Sonnet: 70.3% resolve rate; OpenHands + Claude 3.5 Sonnet: 53.0%; GPT-4o: 33.2%. SWE-bench represents the current frontier for agentic coding benchmarks.

Vision and Multimodal

  • ImageNet top-1 accuracy (image classification): ViT-G/14 (CLIP pre-trained): 90.45%; Swin-V2-G: 90.17%; CoCa ViT-G: 91.0%. COCO object detection mAP: DINO-swin-L: 63.3%; DETA-Swin-L: 63.5%; ViT-Adapter-L: 62.1%. VQAv2 (visual question answering): GPT-4V: 77.2%; Claude 3 Opus: 70.6%; LLaVA-1.6-34B: 67.1%.

Long-Context Evaluation

  • RULER (128K context retrieval and reasoning): Gemini 1.5 Pro: 91.2%; Claude 3 Opus: 87.8%; GPT-4 Turbo: 83.7%; LLaMA 3 70B extended: 73.9%. “Lost in the middle” degradation (Liu et al. 2023): retrieval accuracy of facts placed at positions 0% and 100% of 128K context typically 90-95%; at position 50% drops to 50-70% even for frontier models. Frontier as of Q2 2026 shows improvement but not elimination of this effect.

Inference Efficiency Comparison

  • Throughput on NVIDIA A100 80GB (batch=1, 2K output, BF16): LLaMA 2 7B (vLLM): ~110 tokens/sec; Mistral 7B (vLLM): ~125 tokens/sec; LLaMA 3 8B (vLLM): ~108 tokens/sec; Mixtral 8x7B (2× A100, expert parallel): ~70 tokens/sec at 47B parameter quality. Speculative decoding (LLaMA 3 8B draft + 70B target): ~220-280 tokens/sec effective throughput. Groq LPU (proprietary, deterministic SRAM): LLaMA 3 70B at 300-500 tokens/sec.

Comparisons to Alternatives: RNNs, CNNs, State-Space Models

Why Transformers Won Over RNNs and LSTMs

  • Recurrent Neural Networks create inherently sequential computation — step t depends on step t-1 — and suffer vanishing gradients limiting effective context to ~200-500 tokens even with LSTM gating. Transformers compute all positions simultaneously in O(n²) parallel operations (O(1) sequential complexity), utilising GPU clusters ~100× more efficiently. Direct long-range paths (O(1) vs O(n) for RNNs) explain transformer dominance on Long Range Arena (Tay et al. 2021) benchmarks requiring cross-position dependencies: pathfinder-X, ListOps, retrieval.

Mamba and Selective State Space Models

  • S4 (Gu et al. 2022): Structured SSM with HiPPO state matrix, O(n) time and memory, matches transformers on Long Range Arena. Fixed (input-independent) state matrices limit content-based routing. Mamba (Gu & Dao 2023): Selective SSMs where A, B, C are input-dependent, enabling content-based state retention. Achieves competitive perplexity at 1-3B scale with 5× higher inference throughput vs transformers (O(n) vs O(n²)).
  • Mamba weaknesses: (1) Retrieval accuracy for distant tokens degrades as information must persist in compressed state; attention retrieves directly from all previous KV pairs. (2) In-context learning quality lower than transformers at equivalent scale. (3) No demonstrated GPT-4-class quality at frontier scale as of 2026.
  • Hybrid architectures: Jamba (AI21 Labs, 2024, 52B MoE, ~7:1 SSM:attention); Zamba2-7B (2024, MoE-SSM); Falcon Mamba 7B (TII UAE, 2024, first pure 7B SSM competitive with Mistral on academic benchmarks, 5.5T training tokens).

Transformer Interpretability and Mechanistic Analysis

The Residual Stream Framework

  • Elhage et al. (2021, “A Mathematical Framework for Transformer Circuits”, Anthropic) formalises the transformer computation as a sequence of additive updates to a shared residual stream. Each layer reads from the residual stream via attention/FFN and writes back an additive update: x_{l+1} = x_l + Attn_l(x_l) + FFN_l(x_l). This framing enables independent analysis of each component’s contribution — any head or FFN layer can be ablated (zeroed out) or patched to isolate its causal effect on output logits.
  • The residual stream decomposition reveals that attention heads implement low-rank operations (rank bounded by d_k per head) while FFN layers implement higher-rank transformations. Information composition occurs when one head reads from stream components written by earlier heads (Q-K circuit composition enabling higher-order reasoning).

Induction Heads and In-Context Learning

  • Olsson et al. (2022): Identified “induction heads” — two-attention-layer circuits implementing prefix completion. First layer (K-composition) writes a signal “this token followed that token”; second layer (induction head proper) queries for that signal and attends to the subsequent token. Net effect: whenever token A appears, the induction head copies the token that followed A in any previous occurrence in the context. This implements “fuzzy” n-gram retrieval from context.
  • Induction heads form during a sharp phase transition at approximately 100M training tokens, correlated with a sudden improvement in in-context learning performance. Ablating induction heads degrades in-context learning performance 30-50% in 2-layer models. This provides a mechanistic account of how transformers learn to use their context window — not via explicit memory retrieval but via attention patterns implementing local pattern completion.

Sparse Autoencoders and Feature Decomposition

  • Bricken et al. (2023, Anthropic “Towards Monosemanticity”): Trained a one-layer transformer with d_model=512 and discovered that individual neurons are “polysemantic” — each neuron activates for multiple unrelated concepts (e.g. “legal text”, “DNA base pairs”, and “Italian words”). Proposed sparse autoencoders (SAEs): encoder maps residual stream activations to a high-dimensional sparse representation (d_SAE ≫ d_model) using an L1 penalty; decoder reconstructs the original activations. Features in the SAE representation are monosemantic — each corresponds to a single interpretable concept.
  • Marks et al. (2024): Scaled SAE analysis to Claude 3 Sonnet, identifying thousands of features corresponding to specific concepts (countries, professions, emotional states, programming constructs). Found multi-step reasoning circuits: features representing an intermediate reasoning step causally influencing later features representing a conclusion. Demonstrated that activating a “Golden Gate Bridge” feature via SAE intervention causes the model to obsessively reference the bridge regardless of prompt context — validating that SAE features are causally active in model computation, not merely correlated.

Ethical, Safety, and Societal Considerations

Bias and Representation Harms

  • Transformers trained on internet-scale text inherit statistical biases present in that data. Bolukbasi et al. (2016) identified gender stereotypes in word embeddings (Word2Vec era); these persist in transformer representations. Blodgett et al. (2020) documented systematically higher error rates for African-American Vernacular English (AAVE) in NLP systems. Weidinger et al. (2021, DeepMind) catalogued 21 harm categories from LLMs spanning discrimination, information hazards, misinformation, and human-computer interaction harms.
  • Token fertility disparities disadvantage languages not well-represented in training corpora. A Swahili prompt requiring 10× more tokens than an English prompt of equivalent meaning is penalised in context-length-limited deployments and incurs 10× higher API costs. This systematically disadvantages speakers of under-resourced languages in LLM-mediated services.

Hallucination and Factual Reliability

  • Autoregressive transformers generate statistically plausible continuations, not verified facts. “Hallucination” — generating confident-sounding false statements — arises because next-token prediction maximises perplexity reduction, not factual accuracy. Frontier models (Claude 3.5, GPT-4o) achieve ~95% factual accuracy on verifiable QA benchmarks (TriviaQA, NaturalQuestions) in closed-book settings, but hallucinate biographical details, citation data, and legal specifics at meaningful rates.
  • Mitigation approaches: Retrieval-Augmented Generation (RAG) grounding outputs in retrieved document chunks; factuality-focused RLHF reward models penalising verifiable falsehoods; chain-of-thought reasoning making intermediate steps explicit and checkable; calibrated uncertainty expression teaching models to say “I don’t know” rather than confabulate; and model-generated citations pointing to source documents.

Environmental and Compute Costs

  • Training frontier transformers consumes significant energy: GPT-3 training estimated at 1,287 MWh (Patterson et al. 2021); LLaMA 3 405B at ~11,000 MWh (Meta, disclosed); Gemini Ultra training not publicly disclosed but estimated 30,000-100,000 MWh from cluster size and duration. Carbon intensity depends heavily on grid mix: Google’s TPU clusters in Iowa (40-50% renewable) emit less CO₂ than equivalent clusters in data centres with fossil-dominated grids.
  • Inference at scale: ChatGPT processes ~500M+ daily queries in 2024 (estimated). If each query costs 0.001 kWh (reasonable for a 4K-token response on shared hardware), total daily inference energy exceeds 500 MWh — comparable to training a medium-scale model every day. Efficient architecture choices (quantisation reducing energy ~4×, MoE reducing active compute ~5×, distillation producing smaller models) provide the most direct path to reducing inference energy per useful output.

Key Glossary

  • Scaled Dot-Product Attention: The core attention operation computing Attention(Q,K,V) = softmax(QK^T/√d_k)·V. “Scaled” refers to the √d_k denominator preventing gradient vanishing. O(n²d_k) time and O(n²) memory per head. Introduced in Vaswani et al. 2017.
  • Multi-Head Attention (MHA): Running h parallel scaled dot-product attention operations on d_k = d_model/h dimensional projections, concatenating and projecting outputs. Each head can attend to different representation subspaces simultaneously. Standard since Vaswani et al. 2017.
  • Multi-Query Attention (MQA): Sharing a single set of K,V projections across all query heads. Reduces KV cache h-fold. Used in Falcon, early PaLM 2. First proposed by Shazeer 2019.
  • Grouped-Query Attention (GQA): Dividing query heads into G groups each sharing one K,V pair. Interpolates between MHA quality and MQA efficiency. Adopted by LLaMA 2 70B, Mistral, Gemma, Claude 3, Phi-3. Formalised in Ainslie et al. EMNLP 2023.
  • Flash Attention: IO-aware exact attention algorithm tiling computation into SRAM blocks, avoiding full n×n materialisation in HBM. Memory O(n), runtime 2-10× faster. Dao et al. NeurIPS 2022; v2 (2023); v3 (2024, H100 optimised).
  • Rotary Position Embedding (RoPE): Encodes relative positions by rotating Q,K vectors in 2D subspaces. Q_m^T K_n = f(m-n) — relative position enters dot product naturally. Su et al. 2021. Adopted by LLaMA, Mistral, Gemini Multimodal Language Model, Phi, Qwen.
  • ALiBi: Attention with Linear Biases. Adds −m_h|i-j| penalty to attention logits where m_h is a head-specific slope. No position embeddings. Strong length extrapolation. Press et al. ICLR 2022. Used in BLOOM, MPT.
  • YaRN: Yet Another RoPE extensioN. Frequency-aware interpolation of RoPE dimensions to extend context window 32× without full retraining. Peng et al. 2023. Extends LLaMA 2 to 128K context.
  • SwiGLU: Swish-Gated Linear Unit. FFN variant: SwiGLU(x,W,V) = Swish(xW) ⊙ (xV). Outperforms ReLU/GELU at matched compute. Shazeer 2020. Adopted by LLaMA, PaLM, Gemini Multimodal Language Model, T5 v1.1.
  • RMSNorm: Root Mean Square Layer Normalisation. Normalises by RMS(x) = √(mean(x²)); no mean-centring or bias. ~10% faster than Layer Normalization (LayerNorm). Zhang & Sennrich NeurIPS 2019. Adopted by LLaMA, T5, Mistral, Gemma.
  • Pre-Norm: Normalisation applied before each sub-layer. x’ = x + Sublayer(Norm(x)). Standard in all modern large models. More stable gradient flow than Post-Norm (Vaswani 2017 original) for deep networks.
  • Mixture of Experts (MoE): Replacing dense FFN layers with N expert networks, routing each token to top-k (k=1 or 2) via a gating network. Only k/N parameters active per token. Enables trillion-scale capacity at proportional inference cost. Switch Transformer 2021, Mixtral 2024.
  • KV Cache: Storing computed keys and values from previous tokens during autoregressive generation to avoid recomputation. Memory scales as O(n·L·h·d_k) for context length n, layers L, heads h. GQA/MQA reduce by group factor. PagedAttention manages fragmented cache memory.
  • Speculative Decoding: Small draft model proposes k tokens; target model verifies all k in parallel in one forward pass. Accepted tokens (matching target distribution) are kept; first rejected token is resampled. Leviathan et al. 2023. 2-4× throughput improvement with identical output distribution.
  • LoRA: Low-Rank Adaptation. Adds trainable low-rank matrices ΔW = BA to frozen model weights. Typical rank r=8-64, training 0.1-1% of parameters. Hu et al. ICLR 2022. Standard PEFT method for instruction tuning and domain adaptation.
  • RLHF: Reinforcement Learning from Human Feedback. Pipeline: SFT → reward model from preference data → PPO policy optimisation with KL penalty. Ouyang et al. 2022. Used for ChatGPT, GPT-4, LLaMA 2 Chat, Claude.
  • Constitutional AI (CAI): Anthropic’s alignment approach using a set of principles. Model critiques/revises outputs using principles (RLAIF); trains on revised outputs. Reduces dependence on human harm labelling. Bai et al. 2022. Core to Claude training.
  • Byte-Pair Encoding (BPE): Subword tokenisation merging the most frequent symbol pair iteratively until target vocabulary size. Used by GPT-2/3/4 (50K-128K vocab) and LLaMA 3. Sennrich, Haddow & Birch, ACL 2016 (Edinburgh NLP group).
  • Induction Heads: Two-layer attention circuits implementing prefix completion: when token A appears, attend to the token following any prior occurrence of A in context. Primary mechanism for in-context learning. Olsson et al. 2022.
  • Hallucination: Confident-sounding generation of factually incorrect content by autoregressive transformers. Arises from next-token prediction objective not penalising factual errors. Mitigated by Retrieval-Augmented Generation, factuality-focused RLHF, and chain-of-thought transparency.
  • Chinchilla Scaling Laws: Hoffmann et al. 2022 finding that model parameters and training tokens should scale equally for compute-optimal training. A 70B model trained on 1.4T tokens outperforms GPT-3 175B trained on 300B tokens. Shifted the field toward smaller, heavily-trained models.

Cross-Cutting Applications

Agentic Systems

  • Transformers are the reasoning backbone of Agents that use tools, plan multi-step tasks, and operate in agentic loops. Large Language Model agents invoke web search, code execution, file systems, and APIs via structured function calling (tool use, OpenAI function calling API, Anthropic tool use API). Agents generate action sequences, observe environment responses, and update plans — all processed as tokens in the transformer’s context window.
  • ReAct (Yao et al. 2022): Interleaves reasoning (chain-of-thought) and acting (tool invocation) in a single transformer forward pass. SWE-bench agents (OpenDevin, SWE-agent, Claude Agents) coordinate file editing, test execution, and debugging in software engineering tasks — achieving up to 70% GitHub issue resolution as of 2025.
  • Multi-agent transformer systems (as studied by DeepMind, Anthropic, and Microsoft Research) coordinate multiple LLM instances in producer-consumer, debate, or hierarchical orchestration patterns, improving reliability and specialisation while sharing transformer backbone weights.

Code Intelligence and Software Engineering

  • Transformers have transformed software development tooling. Code Generation via decoder-only models (Codex/GPT-4o, Claude 3.5 Sonnet, Code LLaMA, StarCoder 2) handles function synthesis, docstring generation, test generation, and code review. GitHub Copilot (powered by OpenAI Codex then GPT-4o) reports 30-40% of code written by AI assistance in projects where it is enabled.
  • AlphaCode 2 (DeepMind, 2023): Transformer-based competitive programming system reaching top 15th percentile on Codeforces — surpassing 85% of human contestants. Uses an ensemble of specialised models for problem understanding, solution generation, and test-based filtering.
  • Formal verification integration (2025-2026): Lean 4 and Coq proof assistants receive LLM-assisted tactic suggestions. AlphaProof (DeepMind, 2024) proved IMO-level mathematical problems by combining LLM-guided proof search with Lean’s formal verification engine.

Healthcare and Clinical NLP

  • BERT-based clinical NLP models (ClinicalBERT, BioBERT, PubMedBERT) fine-tuned on clinical notes, discharge summaries, and medical literature achieve state-of-the-art on named entity recognition (medications, conditions, procedures), relation extraction (drug-disease, gene-phenotype), and clinical event classification.
  • The University of Manchester’s NACTEM group (Sophia Ananiadou) deploys transformer-based text mining across NHS data, UK Biobank, and CPRD, processing millions of clinical documents for pharmacovigilance and epidemiology. Their work on PubMedBERT variants is integrated into HDRUK Gateway annotation pipelines.
  • GPT-4 and Claude 3 achieve physician-level performance on USMLE (US Medical Licensing Examination) — GPT-4 scoring 86.7% vs 60% passing threshold — raising questions about clinical decision support, medical education, and regulatory oversight of LLM-based diagnostic tools.

Robotics and Embodied AI

  • RT-2 (Google DeepMind, 2023): 562B-parameter vision-language-action model fine-tuned from PaLM-E. Maps images + text instructions to robot arm actions, transferring web knowledge to physical manipulation. Achieves 2× improvement over RT-1 on novel tasks via internet pre-training. Demonstrates that transformer pre-training on web data generalises to physical-world affordances.
  • Reinforcement Learning from large pre-trained transformers: Decision Transformer (Chen et al. 2021) frames RL as sequence modelling — conditioning on desired return generates actions achieving that return. Gato (DeepMind, 2022) handles 600+ tasks (Atari, robotics, image captioning, dialogue) in a single 1.2B-parameter transformer via multi-task training on heterogeneous data.

Quantisation and Edge Inference

Post-Training Quantisation

  • INT8 Weight Quantisation: Quantise model weights to 8-bit integers post-training. Reduces model size ~2× with minimal quality loss on most benchmarks. Supported natively in NVIDIA cuBLAS INT8 GEMM and Intel oneDNN for CPU. LLM.int8() (Dettmers et al. 2022): Mixed-precision decomposition preserving the 0.1% of outlier features in FP16 while quantising the remainder to INT8.
  • GPTQ (Frantar et al. 2022): Post-training quantisation using approximate second-order (Optimal Brain Quantisation) method. Quantises layer weights block-by-block, compensating quantisation error within each block via inverse Hessian updates. Achieves 3-4 bit weight quantisation of LLaMA 65B with ~1.5 perplexity increase — enabling 65B models on 2× A100 80GB. Supported in AutoGPTQ and HuggingFace.
  • AWQ (Activation-aware Weight Quantisation, Lin et al. 2023): Observes that 1% of weights (those multiplied by large activations) are disproportionately important for model quality. Protects these salient weights at higher precision; quantises remaining 99% to INT4/INT3. Achieves lower perplexity than GPTQ at matched bit-width. Supported in AutoAWQ, TensorRT-LLM, llama.cpp.
  • GGUF/GGML (llama.cpp): File format supporting block-quantised weights. Q4_K_M (4-bit quantised with k-quant mixed precision on attention layers): LLaMA 3 8B at 4.6GB RAM — runs on MacBook Air. Q8_0 (8-bit, near-lossless): 8.5GB. Enables local inference on consumer hardware without GPU. Foundation for Ollama, LM Studio, Jan desktop apps.

Quantisation-Aware Training

  • BitNet b1.58 (Ma et al. 2024, Microsoft Research): Train transformers with ternary weights {-1, 0, +1} from scratch. Uses STE (straight-through estimator) for gradients through the quantisation step function. At 7B scale, BitNet b1.58 matches LLaMA-3 7B on perplexity while enabling binary multiply-accumulate operations on dedicated hardware (no FP GEMM needed). Energy reduction estimated at 3-6× vs FP16 equivalents.
  • LLM.int8() and SmoothQuant: Post-training activation quantisation (weights + activations both INT8) for Tensor Core utilisation. SmoothQuant (Xiao et al. 2022) migrates quantisation difficulty from activations (which have outliers) to weights (which are smoother) via a mathematically equivalent per-channel smoothing transformation, achieving W8A8 without quality degradation.

Scaling Laws and Data Efficiency

Power-Law Scaling

  • Kaplan et al. (2020): Cross-entropy loss L scales as power laws in parameters N, tokens D, and compute C. Compute-optimal allocation: equal scaling of N and D with C budget. Hoffmann et al. (2022, Chinchilla): Corrected Kaplan et al. with finer-grained sweep; optimal N_opt ∝ C^(0.5), D_opt ∝ C^(0.5). Showed GPT-3 (175B, 300B tokens) substantially compute-suboptimal; Chinchilla 70B on 1.4T tokens outperforms GPT-3 on all benchmarks at equal compute cost.
  • Post-Chinchilla practice: LLaMA 3 (8B on 15T tokens) and Mistral 7B significantly over-train relative to compute-optimality for training cost, optimising for inference efficiency — a smaller heavily-trained model serves more queries per GPU-hour than a larger lightly-trained one. This inference-economics rationale now dominates model training decisions.

Data Quality vs Quantity

  • Muennighoff et al. (2023, “Scaling Data-Constrained Language Models”): Shows that repeating training data up to 4 epochs is beneficial when unique data is scarce. Beyond 4 epochs, gains plateau. The study found that FineWeb (200B unique tokens, carefully filtered Common Crawl) yields better model quality than 2× as many tokens of noisier Common Crawl.
  • Phi series (Microsoft Research, 2023-2024): Phi-1 (1.3B, “textbook quality” synthetic data) achieves 50.6% HumanEval pass@1 — outperforming much larger models trained on raw web data. Phi-2 (2.7B) matches LLaMA 2 7B on reasoning benchmarks. Demonstrates that training data quality can substitute for scale within a domain. Phi-3 (3.8B-14B, April 2024) extends this approach to achieve frontier small-model performance.
  • The FineWeb dataset (HuggingFace, 2024) and its FineWeb-edu subset represent the current state of the art in web data curation: aggressive filtering via educational quality classifier, near-deduplication via MinHash LSH, language identification, and heuristic quality filters. Models trained on FineWeb-edu consistently outperform comparable models on educational and reasoning benchmarks.

Research & Literature

    1. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017). “Attention Is All You Need.” NeurIPS 2017. arXiv:1706.03762.
    1. Devlin J, Chang MW, Lee K, Toutanova K (2019). “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” NAACL 2019. arXiv:1810.04805.
    1. Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I (2019). “Language Models are Unsupervised Multitask Learners.” OpenAI Blog.
    1. Brown TB, Mann B, Ryder N et al. (2020). “Language Models are Few-Shot Learners.” NeurIPS 2020. arXiv:2005.14165.
    1. Raffel C, Shazeer N, Roberts A et al. (2020). “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.” JMLR 21(140). arXiv:1910.10683.
    1. Dosovitskiy A, Beyer L, Kolesnikov A et al. (2021). “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.” ICLR 2021. arXiv:2010.11929.
    1. Su J, Lu Y, Pan S, Murtadha A, Wen B, Liu Y (2021). “RoFormer: Enhanced Transformer with Rotary Position Embedding.” arXiv:2104.09864.
    1. Press O, Smith NA, Lewis M (2022). “Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.” ICLR 2022. arXiv:2108.12409.
    1. Dao T, Fu DY, Ermon S, Rudra A, Re C (2022). “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.” NeurIPS 2022. arXiv:2205.14135.
    1. Fedus W, Zoph B, Shazeer N (2022). “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” JMLR 2022. arXiv:2101.03961.
    1. Kaplan J, McCandlish S, Henighan T et al. (2020). “Scaling Laws for Neural Language Models.” arXiv:2001.08361.
    1. Hoffmann J, Borgeaud S, Mensch A et al. (2022). “Training Compute-Optimal Large Language Models.” NeurIPS 2022. arXiv:2203.15556.
    1. Touvron H, Lavril T, Izacard G et al. (2023). “LLaMA: Open and Efficient Foundation Language Models.” arXiv:2302.06825.
    1. Touvron H, Martin L, Stone K et al. (2023). “Llama 2: Open Foundation and Fine-Tuned Chat Models.” arXiv:2307.09288.
    1. Jiang AQ, Sablayrolles A, Mensch A et al. (2023). “Mistral 7B.” arXiv:2310.06825.
    1. Jiang AQ, Sablayrolles A, Roux A et al. (2024). “Mixtral of Experts.” arXiv:2401.04088.
    1. Ainslie J, Lee-Thorp J, de Jong M et al. (2023). “GQA: Training Generalised Multi-Query Transformer Models from Multi-Head Checkpoints.” EMNLP 2023. arXiv:2305.13245.
    1. Zhang B, Sennrich R (2019). “Root Mean Square Layer Normalization.” NeurIPS 2019. arXiv:1910.07467.
    1. Ba JL, Kiros JR, Hinton GE (2016). “Layer Normalization.” arXiv:1607.06450.
    1. Shazeer N (2020). “GLU Variants Improve Transformer.” arXiv:2002.05202.
    1. Peng B, Quesnelle J, Fan H, Shippole E (2023). “YaRN: Efficient Context Window Extension of Large Language Models.” arXiv:2309.00071.
    1. Gu A, Dao T (2023). “Mamba: Linear-Time Sequence Modeling with Selective State Spaces.” arXiv:2312.00752.
    1. Jumper J, Evans R, Pritzel A et al. (2021). “Highly accurate protein structure prediction with AlphaFold.” Nature 596, 583-589.
    1. Olsson C, Elhage N, Nanda N et al. (2022). “In-context Learning and Induction Heads.” Anthropic Transformer Circuits Thread. arXiv:2209.11895.
    1. Radford A, Kim JW, Xu T et al. (2022). “Robust Speech Recognition via Large-Scale Weak Supervision.” arXiv:2212.04356.
    1. Liu Y, Ott M, Goyal N et al. (2019). “RoBERTa: A Robustly Optimized BERT Pretraining Approach.” arXiv:1907.11692.
    1. Sennrich R, Haddow B, Birch A (2016). “Neural Machine Translation of Rare Words with Subword Units.” ACL 2016. arXiv:1508.07909.
    1. Dao T (2023). “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.” ICLR 2024. arXiv:2307.08691.
    1. Elhage N, Nanda N, Olsson C et al. (2021). “A Mathematical Framework for Transformer Circuits.” Anthropic Transformer Circuits Thread. transformercircuits.pub.
    1. Bai Y, Jones A, Ndousse K et al. (2022). “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.” Anthropic. arXiv:2204.05862.
    1. Hu EJ, Shen Y, Wallis P et al. (2022). “LoRA: Low-Rank Adaptation of Large Language Models.” ICLR 2022. arXiv:2106.09685.
    1. Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L (2023). “QLoRA: Efficient Finetuning of Quantized LLMs.” NeurIPS 2023. arXiv:2305.14314.
    1. Ouyang L, Wu J, Jiang X et al. (2022). “Training language models to follow instructions with human feedback.” NeurIPS 2022. arXiv:2203.02155.
    1. Rafailov R, Sharma A, Mitchell E et al. (2023). “Direct Preference Optimization: Your Language Model is Secretly a Reward Model.” NeurIPS 2023. arXiv:2305.18290.
    1. Lin Z, Akin H, Rao R et al. (2023). “Evolutionary-scale prediction of atomic-level protein structure with a language model.” Science 379(6637):1123-1130.

Deployment Ecosystem and Tooling

Training Frameworks

  • PyTorch (Meta AI): Dominant research framework (~80% of ML papers as of 2024). Dynamic computation graph (define-by-run) enabling flexible debugging. torch.compile (2.0+) applies TorchDynamo graph capture and TorchInductor backend to JIT-compile model operations, achieving 2-3× training speedup on CUDA. FSDP2 (Fully Sharded Data Parallel 2.0) enables ZeRO-3-equivalent sharding with cleaner API and better composability with TP/PP.
  • JAX (Google DeepMind): Functional framework with XLA (Accelerated Linear Algebra) compiler. jit() transforms pure Python functions to XLA computations; vmap() enables automatic vectorisation; grad() computes exact gradients. Used for Gemini Multimodal Language Model, Gemma, and most Google research models. TPU-native performance advantage over PyTorch due to XLA’s whole-program optimisation and static shapes.
  • Megatron-LM (NVIDIA): Optimised 3D-parallel transformer training for GPU clusters. Provides tensor parallelism primitives, pipeline parallelism schedules (1F1B, V-schedule), and flash attention integration. Used for GPT-3 and LLaMA pre-training recipes at NVIDIA scale. Megatron-Core (2024) provides a more modular version supporting custom architectures.
  • DeepSpeed (Microsoft): ZeRO optimiser states/gradients/parameters partitioning; pipeline parallelism; activation checkpointing; FP16/BF16 mixed precision; quantisation tooling. Integrates with HuggingFace Transformers. Used widely in academic settings for training 1B-70B models on limited GPU clusters.

Inference Servers

  • vLLM (Kwon et al. 2023, UC Berkeley): PagedAttention — KV cache managed in non-contiguous memory pages (analogous to OS virtual memory), allowing dynamic memory allocation and sharing across concurrent requests. Continuous batching inserts new requests into running batches as slots free. Achieves 24× higher throughput over naive HuggingFace inference. OpenAI-compatible REST API. Supports LLaMA, Mistral, Falcon, Claude model formats via GGUF/safetensors.
  • TensorRT-LLM (NVIDIA): Highly optimised CUDA kernels for transformer inference. INT8/FP8/INT4 quantisation via AMQ (automatic mixed quantisation). FP8 GEMM on H100 doubles effective throughput. Deployed in NVIDIA NIM (cloud microservices) for production LLM serving.
  • llama.cpp (Georgi Gerganov, 2023): C/C++ transformer inference supporting GGUF quantisation (Q2_K through Q8_0). Runs LLaMA 3 8B at 20-25 tokens/sec on Apple M3 Pro CPU-only. Enables local inference on consumer hardware. Foundation for Ollama (one-command model management and serving), LM Studio (GUI), and Jan.
  • Ollama (2023): Cross-platform LLM serving with Modelfile abstraction. 8M+ downloads, 1M+ daily active users. Supports 100+ model families via GGUF backend. API compatible with OpenAI client libraries. Key for local Agents and tooling.

Hardware Landscape (2026)

  • NVIDIA H100 SXM5: 80GB HBM3 (3.35TB/s bandwidth), 3.958 TFLOPS BF16 Tensor Core, NVLink 4.0 (900GB/s all-reduce within 8-GPU NVL system). Transformer Engine with FP8 mixed precision auto-scaling. Standard unit for frontier model training and serving.
  • Google TPU v5p: 459 TFLOPS BF16 per chip; 4096-chip pods with 4.8TB/s inter-chip interconnect (ICI) per pod. Used for Gemini Multimodal Language Model 1.5 and Gemma training. TPU v5e (efficient): lower-cost chips for inference.
  • Cerebras WSE-3: 4 trillion transistors on wafer-scale die; 44GB on-chip SRAM (no HBM bottleneck); 125 PFLOPS dense BF16. Eliminates memory bandwidth bottleneck for models fitting on-chip. Achieves 1,000-2,000 tokens/sec for 13B-class models — 10-20× H100 for bandwidth-bound inference.
  • Groq LPU (Language Processing Unit): Deterministic SRAM-based execution model; no caching, no speculative execution; predictable low-latency. Achieves 300-500 tokens/sec for LLaMA 3 70B with sub-100ms first-token latency. Uses compiler-scheduled streaming execution rather than hardware dynamic scheduling.
  • Graphcore IPU (Intelligence Processing Unit): Bulk-synchronous parallel execution; 1472 independent processor tiles per chip; 900MB in-processor memory. Suited to sparse and irregular computation patterns in MoE routing and graph neural networks. Used at Oxford and several UK research institutions.

Cloud Serving Platforms

  • AWS SageMaker JumpStart: One-click deployment of LLaMA, Mistral, Falcon, and Stable Diffusion. SageMaker Endpoints with auto-scaling and managed endpoints. Supports TGI (Text Generation Inference, HuggingFace) and TensorRT-LLM backends.
  • Google Cloud Vertex AI Model Garden: Gemma, LLaMA, Claude via Anthropic partnership, and PaLM API. Vertex AI Pipelines for fine-tuning workflows. Google AI Studio (free tier for experimentation with Gemini Multimodal Language Model models).
  • Azure OpenAI Service: GPT-4, GPT-4o, Ada embeddings, Whisper, DALL-E 3. Enterprise SLAs, private networking, and compliance certifications (ISO 27001, SOC 2, HIPAA). Data residency controls for EU and UK customers (UK South region).
  • Anthropic API: Claude 3.5, 3.7 series. 200K context window. Vision, tool use, and extended thinking mode APIs. Batch API for async workloads at 50% cost reduction. Available via AWS Bedrock and Google Cloud Vertex.
  • Hugging Face Inference Endpoints: Deploy any model from the Hub to managed dedicated endpoints on AWS/Azure/GCP. Supports TGI for LLM serving with token streaming, batching, quantisation.

Metadata

Provenance

  • domain-correction: infrastructure → artificial-intelligence (IRI, URI, same-as, owl-class all corrected)