Deep Learning is a subset of Machine Learning based on Artificial Neural Networks with multiple layers (depth) that learn hierarchical representations of data through Backpropagation and Gradient Descent.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:NeuralNetwork))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:HiddenLayer))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:ActivationFunction))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:LossFunction))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:Optimizer))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:Backpropagation))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:hasPart ai:ModelArchitecture))

Dependency Relationships

SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:requires ai:TrainingData))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:requires ai:GPUCompute))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:requires ai:GradientDescent))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:requires ai:DeepLearningFramework))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:dependsOn ai:Backpropagation))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:dependsOn ai:StochasticGradientDescent))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:dependsOn ai:BatchNormalisation))

Capability Relationships

SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:enables ai:ComputerVision))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:enables ai:NaturalLanguageProcessing))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:enables ai:SpeechRecognition))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:enables ai:ImageClassification))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:enables ai:ObjectDetection))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:enables ai:LargeLanguageModel))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:supports ai:FederatedLearning))

Implementation Relationships

SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:implements ai:SupervisedLearning))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:implements ai:UnsupervisedLearning))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:implements ai:SelfSupervisedLearning))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:implements ai:ReinforcementLearning))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:implements ai:TransferLearning))

Reduction Relationships

SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:reducesTo ai:MachineLearning))
SubClassOf(ai:DeepLearning
  ObjectSomeValuesFrom(ai:reducesTo ai:NumericalOptimisation))

About

Deep learning emerged as a coherent sub-discipline of Machine Learning Discipline in the period 2006–2012, though its intellectual roots extend to McCulloch and Pitts’s 1943 model of the artificial neuron, Rosenblatt’s 1958 Perceptron, and Rumelhart, Hinton & Williams’s 1986 discovery that Backpropagation could train multi-layer Neural Network architectures. The “deep learning” terminology was popularised by Geoffrey Hinton, Yoshua Bengio, and Yann LeCun to distinguish approaches using many layers from the shallow architectures (one or two layers) that dominated the 1990s and 2000s.

The canonical formulation of a deep learning system consists of a Neural Network with L layers, each parameterised by a weight matrix W_l and bias b_l, with a non-linear Activation Function sigma applied element-wise after the affine transformation. Data propagates forward through successive layers: h_l = sigma(W_l h_{l-1} + b_l), with h_0 = x (the input) and h_L = y-hat (the prediction). A scalar Loss Function L(y-hat, y) — cross-entropy for classification, mean squared error for regression — measures discrepancy between prediction and ground truth. Backpropagation computes partial derivatives dL/dW_l and dL/db_l for all l simultaneously via reverse-mode automatic differentiation, and an Optimizer applies these gradients to update parameters in the direction that reduces the loss. This training loop, iterated over many minibatches of Training Data for multiple epochs, is implemented by a Deep Learning Framework executing on GPU Compute hardware.

The qualitative breakthrough that distinguished deep learning from prior neural network work was the discovery that very deep networks, trained at scale with large datasets on GPUs, could learn hierarchical feature representations of unprecedented quality. Convolutional networks for vision learned edge detectors in early layers, shape detectors in intermediate layers, and object-part detectors in deep layers — without any hand-coding of features. This property of compositional representation learning generalises across modalities: Transformer Architecture applied to text learns syntactic structure in lower layers and semantic relationships in higher layers; diffusion models applied to images learn low-frequency structure before high-frequency detail.

Major Architectural Families

Convolutional Neural Networks (CNNs)

Convolutional Neural Network architectures exploit spatial locality and translational equivariance for image and audio processing. Key milestones: LeNet-5 (Yann LeCun, 1998) for handwritten digit recognition; AlexNet (Krizhevsky, Sutskever, Hinton, 2012) which won ImageNet LSVRC with a 16.4% top-5 error, decisively beating the second-place classical method at 26.2%; VGGNet (Simonyan & Zisserman, 2014) demonstrating depth with 3×3 convolutions; ResNet (He et al., 2015) introducing skip/residual connections that solved the Vanishing Gradient Problem for networks of 50–152 layers; EfficientNet (Tan & Le, 2019) systematic compound scaling. Modern CNNs in 2024–2026 include ConvNeXt (Liu et al., 2022) and MaxViT hybrid architectures.

Recurrent Neural Networks (RNNs)

Recurrent Neural Network architectures model sequential data by maintaining a hidden state across time steps. Vanilla RNNs suffer from the Vanishing Gradient Problem over long sequences. Long Short-Term Memory (LSTM; Hochreiter & Schmidhuber, 1997) introduced gating mechanisms to preserve long-range dependencies. Gated Recurrent Units (GRU; Cho et al., 2014) simplified the LSTM gate structure. By 2018–2019, Transformer Architecture models outperformed RNNs on most sequence tasks by processing all positions in parallel rather than sequentially, rendering LSTM largely obsolete for Natural Language Processing though it persists in time-series forecasting, speech synthesis, and real-time control.

Transformer Architecture

The Transformer Architecture (Vaswani et al., “Attention Is All You Need”, 2017) replaced recurrence with the Attention Mechanism — multi-head scaled dot-product attention — operating over all pairs of positions in a sequence in parallel. This architectural choice enabled orders-of-magnitude larger training runs by exploiting GPU parallelism fully, and produced dramatically better representations of long-range dependencies. Transformers underpin all major Large Language Model families: BERT (Devlin et al., 2018; encoder-only, masked language modelling), GPT series (Radford et al., 2018 onward; decoder-only, autoregressive language modelling), T5 (Raffel et al., 2019; encoder-decoder), and multimodal variants (CLIP, Flamingo, GPT-4V). Vision Transformers (ViT; Dosovitskiy et al., 2020) apply patch-based transformers to images, superseding CNNs on large-scale benchmarks when pre-trained at sufficient scale.

Generative Models

Generative Adversarial Network (GAN; Goodfellow et al., 2014) architecture pits a generator network against a discriminator in a minimax game, producing high-fidelity samples. Variational Autoencoders (VAE; Kingma & Welling, 2013) learn a probabilistic latent space. Diffusion Model architectures (Ho et al., 2020; Song et al., 2021) iteratively denoise Gaussian noise, producing state-of-the-art image quality (Stable Diffusion, DALL-E 3, Imagen) and dominating generative AI by 2022–2026. Score-based and flow-matching variants (Lipman et al., 2022) offer improved training stability.

Graph Neural Networks

Graph Neural Networks (GNN; Scarselli et al., 2009; extended by Kipf & Welling, 2016 with GCN; Velickovic et al., 2017 with GAT) extend deep learning to graph-structured data, enabling molecular property prediction, social network analysis, protein structure modelling, and knowledge graph reasoning. AlphaFold 2/3 (DeepMind, 2020/2024) combines Transformer Architecture and invariant point attention on protein sequence and structure graphs, achieving near-experimental accuracy in protein structure prediction and winning the Nobel Prize in Chemistry 2024.

Training Methodology and Regularisation

Deep learning training involves choices across multiple dimensions. Initialisation strategies (Xavier/Glorot, He/Kaiming) set the scale of initial weights to preserve signal variance across layers. Batch Normalisation (Ioffe & Szegedy, 2015) normalises intermediate activations within minibatches, stabilising training and allowing higher learning rates. Dropout (Srivastava et al., 2014) randomly zeroes activations with probability p during training, acting as an ensemble regulariser that prevents co-adaptation of neurons. Residual connections (He et al., 2015) short-circuit the gradient path to earlier layers, eliminating the Vanishing Gradient Problem for very deep networks. Layer Normalisation (Ba et al., 2016) normalises across feature dimensions rather than batch dimensions, preferred in Transformer Architecture models. Weight decay (L2 regularisation) and learning rate warmup-then-decay schedules are standard components of training recipes.

Hyperparameter Tuning involves selecting the learning rate, batch size, weight decay coefficient, dropout rate, number of layers and hidden units, and schedule parameters. Modern approaches use automated methods (Bayesian optimisation via Optuna; hyperparameter sweeps in Weights & Biases) and standardised recipes derived from empirical scaling law studies (Kaplan et al., 2020; Hoffmann et al., 2022 “Chinchilla”).

Scaling Laws and Foundation Models

A transformative insight of 2020–2022 was that Large Language Model performance follows smooth power-law scaling with model parameters N, dataset size D, and compute budget C (Kaplan et al., 2020). The Chinchilla study (Hoffmann et al., 2022) established optimal compute allocation ratios (roughly equal tokens per parameter for training efficiency), reshaping model training practices. This scaling behaviour enabled the development of foundation models — very large models pre-trained on broad data (GPT-4, Llama 3, Gemini, Claude, Mistral, DeepSeek-V3) that are subsequently adapted to specific tasks via fine-tuning, instruction tuning, or Reinforcement Learning from Human Feedback (RLHF).

Transfer Learning — taking a pre-trained model’s learned representations and adapting them to a new task — became the dominant paradigm across Computer Vision and Natural Language Processing by 2018–2019. Fine-tuning pre-trained models on task-specific datasets dramatically reduces the amount of labelled Training Data required. Parameter-efficient fine-tuning methods (LoRA; Hu et al., 2021) update only a small fraction of parameters, enabling foundation model adaptation on consumer hardware.

Use Cases

Computer Vision

Image Classification (ImageNet top-1 accuracy exceeding 91% for EfficientNet-L2 by 2021), Object Detection (YOLO series, Deformable DETR), semantic segmentation (Segment Anything Model, SAM 2), and image generation (Diffusion Model — Stable Diffusion XL, DALL-E 3, Imagen 3) represent the primary deep learning application domains in vision. Medical imaging AI (radiology, pathology, ophthalmology) is a major deployment domain, with 2025 NHS deployments AI-enabling one-third of chest X-ray reads.

Natural Language Processing

Large Language Model pre-training and instruction tuning now accounts for the majority of frontier AI compute expenditure. Applications include code generation (GitHub Copilot, Cursor), text summarisation, translation, question answering, reasoning, and agentic task completion. Retrieval-augmented generation (RAG) combines Large Language Model generation with external knowledge retrieval.

Speech Recognition

Transformer-based speech models (Whisper, wav2vec 2.0) have achieved near-human word-error rates on standard benchmarks. Deep learning-based text-to-speech (TTS) systems (WaveNet, SoundStorm, Voicebox) produce natural-sounding synthesis.

Scientific AI

AlphaFold (2020, 2021, 2022), AlphaFold 3 (2024) for protein structure prediction; Graph Neural Networks for materials discovery (GNoME, DeepMind, 2023 identifying 2.2 million new stable crystal structures); physics-informed neural networks (PINNs) for solving partial differential equations; climate modelling (GraphCast for global weather prediction).

Reinforcement Learning

Reinforcement Learning combined with deep learning (Deep Q-Networks, Proximal Policy Optimisation, AlphaZero) achieved superhuman performance in Atari games, Go, Chess, and StarCraft II. RLHF has become a critical post-training step for aligning Large Language Model outputs with human preferences.

Academic Context

Deep learning’s intellectual lineage begins with the formal theory of artificial neurons (McCulloch & Pitts, 1943), the Perceptron (Rosenblatt, 1958), and the Minsky & Papert (1969) critique that catalysed the first AI winter. The backpropagation rediscovery (Rumelhart, Hinton & Williams, 1986 in Nature; independently Werbos 1974) enabled multi-layer network training but was limited by compute and data until the 2000s. Hinton & Salakhutdinov’s 2006 Science paper on “Reducing the Dimensionality of Data with Neural Networks” — demonstrating deep belief networks for pre-training — reignited the field. Yann LeCun’s convolutional networks applied at scale to MNIST and later ImageNet showed the practical promise. The AlexNet result (Krizhevsky, Sutskever & Geoffrey Hinton, NeurIPS 2012) is widely considered the field’s inflection point, winning ImageNet with a 10.9-percentage-point margin over the second-place classical pipeline.

The 2014–2018 period saw rapid architectural innovation: GANs (Goodfellow et al., NeurIPS 2014), residual networks (He et al., CVPR 2016), attention mechanisms (Bahdanau et al., ICLR 2015), and the Transformer (Vaswani et al., NeurIPS 2017). BERT (Devlin et al., 2018) established the pre-train/fine-tune paradigm for NLP. The GPT series (Radford et al., 2018, 2019; Brown et al., NeurIPS 2020) demonstrated emergent few-shot capability at scale. Yann LeCun, Geoffrey Hinton, and Yoshua Bengio received the ACM Turing Award in 2019 for their foundational contributions to deep learning.

Key research institutions globally include: MILA (Montréal; Yoshua Bengio’s group), Vector Institute (Toronto; Geoffrey Hinton affiliate), FAIR (Meta AI Research; Yann LeCun), Google DeepMind (London), Google Brain (merged with DeepMind 2023), OpenAI, Microsoft Research, and academic groups at Stanford, MIT, CMU, Berkeley, Oxford, Cambridge, Edinburgh, and UCL.

Current Landscape (2026)

As of mid-2026, deep learning has become the dominant paradigm across applied AI. The frontier is defined by large-scale Transformer Architecture models trained with self-supervised objectives on multimodal data (text, images, code, audio, video). The principal frontier models include GPT-4o and o3 (OpenAI), Claude Sonnet/Opus (Anthropic), Gemini 2.0 Ultra (Google DeepMind), Llama 3.3 (Meta), Mistral Large, and DeepSeek-V3/R1. The 2024 Nobel Prize in Chemistry was awarded to David Baker, Demis Hassabis, and John Jumper for protein structure prediction via deep learning (AlphaFold), the first Nobel recognising a deep learning contribution.

Industry spending on AI infrastructure is projected to reach $2.5 trillion globally in 2026 (Gartner). AI model sizes continue to grow — frontier models exceed one trillion parameters for mixture-of-expert architectures — though the research community is simultaneously pushing for efficiency, with models like DeepSeek-R1 and Mistral achieving near-frontier performance at dramatically lower training cost. The paradigm has shifted from pure pre-training scale to techniques combining efficient training (LoRA/QLoRA fine-tuning), inference-time compute scaling (chain-of-thought, GRPO reinforcement learning), and retrieval augmentation. Diffusion Model architectures have consolidated their position as the leading framework for image, video, and audio generation.

In the UK, 23% of businesses use some form of AI (up from 9% in 2023), with deep learning models deployed in NHS radiology (one third of chest X-ray reads AI-enabled), financial fraud detection, autonomous vehicle research, and industrial quality control. Investment in UK AI firms grew from £2.6 billion (2023) to £4.7 billion (2025), with AI companies accounting for a record share of UK venture funding. The government’s AI Opportunities Action Plan (January 2025) and associated infrastructure investments (Isambard-AI supercomputer at Bristol, Cambridge capacity expansion) are building the compute substrate for deep learning research at national scale.

UK Context

The UK has been a foundational contributor to deep learning’s development. Geoffrey Hinton held a professorship at Edinburgh and King’s College London before moving to Toronto, and his backpropagation and deep belief network work was partly conducted in UK settings. DeepMind (founded London, 2010; acquired by Google 2014; merged into Google DeepMind 2023) is the most prominent UK deep learning research organisation, responsible for AlphaGo (2016), AlphaFold (2020/2021/2022/2024), AlphaCode, and Gemini. DeepMind London employs over 1,500 researchers and engineers.

UCL leads the UKRI national generative AI hub, a multi-institution collaboration (with Imperial, Cardiff, Cambridge, Manchester, Edinburgh, Surrey) conducting deep learning research with industry partners including IBM, BT, Google DeepMind, and Cisco. UCL’s machine learning group (Gatsby Computational Neuroscience Unit, Centre for Artificial Intelligence) hosts foundational deep learning theory researchers. Edinburgh’s Bayes Centre and ELIAI group are active in transformers, efficient training, and Bayesian deep learning. Cambridge’s Computer Laboratory contributes to program synthesis, differentiable programming, and large model efficiency. Imperial College London’s Data Science Institute applies deep learning to scientific computing and healthcare.

Manchester’s £120 million AI research hub (opened 2024) — among the UK’s largest academic AI investments — combines GPU computing infrastructure with research into Distributed Training, efficient architectures, and industrial ML. Leeds’ Institute for Data Analytics applies deep learning to manufacturing, supply chain, and healthcare analytics. Newcastle’s Digital Institute conducts deep learning research for autonomous systems and smart cities. Sheffield’s Advanced Manufacturing Research Centre deploys Convolutional Neural Network and deep learning models for industrial quality control and predictive maintenance.

The UK Government’s NHS Long-Term Plan and the Spending Review 2025 (£10 billion for NHS technology by 2028–29) are driving deployment of deep learning-based clinical AI at scale: AI-assisted triage, diagnostic imaging, administrative automation, and clinical decision support are moving from pilots to system-wide deployment. The Health Data Research Service (£600 million commitment, government and Wellcome) will make national health datasets available for deep learning research under controlled access, potentially making the UK the world’s leading site for health AI research.

Future Directions (2026–2030)

Several trajectories will shape deep learning’s development in the near term. Multimodal unification — training single models over text, image, audio, video, code, and action — is advancing rapidly, with frontier models (Gemini, GPT-4o) already demonstrating this capability; by 2028–2030 most production models are expected to be natively multimodal. Inference-time compute scaling (test-time training, chain-of-thought reasoning, MCTS-based search as in o3) is shifting the frontier from pre-training scale to the efficiency of reasoning at deployment time.

Efficient model training and deployment methods will continue to advance: parameter-efficient fine-tuning (LoRA variants), quantisation (GGUF, AWQ, SmoothQuant), and distillation enable deployment of powerful models on consumer hardware and edge devices. Federated Learning will become more prevalent for privacy-preserving deep learning over distributed health, financial, and personal data under UK GDPR and NHS data governance frameworks. Neuromorphic computing (Intel Loihi, IBM NorthPole) offers the prospect of spiking neural network implementations that dramatically reduce inference energy consumption, relevant to sustainability targets.

The theoretical foundations of deep learning remain incompletely understood: why overparameterised networks generalise, the nature of loss landscape geometry, and mechanistic interpretability (understanding what individual circuits in a network compute) are active research frontiers. Mechanistic interpretability research (Anthropic, DeepMind) aims to identify the computational circuits underlying model behaviour, with implications for AI safety and alignment. Constitutional AI, RLHF, and reinforcement learning from AI feedback (RLAIF) are maturing into standard post-training pipelines for aligning Large Language Model behaviour with human preferences and safety constraints.

Research and Literature

  1. McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. Bulletin of Mathematical Biophysics, 5(4), 115–133.
  2. Rosenblatt, F. (1958). The Perceptron: A probabilistic model for information storage and organisation in the brain. Psychological Review, 65(6), 386–408.
  3. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536.
  4. Hinton, G. E., Osindero, S., & Teh, Y. W. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18(7), 1527–1554.
  5. Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507.
  6. LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278–2324.
  7. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. NeurIPS 2012.
  8. Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780.
  9. Goodfellow, I., Pouget-Abadie, J., Mirza, M., et al. (2014). Generative adversarial nets. NeurIPS 2014.
  10. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. CVPR 2016.
  11. Ioffe, S., & Szegedy, C. (2015). Batch normalisation: Accelerating deep network training by reducing internal covariate shift. ICML 2015.
  12. Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929–1958.
  13. Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. NeurIPS 2017.
  14. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL-HLT 2019.
  15. Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. NeurIPS 2020 (GPT-3).
  16. LeCun, Y., Bengio, Y., & Hinton, G. E. (2015). Deep learning. Nature, 521(7553), 436–444.
  17. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.
  18. Kaplan, J., McCandlish, S., Henighan, T., et al. (2020). Scaling laws for neural language models. arXiv:2001.08361.
  19. Hoffmann, J., Borgeaud, S., Mensch, A., et al. (2022). Training compute-optimal large language models. NeurIPS 2022 (Chinchilla).
  20. Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. (2021). An image is worth 16×16 words: Transformers for image recognition at scale. ICLR 2021 (ViT).
  21. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. NeurIPS 2020.
  22. Jumper, J., Evans, R., Pritzel, A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.
  23. Hu, E. J., Shen, Y., Wallis, P., et al. (2022). LoRA: Low-rank adaptation of large language models. ICLR 2022.
  24. Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimisation. ICLR 2015.
  25. LeCun, Y., Bengio, Y., & Hinton, G. E. (2019). ACM Turing Award announcement. Communications of the ACM.
  26. UK Government (2025). AI Opportunities Action Plan: One Year On. DSIT. https://delivery.ai.gov.uk/
  27. Gartner (2026). Worldwide AI spending forecast. Gartner Research.
  28. Nobel Prize Committee (2024). Nobel Prize in Chemistry 2024. The Royal Swedish Academy of Sciences.

Key Terminology

  • Depth: The number of successive processing layers in a neural network; depth enables hierarchical feature learning
  • Representation Learning: Automatic discovery of useful feature representations from raw data, the defining capability of deep learning
  • Feedforward Pass / Forward Propagation: Left-to-right computation of the network output from input through successive layers
  • Backpropagation: Reverse-mode automatic differentiation through the network to compute gradients of the loss with respect to all parameters
  • Gradient Descent: Iterative parameter update rule that moves weights in the direction that reduces the loss
  • Stochastic Gradient Descent (SGD): Gradient descent applied to random minibatches rather than the full dataset
  • Loss Function: Scalar function measuring discrepancy between predictions and ground truth
  • Activation Function: Non-linear function (ReLU, GELU, SiLU, tanh, sigmoid) applied after each affine transformation
  • Hidden Layer: Any intermediate layer between input and output; the “depth” count
  • Vanishing Gradient Problem: Phenomenon in which gradients shrink exponentially as they propagate backward through many layers
  • Batch Normalisation: Technique normalising activations within a minibatch to stabilise and accelerate training
  • Dropout: Regularisation method randomly zeroing activations with probability p during training
  • Transfer Learning: Adapting a model pre-trained on a large dataset to a new task with less data
  • Foundation Model: A large model pre-trained at scale on broad data and adapted to many downstream tasks
  • Scaling Law: Empirical power-law relationship between model performance and parameter count, dataset size, and compute
  • RLHF (Reinforcement Learning from Human Feedback): Post-training alignment technique using human preferences to guide model behaviour

Provenance