Self-Supervised Learning (SSL) is a machine learning paradigm in which a model learns rich representations of data by solving pretext tasks whose supervisory signal is derived automatically from the input data itself, requiring no human-provided labels. The model learns to predict masked or hidden portions of an input, to match different views of the same data, or to distinguish positive from negative data pairs, developing features that transfer effectively to downstream supervised tasks with limited labelled data. Self-supervised learning has become the dominant pre-training strategy for large language models, visual foundation models, and multimodal systems, enabling training at scales that would be infeasible with manually annotated datasets.

Content

  • The conceptual roots of self-supervised learning lie in the observation that raw, unlabelled data already contains rich structural information that a model can learn to predict. Early applications included word2vec (2013), which trained word embeddings by predicting context words from a target word or vice versa. However, the term “self-supervised learning” was popularised by Yann LeCun around 2018 as he advocated for it as the path toward human-level AI, arguing that humans and animals learn primarily through self-supervised interaction with the world rather than labelled instruction.
  • Language modelling is the most successful application of self-supervised learning. The BERT model (2018) introduced masked language modelling: tokens in an input sentence are randomly replaced with a [MASK] token, and the model is trained to predict the original tokens from their context. GPT-series models use causal (autoregressive) language modelling, predicting each next token given all preceding tokens. Both approaches produce representations that capture syntactic and semantic structure of language without any labelled training data, and both can be fine-tuned to achieve state-of-the-art performance on a wide range of downstream NLP tasks.
  • In computer vision, self-supervised learning techniques have closed the gap with supervised pre-training on ImageNet. Contrastive methods such as SimCLR and MoCo learn by generating two augmented views of each image and training an encoder to produce similar representations for the two views whilst dissimilar representations for different images. Masked image modelling approaches such as MAE (Masked Autoencoder) borrow the language modelling masking strategy, masking patches of an image and training a ViT encoder-decoder to reconstruct the masked pixels.
  • The success of self-supervised pre-training for text has been extended to multimodal settings. CLIP (Contrastive Language-Image Pre-Training) uses natural language descriptions paired with images as free training signal, training image and text encoders to align their representations through contrastive loss. This produces an image encoder with zero-shot classification capability and has become a widely used backbone for vision-language models.
  • Self-supervised learning raises interesting theoretical questions about what representations are learned and why they transfer. Recent work has connected self-supervised objectives to information-theoretic frameworks, showing that good pretext tasks are those that maximise mutual information between learned representations and task-relevant features whilst discarding irrelevant variation. This perspective informs the design of new pretext tasks for emerging modalities including audio, video, point clouds, and biological sequence data.