Meta-Learning, colloquially described as ‘learning to learn’, is the study and design of machine learning systems that improve their own learning algorithms or initialisation through experience across multiple tasks, enabling rapid adaptation to new tasks with minimal data. Rather than learning a task directly, a meta-learning algorithm learns a prior or inductive bias that facilitates fast generalisation. Key paradigms include model-agnostic meta-learning (MAML), which optimises for a parameter initialisation that is close to good solutions for many tasks, and metric-based approaches such as prototypical networks that learn a task-agnostic embedding space. Meta-learning is central to few-shot learning, continual learning, and automated machine learning research.

Content

  • The meta-learning problem was first formalised in the 1990s by Thrun, Pratt, and Schmidhuber, but gained widespread practical attention from the 2017 publication of Model-Agnostic Meta-Learning (MAML) by Finn, Abbeel, and Levine. MAML’s key insight is elegantly simple: rather than training a model to solve a specific task, train it such that a small number of gradient steps on any new task leads to good performance. This requires differentiating through the inner optimisation loop, using second-order gradients that are computationally expensive but produce remarkably generalisable initialisations.
  • Three main families of meta-learning algorithms have emerged. Optimisation-based methods (MAML, Reptile) directly learn parameter initialisations or update rules. Model-based methods use recurrent or memory-augmented architectures (SNAIL, Neural Turing Machines) that encode task context in activations, enabling implicit adaptation without explicit gradient steps. Metric-based methods (Prototypical Networks, Matching Networks) learn embedding spaces where class prototypes can be computed from few examples using nearest-neighbour inference. Each family has distinct computational trade-offs and generalisation characteristics.
  • Meta-learning has had significant practical impact on few-shot image classification, few-shot natural language processing, and drug discovery. In computational drug discovery, meta-learning enables models trained on chemical property prediction tasks to adapt quickly to novel assay targets with sparse measurements—a critical capability when experimental data is expensive. In NLP, pre-trained language models can be interpreted through a meta-learning lens: their pre-training exposes them to diverse linguistic tasks, creating an initialisation from which fine-tuning to specific tasks is efficient, directly paralleling the MAML framework.
  • The relationship between meta-learning and in-context learning in large language models is an active research area. Large language models appear to perform few-shot adaptation purely through forward-pass inference on demonstration examples, without gradient updates—a form of implicit meta-learning encoded in the attention mechanism. Understanding whether this constitutes genuine meta-learning or sophisticated pattern matching has implications for how AI systems should be designed for rapid, safe deployment in novel domains with limited supervision.