The softmax function is a normalising transformation that maps a vector of real-valued scores (logits) into a probability distribution, where each output lies in the open interval (0, 1) and the outputs sum to one. It exponentiates each input and divides by the sum of all exponentials, amplifying larger scores while preserving rank order. Softmax is ubiquitous in machine learning as the final layer of multi-class classifiers and as the normalisation step inside attention mechanisms, and it pairs naturally with the cross-entropy loss whose gradient simplifies to the difference between predicted and target distributions.

Overview

  • Softmax generalises the logistic sigmoid to more than two classes. Given inputs z₁ … zₙ, it computes exp(zᵢ) divided by the sum over all j of exp(zⱼ). The exponential makes every output positive, and the normalisation forces the outputs to form a valid probability distribution.
  • Because the exponential amplifies differences between scores, softmax tends to assign most of the probability mass to the largest logit while still retaining a smooth, differentiable gradient. A temperature parameter can sharpen or soften this distribution, which is useful for sampling and knowledge distillation.
  • Numerical stability is a practical concern: naive computation of exp(z) can overflow. The standard implementation subtracts the maximum logit before exponentiating, which leaves the result unchanged mathematically but keeps the values in a safe range.

Mechanisms

  • Exponentiation and normalisation — Each logit is exponentiated and divided by the sum of all exponentials, producing a normalised probability vector. This is the defining operation.
  • Rank preservation — Softmax is monotonic in each input relative to the others, so the index of the maximum logit equals the index of the maximum probability; it never reorders the classes.
  • Temperature scaling — Dividing logits by a temperature T before softmax controls entropy: high T yields a near-uniform distribution, low T approaches a one-hot vector. This is central to sampling strategies and to calibration.
  • Gradient with cross-entropy — When combined with Cross-Entropy Loss, the gradient with respect to the logits reduces to the predicted probability minus the target, a simple and numerically benign expression that accelerates Backpropagation.
  • Attention normalisation — Inside an Attention Mechanism, softmax converts raw query-key similarity scores into attention weights that sum to one, determining how much each value contributes to the output.

Applications

  • Multi-class classification — Softmax is the canonical output layer for Classification networks, turning logits into class probabilities for tasks from image recognition to language modelling.
  • Transformer attention — Every Transformer layer applies softmax to scaled dot-product scores, the mechanism that lets the model focus on relevant tokens.
  • Language model decoding — Softmax over the vocabulary produces the next-token distribution; temperature and top-k/top-p sampling operate on this distribution.
  • Reinforcement learning — Policy networks often use softmax to convert action preferences into a stochastic policy.
  • Knowledge distillation — Temperature-scaled softmax produces soft targets that transfer richer information from a teacher to a student model.

Provenance