Universal approximation refers to the theoretical property, formalised in the universal approximation theorem, that a feedforward neural network with at least one hidden layer of sufficient width and a suitable non-linear activation function can approximate any continuous function on a compact domain to arbitrary precision. It provides the mathematical justification for using neural networks as general-purpose function approximators. The property depends on the activation function being non-polynomial and does not by itself guarantee that such a network can be learned efficiently from data.