A cost function is a scalar-valued mathematical function that maps a model’s parameters or a system’s state to a real number representing the magnitude of error, resource expenditure, or divergence from a desired outcome. In supervised machine learning, the cost function (also called a loss function) measures the aggregate discrepancy between predicted and ground-truth outputs across a training set, providing the objective that optimisation algorithms such as gradient descent minimise. In control theory and robotics, cost functions encode trajectory quality criteria including path length, energy consumption, and collision risk, enabling optimal control policies via formulations such as LQR and model predictive control. The design of the cost function is one of the most consequential decisions in any learning or optimisation system, as misspecified objectives lead to reward hacking, degenerate solutions, or physically unrealisable behaviours.

Semantic Classification

Content

Compositional Relationships (Components)

SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:DataFidelityTerm))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:RegularisationTerm))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:PenaltyTerm))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:WeightingScheme))
SubClassOf(ai:RegressionCostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:MeanSquaredError))
SubClassOf(ai:ClassificationCostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:CrossEntropyLoss))
SubClassOf(ai:GenerativeCostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:KullbackLeiblerDivergence))
SubClassOf(ai:ControlCostFunction
  ObjectSomeValuesFrom(ai:hasPart ai:QuadraticPerformanceIndex))

Dependency Relationships

SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:requires ai:ParametricModel))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:requires ai:TrainingData))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:requires ai:Differentiability))
SubClassOf(ai:NeuralCostFunction
  ObjectSomeValuesFrom(ai:dependsOn ai:LossLandscape))
SubClassOf(ai:NeuralCostFunction
  ObjectSomeValuesFrom(ai:dependsOn ai:ActivationFunction))
SubClassOf(ai:BayesianCostFunction
  ObjectSomeValuesFrom(ai:dependsOn ai:PriorDistribution))
SubClassOf(ai:MetricCostFunction
  ObjectSomeValuesFrom(ai:dependsOn ai:EmbeddingSpace))

Capability Relationships

SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:enables ai:GradientDescent))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:enables ai:ModelTraining))
SubClassOf(ai:ControlCostFunction
  ObjectSomeValuesFrom(ai:enables ai:OptimalControl))
SubClassOf(ai:ControlCostFunction
  ObjectSomeValuesFrom(ai:enables ai:MotionPlanning))
SubClassOf(ai:GenerativeCostFunction
  ObjectSomeValuesFrom(ai:enables ai:GenerativeModelling))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:enables ai:NeuralArchitectureSearch))
SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:enables ai:InverseReinforcementLearning))

Implementation Relationships

SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:implements ai:MaximumLikelihoodEstimation))
SubClassOf(ai:RegularisedCostFunction
  ObjectSomeValuesFrom(ai:implements ai:BayesianInference))
SubClassOf(ai:SurrogateLoss
  ObjectSomeValuesFrom(ai:implements ai:ContrastiveLearning))
SubClassOf(ai:RLCostFunction
  ObjectSomeValuesFrom(ai:implements ai:BellmanEquation))
SubClassOf(ai:MetricLoss
  ObjectSomeValuesFrom(ai:implements ai:MetricLearning))

Reduction Relationships

SubClassOf(ai:CostFunction
  ObjectSomeValuesFrom(ai:reducesTo ai:ScalarObjective))
SubClassOf(ai:MultiObjectiveCost
  ObjectSomeValuesFrom(ai:reducesTo ai:WeightedScalarisation))
SubClassOf(ai:NeuralCostFunction
  ObjectSomeValuesFrom(ai:reducesTo ai:GradientSignal))

About

The cost function is the conceptual fulcrum around which the entire enterprise of machine learning, optimal control, and algorithmic decision-making pivots. Its role is to convert the informal, domain-specific notion of “the system should perform well” into a precise mathematical quantity that algorithmic machinery can minimise through iterative parameter adjustment. Without a cost function, gradient-based learning is impossible — it is the cost function that renders the training problem well-posed and differentiable, creating the slope that gradient descent descends. The relationship between cost function design and model behaviour is not merely technical but fundamentally normative: what we choose to minimise defines what the system values, and therefore what the system does. This has become increasingly recognised as a central concern of AI safety and alignment research, since even small mismatches between the cost function and the true objective can lead to degenerate or harmful outcomes at deployment scale.

In Supervised Learning, the cost function aggregates a per-sample Loss Function over the training dataset, typically averaging squared errors for regression or cross-entropy for classification. This aggregate provides the signal that Backpropagation distributes through the Neural Network via the chain rule, allowing every weight in even the largest models to receive a gradient proportional to its contribution to total error. The scalar output is also the objective function for the optimiser: Stochastic Gradient Descent, Adam, RMSProp, and AdaGrad all navigate the high-dimensional Loss Landscape defined by the cost function, seeking parameter configurations that generalise to unseen data. The stochastic nature of mini-batch gradient estimation introduces noise into the gradient signal that is, counterintuitively, often beneficial: it acts as an implicit regulariser, helping the optimiser escape sharp minima and settle in flat basins that generalise better to unseen data. This insight — that noise in the cost gradient can improve generalisation — has driven research into the optimal batch size, gradient noise annealing schedules, and connections between the curvature of the cost landscape and the generalisation gap between training and test performance.

In Control Theory, the cost function serves a formally analogous but physically distinct role. Here it encodes the performance criteria of a dynamical system: for a Linear-Quadratic Regulator the cost is a quadratic integral penalising deviations from a target state and the magnitude of control effort; for Model Predictive Control it is minimised over a finite horizon at each timestep, with the process repeated as the system evolves. The elegant mathematical duality between RL (maximising reward) and control (minimising cost) means that modern deep RL systems often blur the boundary entirely, particularly in robotics settings where Deep Learning policy networks are trained against composite cost functions that include physics-based trajectory costs. The Bellman principle of optimality provides the theoretical foundation: under certain regularity conditions, the optimal cost-to-go (value function) satisfies a recursive equation V(x) = min_u [c(x,u) + γV(f(x,u))], connecting the control-theoretic and RL perspectives in a unified mathematical framework.

The design process for a cost function in practice is rarely straightforward. Practitioners begin with a primary objective — prediction accuracy, task completion, energy efficiency — and iteratively discover secondary constraints as the optimiser finds unexpected solutions that minimise the cost without satisfying unstated requirements. This phenomenon, often called Reward Hacking in the RL literature or Goodhart’s Law in economics, drives an ongoing cycle of cost function refinement. Modern machine learning workflows therefore treat cost function design as an empirical process: proposing a candidate, training to convergence, evaluating on held-out data using human-interpretable metrics, identifying failure modes, and revising the cost accordingly. The distance between the designed cost function and the true evaluation criterion — called the proxy gap — is a central concern of AI alignment research, particularly as systems become sufficiently capable to exploit this gap in non-obvious ways.

An important but often overlooked dimension of cost function design is its relationship to the Covariance Matrix of prediction errors. For regression problems under Gaussian noise, the optimal cost function is proportional to the inverse of the noise covariance: J(θ) = Σᵢ (yᵢ - f(xᵢ; θ))ᵀ Σ⁻¹_noise (yᵢ - f(xᵢ; θ)). This weighted MSE arises naturally from the maximum likelihood principle and accounts for correlations and differing reliabilities across output dimensions. In Gaussian Process regression, the cost function involves the log-determinant of the covariance matrix as a normalising constant, making covariance matrix operations central to probabilistic cost function optimisation. Understanding this connection illuminates why regularisation techniques (weight decay, dropout, spectral normalisation) can be derived as modifications to an implicit cost function that accounts for parameter uncertainty, drawing a direct line between statistical estimation theory and practical deep learning regularisation.

Components / Architecture

A well-designed cost function typically decomposes into four components, each serving a distinct role in shaping the optimisation problem:

  • Data-fidelity term — measures the discrepancy between model predictions and ground-truth labels. This is the dominant driver of parameter updates during training and defines the core supervised signal. Examples include Mean Squared Error for regression (penalises squared deviations; scale-sensitive; optimal under Gaussian noise), Cross-Entropy Loss for classification (derived from Maximum Likelihood Estimation under categorical distributions; decomposes into entropy of the true distribution plus KL divergence from model to truth), Huber Loss for robust regression (quadratic for small errors, linear for large ones; less sensitive to outliers than MSE), and Focal Loss for class-imbalanced object detection (down-weights confidently-classified easy examples to focus gradient on hard cases). The choice of data-fidelity term encodes implicit assumptions about the noise distribution of the targets: MSE assumes Gaussian noise, cross-entropy assumes categorical noise (the labels are draws from the true class distribution), and mean absolute error assumes Laplacian noise. This Bayesian interpretation — that a loss function is the negative log-likelihood of a noise model — provides a principled framework for choosing among loss functions by asking what noise model is appropriate for the data-generating process.

  • Regularisation term — penalises model complexity, typically through the L1 norm (LASSO, promoting sparsity by driving small weights to exactly zero; useful for feature selection), the L2 norm (Ridge/weight decay, shrinking all weights toward zero; corresponds to a Gaussian prior), or Elastic Net (combining both L1 and L2 with separate coefficients; useful when groups of correlated features should be selected together). In Bayesian Inference interpretation, the regularisation term is the negative log-prior: J_reg(θ) = -log p(θ). A Gaussian prior N(0, σ²I) corresponds to L2 regularisation with λ = 1/(2σ²); a Laplace prior corresponds to L1 regularisation. Beyond weight norms, modern regularisation includes dropout (randomly zeroing activations during training, equivalent to a geometrically distributed prior over neural network architectures), label smoothing (replacing hard one-hot targets with a soft distribution, equivalent to mixing with the uniform prior), spectral normalisation (constraining the Lipschitz constant of each layer), and Batch Normalisation (which implicitly regularises by normalising the activations and reducing the sensitivity of the loss to individual weight magnitudes). Mixup (Zhang et al., 2018) is a data-level regulariser that creates interpolated training examples and trains the model on interpolated labels, encouraging linear behaviour between examples and improving calibration.

  • Penalty Term — encodes hard or soft constraints that the solution must satisfy: collision-avoidance distances in Robotics and autonomous driving, actuator limits (torque, velocity, force) in Control Theory, fairness budgets in responsible ML (demographic parity, equalised odds), safety constraints in high-risk AI (maximum refusal rate, minimum detection precision), or output-length constraints in language generation. Soft penalties are added as additive terms proportional to the constraint violation: J_penalty(θ) = J(θ) + λ·max(0, g(θ))² for an inequality constraint g(θ) ≤ 0. Hard constraints use Lagrangian multipliers or augmented Lagrangian methods: L(θ, λ) = J(θ) + λ·g(θ), solved via alternating updates of θ and λ (the dual variable). The penalty framework is also used in constrained RL (CPO, PCPO, RCPO) where safety constraints are maintained during policy optimisation by augmenting the reward-based cost with penalty terms for constraint violations. In practice, the choice of penalty coefficient λ is a hyperparameter that controls the trade-off between primary objective and constraint satisfaction.

  • Surrogate components — differentiable proxies for non-differentiable target metrics. Since accuracy, BLEU score, ROUGE, AUC, and most practical evaluation metrics are piecewise constant and unsuitable for gradient optimisation, surrogate losses serve as training objectives while the true metric is evaluated at inference. Hinge Loss is a differentiable (except at the margin boundary) surrogate for the 0-1 classification error that achieves the same Bayes-optimal classifier while being amenable to convex optimisation. Kullback-Leibler Divergence is a differentiable surrogate for statistical distance between distributions; it appears as a regulariser in Variational Inference (ELBO = reconstruction - KL) and in RLHF (penalty for diverging from the reference policy). Wasserstein Distance provides geometrically meaningful gradients even when distributions have non-overlapping supports, enabling training of generative models without mode collapse. Contrastive Loss and InfoNCE are differentiable surrogates for mutual information maximisation, enabling self-supervised learning without explicit probability density estimation. The gap between surrogate training loss and true evaluation metric — called the proxy gap or surrogate regret — is a fundamental challenge: a model can achieve near-zero surrogate loss while performing poorly on the true metric, motivating work on tighter surrogates and direct optimisation of evaluation metrics through reinforcement learning or task-specific loss design.

    Cost Function Families

    The diversity of cost functions used in practice reflects the diversity of learning tasks and the distinct mathematical properties required by different problem structures. The following taxonomy organises the principal families by the type of prediction task and the statistical assumptions embedded in the cost.

    Regression Losses

    Regression losses measure deviation between a continuous predicted value and a continuous target. The choice among them encodes assumptions about the noise distribution of the targets and the desired robustness properties of the estimator.

  • Mean Squared Error (MSE / L2 loss) — J(θ) = (1/n) Σᵢ (yᵢ - ŷᵢ)². Quadratic, globally differentiable, sensitive to outliers; standard in linear regression and Linear-Quadratic Regulator design. Minimising MSE under Gaussian noise is equivalent to maximum likelihood estimation. The squared penalty means that large errors (outliers or labelling mistakes) receive disproportionately large gradients, pulling the model towards them; this is beneficial when all deviations should be penalised severely, but detrimental when training data contains noise or outliers.

  • Mean Absolute Error (MAE / L1 loss) — J(θ) = (1/n) Σᵢ |yᵢ - ŷᵢ|. Robust to outliers (gradient magnitude does not grow with error magnitude), but non-differentiable at zero, requiring subgradient methods or smooth approximations. MAE is the MLE estimator under Laplace noise and produces sparser gradients — zero gradient for correctly-ranked predictions — which can slow convergence relative to MSE for well-fitting regions.

  • Huber Loss — J_δ(θ) = Σᵢ hᵢ where hᵢ = (1/2)(yᵢ - ŷᵢ)² if |yᵢ - ŷᵢ| ≤ δ, else δ|yᵢ - ŷᵢ| - δ²/2. Smooth quadratic near zero, linear in the tails with a transition at threshold δ; balances robustness and differentiability. Used in DQN (temporal difference error), Faster-RCNN (bounding box regression), and many regression tasks. The threshold δ is a hyperparameter that controls the transition between quadratic and linear regimes; δ → ∞ recovers MSE, δ → 0 recovers MAE.

  • Log-Cosh loss — J(θ) = Σᵢ log(cosh(yᵢ - ŷᵢ)). Twice-differentiable approximation to MAE; superior numerical stability for second-order optimisers (Newton’s method, L-BFGS). The gradient tanh(yᵢ - ŷᵢ) is bounded, providing implicit robustness to outliers.

  • Quantile loss (Pinball loss) — J_q(θ) = Σᵢ max(q(yᵢ - ŷᵢ), (q-1)(yᵢ - ŷᵢ)). Enables calibrated interval prediction by penalising under- and over-prediction asymmetrically, where q ∈ (0,1) is the target quantile. Minimising the q-quantile loss produces an estimator of the q-th quantile of the conditional distribution P(Y|X); training models for multiple quantiles provides calibrated prediction intervals. Critical in supply chain and finance applications where the asymmetric cost of over- versus under-prediction must be modelled explicitly.

    Classification Losses

    Classification losses measure the quality of categorical predictions. The dominant paradigm is the probabilistic classifier that outputs a distribution over classes; the cost function measures the distance between this predicted distribution and the true label.

  • Cross-Entropy Loss (categorical / binary) — J(θ) = -Σᵢ Σₖ yᵢₖ log(p̂ᵢₖ). Gold standard for probabilistic classifiers; derived from negative log-likelihood under categorical distribution. Minimised when the predicted distribution matches the true class-conditional distribution. Binary cross-entropy is the special case for two-class problems. The gradient of cross-entropy with respect to Softmax logits equals (predicted probability - true probability), a remarkably clean signal that drives the predicted probability toward the true label. Cross-entropy penalises confident wrong predictions extremely heavily (log(0) → ∞), which motivates label smoothing.

  • Hinge Loss — J(θ) = Σᵢ max(0, 1 - yᵢ · f(xᵢ; θ)). Margin-based; used in support vector machines; not differentiable at the margin boundary (yᵢ · f(xᵢ) = 1). Encourages a margin of at least 1 between the score of the correct class and the maximum competing class score. Extended to multiclass via Crammer-Singer formulation. Computationally sparse gradients — zero for correctly-classified examples beyond the margin — enable efficient training with subgradient methods.

  • Focal Loss — J_FL(θ) = -Σᵢ (1 - p̂ᵢ)^γ log(p̂ᵢ). Variant of binary cross-entropy that down-weights easy examples via a modulating factor (1-p̂)^γ where γ > 0; introduced by Lin et al. (2017) for single-stage object detection in RetinaNet. The modulating factor reduces the loss contribution from examples where the model is already confident, focusing gradient updates on the hard misclassified examples. With γ = 0 recovers standard cross-entropy; γ = 2 is the default setting. Widely adopted in imbalanced classification, medical imaging, and satellite image analysis.

  • Label-smoothed cross-entropy — replaces hard one-hot targets yᵢ with soft distribution y_smooth = (1-ε)·y_one-hot + ε/K (where K is the number of classes). Reduces overconfidence and improves calibration in Transformer Architecture models including BERT, GPT, and ViT. Equivalent to adding the KL divergence between the model distribution and the uniform distribution as a regulariser.

  • Pairwise ranking losses — BPR (Bayesian Personalised Ranking) and RankNet are cost functions that train models to assign higher scores to preferred items than to non-preferred items; they power recommendation systems by directly optimising the ranking objective rather than a pointwise prediction error.

    Generative and Probabilistic Losses

    Generative model training requires cost functions that measure the discrepancy between learned and true data distributions, operating in the distribution space rather than the prediction-target space.

  • Kullback-Leibler Divergence (KL divergence) — KL(p||q) = ∫ p(x) log(p(x)/q(x)) dx. Measures the information cost of approximating distribution p with distribution q; asymmetric (KL(p||q) ≠ KL(q||p)); zero when p = q, infinite when q assigns probability 0 to a region where p > 0 (mode-seeking behaviour). Core to Variational Autoencoder training (ELBO = -KL[q(z|x)||p(z)] + E[log p(x|z)]) and RLHF (KL penalty against the reference policy prevents the fine-tuned model from diverging to a degenerate solution).

  • Evidence Lower Bound (ELBO) — ELBO(θ, φ) = E_{q_φ(z|x)}[log p_θ(x|z)] - KL[q_φ(z|x)||p(z)]. Negated and minimised in Variational Inference; the first term (reconstruction) drives the decoder to reconstruct inputs accurately, while the KL term (regularisation) keeps the approximate posterior close to the prior. The ELBO is a lower bound on the marginal log-likelihood log p(x) by Jensen’s inequality. Variants include the β-VAE (Higgins et al., 2017) which scales the KL term by β > 1 to encourage disentangled representations.

  • Wasserstein Distance (Earth Mover’s Distance) — W₁(p, q) = inf_{γ ∈ Π(p,q)} E_{(x,y)~γ}[||x-y||]. Geometrically meaningful probability metric providing smoother gradients than JS divergence, even when distributions have non-overlapping supports; central to Wasserstein GAN (WGAN) and optimal transport-based generative models. The dual form W₁(p,q) = sup_{||f||_L ≤ 1} [E_p[f(x)] - E_q[f(x)]] (Kantorovich-Rubinstein duality) enables practical computation using a discriminator network with Lipschitz constraint enforced via gradient penalty (WGAN-GP, Gulrajani et al., 2017).

  • Diffusion loss (denoising score matching) — J(θ) = E_{t,x₀,ε}[||ε - ε_θ(√ᾱₜ·x₀ + √(1-ᾱₜ)·ε, t)||²]. The simplified training objective for denoising diffusion probabilistic models; minimises the expected squared difference between the added noise ε and the model’s prediction ε_θ of that noise. Equivalent to weighted score matching; underpins Stable Diffusion, DALL-E 3, Sora, and all modern diffusion-based generative models. Each timestep t has a different noise level, and the expectation over t with a specific weighting corresponds to an ELBO under the forward diffusion process.

  • Contrastive Loss — pushes representations of similar pairs together and dissimilar pairs apart in an embedding space. NT-Xent (normalised temperature cross-entropy, SimCLR): for a batch of 2N examples with N augmented pairs, the loss for example i is -log(exp(sim(zᵢ, zⱼ)/τ) / Σₖ≠ᵢ exp(sim(zᵢ, zₖ)/τ)). InfoNCE generalises this to multiple negative samples. Foundational to CLIP (image-text contrastive pre-training), DINOv2 (self-distillation with no labels), and SigLIP (sigmoid loss for vision-language).

    Control and Planning Losses

    Control cost functions formalise the notion of trajectory quality, enabling the computation of optimal policies through mathematical programming.

  • Quadratic performance index (LQR) — J = ∫₀^∞ (xᵀQx + uᵀRu)dt where Q ≥ 0 penalises state deviation and R > 0 penalises control effort. For a linear system ẋ = Ax + Bu, minimising J yields a closed-form optimal gain matrix K = R⁻¹BᵀP where P solves the algebraic Riccati equation AᵀP + PA - PBR⁻¹BᵀP + Q = 0. The relative weighting Q/R trades off state precision against control economy; large Q/R produces aggressive control, small Q/R produces sluggish control.

  • Receding-horizon cost (MPC) — J_N(x(t), u(·)) = Σₖ₌₀^{N-1} [xₖᵀQxₖ + uₖᵀRuₖ] + xₙᵀPxₙ. Finite-horizon quadratic or nonlinear cost minimised at each timestep in Model Predictive Control; the terminal cost P stabilises the closed-loop system. Solved as a quadratic programme (for linear systems) or nonlinear programme (for nonlinear MPC) at each timestep; powerful for constrained control, autonomous driving, and chemical process optimisation where safety constraints must be enforced explicitly.

  • Trajectory smoothness costs — Σₜ ||xᵢ₊₁ - 2xᵢ + xᵢ₋₁||² (acceleration penalty) or higher-order jerk/snap penalties along a planned path. Used in CHOMP (Covariant Hamiltonian Optimisation for Motion Planning) and TrajOpt planners for Motion Planning in Robotics. Smooth trajectories are mechanically less stressful, consume less energy, and are more predictable to surrounding agents; balancing smoothness against obstacle avoidance defines the characteristic cost structure of motion planning.

  • Energy-optimal cost — J = ∫ uᵀ(t)u(t) dt minimises total actuator effort; critical for battery-constrained mobile robots, spacecraft manoeuvring, and humanoid locomotion where energy budget is a primary constraint.

    Metric and Ranking Losses

  • Triplet loss — L = Σᵢ max(0, d(aᵢ, pᵢ) - d(aᵢ, nᵢ) + margin). Learns an embedding where the distance to anchor-positive is smaller by a margin than anchor-negative; used in face recognition (FaceNet), image retrieval, and speaker verification. Hard negative mining (selecting the most confusable negatives from the batch) is critical for efficient triplet loss training.

  • NT-Xent (Normalised Temperature-scaled Cross-Entropy) — contrastive loss used in SimCLR; maximises agreement between augmented views of the same image within a batch of negatives. The temperature parameter τ controls the hardness of the contrastive distribution.

  • InfoNCE loss — mutual information lower bound I(X;Y) ≥ E[log softmax(f(x,y)/τ)] derived by Van den Oord et al. (2018); used in contrastive predictive coding (CPC) for speech and video; key to representation learning that captures high-level semantic structure.

    Formal Analysis

    The cost function J(θ) is formally a functional over parameter space Θ. Its gradient ∇_θ J(θ) is the vector of partial derivatives with respect to every parameter — computed via Automatic Differentiation in frameworks such as PyTorch, TensorFlow, and JAX. The optimisation problem:

    θ* = argmin_{θ ∈ Θ} J(θ)

    is unconstrained for standard training and constrained when fairness or safety requirements impose bounds on θ or on model outputs. Key analytical properties:

  • Convexity — MSE for linear models and cross-entropy for logistic regression are convex in θ, guaranteeing a unique global minimum reachable by any gradient method. Deep network losses are generically non-convex; however, overparameterisation (width >> depth) creates a benign geometric structure where local minima cluster near the global minimum in terms of generalisation performance. Neural tangent kernel (NTK) theory (Jacot et al., 2018) provides a linearisation of deep network cost functions in the infinite-width limit, showing convergence to a global minimum under gradient descent in that regime.

  • Smoothness (L-Lipschitz gradient) — determines the maximum stable learning rate for gradient descent: α < 2/L. Batch Normalisation and gradient clipping improve effective smoothness during training. Layer-wise adaptive rate scaling (LARS) and its successors adaptively set per-layer learning rates based on the ratio of weight norms to gradient norms, implicitly accounting for layer-dependent curvature of the cost landscape.

  • Variance of stochastic gradients — determines the noise in Stochastic Gradient Descent updates; high variance slows convergence but provides implicit regularisation that can help escape sharp minima. The variance-bias tradeoff in stochastic gradient estimation has been extensively studied; variance reduction methods (SVRG, SAGA, SARAH) reduce gradient noise while preserving the beneficial annealing effect of decreasing learning rates.

  • Saddle points — in high dimensions, most critical points are saddle points rather than local minima; first-order methods with noise escape these efficiently, making SGD practically superior to full-batch gradient descent for deep Neural Network training. Perturbed gradient descent (Jin et al., 2017) provides theoretical guarantees for escaping strict saddle points in polynomial time, formalising the empirical observation that deep networks train successfully despite non-convexity.

  • Loss landscape geometry — the Loss Landscape of deep networks has been studied extensively through random matrix theory and differential geometry. Sharp minima (high Hessian curvature) correlate with poor generalisation; flat minima (low curvature) correlate with good generalisation, motivating sharpness-aware minimisation (SAM) which explicitly penalises the maximum cost in a neighbourhood of the current parameters. This can be interpreted as a cost function augmentation: J_SAM(θ) = max_{||ε|| ≤ ρ} J(θ + ε).

  • Duality theory — constrained cost minimisation problems (e.g., minimise J(θ) subject to fairness constraint g(θ) ≤ c) are addressed via Lagrangian duality: L(θ, λ) = J(θ) + λ·g(θ). The dual problem max_λ min_θ L(θ, λ) is the foundation of constrained deep learning, safety-constrained RL (CPO, PCPO), and fairness-aware training. Strong duality holds when the constraint set and cost are convex; in non-convex settings, the duality gap requires special handling.

  • Information-geometric perspective — the natural gradient (Amari, 1998) preconditions gradient updates by the inverse Fisher information matrix F(θ)⁻¹, which equals the inverse Covariance Matrix of the score function ∇_θ log p(y|x;θ). This accounts for the Riemannian geometry of the parameter space under the cost function’s induced probability model, giving steeper descent in directions of high parameter sensitivity. K-FAC (Kronecker-factored approximate curvature) approximates F(θ)⁻¹ efficiently for deep networks, achieving faster convergence than first-order methods.

    Use Cases / Major Families

    The cost function is not a component of a system but rather the specification of what the system is trying to achieve. Its design varies substantially across application domains, reflecting the different learning objectives, data characteristics, and deployment constraints of each field.

  • Deep Learning model training — Backpropagation computes the gradient of the cost with respect to all parameters via the chain rule; Adam and SGD use these gradients to update weights across billions of parameters in Transformer Architecture language models and vision systems. The cost function for pre-training large language models is typically autoregressive next-token cross-entropy: J(θ) = -E_{x~D}[Σₜ log p_θ(xₜ|x_{<t})], which trains the model to predict each token given all preceding tokens. This apparently simple cost function produces surprisingly general representations; scaling laws (Kaplan et al., 2020; Hoffmann et al., 2022) characterise how cross-entropy loss decreases predictably with model size, dataset size, and compute budget following power laws.

  • Reinforcement Learning agent training — Policy gradient methods (PPO, GRPO, TRPO) and value-based methods (DQN, SAC) optimise a cost defined on accumulated reward signals; the RL cost function is typically J(π) = -E_{τ~π}[Σₜ γᵗ rₜ] (negated discounted return). DeepSeek-R1’s final training stage used both RLVR (for reasoning tasks — a verifiable binary reward) and RLHF (for helpfulness/harmlessness) with separate reward models scored against different cost functions: helpfulness was scored based only on the final answer, while harmlessness was evaluated on the entire reasoning chain. This multi-component cost reflects the recognition that different aspects of model quality require different measurement approaches.

  • Inverse Reinforcement Learning — infers the cost function from expert demonstrations; the problem is ill-posed (many cost functions can rationalise any behaviour), resolved by maximum entropy IRL (Ziebart et al., 2008) which selects the cost function that makes the demonstrated behaviour the most probable among all behaviours while remaining maximally uncertain about everything else. Enables robots to acquire task objectives from human or animal demonstrations without hand-coded reward engineering; used in autonomous driving (learning driving cost functions from expert logs) and robot manipulation (learning from videos).

  • Model Predictive Control — a receding-horizon optimiser solves finite-horizon cost minimisation at each timestep J_MPC = Σₖ₌₀^{N-1} l(xₖ, uₖ) + V_f(x_N), replanning continuously as new state measurements arrive; standard in autonomous vehicles, chemical process optimisation, building energy management, and humanoid locomotion (where the cost balances locomotion speed, stability, energy efficiency, and collision avoidance over a horizon of 0.5-2 seconds).

  • Neural Architecture Search — validation loss serves as the cost function driving architecture optimisation via evolutionary algorithms or differentiable NAS; the upper-level cost penalises both validation loss and resource constraints (FLOPs, latency, memory footprint) via a multi-objective scalarisation; one-shot NAS frameworks (DARTS, SNAS) use architecture parameters α differentiably coupled to a shared-weight supernet, computing architecture gradients by chain rule through the training cost.

  • Large Language Models fine-tuning — RLHF uses a KL-regularised reward-model score as the cost J_RLHF(π) = E_{xD,yπ(·|x)}[r_θ(y|x) - β·KL[π(·|x)||π_ref(·|x)]], balancing human preference alignment against divergence from the base pre-trained distribution. The Bradley-Terry pairwise preference model defines the reward-model training objective: J_RM = -E[(r(y_w) - r(y_l))] where y_w is the preferred response and y_l the rejected response. Modern approaches (DPO, RLVR, GRPO) directly optimise variants of this cost without a separate reward model.

  • Medical image segmentation — Dice loss J_Dice = 1 - 2|P∩G|/(|P|+|G|) and combined cross-entropy/Dice composite costs are minimised to train pixel-level classifiers in CT, MRI, and histology imaging; Dice loss handles class imbalance inherently through its F1-based formulation (the denominator normalises by total predicted and true positive counts, not by all pixels), making it far superior to cross-entropy alone when foreground structures occupy < 1% of image volume.

  • Financial portfolio optimisation — mean-variance J_MV(w) = -wᵀμ + λ·wᵀΣw and Conditional Value-at-Risk objectives define cost functions over portfolio weights w; solved via quadratic programming (for mean-variance) or linear programming (for CVaR); Ledoit-Wolf shrinkage of the Covariance Matrix is standard preprocessing to ensure the quadratic term wᵀΣw uses a reliable estimate.

  • Autonomous vehicle planning — composite costs J_AV = w₁·J_collision + w₂·J_lane + w₃·J_comfort + w₄·J_progress combining collision risk, lane-keeping error, ride comfort (jerk minimisation), and forward progress are minimised in real-time MPC loops at 20-50 Hz. The weighting coefficients w encode the relative importance of safety versus efficiency; safety-critical violations (collision cost) use hard constraints rather than soft penalties.

  • Climate and weather modelling — NeuralGCM (Google DeepMind, 2024) and GraphCast (DeepMind, 2023) are trained against combinations of spectral coefficient loss (weighted MSE in frequency space) and physical conservation law penalties (conservation of total energy, mass, and angular momentum). These physics-informed cost functions produce physically consistent forecasts while benefiting from data-driven learning, achieving state-of-the-art accuracy on 10-day global weather prediction at fraction of the computational cost of traditional numerical weather prediction models.

  • Protein structure prediction — AlphaFold2 (DeepMind, 2021) uses the Frame Aligned Point Error (FAPE) loss — an average Euclidean distance between predicted and true atom positions computed in local coordinate frames — combined with a violation loss penalising physically impossible bond lengths and angles. The composite cost function reflects the specific geometry of 3D protein structures and the need for rotational/translational equivariance.

  • Drug discovery and molecular optimisation — graph neural networks for molecular property prediction use MSE or cross-entropy on binding affinity or toxicity labels; generative models for de novo drug design use ELBO on molecular graphs. Reinforcement Learning-based molecular optimisation uses reward functions (negative cost) that balance predicted activity against synthesisability, drug-likeness (QED score), and diversity to avoid trivial solutions.

  • Recommendation systems — implicit feedback recommendation uses Bayesian Personalised Ranking (BPR) loss, which optimises pairwise ranking: J_BPR = -Σ_{(u,i,j) ∈ DS} log σ(ŷ_ui - ŷ_uj) where u is a user, i a positive item, j a negative item. This directly optimises the ranking objective (AUC equivalent) rather than a surrogate point-prediction error.

    Academic Context

    The concept of a cost function traces to the calculus of variations (Euler, 1744; Lagrange, 1788), which formulated the problem of finding functions that extremise integral functionals — the direct ancestor of modern variational losses. Lagrange introduced the method of multipliers for constrained optimisation; Hamilton reformulated mechanics as extremisation of the action functional; these are all special cases of the general cost function minimisation framework that underlies modern machine learning. Legendre and Fenchel developed convex conjugate theory, which provides the duality structure used in max-margin learning and constrained optimisation. The term “cost function” in control theory was formalised by Bellman (1957) in dynamic programming and by Kalman (1960) in the LQR formulation. Wiener’s (1949) minimum mean-squared-error estimation is an early statistical cost function, derived from minimising the expected quadratic discrepancy between estimator and true signal — the same MSE loss that remains ubiquitous in regression tasks.

    In machine learning, Rosenblatt’s (1958) perceptron used a hinge-like cost; the breakthrough formalisation of differentiable cost + backpropagation came with Rumelhart, Hinton and Williams (1986), who showed that a smooth differentiable cost enables efficient computation of parameter gradients through the chain rule. The cross-entropy loss was derived from information-theoretic principles (Shannon, 1948) and its connection to maximum likelihood estimation under categorical distributions was made explicit by Duda and Hart (1973). The kernel SVM hinge loss (Boser, Guyon, Vapnik, 1992) introduced the margin-based perspective, providing a geometric interpretation of the cost function as measuring distance from the decision boundary. Modern deep learning practice was codified by LeCun et al. (1998) for convolutional networks and by Goodfellow, Bengio and Courville (2016) in the canonical textbook, which contains a comprehensive treatment of cost function design from both frequentist (MLE) and Bayesian perspectives.

    Key recent academic milestones include: the Wasserstein GAN loss (Arjovsky et al., 2017), showing superior training stability for generative models through a cost function with better gradient properties than the Jensen-Shannon divergence; the Focal Loss (Lin et al., 2017) for dense object detection, which modulated cross-entropy to down-weight easy examples and focus learning on hard negatives; the SimCLR contrastive loss (Chen et al., 2020) which scaled contrastive learning to vision and demonstrated that a well-designed cost function enables self-supervised pre-training competitive with supervised learning; the diffusion denoising score matching objective (Ho et al., 2020) which frames generative modelling as minimising a cost over time-indexed noise predictions rather than directly modelling data density; and the RLHF reward-model objective using Bradley-Terry preferences (Ouyang et al., 2022; Bai et al., 2022) which operationalised human preference signals as a differentiable cost function for fine-tuning large language models. A 2025 Springer Artificial Intelligence Review provided a comprehensive taxonomy of loss functions in deep learning covering regression, classification, metric, generative, energy-based, and ranking categories, noting that loss function design remains one of the most active research areas in the field. A 2025 MDPI Mathematics survey surveyed 90+ loss functions, systematically comparing their properties, gradients, and suitable application domains.

    The relationship between cost function choice and the implicit bias of gradient descent has emerged as a major theoretical question in the last decade. Different cost functions, even when minimised to the same value, create different inductive biases through the trajectory of the optimiser: gradient descent on cross-entropy for linear classifiers converges to the maximum-margin solution (Soudry et al., 2018); for matrix factorisation problems it converges to low-rank solutions (Gunasekar et al., 2017). These implicit regularisation effects — invisible in the cost function itself but manifest in the optimisation dynamics — have profound implications for generalisation and are an active area of theoretical investigation.

    Current Landscape (2026)

    Cost function research is experiencing significant diversification in 2025-2026. Several major trends are reshaping the field:

    Meta-learning of loss functions — systems that automatically discover loss functions optimised for a given task and data distribution, using bilevel optimisation or evolutionary search. Meta-Loss (2024, arXiv:2406.09713) demonstrated that learned loss functions can outperform hand-designed counterparts on few-shot learning benchmarks.

    Physics-informed cost functions — frontier AI systems for scientific computing embed physical conservation laws (mass, energy, momentum) directly into the cost function as equality constraints or soft penalty terms. GraphCast (DeepMind) and NeuralGCM use these for weather prediction at unprecedented accuracy.

    Composite RLHF objectives — the 2025 landscape of frontier LLM training (GPT-o3, Claude 3.7, Gemini 2.5, DeepSeek-R1) uses increasingly sophisticated multi-objective cost functions that simultaneously optimise for helpfulness, harmlessness, honesty, and reasoning quality. DeepSeek-R1 separates format reward, language-consistency reward, and reasoning-verification reward into distinct sub-costs.

    Certified-safe cost functions — following the EU AI Act (effective August 2024) and NIST AI RMF 1.0, high-risk AI applications require that cost functions demonstrably encode safety and fairness constraints. Certifiable Safe RLHF (2025, arXiv:2510.03520) proposed fixed-penalty constraint optimisation to ensure language model outputs satisfy safety constraints with provable guarantees.

    Contrastive and metric losses at scale — CLIP, ALIGN, and their successors use contrastive cost functions at scales of billions of image-text pairs; the InfoNCE and NT-Xent objectives have become standard for vision-language pre-training.

    UK Context

    UK academic institutions have made significant contributions across multiple aspects of cost function research. The Alan Turing Institute (London) has active programmes in robust loss function design for safety-critical AI, including medical imaging and financial risk applications. Imperial College London (Computing and Electrical Engineering departments) has contributed to optimal control cost function design and robust MPC; researchers there have collaborated on automotive and aerospace MPC applications with industry partners including Rolls-Royce and BAE Systems.

    University of Edinburgh (Informatics, notably the Bayesian and Neural Computing groups) has long-standing research on variational inference objectives including ELBO variants, and contributed foundational work on normalising flows and latent variable model training objectives. UCL (Gatsby Computational Neuroscience Unit, Machine Learning group) has been central to meta-learning and Bayesian deep learning, both of which depend critically on well-specified cost functions; Arthur Gretton’s group at UCL Gatsby contributed foundational work on kernel maximum mean discrepancy as a cost function for generative models.

    University of Cambridge (Machine Learning Group, Engineering Department) has produced research on probabilistic cost functions for Gaussian process models and deep kernel learning. Richard Turner’s group has contributed to variational objectives in deep generative models.

    Northern England institutions also contribute: University of Manchester (AI and Data Science, working with The Alan Turing Institute’s Manchester node), University of Leeds (industrial AI applications including process control MPC), and University of Sheffield (Gaussian processes and probabilistic machine learning under Neil Lawrence, now at Cambridge).

    In industry, DeepMind (London) has been a prolific producer of novel cost functions, including the AlphaFold2 FAPE loss for protein structure, the AlphaCode competitive programming objectives, and physics-informed weather losses in GraphCast. Wayve (London) develops composite cost functions for end-to-end autonomous driving from video. Stability AI (London) pioneered diffusion model training losses at scale.

    Future Directions (2026-2030)

  • Learned and adaptive cost functions — moving beyond manually designed objectives toward systems that discover their own loss functions via meta-learning or neural architecture search. This raises alignment concerns: an AI that modifies its own training cost may diverge from human intent.

  • Causal cost functions — embedding causal structure into the loss (do-calculus constraints, counterfactual invariance) to produce models that learn genuinely causal relationships rather than spurious correlations. CausalRM (2026, arXiv:2603.18736) applies causal reward modelling to RLHF.

  • Multi-objective Pareto-optimal costs — rather than scalarising multiple objectives into a single cost, Pareto-front methods learn a family of solutions; this is expected to dominate in healthcare AI where safety and accuracy trade-offs must be explicitly represented.

  • Neurosymbolic cost functions — hybrid losses combining differentiable neural components with symbolic logical constraints; enabling formal verification of model properties during training rather than post hoc.

  • Energy-efficient cost functions — as compute cost becomes a first-class concern, cost functions for model compression, quantisation, and sparsification will increasingly include FLOPs and memory footprint as regularisation terms.

  • Continual learning objectives — cost functions that minimise catastrophic forgetting while enabling plasticity; regularisation-based (EWC, SI) and replay-based approaches compete in this space.

  • Regulatory-driven cost function auditing — the EU AI Act’s requirements for high-risk AI documentation will likely drive standardisation of cost function specification formats, creating a new area of AI governance and auditing.

    Research & Literature

    1. Rumelhart, D.E., Hinton, G.E. & Williams, R.J. (1986). “Learning representations by back-propagating errors.” Nature, 323, 533-536.
    2. LeCun, Y., Bottou, L., Bengio, Y. & Haffner, P. (1998). “Gradient-based learning applied to document recognition.” Proceedings of the IEEE, 86(11), 2278-2324.
    3. Bishop, C.M. (2006). Pattern Recognition and Machine Learning. Springer, New York.
    4. Goodfellow, I., Bengio, Y. & Courville, A. (2016). Deep Learning. MIT Press, Cambridge MA.
    5. Sutton, R.S. & Barto, A.G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
    6. Bertsekas, D.P. (2012). Dynamic Programming and Optimal Control, Vol. 1 (4th ed.). Athena Scientific.
    7. Bellman, R. (1957). Dynamic Programming. Princeton University Press.
    8. Kalman, R.E. (1960). “A new approach to linear filtering and prediction problems.” Journal of Basic Engineering, 82(1), 35-45.
    9. Arjovsky, M., Chintala, S. & Bottou, L. (2017). “Wasserstein Generative Adversarial Networks.” ICML 2017. PMLR 70, 214-223.
    10. Lin, T.-Y., Goyal, P., Girshick, R., He, K. & Dollar, P. (2017). “Focal Loss for Dense Object Detection.” ICCV 2017, 2980-2988.
    11. Chen, T., Kornblith, S., Norouzi, M. & Hinton, G. (2020). “A Simple Framework for Contrastive Self-Supervised Learning.” ICML 2020. PMLR 119, 1597-1607.
    12. Ho, J., Jain, A. & Abbeel, P. (2020). “Denoising Diffusion Probabilistic Models.” NeurIPS 2020, 33, 6840-6851.
    13. Ouyang, L. et al. (2022). “Training language models to follow instructions with human feedback.” NeurIPS 2022, 35, 27730-27744.
    14. Bai, Y. et al. (2022). “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.” arXiv:2204.05862.
    15. Kingma, D.P. & Welling, M. (2013). “Auto-Encoding Variational Bayes.” ICLR 2014. arXiv:1312.6114.
    16. Radford, A. et al. (2021). “Learning Transferable Visual Models from Natural Language Supervision.” ICML 2021. PMLR 139, 8748-8763. [CLIP contrastive loss]
    17. Lafferty, J., McCallum, A. & Pereira, F. (2001). “Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data.” ICML 2001, 282-289.
    18. Tibshirani, R. (1996). “Regression shrinkage and selection via the lasso.” JRSS-B, 58(1), 267-288.
    19. Huber, P.J. (1964). “Robust estimation of a location parameter.” Annals of Mathematical Statistics, 35(1), 73-101.
    20. Lyu, S. & Simoncelli, E.P. (2009). “Nonlinear image representation using divisive normalization.” CVPR 2009.
    21. Wright, S.J. (1997). Primal-Dual Interior-Point Methods. SIAM.
    22. Rawlings, J.B., Mayne, D.Q. & Diehl, M.M. (2017). Model Predictive Control: Theory, Computation, and Design (2nd ed.). Nob Hill Publishing.
    23. Lambert, N. (2024). RLHF Book: Reward Modeling. rlhfbook.com/c/05-reward-models.
    24. Wang, Z. et al. (2025). “A Survey of Loss Functions in Deep Learning.” Mathematics, 13(15), 2417. MDPI. doi:10.3390/math13152417.
    25. Le Gall, F. & Sedrakyan, G. (2025). “A Survey and Taxonomy of Loss Functions in Machine Learning.” AI, 7(4), 128. MDPI. doi:10.3390/ai7040128.
    26. Gretton, A., Borgwardt, K.M., Rasch, M.J., Scholkopf, B. & Smola, A. (2012). “A Kernel Two-Sample Test.” JMLR, 13, 723-773. [MMD as cost function]
    27. DeepMind (2023). “GraphCast: Learning skillful medium-range global weather forecasting.” Science, 382(6677). [Physics-informed loss]
    28. DeepSeek-AI (2025). “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” arXiv:2501.12948. [Multi-objective RLHF costs]

    Standards and Implementation

    Cost function design and implementation is supported by all major deep learning frameworks. In PyTorch, torch.nn.functional and torch.nn expose all standard losses (MSE, cross-entropy, hinge, NLL, KL divergence, cosine embedding) with automatic gradient computation via autograd. TensorFlow provides tf.keras.losses with equivalent coverage. JAX via the optax library provides composable loss functions with functional transforms. The frameworks handle the numerical subtleties automatically: log-sum-exp tricks to avoid overflow in cross-entropy, numerically stable KL computation, and half-precision support for large-scale training.

    The OpenAI Gym and DeepMind Control Suite standardise reward (negative cost) signals for Reinforcement Learning benchmarking, enabling consistent comparison of algorithms across cost function formulations. For control applications, the LQR cost formulation is specified in IEEE and IEC control standards and implemented in MATLAB’s Control System Toolbox. IEC 22989:2022 (AI concepts and terminology) references objective functions and loss functions as core AI vocabulary. The EU AI Act (effective August 2024) and NIST AI RMF 1.0 require documentation of cost function choices in high-risk AI systems; this is driving standardisation of cost function specification formats including machine-readable documentation of what is and is not captured in the training objective.

    Key Terminology

  • Cost function — the aggregate training objective; sum or mean of per-sample losses across the dataset; the quantity being minimised by the optimiser

  • Loss function — per-sample measure of prediction error; the cost function is typically the empirical expectation (average) of the loss over training data

  • Objective function — generic term for any mathematical function being optimised; cost/loss functions are special cases; also used in optimisation theory for functions whose extremum is sought

  • Data-fidelity term — the component of the cost that measures fit to training data; drives parameter updates toward better prediction; dominant signal during training

  • Regularisation — penalty term added to the cost to discourage model complexity; acts as a Bayesian prior on parameters; L1 induces sparsity, L2 shrinks weights, dropout and batch normalisation are data-level regularisers

  • Surrogate loss — differentiable proxy for a non-differentiable evaluation metric; enables gradient-based optimisation of complex metrics; the gap between surrogate and true metric is the proxy gap

  • Reward hacking — when an agent finds a policy that scores well on the proxy cost without achieving the intended goal; Goodhart’s Law applied to cost functions: “When a measure becomes a target, it ceases to be a good measure”

  • Loss landscape — the hypersurface J(θ) over parameter space Θ; its geometry (convexity, smoothness, saddle points, basin widths) determines optimiser behaviour and generalisation

  • ELBO — Evidence Lower BOund; the variational inference training objective; equals E[log p(x|z)] - KL[q(z|x)||p(z)]; tight lower bound on log p(x)

  • Bradley-Terry model — pairwise preference model underlying reward model training in RLHF; models the probability of preferring response a over b as σ(r(a) - r(b)) where r is the reward model

  • KL regularisation — penalty term in RLHF cost function: J_RLHF = r(y) - β·KL[π_θ||π_ref]; prevents the fine-tuned model from diverging too far from the base distribution

  • Proxy gap — the difference between the training cost and the true evaluation criterion; source of reward hacking and out-of-distribution failures; central concern of AI alignment

  • Convexity — property of cost functions where J(λθ₁ + (1-λ)θ₂) ≤ λJ(θ₁) + (1-λ)J(θ₂); guarantees global minimisability; rare in deep learning but common in classical ML

  • L-smoothness — property where ||∇J(θ₁) - ∇J(θ₂)|| ≤ L||θ₁ - θ₂||; determines stable step sizes; improved by gradient clipping and batch normalisation

Provenance