Statistical modelling is the practice of representing data-generating processes with mathematical structures built on probability theory, in order to describe relationships, quantify uncertainty, test hypotheses, and make inferences or predictions. It encompasses approaches such as regression, generalised linear models, time-series models, and Bayesian methods, emphasising interpretable parameters and explicit assumptions. It provides the formal foundation on which much of machine learning and data analysis is built.

Overview

  • A statistical model encodes assumptions about how observed data arise, often via parameters with interpretable meaning.
  • Fitting estimates those parameters from data, and inference characterises their uncertainty.
  • The emphasis is on explanation, calibrated uncertainty, and hypothesis testing as much as raw prediction.
  • It contrasts with black-box Deep Learning in its transparency and explicit assumptions.

Mechanisms

  • Model specification: choosing a family (linear, GLM, mixed, time-series) and link structure.
  • Estimation: maximum likelihood, least squares, or Bayesian posterior inference.
  • Inference: confidence intervals, hypothesis tests, and credible intervals.
  • Diagnostics: residual analysis, goodness-of-fit, and model selection.

Key aspects

  • Interpretability: parameters carry domain meaning.
  • Uncertainty quantification: a first-class output, not an afterthought.
  • Assumptions: explicit and testable, governing validity.
  • Parsimony: preferring simpler models that generalise.

Applications

Provenance