Hyperparameter optimisation is the automated search process for the configuration values — such as learning rate, regularisation strength, architecture depth, and batch size — that govern a machine learning model’s training dynamics but are not learned directly from data, with the aim of maximising held-out validation performance. It encompasses grid search, random search, Bayesian optimisation, and gradient-based meta-learning, operating over an outer loop that wraps the inner model training procedure.

Content

  • Early machine learning practitioners performed hyperparameter selection manually through domain intuition and grid search, a computationally expensive exhaustive enumeration of discrete value grids. Bergstra and Bengio’s 2012 paper on random search demonstrated that random sampling over continuous ranges empirically outperforms grid search at equal compute budgets, because relevant hyperparameters typically form a low-dimensional manifold within the full configuration space. This finding shifted standard practice toward random search as the default baseline.
  • Bayesian optimisation formalised principled sequential search by fitting a surrogate probabilistic model (Gaussian Process or Tree-structured Parzen Estimator) to observed (configuration, validation score) pairs and selecting the next configuration via an acquisition function that balances exploration and exploitation. Spearmint, Hyperopt, BOHB, and Optuna are widely adopted implementations. Asynchronous Hyperband (ASHA) and Population-Based Training (PBT) address the compute cost of evaluating each configuration to convergence by early-stopping poor performers and recycling resources to promising configurations.
  • In practice, hyperparameter optimisation is integrated into MLOps platforms — Weights & Biases Sweeps, MLflow, AWS SageMaker Automatic Model Tuning, Google Vertex AI Vizier — making it accessible without manual orchestration. Neural architecture search tools such as Keras Tuner and AutoGluon extend the search to architectural choices, approaching the vision of fully automated machine learning (AutoML). The commercial impact is significant: automated HPO routinely improves model accuracy by 2–10 percentage points over hand-tuned baselines on structured data tasks.
  • In 2024–2025, hyperparameter optimisation for large language models presents unique challenges: the cost of a single training run precludes the hundreds of evaluations assumed by classical Bayesian methods. Proxy tasks, scaling laws extrapolation, and transfer learning from smaller models to inform large-model configuration are active research directions. Curriculum learning schedules, mixed-precision training configurations, and RLHF reward model hyperparameters are new search axes that existing optimisation libraries are being extended to handle.