Capacity planning is the process of determining the production, infrastructure, or service capacity required by an organisation to meet changing demand over a defined time horizon, balancing the cost of over-provisioned resources against the risk of under-provisioned systems that cannot meet service-level objectives. In technology contexts, capacity planning encompasses compute, storage, network bandwidth, and human resources, employing demand forecasting models, utilisation metrics, and growth projections to derive procurement and scaling roadmaps. In manufacturing, the same principles apply to machine-hours, floor space, and workforce shifts.

Content

  • Capacity planning has roots in operations research and manufacturing scheduling, where linear programming models were applied from the 1950s onwards to optimise factory output against machine and labour constraints. In IT, the discipline emerged formally in the 1980s as mainframe compute time was an expensive, constrained resource that had to be carefully allocated across business units. Queuing theory models (particularly M/M/1 and M/G/1 queue analysis) were applied to predict response times at various utilisation levels, forming the mathematical core of IT capacity planning practice.
  • In the era of virtualisation and cloud computing, capacity planning shifted from predicting hardware procurement cycles (measured in months or years) to configuring auto-scaling policies measured in minutes. Tools such as Kubernetes Horizontal Pod Autoscaler and Vertical Pod Autoscaler automate reactive scaling, while capacity planning now focuses on the medium-term (90-day to 2-year) horizon: reserved instance commitments, GPU cluster procurement for AI workloads, and data centre power and cooling constraints. Statistical models — including time-series decomposition, regression against business metrics, and percentile-based SLO budgeting — are standard inputs to planning cycles.
  • For AI inference infrastructure in particular, capacity planning presents novel challenges. GPU memory is a hard constraint that determines the maximum model size and batch size concurrently serveable on a given instance. Demand patterns for AI APIs exhibit high burstiness and unpredictability compared to traditional web traffic, making over-provisioning costly and under-provisioning immediately visible as latency degradation. Organisations operating large language model services run dedicated capacity planning teams that model token throughput, KV cache utilisation, and request queue depth alongside conventional compute and network metrics.
  • By 2024-2025, capacity planning has become tightly coupled with FinOps (cloud financial operations) disciplines, with organisations using real-time cost attribution alongside utilisation data to make economically optimal scaling decisions. The proliferation of spot and preemptible instances has introduced probabilistic capacity planning, where workloads must be designed to tolerate instance reclamation events. Carbon-aware capacity planning — scheduling flexible workloads to times and regions with lower grid carbon intensity — is an emerging practice aligned with corporate sustainability commitments.