Cloud computing is the on-demand delivery of computing resources — servers, storage, databases, networking, software, analytics, and AI accelerators — over the internet via provider-managed data centres, abstracting physical infrastructure into programmable APIs with pay-per-use economics. Service models (IaaS, PaaS, SaaS) and deployment models (public, private, hybrid, multi-cloud) define the boundary of managed responsibility between provider and consumer. Hyperscale providers such as AWS, Microsoft Azure, and Google Cloud Platform underpin modern AI training, inference serving, and distributed application deployment at global scale. The paradigm enables elastic provisioning — scaling from zero to thousands of compute nodes in seconds — transforming both software engineering and machine learning operations.
Overview
- Cloud computing emerged from the insight that large-scale data centre operators (initially Amazon, then Microsoft and Google) could expose spare capacity as rentable compute units via standardised web APIs. The NIST SP 800-145 definition codified five essential characteristics: on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured service.
- Why it matters: it democratises access to supercomputing-class infrastructure, allowing a startup to train a large neural network or serve millions of API requests with the same infrastructure primitives available to the world’s largest enterprises — billed only for active usage.
- How it works: physical resources in geographically distributed data centres are partitioned via Virtualisation and Containerisation into isolated tenant environments. A global control plane (the cloud provider’s management layer) handles provisioning, billing, identity, and observability, exposing these through REST/gRPC APIs, CLIs, and infrastructure-as-code tooling.
Key Components
Service Models
- Infrastructure as a Service (IaaS) — raw virtual machines, block and Object Storage, and virtual networks (e.g. AWS EC2, Azure Virtual Machines, Google Compute Engine). Consumer manages OS and above.
- Platform as a Service (PaaS) — managed runtimes, databases, and middleware; consumer manages only application code (e.g. Google App Engine, AWS Elastic Beanstalk, Azure App Service).
- Software as a Service (SaaS) — fully managed applications delivered via browser or API; consumer configures, not operates (e.g. Salesforce, Google Workspace, Microsoft 365).
- Serverless Computing — event-driven functions billed per invocation with no persistent server management (AWS Lambda, Azure Functions, Google Cloud Functions).
Deployment Models
- Public Cloud — infrastructure owned and operated by a third-party hyperscaler, shared among tenants with logical isolation.
- Private Cloud — dedicated infrastructure operated for a single organisation, on-premises or collocated.
- Hybrid Cloud — orchestrated integration of public and private cloud resources, connected via secure network fabric, enabling data-sovereignty compliance and burst capacity.
- Multi-Cloud — use of services from multiple providers simultaneously to avoid vendor lock-in and optimise cost or capability.
Core Infrastructure Primitives
- Virtualisation — hypervisors (KVM, Hyper-V, Xen) partition physical servers into isolated VMs.
- Containerisation — lightweight process isolation via namespaces and cgroups (Docker, OCI images) run on shared OS kernels.
- Kubernetes — de-facto orchestration layer for scheduling, scaling, and managing containerised workloads across cloud and hybrid environments.
- Object Storage — massively scalable, durable blob storage (S3, Azure Blob, GCS) underpinning data lakes and ML training datasets.
- Content Delivery Network — geographically distributed cache layer that moves static and dynamic content close to users, reducing latency.
- Network Infrastructure — software-defined networking (SDN), virtual private clouds (VPCs), load balancers, and global backbone connectivity.
- Data Centre — physical facilities housing hyperscale server racks, cooling, power, and physical security.
Mechanisms
- Elasticity and Auto-Scaling: workloads can burst to thousands of compute nodes on demand and release them automatically based on CPU, memory, or custom metrics — critical for AI training and inference traffic spikes.
- Measured Service: metering at fine granularity (per-second, per-request, per-GB) enables pay-per-use billing and FinOps Cost Optimisation practices.
- Resource Pooling (Multi-Tenancy): physical resources are shared among multiple customers with logical isolation enforced via Virtualisation, network segmentation, and IAM policies.
- High Availability and Fault Tolerance: High Availability is achieved through geographically redundant availability zones and regions, automatic failover, and replication at the storage layer.
- Managed Services Ecosystem: providers offer hundreds of managed services (databases, message queues, ML pipelines, monitoring) reducing operational burden and enabling DevOps teams to focus on product logic.
Applications / Use Cases
- AI and Machine Learning: Distributed Training of large language models requires ephemeral access to thousands of GPUs/TPUs for days; cloud eliminates the need to own these accelerators. Inference serving scales elastically with traffic.
- MLOps Pipelines: managed experiment tracking, feature stores, model registries, and CI/CD for models are increasingly cloud-native services.
- Big Data Analytics: petabyte-scale data processing via cloud-managed Spark, BigQuery, Snowflake, and Databricks clusters.
- DevOps and CI/CD: cloud-hosted build pipelines, container registries, and deployment targets accelerate software delivery.
- Digital Twin Simulation: real-time simulation of physical systems (manufacturing lines, urban environments) leverages cloud burst compute for high-fidelity models.
- Spatial Computing and XR: cloud rendering and streaming reduces client-side hardware requirements for augmented and virtual reality applications.
- Federated Learning: cloud orchestrates distributed model training across decentralised data silos without raw data leaving each site, bridging privacy and AI capability.
- Disaster Recovery and Business Continuity: cloud-based replication and failover replace expensive secondary data centres.
- Regulated Industries: Hybrid Cloud and sovereign cloud configurations (e.g. AWS GovCloud, Azure Government) meet data-residency requirements in healthcare, finance, and defence.
Standards & Context
- NIST SP 800-145 — the authoritative US government definition of cloud computing, establishing the five essential characteristics, three service models, and four deployment models. Published by the National Institute of Standards and Technology.
- ISO/IEC 17788:2014 — international standard providing the overview and vocabulary for cloud computing, harmonising terminology across the industry.
- ISO/IEC 17789:2014 — cloud computing reference architecture, defining the roles and activities of cloud service customers, providers, and partners.
- CSA Cloud Controls Matrix (CCM) — the Cloud Security Alliance’s cybersecurity control framework for Cloud Security assessment and compliance, widely referenced alongside ISO 27001 and SOC 2.
- GDPR and Data Residency Regulation — European data protection law constrains where personal data may be processed; drives Hybrid Cloud and sovereign cloud architectures for EU-based workloads.
- FinOps Foundation — open practitioner community standardising Cost Optimisation practices for cloud spend; FinOps framework defines crawl/walk/run maturity for cloud financial management.
- Green Software Foundation — develops standards for measuring and reducing the carbon footprint of cloud workloads; relevant to sustainable AI training.
Current Landscape (2026)
- AI has become the dominant growth engine: global cloud infrastructure services spending hit US$129bn in Q1 2026, up 35% year on year (the steepest since late 2021) on a run-rate above half a trillion dollars, per Synergy Research Group, driven by generative-AI workloads rather than the earlier pandemic-era migration wave.
- Growth-rate divergence among the “Big Three” widened in Q1 2026: AWS grew 28% (its fastest in 15 quarters, ~28% share), Microsoft Azure 40% (~21-25% share) and Google Cloud 63% to US460bn as enterprises commit to multi-year capacity.
- Hyperscaler capital expenditure is exploding: the big five (Amazon, Alphabet, Microsoft, Meta, Oracle) are forecast to spend over US450bn) tied directly to AI infrastructure (GPUs, servers, data centres); AWS alone guided to ~US$200bn capex in 2026.
- “Neocloud” GPU-specialist providers (CoreWeave and peers) collectively crossed 5% of the cloud market by late 2025, with several second-tier names entering the top thirty, reshaping supply for scarce accelerated-compute capacity.
- The EU Data Act’s cloud-switching regime (Chapter VI) became applicable on 12 September 2025, forcing providers to remove technical and contractual lock-in and complete switches within 30 days; all switching and data-egress fees must fall to zero by 12 January 2027, with national enforcers (e.g. Germany’s Bundesnetzagentur) empowered to fine up to 4% of global turnover.
- Sovereign cloud moved from preference to mandate: the Commission proposed the Cloud and AI Development Act (CADA) on 3 June 2026, introducing a four-level Union sovereignty-assurance framework and aiming to at least triple EU data-centre capacity within 5-7 years, layered on DORA (full enforcement January 2025), NIS2 and the EU AI Act (full application 2 August 2026).
- Open challenges as of 2026 include AI-driven power and data-centre capacity constraints, GPU scarcity and capital intensity straining ROI, spiralling FinOps/GreenOps pressure, and reconciling US CLOUD Act extraterritorial exposure with tightening EU sovereignty and switching obligations.
References
-
- Omdia / Informa (2026). Global cloud infrastructure spending rose 29% in Q4 2025 as hyperscalers scaled AI infrastructure investment. https://omdia.tech.informa.com/pr/2026/mar/global-cloud-infrastructure-spending-rose-29percent-in-q4-2025-as-hyperscalers-scaled-ai-infrastructure-investment
-
- Synergy Research Group / VoxBooster (2026). Cloud Computing Statistics 2026: 55+ Data Points on Market Growth, Provider Share, and Spending. https://voxbooster.com/blog/cloud-computing-statistics-2026/
-
- TechBlog / ComSoc (2025). Hyperscaler capex > $600bn in 2026, a 36% increase over 2025, while global spending on cloud infrastructure services skyrockets. https://techblog.comsoc.org/2025/12/22/hyperscaler-capex-600-bn-in-2026-a-36-increase-over-2025-while-global-spending-on-cloud-infrastructure-services-skyrockets/
-
- Greenberg Traurig (2025). Cloud Switching Under the EU Data Act. https://www.gtlaw.com/en/insights/2025/9/cloud-switching-under-the-eu-data-act
-
- Cloud Security Alliance (2026). EU Tech Sovereignty: Cloud Concentration Risk and the Cloud and AI Development Act (CADA). https://labs.cloudsecurityalliance.org/research/eu-tech-sovereignty-cloud-ai-enterprise-risk-v1-0-csa-styled/