A computer system engineered to provide services — computation, storage, or content — to other machines over a network, typically built for sustained load, remote management, and high availability with error-correcting memory, redundant power supplies, and hot-swappable storage. Racked in their thousands inside data centres, servers are the physical substrate of cloud computing, and the term equally names the software process that answers client requests in the client-server model.
Semantic Classification
Content
Definition
A server is the answering half of the client-server model: a machine (or process) that waits for requests over a network and fulfils them — serving web pages, executing database queries, storing files, brokering messages, or running the inference and rendering workloads that underpin this graph’s AI and spatial-computing systems. The word covers both the physical Hardware and the software daemon; a single physical server routinely hosts dozens of logical servers through virtualisation and containers.
Server hardware differs from consumer computing in its bias towards sustained, unattended operation. Typical features include multi-socket Xeon or EPYC processors (and increasingly Arm parts such as Graviton and NVIDIA Grace), error-correcting (ECC) memory, redundant hot-swappable Power Supply units and fans, NVMe storage arrays, out-of-band management controllers (BMC/IPMI, iDRAC, iLO) for lights-out administration, and high-bandwidth network interfaces from 25 to 800 Gb/s. Form factors span 1U/2U rack units, blades, and open-standard designs from the Open Compute Project; GPU servers for AI training add 4-8 accelerators with their own interconnect fabrics and multi-kilowatt power draw.
Aggregation is what gives servers their significance: racked and networked by the thousand inside a Data Centre, provisioned by orchestration software, and rented by the second, they become the elastic pool that Cloud Computing abstracts into instances, functions, and managed services. The engineering agenda has accordingly shifted from the single machine to the fleet — failure is handled by replication rather than repair, capacity by horizontal scaling, and efficiency by driving utilisation and power effectiveness across the whole facility, with liquid cooling spreading as AI-class racks exceed 100 kW.
Technical Details
-
Roles: web/application servers (nginx, Apache), database servers (PostgreSQL, MySQL), file and object storage, mail, DNS, game and real-time media servers, GPU inference/training nodes.
-
Operating systems: Linux dominates the installed base; Windows Server persists in enterprise directories and .NET estates; the Operating System is increasingly wrapped by hypervisors (KVM, ESXi) and Kubernetes.
-
Reliability vocabulary: availability tiers expressed in “nines”; redundancy patterns N+1 and 2N for power and cooling; RAID/erasure coding for storage; live migration for maintenance.
-
Management plane: PXE/Redfish provisioning, infrastructure-as-code (Terraform, Ansible), telemetry and remote KVM via the BMC — itself a notable attack surface requiring firmware hygiene.
-
Trends: heterogeneous compute (GPUs, DPUs/SmartNICs offloading Networking and storage), Arm adoption for performance-per-watt, confidential-computing enclaves, and rack-scale liquid cooling for AI density.
Current Landscape
-
The rack is the new server: NVIDIA’s GB200 NVL72 fuses 72 Blackwell GPUs and 36 Grace CPUs into a single liquid-cooled rack behaving as one accelerator, linked by an NVLink domain delivering 130 TB/s of aggregate GPU bandwidth — an “exascale computer in a single rack”.
-
Power density has jumped an order of magnitude: a GB200 NVL72 draws roughly 120 kW nominal (130-132 kW observed under full load) per rack, versus the ~7.6 kW industry-average rack reported by the Uptime Institute in 2025, forcing purpose-built or modular facilities.
-
Liquid cooling is now mandatory, not optional: NVIDIA ships the NVL72 only with direct-to-chip cold plates on GPUs, CPUs and NVLink switches (roughly 2 L/s coolant flow at a ~20-25°C inlet); air cooling is not viable at this density.
-
Open hardware for AI racks: NVIDIA contributed the GB200 NVL72 rack and its liquid-cooled compute/switch tray designs to the Open Compute Project at the 2024 OCP Global Summit, and the successor GB300 NVL72 (Blackwell Ultra) is positioned for test-time-scaling inference and reasoning workloads.
Sources: