A single point of failure is a component, service, or dependency within a system whose failure would cause the entire system to stop functioning, because no redundant alternative exists to take over its role. Identifying and eliminating single points of failure is a central goal of high-availability and fault-tolerant design. Mitigation strategies include redundancy, replication, clustering, and load balancing so that no individual element is indispensable.
- A single point of failure is an indispensable Infrastructure element whose loss halts the whole system.
- It is the negative condition that Fault Tolerance and High Availability design seek to eliminate.
- It is directly contrasted with Redundancy and mitigated by Clustering.
- Detection relies on Health Monitoring and architectural review.
Overview
- The term denotes any node, link, process, or dependency that lacks a backup capable of assuming its function.
- Common examples include a sole database primary, a single network gateway, a lone power feed, or an unreplicated service.
- Eliminating single points of failure is a structured exercise in resilience and reliability engineering.
- Residual single points often hide in shared dependencies such as DNS, authentication, or configuration stores.
Key aspects
- Dependency mapping to surface critical, non-redundant components.
- Redundancy and replication to provide hot or warm standbys.
- Failover orchestration so traffic shifts away from a failed element.
- Continuous monitoring to verify that redundancy remains intact.
Applications
- Database clusters with replica promotion to avoid a single primary.
- Multi-zone deployments removing reliance on one data centre.
- Redundant load balancers and network paths.
- Resilience reviews and chaos testing to expose hidden dependencies.