Disaster recovery (DR) is the set of policies, tools, and procedures enabling an organisation to restore its IT systems, data, and operations following a disruptive event such as hardware failure, cyberattack, natural disaster, or human error. It is quantified by Recovery Time Objective (RTO) and Recovery Point Objective (RPO), and encompasses backup strategies, replication architectures, and tested failover procedures.
Content
- Disaster recovery as a formalised discipline emerged in the 1970s alongside the growing dependence of financial institutions and government agencies on mainframe computing. Early DR programmes relied on hot-site agreements—contracted access to preconfigured hardware at a third-party facility—and daily tape backups transported off-site by courier. The 9/11 attacks (2001) and Hurricane Katrina (2005) exposed critical gaps in many organisations’ DR capabilities, triggering regulatory mandates across financial services (FFIEC, Basel II) and healthcare (HIPAA) sectors.
- Modern DR architectures are tiered by RTO and RPO requirements. Tier 0 (no DR) through Tier 7 (zero data loss, automated recovery) structures guide technology selection. Active-active architectures—where multiple sites concurrently serve production workloads—represent Tier 7 implementations, eliminating both RTO and RPO at significant infrastructure cost. Storage replication technologies (synchronous for near-zero RPO, asynchronous for cost-efficient longer RPO) combined with orchestration tools (Zerto, Veeam, AWS CloudEndure) automate failover sequencing. Regular DR testing—full failover exercises, tabletop simulations, chaos engineering—is essential to validate recovery runbooks.
- Cloud-native DR has fundamentally altered the economics and architecture of the discipline. DR-as-a-Service (DRaaS) providers use cloud elasticity to spin up recovery environments on demand rather than maintaining permanently provisioned hardware. Infrastructure-as-Code tools (Terraform, CloudFormation) enable entire environment configurations to be version-controlled and redeployed automatically. Kubernetes-based workloads can migrate between regions by reapplying manifests against pre-replicated data volumes, drastically reducing manual intervention during recovery.
- In 2024-2025, ransomware resilience has become the dominant DR design driver: immutable backup storage (object lock, air-gapped vaults), rapid detection of encryption events, and clean recovery point identification are now core DR requirements. AI-driven DR tools are emerging that automatically assess blast radius during incidents and prioritise recovery sequences by business criticality. Regulatory frameworks including DORA (Digital Operational Resilience Act) in the EU now mandate DR testing and incident reporting for financial institutions, elevating DR from operational best practice to compliance obligation.