Root cause analysis (RCA) is a structured problem-solving discipline that identifies the underlying origin of a fault, failure, or incident rather than merely treating its visible symptoms. In infrastructure and reliability engineering it traces a chain of contributing causes back to the conditions that, if corrected, would have prevented the event. RCA produces durable, systemic fixes and feeds learning back into operational practice.
Overview
- RCA begins after detection and containment of an incident, when responders shift from restoring service to understanding why the failure occurred. Investigators reconstruct a timeline from telemetry, logs, and traces, then iteratively ask why each contributing condition existed until the analysis reaches systemic factors that are actionable. Common techniques include the Five Whys, fault tree analysis, fishbone (Ishikawa) diagrams, and causal change analysis. The output is a set of corrective and preventive actions tracked to completion.
Mechanisms
- Causal chaining: tracing proximate symptoms back through intermediate failures to systemic origins.
- Blameless culture: focusing on process and system weaknesses rather than individual fault to encourage honest disclosure.
- Evidence gathering: correlating audit trails, distributed traces, and time-series telemetry to establish a defensible timeline.
- Categorisation: classifying causes as technical, procedural, or organisational to target the correct remediation layer.
- Corrective action tracking: assigning, prioritising, and verifying fixes so the same root cause cannot recur.
Applications
- Production outage post-mortems in cloud and on-premises infrastructure.
- Quality and safety investigations in manufacturing and regulated industries.
- Recurring-defect elimination in software delivery pipelines.
- Security incident analysis to identify the initial compromise vector.