Alerting is the observability capability that evaluates monitored signals against defined conditions and notifies responsible humans or automated systems when those conditions indicate a problem or an impending one. It converts continuous telemetry into discrete, actionable notifications routed to the appropriate on-call recipient. Effective alerting balances sensitivity against noise so that every alert is meaningful and timely.
Overview
- Alerting sits atop a monitoring stack, consuming metrics, logs, and traces and applying rules, thresholds, or statistical models to decide when a condition warrants attention. When a rule fires, the alerting system deduplicates, groups, and routes the notification through escalation policies to an on-call engineer or an automated remediation workflow. Mature practice ties alerts to symptoms that affect users — typically service-level objectives — rather than to every low-level metric, reducing fatigue and improving signal quality.
Key aspects
- Threshold and rule evaluation: static thresholds, rate-of-change rules, and dynamic baselines derived from anomaly detection.
- Routing and escalation: directing alerts to the correct team via on-call schedules and escalating unacknowledged alerts.
- Deduplication and grouping: collapsing related alerts to prevent storms during widespread failures.
- Severity classification: distinguishing pages that demand immediate action from informational warnings.
- Alert quality management: tuning rules to minimise false positives and combat alert fatigue.
Applications
- On-call paging for production service degradation.
- SLO burn-rate alerts in site reliability engineering.
- Security event notification from intrusion detection systems.
- Capacity and saturation warnings in infrastructure operations.