Alerting is the observability capability that evaluates monitored signals against defined conditions and notifies responsible humans or automated systems when those conditions indicate a problem or an impending one. It converts continuous telemetry into discrete, actionable notifications routed to the appropriate on-call recipient. Effective alerting balances sensitivity against noise so that every alert is meaningful and timely.

Overview

  • Alerting sits atop a monitoring stack, consuming metrics, logs, and traces and applying rules, thresholds, or statistical models to decide when a condition warrants attention. When a rule fires, the alerting system deduplicates, groups, and routes the notification through escalation policies to an on-call engineer or an automated remediation workflow. Mature practice ties alerts to symptoms that affect users — typically service-level objectives — rather than to every low-level metric, reducing fatigue and improving signal quality.

Key aspects

  • Threshold and rule evaluation: static thresholds, rate-of-change rules, and dynamic baselines derived from anomaly detection.
  • Routing and escalation: directing alerts to the correct team via on-call schedules and escalating unacknowledged alerts.
  • Deduplication and grouping: collapsing related alerts to prevent storms during widespread failures.
  • Severity classification: distinguishing pages that demand immediate action from informational warnings.
  • Alert quality management: tuning rules to minimise false positives and combat alert fatigue.

Applications

  • On-call paging for production service degradation.
  • SLO burn-rate alerts in site reliability engineering.
  • Security event notification from intrusion detection systems.
  • Capacity and saturation warnings in infrastructure operations.

Provenance