Metrics are numeric, time-series measurements that quantify the state and behaviour of systems, services and infrastructure over time. As one of the three pillars of observability alongside logs and traces, they are typically aggregated counters, gauges and histograms scraped or pushed at regular intervals. Metrics enable efficient trend analysis, alerting and capacity planning at scale because they compress system behaviour into compact, queryable numeric series.

Overview

  • Metrics provide the low-cardinality, numeric backbone of observability. Because each data point is a small number with a timestamp and a set of labels, metrics can be retained for long periods and queried cheaply to reveal trends, seasonality and anomalies. They are the natural substrate for dashboards, alert thresholds and autoscaling decisions.

Key aspects

  • Metric types: counters (monotonic), gauges (point-in-time values), histograms and summaries (distribution of observations).
  • Dimensionality: labels or tags that partition a metric into related series for slicing and aggregation.
  • Collection model: pull-based scraping or push-based emission at fixed intervals into a time-series database.
  • Aggregation: rollups, rate calculations and percentile estimation that summarise raw samples for analysis.
  • Retention and downsampling: storing high-resolution recent data and coarser long-term history to bound storage cost.

Applications

  • Powering real-time dashboards and service-level indicator tracking.
  • Triggering threshold-based and anomaly-based alerts.
  • Informing autoscaling and capacity-planning decisions.
  • Quantifying the impact of deployments and incidents on system health.

Provenance