Change data capture is a set of techniques for identifying and propagating row-level changes — inserts, updates, and deletes — from a source database to downstream systems in near real time. The most robust approach reads the database transaction log, turning committed mutations into an ordered stream of change events without burdening the source with polling. It underpins data replication, event streaming, and incremental data integration, keeping analytical stores, caches, and microservices consistent with operational systems.

Overview

  • Traditional batch ETL re-reads whole tables on a schedule, which is slow, stale, and heavy on the source. CDC instead emits only what changed, as it changes.
  • Log-based CDC tails the database’s write-ahead or binary log, reconstructing each insert, update, and delete with before and after images and preserving commit order.
  • Trigger-based and query-based variants exist but impose overhead or miss intermediate states; log-based capture is the preferred low-impact method.
  • Captured changes are typically published to a durable log such as Apache Kafka, where many consumers can replay and process them independently.

Mechanisms

  • Read the transaction log to obtain an ordered, lossless change stream.
  • Serialise each change with metadata (operation type, source table, transaction, timestamp).
  • Publish to a streaming platform for fan-out to multiple sinks.
  • Apply exactly-once or idempotent delivery semantics to keep targets consistent.

Applications

Provenance