Apache Kafka is an open-source distributed event streaming platform originally developed at LinkedIn and donated to the Apache Software Foundation in 2011. It provides a high-throughput, low-latency, fault-tolerant publish-subscribe messaging system built around an immutable, ordered, partitioned commit log. Kafka decouples producers and consumers of data streams, enabling real-time data pipelines, event-driven architectures, and stream processing applications at scale across thousands of nodes handling trillions of events per day.

Content

  • Apache Kafka was created at LinkedIn by Jay Kreps, Neha Narkhede, and Jun Rao to solve the problem of integrating heterogeneous data systems at LinkedIn’s scale. The name references Franz Kafka as a nod to the system being “a system optimised for writing.” After open-sourcing in 2011 and entering the Apache incubator, Kafka became the de facto standard for enterprise event streaming, with adoption at companies including Netflix, Uber, Airbnb, and Twitter.
  • Kafka’s core abstraction is the topic: an ordered, immutable sequence of records (events) that is partitioned across a cluster of brokers for parallel throughput and replicated for fault tolerance. Producers append records to topic partitions; consumers subscribe to partitions and read at their own pace, with offset tracking enabling both real-time consumption and historical replay. This durable, replayable log model separates Kafka from traditional message queues that delete messages upon consumption.
  • The Kafka Streams API and the companion Apache Flink and Apache Spark frameworks enable stateful stream processing directly on Kafka topics. Kafka Connect provides a standardised connector ecosystem for ingesting data from databases (via Change Data Capture), file systems, cloud storage, and SaaS platforms into Kafka topics, and for sinking processed data to downstream stores. The Schema Registry (from Confluent) enforces data contract compatibility across producer-consumer boundaries using Avro, Protobuf, or JSON Schema.
  • KRaft mode (Kafka Raft metadata), reaching production readiness in Kafka 3.3 (2022), eliminated the historical dependency on Apache ZooKeeper for cluster metadata management. ZooKeeper had been a complexity and scalability bottleneck; replacing it with an internal Raft-based consensus protocol simplified deployment significantly and improved cluster startup and failover performance.
  • In machine learning contexts, Kafka serves as the real-time feature pipeline for online feature stores, delivering model input features with millisecond latency. It also underpins model monitoring infrastructure by streaming prediction logs and ground-truth labels to detection systems that watch for data drift and model degradation. Confluent’s managed Kafka-as-a-service and the availability of Kafka on all major cloud providers have made it foundational infrastructure in both traditional enterprise data architectures and modern AI/ML platforms.