/ BigData, RealTimeData, DataStreaming, Kafka, Confluent, OpenSource, DataEngineering

Apache Kafka's Evolution: Timeline, Fun Facts & Key Terminology

A peek into Kafka's Evolution

As someone following the incredible journey of Apache Kafka, I wanted to share a comprehensive timeline, some fun facts, and key terminology you can't miss. Kafka has revolutionised real-time data processing and continues to shape the future of data infrastructure.

Kafka Timeline

| Year | Milestone | |------|-----------| | 2008 | Conception at LinkedIn | | 2011 | Open-sourced under Apache | | 2013 | Kafka 0.8 introduces replication | | 2014 | Confluent founded for adoption | | 2015 | Multiple security features released | | 2017 | Kafka Streams launched in Kafka 0.10 | | 2018 | Kafka 1.0 released, marking maturity | | 2021 | K8s support and KRaft early access | | 2023 | 3.0 tiered storage and security |

Fun Facts

  • Developed originally by LinkedIn
  • Named after novelist Franz Kafka
  • Joined the Apache Incubator in 2011
  • LinkedIn processes over a trillion messages per day
  • 80% of Fortune 100 companies use Kafka

Key Terminology

  • Broker: Server that stores data and serves clients
  • Topic: Address to which data and records are sent
  • Partition: A division of a topic, allowing parallelism
  • Producer: Application (source of data) that writes data to topics
  • Consumer: Application (destination) that reads data from topics
  • ZooKeeper: A centralised service for configurations
  • Kafka Streams: A client library for stream-processing applications