Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #113Advanced 10 min read

Kafka Deep Dive: Topics, Partitions, Brokers, & Consumer Groups

💡
Core Architecture Summary

Master Kafka internals: Topic partitioning, segment storage, leader/follower replication, ISR (In-Sync Replicas), acks=all durability, and consumer group rebalancing.

Key Glossary Concepts in this TopicAll Glossary Terms

Kafka Partition Topology, Replication, & Consumer Group Assignment

Partition leader/follower replication across brokers and 1-to-1 consumer group partition mapping.

Kafka Partition Topology, Replication, & Consumer Group Assignment
100%
Rendering visual architecture flowchart...

01.1. Topics & Partitions: The Core Scalability Primitive

In Apache Kafka, a Topic is a logical category or feed name to which records are published. Topics are divided into one or more Partitions:

  • Physical Storage: A partition is an ordered, immutable sequence of records continuously appended to structured log files on disk (stored in ≈ 1GB segment files such as 00000000000000000000.log and indexed by .index and .timeindex).
  • Unit of Parallelism: A single partition cannot be split across multiple broker nodes. Therefore, the number of partitions in a topic represents the maximum theoretical concurrency for write ingestion and consumer processing.
  • Partition Keys & Deterministic Hashing: When a producer publishes a record with a key (user_12345), Kafka computes MurmurHash2(key) % num_partitions to assign the partition. All records sharing the same key are guaranteed to land on the identical partition in strict chronological order.

02.2. Leader-Follower Replication & In-Sync Replicas (ISR)

To guarantee high availability and prevent data loss upon hardware failure, Kafka partitions are replicated across multiple broker nodes according to the Replication Factor (typically 3 in production):

  • Partition Leader: One broker is elected as the Leader for each partition. All producer writes and consumer reads route directly through the Leader.
  • Follower Replicas: The remaining replicas actively pull records from the Leader over the network to replicate the log sequentially.
  • In-Sync Replicas (ISR): The subset of follower replicas that are actively caught up with the Leader's log end offset. If a follower lags behind (controlled by replica.lag.time.max.ms = 30000), the Leader ejects it from the ISR pool.
  • Producer Durability (acks=all):
    • acks=0: Producer does not wait for any broker acknowledgment (< 1ms latency, high data loss risk).
    • acks=1: Producer waits only for the local Leader to write to disk (2 - 5ms).
    • acks=all (or -1) with min.insync.replicas=2: The Leader writes to disk AND confirms replication to at least one ISR follower before returning success to the producer (10 - 20ms), guaranteeing zero data loss even if the leader crashes.

03.3. Consumer Groups & The Cardinality Rule

A Consumer Group represents a coordinated fleet of consumers pooling together to ingest data from a topic:

The Fundamental Cardinality Rule:

Within any single Consumer Group, each topic partition is assigned to AT MOST ONE active consumer instance.

  • Scenario 1 (Partitions = Consumers): A topic has 6 partitions and the group has 6 consumers. Each consumer processes exactly 1 partition (100\% efficiency).
  • Scenario 2 (Consumers < Partitions): A topic has 6 partitions and the group has 3 consumers. Each consumer is assigned 2 partitions.
  • Scenario 3 (Consumers > Partitions): A topic has 6 partitions and the group has 10 consumers. 6 consumers process data, while 4 consumers sit completely idle as hot standbys.
  • Independent Groups: Multiple distinct consumer groups (e.g., order-fulfillment and fraud-analytics) each receive their own independent stream of all messages from all partitions without competing with each other.

04.4. Consumer Group Rebalancing Protocols

When a consumer joins or leaves a group (or crashes due to exceeding max.poll.interval.ms), Kafka triggers a Group Rebalance to reassign partition ownership:

  • Eager Rebalancing (Legacy): All consumers stop processing, drop all partition assignments, rejoin the group coordinator, and receive new assignments. This causes a "stop-the-world" latency spike across the entire processing pipeline.
  • Cooperative Sticky Rebalancing (Modern Kafka): Consumers only give up partitions that must be migrated to rebalance the load. Surviving consumers continue processing unaffected partitions seamlessly without global pipeline pauses.
  • Heartbeat Thread: Consumers run a background heartbeat thread (heartbeat.interval.ms = 3000) sending liveness pings to the broker Group Coordinator. If no heartbeat is received within session.timeout.ms = 45000, the broker assumes the consumer is dead and initiates a rebalance.

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Massive linear write and read scaling by adding partitions and consumer instances
  • Configurable durability guarantees via In-Sync Replicas (ISR) and acks=all settings
  • Supports multiple independent consumer groups reading at independent speeds from the same topic

Trade-offs & Constraints

  • Consumer group maximum concurrency is strictly bounded by the topic partition count
  • Partition count cannot be easily decreased after creation; increasing partitions breaks key-to-partition hash mapping
Production Implementation in Big Tech
Netflix• Keystone Real-Time Telemetry Pipeline

Netflix operates thousands of Kafka brokers processing over 1.5 trillion messages and several petabytes of telemetry daily. Topics are provisioned with 64 to 256 partitions each. Consumer groups scale across AWS EC2 fleets to calculate viewing metrics, deliver personalized video recommendations, and detect network streaming anomalies in sub-second timeframes.

🎯 Staff+ Engineering Takeaways

  • Partitions are the fundamental unit of storage, throughput, and parallelism in Kafka.
  • Max active consumers in a consumer group cannot exceed the number of partitions in the topic.
  • In-Sync Replicas (ISR) and acks=all guarantee durability against broker hardware failure.
  • Cooperative Sticky Rebalancing prevents stop-the-world pauses during consumer scaling.

Topic Knowledge Assessment 🧠

Step through 3 scenario questions to test your staff-level grasp.

Question 1 of 30 answered
#1

A Kafka topic has 8 partitions. A development team deploys 12 consumer instances within the same Consumer Group. How many consumers will actively read messages?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?