Consumer groupshard3-5 years

A consumer joins a group of ten and the whole group stops processing for several seconds. Walk through why, and what changed between Kafka's original rebalance protocol and the cooperative one.

Under the original (eager) protocol, any membership change — a join, a leave, a crash — triggers a rebalance where every member first revokes all of its partitions, then the group re-joins through two coordinator round trips (JoinGroup, then SyncGroup), then the new assignment is handed out. Between revoke and re-assign, nobody in the group is processing anything, so a group of ten pauses entirely just to onboard one new member, even though only a fraction of the partitions actually needed to move. CooperativeStickyAssignor (Kafka 2.4) changes this to incremental: a rebalance revokes only the partitions that must actually move, so the nine members keeping their assignment keep processing the whole time.

The lesson behind it →