A lag alert fires. How do you tell from the lag data alone whether it's a capacity problem, a stuck partition, or a rebalance loop?
Lag is log-end offset minus committed offset per partition, and its shape across partitions is the diagnostic. Rising steadily on every partition means the group is consistently slower than the producers — a capacity problem, fixed by more partitions, a faster handler, or both. Rising on one partition while the rest sit at zero means that one partition is stuck — a poison message retrying forever, or a hot key overloading a single consumer. A sawtooth that resets and climbs again, with evidence of duplicate processing, means the group keeps rebalancing and redoing work rather than falling behind on genuinely new records. Reading which shape you have before reaching for a fix is what separates a five-minute diagnosis from an hour of guessing.