One partition stopped moving at 11:02. Nothing alerted. Support did, at 14:41.
01Symptom
Your payments service publishes events to a 24-partition topic, and a consumer group sends payout notifications to customers. At 14:41, support reports that a small slice of customers, around 4%, have received no payout notifications since late morning, while everyone else is fine. Dashboards are green: consumers are up, CPU is idle, and nothing is crash-looping. Restarting the consumers briefly "fixes" nothing, and scaling the consumer group from 12 to 24 instances changes nothing either.
02Constraints
- The topic has 24 partitions, keyed by account ID, so each account's events land in the same partition and are consumed in order
- 12 consumers in one group, each owning two partitions; a partition is owned by exactly one consumer at a time
- The handler catches every exception and retries the same event after 5 seconds, with no retry limit
- Offsets are committed only after an event is processed successfully
- Total throughput is about 240 events per second, roughly 10 per partition
- Alerts: error rate above 5% of processed events, and total lag summed across all partitions above 250,000
03Evidence
- Lag is zero on 23 partitions. On partition 17 it grows steadily and is at about 131,000 events
- The committed offset on partition 17 has been stuck at 8,412,903 for 3 hours 39 minutes; it has not moved at all
- One log line — `failed to process event, retrying` — repeats every 5 seconds against the same offset since 11:02
- That is about 0.2 errors per second against roughly 230 events per second processed overall, so the error rate is under 0.1%
- The event at offset 8,412,903 was produced by a back-office correction tool, run once, at 11:02. Its amount field is the string "12,50", while regular producers send a number
- After a restart, partition 17 moves to another consumer within seconds. The same offset starts failing again immediately
- With 24 consumers, 12 of them have no partitions assigned at all
→The question
Why did one event stall 4% of customers, why couldn't restarts or extra consumers help, and what should the consumer and its monitoring do differently?
04Your prediction
0 / 600 chars