Platformpartitionsconsumer-laghead-of-line-blockingdead-letter-queueretriesmedium

One partition stopped moving at 11:02. Nothing alerted. Support did, at 14:41.

01Symptom

Your payments service publishes events to a 24-partition topic, and a consumer group sends payout notifications to customers. At 14:41, support reports that a small slice of customers, around 4%, have received no payout notifications since late morning, while everyone else is fine. Dashboards are green: consumers are up, CPU is idle, and nothing is crash-looping. Restarting the consumers briefly "fixes" nothing, and scaling the consumer group from 12 to 24 instances changes nothing either.

02Constraints

  • The topic has 24 partitions, keyed by account ID, so each account's events land in the same partition and are consumed in order
  • 12 consumers in one group, each owning two partitions; a partition is owned by exactly one consumer at a time
  • The handler catches every exception and retries the same event after 5 seconds, with no retry limit
  • Offsets are committed only after an event is processed successfully
  • Total throughput is about 240 events per second, roughly 10 per partition
  • Alerts: error rate above 5% of processed events, and total lag summed across all partitions above 250,000

03Evidence

  • Lag is zero on 23 partitions. On partition 17 it grows steadily and is at about 131,000 events
  • The committed offset on partition 17 has been stuck at 8,412,903 for 3 hours 39 minutes; it has not moved at all
  • One log line — `failed to process event, retrying` — repeats every 5 seconds against the same offset since 11:02
  • That is about 0.2 errors per second against roughly 230 events per second processed overall, so the error rate is under 0.1%
  • The event at offset 8,412,903 was produced by a back-office correction tool, run once, at 11:02. Its amount field is the string "12,50", while regular producers send a number
  • After a restart, partition 17 moves to another consumer within seconds. The same offset starts failing again immediately
  • With 24 consumers, 12 of them have no partitions assigned at all

→The question

Why did one event stall 4% of customers, why couldn't restarts or extra consumers help, and what should the consumer and its monitoring do differently?

04Your prediction

01Why were only about 4% of customers affected?
02Why didn't restarts or scaling to 24 consumers help?
03Why did no alert fire while a partition sat stuck for 3 hours 39 minutes?
0 / 600 chars