A 4-minute brownout became a 29-minute outage. Every layer retried correctly.
01Symptom
Your mobile app calls an edge API, which calls the checkout service, which calls the pricing service, which reads a pricing database. At 10:15 the pricing database begins a heavy maintenance compaction, and about 20% of pricing requests start failing with 503s. That's survivable, and each layer retries. For the first minute the user-facing error rate is nearly zero. At 10:19 the compaction ends and the database returns to normal. Pricing's error rate keeps climbing anyway, from 20% to 80% by 10:22 and to 100% by 10:30. User traffic is flat, so no one is attacking you and nothing new was deployed.
02Constraints
- Call chain: mobile app → edge API → checkout → pricing → pricing database
- User-facing traffic is steady at about 1,000 requests per second
- Pricing is healthy up to about 1,800 rps; above roughly 2,500 rps, latency passes the callers' timeouts, most work is wasted, and goodput falls. During the compaction its effective capacity drops to roughly 800 rps
- Every layer retries up to 3 attempts total on timeouts and 5xx, with exponential backoff and full jitter. All price lookups are idempotent
- Three teams own the three retry configs, and each looked reasonable on its own
- No circuit breakers, no load shedding, and no retry budgets. Pricing returns a bare 503 with no Retry-After or "do not retry" signal
03Evidence
- User-facing request rate is flat at about 1,000 rps for the whole incident
- Inbound rate at pricing is 1,000 rps at 10:14, 1,240 at 10:16, and about 5,000 at 10:22
- Pricing's error rate is about 20% from 10:15 to 10:19, about 80% at 10:22, and 100% at 10:30
- The pricing database's latency is back to normal at 10:20. The 503s continue
- Counting attempts at the pricing service per user request gives about 1.25 at 10:16 and about 5.0 at 10:22
- Pricing's logs can't distinguish first attempts from retries; they look identical
- Checkout p99 latency climbs to about 10 seconds, and its thread pool is full of requests waiting on pricing
→
Recent
Previously diagnosed.
- The nightly job had run for three years. One Sunday in March it didn't, and nothing failed.cron · dst · timezones · scheduler · missed-runs · monitoringmedium
- One partition stopped moving at 11:02. Nothing alerted. Support did, at 14:41.partitions · consumer-lag · head-of-line-blocking · dead-letter-queue · retriesmedium
- Duplicate primary keys on one node, for about two seconds, once a monthclocks · id-generation · ntp · upsert · silent-corruptionmedium
- A 30 ms migration took the orders table down for 17 minutespostgres · locks · migrations · lock-queue · lock-timeoutmedium