Platformretriesretry-budgetamplificationbackoffload-sheddingmedium

A 4-minute brownout became a 29-minute outage. Every layer retried correctly.

01Symptom

Your mobile app calls an edge API, which calls the checkout service, which calls the pricing service, which reads a pricing database. At 10:15 the pricing database begins a heavy maintenance compaction, and about 20% of pricing requests start failing with 503s. That's survivable, and each layer retries. For the first minute the user-facing error rate is nearly zero. At 10:19 the compaction ends and the database returns to normal. Pricing's error rate keeps climbing anyway, from 20% to 80% by 10:22 and to 100% by 10:30. User traffic is flat, so no one is attacking you and nothing new was deployed.

02Constraints

  • Call chain: mobile app → edge API → checkout → pricing → pricing database
  • User-facing traffic is steady at about 1,000 requests per second
  • Pricing is healthy up to about 1,800 rps; above roughly 2,500 rps, latency passes the callers' timeouts, most work is wasted, and goodput falls. During the compaction its effective capacity drops to roughly 800 rps
  • Every layer retries up to 3 attempts total on timeouts and 5xx, with exponential backoff and full jitter. All price lookups are idempotent
  • Three teams own the three retry configs, and each looked reasonable on its own
  • No circuit breakers, no load shedding, and no retry budgets. Pricing returns a bare 503 with no Retry-After or "do not retry" signal

03Evidence

  • User-facing request rate is flat at about 1,000 rps for the whole incident
  • Inbound rate at pricing is 1,000 rps at 10:14, 1,240 at 10:16, and about 5,000 at 10:22
  • Pricing's error rate is about 20% from 10:15 to 10:19, about 80% at 10:22, and 100% at 10:30
  • The pricing database's latency is back to normal at 10:20. The 503s continue
  • Counting attempts at the pricing service per user request gives about 1.25 at 10:16 and about 5.0 at 10:22
  • Pricing's logs can't distinguish first attempts from retries; they look identical
  • Checkout p99 latency climbs to about 10 seconds, and its thread pool is full of requests waiting on pricing

→The question

Why did a failure that ended at 10:19 keep getting worse until 10:30, why didn't backoff and jitter prevent it, and what limits the damage next time?

04Your prediction

01At the peak, roughly how many pricing attempts does one user request cause, and why?
02Why did pricing stay down after its database recovered?
03What is a retry budget, and what does it do here?
0 / 600 chars