Platform · 2026-10-11 · retries / retry-budget / amplification / backoff / load-shedding · medium

A 4-minute brownout became a 29-minute outage. Every layer retried correctly.

01Symptom

Your mobile app calls an edge API, which calls the checkout service, which calls the pricing service, which reads a pricing database. At 10:15 the pricing database begins a heavy maintenance compaction, and about 20% of pricing requests start failing with 503s. That's survivable, and each layer retries. For the first minute the user-facing error rate is nearly zero. At 10:19 the compaction ends and the database returns to normal. Pricing's error rate keeps climbing anyway, from 20% to 80% by 10:22 and to 100% by 10:30. User traffic is flat, so no one is attacking you and nothing new was deployed.

02Constraints

  • Call chain: mobile app → edge API → checkout → pricing → pricing database
  • User-facing traffic is steady at about 1,000 requests per second
  • Pricing is healthy up to about 1,800 rps; above roughly 2,500 rps, latency passes the callers' timeouts, most work is wasted, and goodput falls. During the compaction its effective capacity drops to roughly 800 rps
  • Every layer retries up to 3 attempts total on timeouts and 5xx, with exponential backoff and full jitter. All price lookups are idempotent
  • Three teams own the three retry configs, and each looked reasonable on its own
  • No circuit breakers, no load shedding, and no retry budgets. Pricing returns a bare 503 with no Retry-After or "do not retry" signal

03Evidence

  • User-facing request rate is flat at about 1,000 rps for the whole incident
  • Inbound rate at pricing is 1,000 rps at 10:14, 1,240 at 10:16, and about 5,000 at 10:22
  • Pricing's error rate is about 20% from 10:15 to 10:19, about 80% at 10:22, and 100% at 10:30
  • The pricing database's latency is back to normal at 10:20. The 503s continue
  • Counting attempts at the pricing service per user request gives about 1.25 at 10:16 and about 5.0 at 10:22
  • Pricing's logs can't distinguish first attempts from retries; they look identical
  • Checkout p99 latency climbs to about 10 seconds, and its thread pool is full of requests waiting on pricing
→

Recent

Previously diagnosed.