Cache hit rate fell to 34% and stayed there. Stopping the rollout 18 minutes in changed nothing.
01Symptom
You run a team chat product. Every session's first call is client boot: the client asks who it is and which conversations it belongs to, and nothing works until that call returns. Boot reads group-chat membership — who is in which group DM — and that data is immutable once written and cached per channel with a long TTL, so boot is normally a handful of cache reads and finishes in single-digit milliseconds. You are rolling out a new version of the host agent across the fleet in 25% steps; two earlier steps that week went without incident. Just after 06:00 Pacific, tickets arrive, internal users cannot get in, and several teams are paged. Boot is failing outright or taking tens of seconds. One database keyspace is severely overloaded, and the query filling it is the one that lists group DMs. You pause the rollout. Twenty minutes later the metrics have not moved.
02Constraints
- The cache tier sits behind a proxy that maps each key to a node by consistent hashing over the ordered list of currently healthy nodes — change the list and keys remap
- 200 cache nodes in the ring. The paused step was taking 50 of them out, sequentially, and the proxy's healthy-host list changes with every one
- A provisioning controller watches the service catalog: when a cache node leaves the catalog it promotes an empty spare, and a node that rejoins is flushed before it is promoted. It is designed to keep the ring stable and to flush only when it must
- Group-DM membership is immutable and cached by channel ID with a long TTL, so it is almost always served from cache — 99.2% hit rate in the week before the rollout
- Membership rows live in the table sharded by user, across 200 shards. No shard owns a given channel, so answering 'who is in channel C' means querying every shard — a scatter query, one execution, 200 shard reads
- A table sharded by channel already exists. It holds channel metadata only — name, purpose, created_at — and none of the member rows this query needs
- Client calls time out at 400ms. The cache's fill path uses the same query over the same connection, so a fill fails whenever the read would
- Clients retry three times with exponential backoff and full jitter
- The only mitigation available at the time was throttling client boot, which holds unbooted clients at the door
03Evidence
- Membership cache hit rate falls from 99.2% to 34% across the step and then stays at 34% for the following 40 minutes — the controller log shows the last node replacement at minute 34, and the hit rate does not move after it
- Every entry in the controller log during the step is the same two events: a cold spare promoted to replace a departed node, and a rejoining node flushed before promotion. 50 hosts left the catalog and 50 empty spares were promoted over 22 minutes
- The hot query against the user-sharded table takes a single channel ID as its only parameter, and its shard count is 200 — no execution can be answered by one shard. Measured against the boot rate, database query count tracks boot rate × miss fraction × 200, not boot rate
- Median latency for that query sits at the 400ms client timeout for the whole incident, and the hit rate is flat rather than recovering: the fill path is timing out at the same rate the read path is
- Raising the boot cap from 250/min to 2,500/min put query fan-out back to its pre-throttle level within 90 seconds. The cap had to be cut back to 250 and then raised in roughly 10% steps, holding at each step until fan-out stopped climbing
- Replica lag on the user-sharded table is 200ms at p99, membership rows are never updated after insert, and 100% of the 200-way fan-out went to primaries
Recent
Previously diagnosed.
- One customer's 50,000 req/min spreads across 20 instances. Only shared state sees the whole stream.redis · rate-limiting · distributed-systems · concurrency · slahard
- 100,000 requests in. The counter says 97,832 — every run, a different shortfall.go · concurrency · race-condition · atomics · runtimemedium
- 80% of reads hit one post ID. Redis pins at 100% and every key's p99 goes 1ms → 300ms. Then the TTL expires.redis · caching · hot-key · thundering-herd · postgreshard
- 98% cache hit rate → 40%, every 60 seconds, like clockworkredis · caching · postgresmedium