No deploy, no traffic change — every internal call fails at exactly 00:00:07
01Symptom
Your product runs 38 services behind a public ingress. At 00:00:07 UTC, customers start seeing 502s on nearly every page. Nothing was deployed today, traffic is a normal late-night trickle, and CPU, memory, and database dashboards are all green. The on-call engineer rolls back yesterday's release and restarts the busiest services. Neither changes anything.
02Constraints
- Services talk to each other over mTLS, and each service mounts the same certificate and key from a shared secret
- The public ingress certificate is issued and renewed automatically
- Staging runs the same code and the same architecture
- Your only expiry monitoring is a dashboard panel for the ingress certificates
- The engineer who set up the internal CA has since left the company
03Evidence
- All 38 services start failing in the same second, 00:00:07, with no deploy and no traffic change
- Every internal error line reads x509: certificate has expired or is not yet valid, and the current time in it is 00:00:09, two seconds past the certificate's end date
- The ingress certificate, which renews automatically, has 61 days left and is healthy
- The internal certificate has Not Before: 2025-10-04 and Not After: 2026-10-04 — exactly one year, generated by a one-off command
- Staging is healthy. Its certificate was generated in March
- Restarting any pod changes nothing, and the pod mounts the same secret
→The question
Why did everything fail at once, why did a rollback and restarts not help, and what would you change so it can't happen again?
04Your prediction
0 / 600 chars