Platformtlsmtlscertificatesexpiryobservabilityeasy

No deploy, no traffic change — every internal call fails at exactly 00:00:07

01Symptom

Your product runs 38 services behind a public ingress. At 00:00:07 UTC, customers start seeing 502s on nearly every page. Nothing was deployed today, traffic is a normal late-night trickle, and CPU, memory, and database dashboards are all green. The on-call engineer rolls back yesterday's release and restarts the busiest services. Neither changes anything.

02Constraints

  • Services talk to each other over mTLS, and each service mounts the same certificate and key from a shared secret
  • The public ingress certificate is issued and renewed automatically
  • Staging runs the same code and the same architecture
  • Your only expiry monitoring is a dashboard panel for the ingress certificates
  • The engineer who set up the internal CA has since left the company

03Evidence

  • All 38 services start failing in the same second, 00:00:07, with no deploy and no traffic change
  • Every internal error line reads x509: certificate has expired or is not yet valid, and the current time in it is 00:00:09, two seconds past the certificate's end date
  • The ingress certificate, which renews automatically, has 61 days left and is healthy
  • The internal certificate has Not Before: 2025-10-04 and Not After: 2026-10-04 — exactly one year, generated by a one-off command
  • Staging is healthy. Its certificate was generated in March
  • Restarting any pod changes nothing, and the pod mounts the same secret

→The question

Why did everything fail at once, why did a rollback and restarts not help, and what would you change so it can't happen again?

04Your prediction

01What's the most likely cause?
02Why didn't rollbacks and restarts help?
03What is the best long-term fix?
0 / 600 chars