Platform · 2026-10-04 · tls / mtls / certificates / expiry / observability · easy

No deploy, no traffic change — every internal call fails at exactly 00:00:07

01Symptom

Your product runs 38 services behind a public ingress. At 00:00:07 UTC, customers start seeing 502s on nearly every page. Nothing was deployed today, traffic is a normal late-night trickle, and CPU, memory, and database dashboards are all green. The on-call engineer rolls back yesterday's release and restarts the busiest services. Neither changes anything.

02Constraints

  • Services talk to each other over mTLS, and each service mounts the same certificate and key from a shared secret
  • The public ingress certificate is issued and renewed automatically
  • Staging runs the same code and the same architecture
  • Your only expiry monitoring is a dashboard panel for the ingress certificates
  • The engineer who set up the internal CA has since left the company

03Evidence

  • All 38 services start failing in the same second, 00:00:07, with no deploy and no traffic change
  • Every internal error line reads x509: certificate has expired or is not yet valid, and the current time in it is 00:00:09, two seconds past the certificate's end date
  • The ingress certificate, which renews automatically, has 61 days left and is healthy
  • The internal certificate has Not Before: 2025-10-04 and Not After: 2026-10-04 — exactly one year, generated by a one-off command
  • Staging is healthy. Its certificate was generated in March
  • Restarting any pod changes nothing, and the pod mounts the same secret
→

Recent

Previously diagnosed.