Certificate expiry outages persist because the certificate that expires is almost always one nobody knew about. Discovery has to include probing your actual endpoints, not just reading the systems that were supposed to issue them. Automate renewal wherever an ACME or internal issuer is available, and route the remainder to a team rather than a mailbox.
How do you prevent SSL certificate expiry outages?
Prevent expiry outages by discovering every certificate in the estate (including by probing endpoints directly rather than only reading the systems that issue them) automating renewal wherever an ACME or internal certificate authority is available, and routing escalating expiry warnings to the team that owns the affected service weeks in advance.
Why this still happens
Certificate expiry is completely predictable (the date is printed on the certificate), and it still causes outages at organisations with mature operations. The reason is consistent: the certificate that expires is the one nobody had in an inventory.
Typical culprits are a certificate issued by hand during an urgent migration, one on an internal service that was never in scope for the public certificate process, one on a load balancer configured before the current team arrived, and one whose renewal reminder goes to an individual who has left.
Discovery has to include probing
Reading your certificate authority and your cert-manager resources tells you about the certificates that went through those processes. By definition it cannot tell you about the ones that did not, which are exactly the ones that cause the outages.
- Read Kubernetes Ingress and TLS Secret resources across every cluster.
- Read cloud load balancer and certificate service configuration in every account.
- Probe every endpoint in your infrastructure inventory directly and record what it presents.
- Include internal endpoints. Internal certificates expire too, and internal outages are still outages.
Direct probing is what closes the gap, because it reports what is actually being served rather than what was supposed to be issued.
Automate what you can, escalate the rest properly
Wherever an ACME issuer or an internal certificate authority is available, renewal should be automatic and coordinated with the load balancer or ingress so the new certificate is installed and verified before the old one lapses.
For the remainder, where certificates come from an external authority with a manual process, the fix is routing rather than reminders. Notifications should begin weeks ahead, escalate as the date approaches, and go to the team that owns the service according to a live catalogue, not to an address in a spreadsheet.
Wildcards, chains and the things that break quietly
| Problem | Symptom | Prevention |
|---|---|---|
| Wildcard blast radius | One renewal affects dozens of services at once | Track every service depending on a shared certificate before renewing |
| Incomplete chain | Works in browsers, fails in curl and Java clients | Probe with a strict client, not just a browser |
| Hostname mismatch | Certificate valid but for the wrong name | Verify SANs against the hostnames actually served |
| Clock skew | Certificate rejected as not yet valid | Monitor time synchronisation on hosts and clients |
| Pinned certificates in clients | Renewal breaks a mobile app or partner integration | Record pinning dependencies; pinning to a leaf certificate is fragile by design |
Audit configuration, not just expiry
The same endpoint probing that catches expiry gives you protocol and cipher configuration for free. Deprecated TLS versions, weak cipher suites and incomplete chains persist for years precisely because nobody looks unless there is an audit.
Running this continuously rather than annually means a misconfigured endpoint is caught when it is deployed, which is when it is cheapest to fix and before anyone has come to depend on it.
Key takeaways
- The certificate that expires is almost always the one nobody had in an inventory.
- Discovery must include probing your endpoints directly, not just reading issuing systems.
- Automate renewal wherever ACME or an internal CA is available, and verify by probing afterwards.
- Route the remaining manual renewals to a team from a live catalogue, not to a mailbox.
- Know the blast radius of a shared wildcard certificate before renewing it.
- Continuous probing gives you protocol and cipher auditing at no extra cost.
Frequently asked questions
Combine three sources: Kubernetes Ingress and TLS Secret resources, cloud load balancer and certificate service configuration, and direct probing of every endpoint in your infrastructure inventory. The third is what finds the ones issued outside any process.
Most can, using ACME with a public authority or an internal certificate authority. Extended validation certificates and some partner-issued certificates still involve manual steps, and those need escalating notifications routed to a team rather than an individual.
A single private key covers every subdomain, so a compromise is broad, and a single renewal affects every service using it. They reduce administrative overhead at the cost of concentrating both security and operational risk.
Far enough to accommodate the slowest renewal path you have. Thirty days is a common starting point for automated renewals, with escalation at fourteen and seven; manual processes involving an external authority often need longer.
Almost always an incomplete chain. Browsers frequently fetch or cache missing intermediate certificates; strict clients such as curl and Java HTTP clients do not. Serve the full chain rather than relying on the client to complete it.