Cloud

SSL certificate management: preventing the outage nobody planned for

Why certificate expiry still causes outages, how to find the certificates nobody documented, and how to automate renewal so the problem stops recurring.

DevOpsArk EngineeringEngineering team, DevOpsArkPublished 15 July 2026 · Updated 19 August 20266 min read
TL;DR

Certificate expiry outages persist because the certificate that expires is almost always one nobody knew about. Discovery has to include probing your actual endpoints, not just reading the systems that were supposed to issue them. Automate renewal wherever an ACME or internal issuer is available, and route the remainder to a team rather than a mailbox.

Short answer

How do you prevent SSL certificate expiry outages?

Prevent expiry outages by discovering every certificate in the estate (including by probing endpoints directly rather than only reading the systems that issue them) automating renewal wherever an ACME or internal certificate authority is available, and routing escalating expiry warnings to the team that owns the affected service weeks in advance.

Why this still happens

Certificate expiry is completely predictable (the date is printed on the certificate), and it still causes outages at organisations with mature operations. The reason is consistent: the certificate that expires is the one nobody had in an inventory.

Typical culprits are a certificate issued by hand during an urgent migration, one on an internal service that was never in scope for the public certificate process, one on a load balancer configured before the current team arrived, and one whose renewal reminder goes to an individual who has left.

Discovery has to include probing

Reading your certificate authority and your cert-manager resources tells you about the certificates that went through those processes. By definition it cannot tell you about the ones that did not, which are exactly the ones that cause the outages.

  • Read Kubernetes Ingress and TLS Secret resources across every cluster.
  • Read cloud load balancer and certificate service configuration in every account.
  • Probe every endpoint in your infrastructure inventory directly and record what it presents.
  • Include internal endpoints. Internal certificates expire too, and internal outages are still outages.

Direct probing is what closes the gap, because it reports what is actually being served rather than what was supposed to be issued.

Automate what you can, escalate the rest properly

Wherever an ACME issuer or an internal certificate authority is available, renewal should be automatic and coordinated with the load balancer or ingress so the new certificate is installed and verified before the old one lapses.

For the remainder, where certificates come from an external authority with a manual process, the fix is routing rather than reminders. Notifications should begin weeks ahead, escalate as the date approaches, and go to the team that owns the service according to a live catalogue, not to an address in a spreadsheet.

Verify after renewal, not before. A renewal that succeeded in the certificate authority but was never installed on the load balancer looks complete in the issuing system and fails in production. Probe the endpoint afterwards and confirm what it actually serves.

Wildcards, chains and the things that break quietly

ProblemSymptomPrevention
Wildcard blast radiusOne renewal affects dozens of services at onceTrack every service depending on a shared certificate before renewing
Incomplete chainWorks in browsers, fails in curl and Java clientsProbe with a strict client, not just a browser
Hostname mismatchCertificate valid but for the wrong nameVerify SANs against the hostnames actually served
Clock skewCertificate rejected as not yet validMonitor time synchronisation on hosts and clients
Pinned certificates in clientsRenewal breaks a mobile app or partner integrationRecord pinning dependencies; pinning to a leaf certificate is fragile by design

Audit configuration, not just expiry

The same endpoint probing that catches expiry gives you protocol and cipher configuration for free. Deprecated TLS versions, weak cipher suites and incomplete chains persist for years precisely because nobody looks unless there is an audit.

Running this continuously rather than annually means a misconfigured endpoint is caught when it is deployed, which is when it is cheapest to fix and before anyone has come to depend on it.

Key takeaways

  • The certificate that expires is almost always the one nobody had in an inventory.
  • Discovery must include probing your endpoints directly, not just reading issuing systems.
  • Automate renewal wherever ACME or an internal CA is available, and verify by probing afterwards.
  • Route the remaining manual renewals to a team from a live catalogue, not to a mailbox.
  • Know the blast radius of a shared wildcard certificate before renewing it.
  • Continuous probing gives you protocol and cipher auditing at no extra cost.

Frequently asked questions

TLSCertificatesInfrastructure

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.