Skip to main content

TLS Certificate Expiration Is Becoming an Observability Problem

The one outage you can predict is about to multiply
Meenakshisundaram Ramakrishna Sahadevan
ManageEngine

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around.

Under the CA/Browser Forum's 2025 decision, Ballot SC-081v3, the maximum validity period for a publicly trusted TLS certificate is dropping on a fixed schedule: from 398 days to 200 on March 15, 2026, to 100 on March 15, 2027, and to 47 on March 15, 2029. The limits become part of the requirements for publicly trusted certificate authorities (CAs), with browser root programs enforcing compliance as a condition of trust. Organizations that need publicly trusted certificates cannot retain the old validity periods. The first maximum length certificates issued under the 200-day rule will begin expiring around early October 2026.

A certificate that renewed once a year in 2025 will need to be rotated roughly eight times a year in 2029. Across 200 certificates, that turns roughly 200 annual renewals into approximately 1,600 certificate replacements, or around six every business day. This could vary based on the organization scale and size, and you can check your impact here. For a monitoring team, that shift changes what the monitoring is for.

Why Monitoring Alone Won't Keep Up

If you run an observability practice, some form of a certificate expiration check is probably already in place, so visibility is rarely the whole problem. The harder issue is what happens after the alert. Monitoring can flag that a certificate is due or that a renewal failed, but the certificate still has to be renewed, deployed, and activated, and that work is what the schedule multiplies.

When a certificate lasted a year, that work was easy to absorb. It came up rarely, one certificate at a time, with weeks of slack to handle each one. Shorter lifespans make the same work recur far more often: several times a year across the publicly trusted certificate estate. A process that clears those renewals by hand then falls behind.

The mandate applies specifically to publicly trusted TLS certificates. Private PKI is not directly covered by Ballot SC-081v3. However, the same operational weaknesses appear in internal environments, particularly when organizations are already adopting shorter-lived private certificates as a security practice.

What ACME Solved and Where It Stops

There's already a working template for this, proven for years on web servers that use ACME. On many web servers, an ACME client can request a new certificate, install it, and reload the service to put it in use. That works because the whole chain is automated. A certificate that's issued but never installed protects nothing, so automating renewal without deployment solves very little.

Beyond common web server integrations, however, certificate automation often becomes uneven. Internal services, application key stores, appliances, and the load balancers in front of them can often get a certificate issued automatically over ACME or another CA integration. Yet there's rarely a built-in way to install it, update the relevant binding, and reload the service automatically the way there is on a web server, so that step stays manual or falls to a script someone maintains. With annual renewals, that was an occasional chore, but with eight renewals a year across the estate, it will become the part most likely to fall behind or fail unnoticed.

Automate the Chain and Let Monitoring Measure It

The way forward is to give the rest of the estate what the web tier already has. That means the same chain is automated from end to end, from discovery through renewal and deployment to the reload that puts the certificate in use. It should also include verification that the endpoint is actually serving the renewed certificate. That's a certificate life cycle management job.

A platform built for that job runs the chain: It discovers certificates by scanning for them, renews them across public and private CAs, and deploys the renewed certificates to load balancers, IIS bindings, Linux and Windows hosts, and cloud key stores. Once the required integrations and workflows are configured, a renewed certificate can reach the endpoint without a manual handoff in the middle.

Monitoring doesn't go away in that model, but its job changes. It now confirms the automation is keeping up: that renewals are completing across the estate and that nothing has slipped outside what the automation covers.

That means measuring the automation coverage, renewal success, deployment lag, reload failures, and number of certificates approaching a defined safety threshold. Additionally, monitoring also focuses on sending alerts for critical certificate renewals that can't be automated and require human intervention. That turns certificate management from a set of expiration alerts into an observable deployment pipeline.

Because the schedule is public, that load is countable in advance. Automate the renewal chain, monitor the outcomes, and treat certificate rotation as a continuously running operational pipeline. Then the 2027 and 2029 reductions will become planned capacity changes that the pipeline will handle on schedule, and missed renewal outages will drop from a recurring operational risk to a rare exception.

ManageEngine Key Manager Plus is one platform built to run this chain end to end, from discovery through renewal, deployment, and verification across public and private CAs. See how Key Manager Plus handles it.

Meenakshisundaram Ramakrishna Sahadevan is Product Expert - ManageEngine Key Manager Plus

The Latest

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

TLS Certificate Expiration Is Becoming an Observability Problem

The one outage you can predict is about to multiply
Meenakshisundaram Ramakrishna Sahadevan
ManageEngine

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around.

Under the CA/Browser Forum's 2025 decision, Ballot SC-081v3, the maximum validity period for a publicly trusted TLS certificate is dropping on a fixed schedule: from 398 days to 200 on March 15, 2026, to 100 on March 15, 2027, and to 47 on March 15, 2029. The limits become part of the requirements for publicly trusted certificate authorities (CAs), with browser root programs enforcing compliance as a condition of trust. Organizations that need publicly trusted certificates cannot retain the old validity periods. The first maximum length certificates issued under the 200-day rule will begin expiring around early October 2026.

A certificate that renewed once a year in 2025 will need to be rotated roughly eight times a year in 2029. Across 200 certificates, that turns roughly 200 annual renewals into approximately 1,600 certificate replacements, or around six every business day. This could vary based on the organization scale and size, and you can check your impact here. For a monitoring team, that shift changes what the monitoring is for.

Why Monitoring Alone Won't Keep Up

If you run an observability practice, some form of a certificate expiration check is probably already in place, so visibility is rarely the whole problem. The harder issue is what happens after the alert. Monitoring can flag that a certificate is due or that a renewal failed, but the certificate still has to be renewed, deployed, and activated, and that work is what the schedule multiplies.

When a certificate lasted a year, that work was easy to absorb. It came up rarely, one certificate at a time, with weeks of slack to handle each one. Shorter lifespans make the same work recur far more often: several times a year across the publicly trusted certificate estate. A process that clears those renewals by hand then falls behind.

The mandate applies specifically to publicly trusted TLS certificates. Private PKI is not directly covered by Ballot SC-081v3. However, the same operational weaknesses appear in internal environments, particularly when organizations are already adopting shorter-lived private certificates as a security practice.

What ACME Solved and Where It Stops

There's already a working template for this, proven for years on web servers that use ACME. On many web servers, an ACME client can request a new certificate, install it, and reload the service to put it in use. That works because the whole chain is automated. A certificate that's issued but never installed protects nothing, so automating renewal without deployment solves very little.

Beyond common web server integrations, however, certificate automation often becomes uneven. Internal services, application key stores, appliances, and the load balancers in front of them can often get a certificate issued automatically over ACME or another CA integration. Yet there's rarely a built-in way to install it, update the relevant binding, and reload the service automatically the way there is on a web server, so that step stays manual or falls to a script someone maintains. With annual renewals, that was an occasional chore, but with eight renewals a year across the estate, it will become the part most likely to fall behind or fail unnoticed.

Automate the Chain and Let Monitoring Measure It

The way forward is to give the rest of the estate what the web tier already has. That means the same chain is automated from end to end, from discovery through renewal and deployment to the reload that puts the certificate in use. It should also include verification that the endpoint is actually serving the renewed certificate. That's a certificate life cycle management job.

A platform built for that job runs the chain: It discovers certificates by scanning for them, renews them across public and private CAs, and deploys the renewed certificates to load balancers, IIS bindings, Linux and Windows hosts, and cloud key stores. Once the required integrations and workflows are configured, a renewed certificate can reach the endpoint without a manual handoff in the middle.

Monitoring doesn't go away in that model, but its job changes. It now confirms the automation is keeping up: that renewals are completing across the estate and that nothing has slipped outside what the automation covers.

That means measuring the automation coverage, renewal success, deployment lag, reload failures, and number of certificates approaching a defined safety threshold. Additionally, monitoring also focuses on sending alerts for critical certificate renewals that can't be automated and require human intervention. That turns certificate management from a set of expiration alerts into an observable deployment pipeline.

Because the schedule is public, that load is countable in advance. Automate the renewal chain, monitor the outcomes, and treat certificate rotation as a continuously running operational pipeline. Then the 2027 and 2029 reductions will become planned capacity changes that the pipeline will handle on schedule, and missed renewal outages will drop from a recurring operational risk to a rare exception.

ManageEngine Key Manager Plus is one platform built to run this chain end to end, from discovery through renewal, deployment, and verification across public and private CAs. See how Key Manager Plus handles it.

Meenakshisundaram Ramakrishna Sahadevan is Product Expert - ManageEngine Key Manager Plus

The Latest

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...