Skip to main content

Downtime in a Downturn Could Mean Customer Churn

Phil Tee

The last year has been challenging for Tech. Everyone in the industry, from IT and DevOps leaders to field technicians, grapples with recessionary pressures like inflation and rising interest rates in their personal life. And thanks to a never-ending barrage of stories about high-profile layoffs, they are also keenly aware that Tech is experiencing an aggravated downturn.

For many IT leaders, the well-reasoned response to these stories is to locate cost-cutting opportunities in their organization. Ultimately, an economic softening will encourage managers to audit their ITOps tech stack. This is a reasonable first step since the average engineering team manages more than 16 monitoring tools alone.

However, IT leaders must ensure their tool consolidation process is strategic. After all, many solutions are mission-critical — especially during an economic downturn, when hitting key metrics like revenue and availability becomes necessary for business continuity. The best rule of thumb is to consider which tools provide actionable insights and ROI without wasting technicians' time. This benchmark for success allows leaders to cut ties with superfluous solutions and double down on those that map back to critical KPIs like system performance and operational efficiency.

An array of tools purport to maintain availability — the trick is sorting through the noise to find the right one. Let us discuss why availability is so important and then unpack the ROI of deploying Artificial Intelligence for IT Operations (AIOps) during an economic downturn.

Maintaining Availability Has Become More Important Than Ever

Over half the world's GDP (60%) is digitized as of 2019. That means organizations with improper digital infrastructure will repeatedly lose out on revenue opportunities. And in a downturn, revenue-generating opportunities are not simply competitive differentiators — they are the difference between sinking and swimming.

True, revenue is a guiding KPI regardless of macroeconomic conditions. But the recent economic softening has refocused efforts from a "growth at all costs" mindset to a "generate revenue efficiently" perspective. Now, organizations are buckling down to the basics — and providing consumers with a reliable online destination to interact with a brand and its products is downright critical.

That is where availability comes in. Availability is the glue that binds all digital interfaces together. Defined by maximum system performance and uptime, availability is achieved through rigorous behind-the-scenes engineering work. AIOps are an essential part of this equation because these tools reduce an organization's mean time to detect (MTTD) and mean time to recover (MTTR) by simplifying, collating and escalating data errors before they create downtime.

Let us use an example to illustrate the importance of reduced MTTX. If a top broadcast network experiences an outage during a major sporting event, they stand to lose millions of viewers — and, as a result, millions of dollars in ad revenue. But if that broadcast network has deployed AIOps, they can expediently identify the nature of the error (low MTTD) and resolve it within 30 seconds (low MTTR). Compare that resolution to a network without AIOps, which may experience an outage measured in minutes not seconds. This extended outage could immediately cost the network millions of dollars, not to mention millions more in lost customer loyalty and damaged brand reputation.

In an economically fraught environment, the losses associated with such an outage are more likely to become exacerbated. Hence, maintaining availability is not a luxury but a necessity.

AIOps Goes Beyond Simple Event Management

Availability, uptime and system performance are leading DevOps concerns. Consequently, many vendors advertise that their monitoring tool can improve these vectors in isolation, but this is not so. Monitoring tools are foundational for a tech stack, but they are fundamentally incapable of identifying and escalating data errors across all telemetry points. Only AIOps solutions that ingest disparate data from all devices, networks and tools will provide a complete overhead of the incident lifecycle. Furthermore, top AIOps solutions rely on machine learning (ML) to grow with their system and fill contextual gaps.

AIOps tools are superior to point solutions because their AI-based algorithms can parse thousands of incidents to determine which are relevant. Consider that any data state change creates an incident, yet data is inherently ephemeral, and only a select few changes indicate an actual system error. AIOps reduce the time technicians spend combing over data by eradicating non-harmful events and escalating the rest to the appropriate party — all with minimal supervision.

And when technicians need to step in, AIOps-based systems provide them with context-rich event tickets that explain the data issue in detail. This provides ample time for technicians to address the problem and return to revenue-generating responsibilities like improving the user experience (UX) and driving down technical debt. During an economic softening, the ROI here is even more apparent, especially given the extended tech talent crunch that continues to leave IT and DevOps teams struggling to fill labor-related gaps.

Of course, budget cuts and hiring freezes are only natural responses to concerns about fluctuations in economic stability. But IT and DevOps leaders should carefully consider the ROI behind each solution they cut — and adopt — during an economic softening.

For example, does a solution of interest provide excess data to interpret, or does it also understand and act on that data?

Does a solution reduce monotonous labor needs?

And, most importantly, does it provide revenue-generating opportunities like increased uptime and availability?

This line of questioning will ultimately demonstrate that certain tools are unnecessary during an economic downturn while others are more critical than ever. But, in general, leaders should treat availability as their guiding light when auditing their tech stack. Doing so will leave their organization better positioned to excel in the months ahead.

The Latest

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Downtime in a Downturn Could Mean Customer Churn

Phil Tee

The last year has been challenging for Tech. Everyone in the industry, from IT and DevOps leaders to field technicians, grapples with recessionary pressures like inflation and rising interest rates in their personal life. And thanks to a never-ending barrage of stories about high-profile layoffs, they are also keenly aware that Tech is experiencing an aggravated downturn.

For many IT leaders, the well-reasoned response to these stories is to locate cost-cutting opportunities in their organization. Ultimately, an economic softening will encourage managers to audit their ITOps tech stack. This is a reasonable first step since the average engineering team manages more than 16 monitoring tools alone.

However, IT leaders must ensure their tool consolidation process is strategic. After all, many solutions are mission-critical — especially during an economic downturn, when hitting key metrics like revenue and availability becomes necessary for business continuity. The best rule of thumb is to consider which tools provide actionable insights and ROI without wasting technicians' time. This benchmark for success allows leaders to cut ties with superfluous solutions and double down on those that map back to critical KPIs like system performance and operational efficiency.

An array of tools purport to maintain availability — the trick is sorting through the noise to find the right one. Let us discuss why availability is so important and then unpack the ROI of deploying Artificial Intelligence for IT Operations (AIOps) during an economic downturn.

Maintaining Availability Has Become More Important Than Ever

Over half the world's GDP (60%) is digitized as of 2019. That means organizations with improper digital infrastructure will repeatedly lose out on revenue opportunities. And in a downturn, revenue-generating opportunities are not simply competitive differentiators — they are the difference between sinking and swimming.

True, revenue is a guiding KPI regardless of macroeconomic conditions. But the recent economic softening has refocused efforts from a "growth at all costs" mindset to a "generate revenue efficiently" perspective. Now, organizations are buckling down to the basics — and providing consumers with a reliable online destination to interact with a brand and its products is downright critical.

That is where availability comes in. Availability is the glue that binds all digital interfaces together. Defined by maximum system performance and uptime, availability is achieved through rigorous behind-the-scenes engineering work. AIOps are an essential part of this equation because these tools reduce an organization's mean time to detect (MTTD) and mean time to recover (MTTR) by simplifying, collating and escalating data errors before they create downtime.

Let us use an example to illustrate the importance of reduced MTTX. If a top broadcast network experiences an outage during a major sporting event, they stand to lose millions of viewers — and, as a result, millions of dollars in ad revenue. But if that broadcast network has deployed AIOps, they can expediently identify the nature of the error (low MTTD) and resolve it within 30 seconds (low MTTR). Compare that resolution to a network without AIOps, which may experience an outage measured in minutes not seconds. This extended outage could immediately cost the network millions of dollars, not to mention millions more in lost customer loyalty and damaged brand reputation.

In an economically fraught environment, the losses associated with such an outage are more likely to become exacerbated. Hence, maintaining availability is not a luxury but a necessity.

AIOps Goes Beyond Simple Event Management

Availability, uptime and system performance are leading DevOps concerns. Consequently, many vendors advertise that their monitoring tool can improve these vectors in isolation, but this is not so. Monitoring tools are foundational for a tech stack, but they are fundamentally incapable of identifying and escalating data errors across all telemetry points. Only AIOps solutions that ingest disparate data from all devices, networks and tools will provide a complete overhead of the incident lifecycle. Furthermore, top AIOps solutions rely on machine learning (ML) to grow with their system and fill contextual gaps.

AIOps tools are superior to point solutions because their AI-based algorithms can parse thousands of incidents to determine which are relevant. Consider that any data state change creates an incident, yet data is inherently ephemeral, and only a select few changes indicate an actual system error. AIOps reduce the time technicians spend combing over data by eradicating non-harmful events and escalating the rest to the appropriate party — all with minimal supervision.

And when technicians need to step in, AIOps-based systems provide them with context-rich event tickets that explain the data issue in detail. This provides ample time for technicians to address the problem and return to revenue-generating responsibilities like improving the user experience (UX) and driving down technical debt. During an economic softening, the ROI here is even more apparent, especially given the extended tech talent crunch that continues to leave IT and DevOps teams struggling to fill labor-related gaps.

Of course, budget cuts and hiring freezes are only natural responses to concerns about fluctuations in economic stability. But IT and DevOps leaders should carefully consider the ROI behind each solution they cut — and adopt — during an economic softening.

For example, does a solution of interest provide excess data to interpret, or does it also understand and act on that data?

Does a solution reduce monotonous labor needs?

And, most importantly, does it provide revenue-generating opportunities like increased uptime and availability?

This line of questioning will ultimately demonstrate that certain tools are unnecessary during an economic downturn while others are more critical than ever. But, in general, leaders should treat availability as their guiding light when auditing their tech stack. Doing so will leave their organization better positioned to excel in the months ahead.

The Latest

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...