Skip to main content

Three Strikes and They're Gone: Why Proactive Operations is Now a Customer Retention Strategy

Phil Christianson
Xurrent

IT organizations have historically measured success by how quickly they can respond when something goes wrong. The entire discipline of Incident Management has been optimized around mean time to resolution, first-response SLAs and ticket closure rates. But new research suggests that even though this is a well-executed playbook, it's no longer enough to retain customers.

A recent Xurrent study found that 60% of respondents would seriously consider switching to a competitor by the third service outage. One in five would start shopping around after the very first failure. The takeaway is that customers are keeping score, and the count begins early. Companies need to be thinking about prevention as much as quick resolutions.

The Loyalty Illusion

Our study's findings are the latest in a wider backdrop of eroding customer patience. PwC's 2025 Customer Experience Survey found that 52% of consumers have stopped using or buying from a brand because of a bad experience with its products or services. For digital businesses, the service itself is the experience.

Only 6% of Xurrent survey respondents said they canceled or switched immediately after their last outage. Half just gave up in the moment and tried again later. It may be tempting to read that as tolerance, but it's really attrition in slow motion.

With each problem recurrence, trust weakens, and by the third strike, a clear majority is ready to leave. A third of respondents said repeat failures break down trust fastest, and 17% said nothing damages confidence more than being told an issue is fixed only to watch it resurface. When an outage hits, restoring service and resolving the underlying problem are two different things.

3 Tips to Avoid Three Strikes

Moving from reactive Incident Management to proactively preventing problems requires a structural shift that relies on three capabilities:

1. Connecting incident data across silos

Recurring outages keep happening because the signals that would reveal them are scattered across infrastructure monitoring, service desk queues and customer teams.

A modern IT Service Management (ITSM) platform can unify these workflows and make patterns visible. For example, the API that fails under load every month or the configuration drift that triggers the same cascade. Root cause analysis becomes standard practice when incident records, change history and monitoring data live in one system.

2. Automating detection and response before customers notice

Most enterprises have numerous monitoring tools, contributing to a fragmented environment that generates large volumes of redundant or low-priority alerts. Teams develop alert fatigue, critical signals get buried, and customers are frequently the first to report the issue.

Modern Incident Management platforms can correlate alerts across sources and reduce redundant or distracting alarms. Reducing distractions can help your teams to respond to more of what matters and head off lower-priority alarms before customers notice.

3. Institutionalizing problem management alongside incident closure

Every repeat incident should trigger a root cause investigation with an owner and a deadline. In many scenarios, teams restore service, close the ticket and move on because leadership rewards shipping new features over stabilizing existing systems.

Linking incidents to permanent solutions saves hours of firefighting, but teams must be intentional about making long-term fixes a priority.

Communication Is Part of the Architecture

While prevention won't eliminate every outage, the way organizations communicate during the ones that occur is its own retention lever. There is an opportunity to turn an outage from a breach of trust into a demonstration of competence. A key element is having automated, proactive status communication in place.

There's an important nuance for leaders designing AI Service Desk strategies: Consumers welcome the speed AI enables, but our study found that 76% still prefer human help when they're frustrated. The key is to use automation for velocity and transparency, while routing high-stakes interactions to people.

Reacting quickly to service issues is no longer enough. Treat every repeat incident as a customer-retention emergency. When customers are counting strikes, don't stop at restoring service. Elevate root-cause problem management to a mandatory practice.

Phil Christianson is Chief Product Officer at Xurrent

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

Three Strikes and They're Gone: Why Proactive Operations is Now a Customer Retention Strategy

Phil Christianson
Xurrent

IT organizations have historically measured success by how quickly they can respond when something goes wrong. The entire discipline of Incident Management has been optimized around mean time to resolution, first-response SLAs and ticket closure rates. But new research suggests that even though this is a well-executed playbook, it's no longer enough to retain customers.

A recent Xurrent study found that 60% of respondents would seriously consider switching to a competitor by the third service outage. One in five would start shopping around after the very first failure. The takeaway is that customers are keeping score, and the count begins early. Companies need to be thinking about prevention as much as quick resolutions.

The Loyalty Illusion

Our study's findings are the latest in a wider backdrop of eroding customer patience. PwC's 2025 Customer Experience Survey found that 52% of consumers have stopped using or buying from a brand because of a bad experience with its products or services. For digital businesses, the service itself is the experience.

Only 6% of Xurrent survey respondents said they canceled or switched immediately after their last outage. Half just gave up in the moment and tried again later. It may be tempting to read that as tolerance, but it's really attrition in slow motion.

With each problem recurrence, trust weakens, and by the third strike, a clear majority is ready to leave. A third of respondents said repeat failures break down trust fastest, and 17% said nothing damages confidence more than being told an issue is fixed only to watch it resurface. When an outage hits, restoring service and resolving the underlying problem are two different things.

3 Tips to Avoid Three Strikes

Moving from reactive Incident Management to proactively preventing problems requires a structural shift that relies on three capabilities:

1. Connecting incident data across silos

Recurring outages keep happening because the signals that would reveal them are scattered across infrastructure monitoring, service desk queues and customer teams.

A modern IT Service Management (ITSM) platform can unify these workflows and make patterns visible. For example, the API that fails under load every month or the configuration drift that triggers the same cascade. Root cause analysis becomes standard practice when incident records, change history and monitoring data live in one system.

2. Automating detection and response before customers notice

Most enterprises have numerous monitoring tools, contributing to a fragmented environment that generates large volumes of redundant or low-priority alerts. Teams develop alert fatigue, critical signals get buried, and customers are frequently the first to report the issue.

Modern Incident Management platforms can correlate alerts across sources and reduce redundant or distracting alarms. Reducing distractions can help your teams to respond to more of what matters and head off lower-priority alarms before customers notice.

3. Institutionalizing problem management alongside incident closure

Every repeat incident should trigger a root cause investigation with an owner and a deadline. In many scenarios, teams restore service, close the ticket and move on because leadership rewards shipping new features over stabilizing existing systems.

Linking incidents to permanent solutions saves hours of firefighting, but teams must be intentional about making long-term fixes a priority.

Communication Is Part of the Architecture

While prevention won't eliminate every outage, the way organizations communicate during the ones that occur is its own retention lever. There is an opportunity to turn an outage from a breach of trust into a demonstration of competence. A key element is having automated, proactive status communication in place.

There's an important nuance for leaders designing AI Service Desk strategies: Consumers welcome the speed AI enables, but our study found that 76% still prefer human help when they're frustrated. The key is to use automation for velocity and transparency, while routing high-stakes interactions to people.

Reacting quickly to service issues is no longer enough. Treat every repeat incident as a customer-retention emergency. When customers are counting strikes, don't stop at restoring service. Elevate root-cause problem management to a mandatory practice.

Phil Christianson is Chief Product Officer at Xurrent

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...