Skip to main content

Internet Outages Cost Companies Upwards of $10 Million per Month

97% of companies assert a reliable, resilient Internet Stack is of utmost importance to their business success

Almost all (97%) of respondents state that a reliable, resilient Internet Stack is of the utmost importance to their business success, according to Catchpoint's inaugural Internet Resilience Report.

CEOs and company boards are prioritizing Internet resilience as a critical factor in maintaining and enhancing business operations. The Internet's complexity and its vital role in connecting businesses and customers necessitate a holistic resilience strategy. Companies rely on the Internet for connectivity and digital experiences, and outages can add up to many millions of dollars, making the need for Internet Resilience a board-level discussion. In 2024, any business, whether involved in eCommerce or operating with a distributed or remote workforce, must ensure strong Internet resilience.

Based on insights gathered from 310 digital business leaders, key findings from the new report revealed: 

■ 78% identify improved customer experience as the primary driver for resilience programs. 

■ 77% highlight the critical role of third-party technology providers in their Internet resilience strategies. 

■ 43% estimate a total economic impact or loss of more than $1 million monthly due to Internet outages or degradations and some Internet outages cost some companies upward of $10 million a month. 

■ 40% cite talent and skillset as a major barrier to implementing IPM. 

"Our findings highlight the criticality of Internet Resilience in our always-on, digital-first world," said Mehdi Daoudi, CEO of Catchpoint. "Businesses must expand their observability boundaries beyond Application Performance Monitoring to include Internet Performance Monitoring. This is the best way to proactively address issues impacting operations as it is the only way to get real-time insights and ensure superior performance that can reduce risks to both reputation and revenue."
 

Catchpoint's report provides actionable insights for businesses to enhance their Internet Resilience, and their top suggestions include:

Think about Resilience from the Top Down

Leaders must incorporate Internet resilience into their strategic plans and daily operations to instill an effective, long-term culture of resilience. This includes the implementation of a new role, that of chief reliability or chief resilience officer to the C-suites, something Fortune 2000 companies are increasingly doing.

Align Internally

Break down organizational silos and foster collaboration between IT and business teams to ensure alignment and a focused, constructive response to outages.

Implement Internet Performance Monitoring

Utilize objective, independent IPM data to settle disputes and remove emotion from the equation, while setting Service Level Objectives (SLOs) to guide actions during incidents. The longer mean time to repair (MTTR) takes, the more the risk of payouts and other ramifications increases, including the threat of decreased customer loyalty — be sure to equip your ITOps team with the tools they need to continue to succeed. As the Internet continues to be an indispensable resource for businesses, ensuring its resilience is paramount. 

"Internet resilience should be a critical part of your overall Disaster Recovery / Business Continuity program," said Pete Charlton, IT Vice President, TMNAS. "Ultimately the CIO/CTO is accountable for the organization's digital resilience, but these are not just technology problems. Resilience and business continuity are in fact overall organizational issues that need to be discussed at the organization's highest levels and tested as frequently as possible. Obviously, you cannot simulate every possible outage, but if the past few years have taught us anything, it is that you need to plan for the unexpected."

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...

Internet Outages Cost Companies Upwards of $10 Million per Month

97% of companies assert a reliable, resilient Internet Stack is of utmost importance to their business success

Almost all (97%) of respondents state that a reliable, resilient Internet Stack is of the utmost importance to their business success, according to Catchpoint's inaugural Internet Resilience Report.

CEOs and company boards are prioritizing Internet resilience as a critical factor in maintaining and enhancing business operations. The Internet's complexity and its vital role in connecting businesses and customers necessitate a holistic resilience strategy. Companies rely on the Internet for connectivity and digital experiences, and outages can add up to many millions of dollars, making the need for Internet Resilience a board-level discussion. In 2024, any business, whether involved in eCommerce or operating with a distributed or remote workforce, must ensure strong Internet resilience.

Based on insights gathered from 310 digital business leaders, key findings from the new report revealed: 

■ 78% identify improved customer experience as the primary driver for resilience programs. 

■ 77% highlight the critical role of third-party technology providers in their Internet resilience strategies. 

■ 43% estimate a total economic impact or loss of more than $1 million monthly due to Internet outages or degradations and some Internet outages cost some companies upward of $10 million a month. 

■ 40% cite talent and skillset as a major barrier to implementing IPM. 

"Our findings highlight the criticality of Internet Resilience in our always-on, digital-first world," said Mehdi Daoudi, CEO of Catchpoint. "Businesses must expand their observability boundaries beyond Application Performance Monitoring to include Internet Performance Monitoring. This is the best way to proactively address issues impacting operations as it is the only way to get real-time insights and ensure superior performance that can reduce risks to both reputation and revenue."
 

Catchpoint's report provides actionable insights for businesses to enhance their Internet Resilience, and their top suggestions include:

Think about Resilience from the Top Down

Leaders must incorporate Internet resilience into their strategic plans and daily operations to instill an effective, long-term culture of resilience. This includes the implementation of a new role, that of chief reliability or chief resilience officer to the C-suites, something Fortune 2000 companies are increasingly doing.

Align Internally

Break down organizational silos and foster collaboration between IT and business teams to ensure alignment and a focused, constructive response to outages.

Implement Internet Performance Monitoring

Utilize objective, independent IPM data to settle disputes and remove emotion from the equation, while setting Service Level Objectives (SLOs) to guide actions during incidents. The longer mean time to repair (MTTR) takes, the more the risk of payouts and other ramifications increases, including the threat of decreased customer loyalty — be sure to equip your ITOps team with the tools they need to continue to succeed. As the Internet continues to be an indispensable resource for businesses, ensuring its resilience is paramount. 

"Internet resilience should be a critical part of your overall Disaster Recovery / Business Continuity program," said Pete Charlton, IT Vice President, TMNAS. "Ultimately the CIO/CTO is accountable for the organization's digital resilience, but these are not just technology problems. Resilience and business continuity are in fact overall organizational issues that need to be discussed at the organization's highest levels and tested as frequently as possible. Obviously, you cannot simulate every possible outage, but if the past few years have taught us anything, it is that you need to plan for the unexpected."

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...