Preventing Outages in 2023: What We Can Learn from Recent Failures

February 09, 2023

Learn more about Catchpoint

"What the recent failures from Internet giants demonstrate is that the question of the next outage is not if, but when," says Dritan Suljoti, Chief Product and Technology Officer of Catchpoint, referencing the company's new white paper, Preventing Outages in 2023: What We Learned from Recent Failures. "Moreover, the downstream effect of major outages to essential Internet infrastructure, such as cloud platforms, CDNs or DNS providers, means that no company is immune, no matter how well prepared they think they are. The white paper demonstrates why it's so important for all of us to be proactive to reduce Mean Time to Repair (MTTR) when the next outage occurs."

Key lessons from the past

■ Develop an Internet Performance Monitoring strategy that allows you to monitor precisely what customers, workforce, and other users expect and build an Experience Score.

■ Monitor not only what is under your direct control, map your Internet stack to ensure you are monitoring every component of the Internet Stack relied on to deliver your content (including DNS, CDN, ISP, BGP, TCP configuration, SSL, and other cloud services, etc.). ■ Automate intelligently – design and test automation to ensure there are no bugs hiding in the code.

■ Be prepared to take fast action to remediate outages as they occur, for example, switching to a backup solution or dropping the third-party causing the issue. Develop runbooks and practice recovery.

■ Whenever change is scheduled, ensure your team is ready for any outages that may occur (intentionally or not) with a crisis call plan that includes a communication plan and templates, a plan to mitigate failures from third-parties, and a best practices monitoring and observability plan. ‍

"Given the impact of serious outages to the bottom line, not to mention the long-tail impact to brand and reputation, amidst a landscape of increased Internet reliance alongside ever-growing Internet fragility and greater and great complexity, the need for community learnings from past failures to be shared and practical advice disseminated around stemming future major incidents and ensuring Internet Resilience is imperative," says Gerardo Dada, CMO at Catchpoint. "We believe this white paper offers an invaluable deep dive into recent outages past and key lessons learned that all of us can learn from to prevent (or mitigate the consequences of) the next major outage."

Hot Topics

Automation

The Latest

MEAN TIME TO INSIGHT Podcast - Episode 24: Network Observability Tool Sprawl

May 29, 2026

In MEAN TIME TO INSIGHT Episode 24, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network observability tool sprawl ...

Capacity Isn't a Guess: Observability-Driven Sizing for On-Prem Databases

May 28, 2026

In cloud-native systems, scaling is often as simple as moving a slider. For on-premise databases, the stakes are different. Over-provisioning hardware is expensive. Under-provisioning leads to performance bottlenecks that are difficult to fix once the equipment is in the rack ...

5 Security Principles Every Entrepreneur Should Apply to Leadership

May 27, 2026

When most people think about cybersecurity, they picture firewalls, encryption, and access controls — technical tools designed to protect systems and data. But beneath the technology lies a deeper set of principles about trust, decision-making, and resilience ... The best leaders don't eliminate risk. They manage it intelligently. And in many ways, cybersecurity offers a surprisingly useful playbook for doing exactly that ...

Signs It May Be Time to Reassess Your IT Infrastructure Strategy

May 26, 2026

Many organizations assumed their infrastructure strategy was settled. It had been implemented, optimized and built into long-term plans. Recent changes in technology and vendor consolidation are forcing a second look. Cloud outages and licensing changes have exposed how much dependency exists on a small number of platforms. As a result, organizations are reevaluating whether those decisions still hold up under current conditions ...

Enterprise Edge AI Reaches Inflection Point

May 22, 2026

Edge AI is strategically embedded in core IT and infrastructure spending across industries, according to the 2026 Edge AI Survey from ZEDEDA. The research shows that 83% of C-suite and IT executive respondents say edge AI is important to their core business strategy ...

AI Is Hitting Operational Limits

May 21, 2026

As AI adoption accelerates, operational complexity — not model intelligence — is becoming the primary barrier to reliable AI at scale, according to the State of AI Engineering 2026 from Datadog ... The report highlights a compounding complexity challenge as AI systems scale ... Around 5% of AI model requests fail in production, with nearly 60% of those failures caused by capacity limits ...

Alert Fatigue Is No Longer a Morale Problem, It's a Reliability Risk and a System Failure

May 20, 2026

For years, production operations teams have treated alert fatigue as a quality-of-life problem: something that makes on-call rotations miserable but isn't considered a direct contributor to outages. That framing doesn't capture how these systems fail, and we now have data to show why. More importantly, it's now clear alert fatigue is a symptom of a deeper issue: production systems have outgrown the current operational approaches ...

Most Enterprises Think They Can Switch AI Vendors in a Month ... Most Who've Tried Couldn't

May 19, 2026

I was on a customer call last fall when an enterprise architect said something I haven't been able to shake. Her team had just spent four months trying to swap one AI vendor for another. The original plan said three weeks. "We didn't switch vendors," she told me. "We rebuilt half our integrations and discovered what we'd actually been depending on." Most enterprise leaders don't expect that to be the experience ...

Your Observability Stack Has a Telemetry Pipeline Problem

May 18, 2026

Ask any senior SRE or platform engineer what keeps them up at night, and the answer probably isn't the monitoring tool — it's the data feeding it. The proliferation of APM, observability, and AIOps platforms has created a telemetry sprawl problem that most teams manage reactively rather than architect proactively. Metrics are going to one platform. Traces routed somewhere else. Logs duplicated across multiple backends because nobody wants to be caught without them when something breaks. Every redundant stream costs money ...

Operator to Orchestrator: 80% of IT Pros See Shift in Role as AI Permeates Workflows

May 15, 2026

80% of respondents agree that the IT role is shifting from operators to orchestrators, according to the 2026 IT Trends Report: The Human Side of Autonomous IT from SolarWinds ...

Preventing Outages in 2023: What We Can Learn from Recent Failures

February 09, 2023

Learn more about Catchpoint

Key lessons from the past

■ Develop an Internet Performance Monitoring strategy that allows you to monitor precisely what customers, workforce, and other users expect and build an Experience Score.

Hot Topics

Automation

The Latest

MEAN TIME TO INSIGHT Podcast - Episode 24: Network Observability Tool Sprawl

May 29, 2026

In MEAN TIME TO INSIGHT Episode 24, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network observability tool sprawl ...

Capacity Isn't a Guess: Observability-Driven Sizing for On-Prem Databases

May 28, 2026

5 Security Principles Every Entrepreneur Should Apply to Leadership

May 27, 2026

Signs It May Be Time to Reassess Your IT Infrastructure Strategy

May 26, 2026

Enterprise Edge AI Reaches Inflection Point

May 22, 2026

AI Is Hitting Operational Limits

May 21, 2026

Alert Fatigue Is No Longer a Morale Problem, It's a Reliability Risk and a System Failure

May 20, 2026

Most Enterprises Think They Can Switch AI Vendors in a Month ... Most Who've Tried Couldn't

May 19, 2026

Your Observability Stack Has a Telemetry Pipeline Problem

May 18, 2026

Operator to Orchestrator: 80% of IT Pros See Shift in Role as AI Permeates Workflows

May 15, 2026

80% of respondents agree that the IT role is shifting from operators to orchestrators, according to the 2026 IT Trends Report: The Human Side of Autonomous IT from SolarWinds ...

Featured Free Trial

Featured eBook

Featured Webinar

Featured Webinar

Featured Webinar

Featured White Paper

Featured Webinar

Featured Webinar

Featured Webinar

Featured White Paper

Featured Report

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured eBook

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Free Trial

Featured Webinar

Featured Report

Featured eBook

Featured Webinar

Featured Webinar

Featured Free Trial

Featured eBook

Featured Webinar

Featured Webinar

Featured White Paper

Featured Free Trial

Featured Free Trial

Featured Webinar

Featured eBook

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured White Paper

Featured Report

Featured Webinar

Featured Webinar

Featured Webinar

Featured Free Trial

Featured Webinar

Featured Report

Featured Webinar

Featured Webinar

Featured eBook

Featured Webinar

Featured Webinar

Featured eBook

Featured Webinar

Featured Webinar

Featured Free Trial

Featured Webinar

Featured Webinar