Down Goes the Internet (Again) – Part One: Are You Ready?

October 28, 2013

Learn more about Dynatrace

What is the cost of downtime? The answer obviously depends on various factors such as the size of an organization, the industry, the duration of the outage and the number of people impacted. To provide a ballpark, though, the Uptime Institute Symposium estimates an average cost of $5,600 per minute.

And it's not all about dollars and cents. Reputation, customer retention, employee satisfaction and overall confidence can be shaken by even a short outage.

For these reasons, media are drawn to outage stories like passers-by of a major roadside accident. Recently, a spate of high-profile outages has once again captured headlines around the world:

- On August 14, the New York Times website experienced a two-hour failure, in which the newspaper had to resort to publishing articles on its Facebook page.

- On the same day, Microsoft customers began reporting email failures en masse. The outage was traced to problems with the Exchange ActiveSync service which serves email to many of the world's smartphones. When Exchange hit a glitch, the volume of phones trying to connect triggered a tsunami of traffic that took three days to get under control.

- On August 22, a software bug and other technology issues brought the NASDAQ stock exchange to a standstill. For approximately three hours, trading was halted for Apple, Google and Facebook and others of this ilk. Because other exchanges rely on NASDAQ's pricing, the fault had a ripple effect that seriously undermined market confidence. This grim fallout resulted in a third fewer shares being traded in the US on that day.

- On August 23, Apple's iCloud service, which helps connect iPhones, iPads and other Apple devices to key services, went down for more than six hours. While Apple claimed that the outage impacted less than one percent of iCloud customers, the sheer size of this user base – 300 million users – translated to approximately three million users being disconnected from services for 11 hours.

- On August 26, an Amazon EC2 outage rocked Instagram, Vine, Netflix and several other major customers of this cloud service, inflicting unplanned downtime across all of them. Last year, a similar Amazon EC2 outage caused Netflix to go down on Christmas day, a busy time for the video streaming service. This occurred in spite of the fact that Amazon had just upgraded their servers to make them less likely to collapse.

In response to these events, industry experts have sounded their alarm and issued a stark warning: we are over-reliant on a digital infrastructure that has become far too complex and exceeds our limits of control.

From high volume securities trading to the explosion in social media and the online consumption of entertainment, the amount of data being carried globally over private networks, such as stock exchanges, and the public internet is placing unprecedented strain on websites and the networks that connect them. According to recent statistics from Cisco, by 2017, the amount of data equivalent to all the films ever produced will be transmitted over the internet in just three minutes.

With these trends showing no signs of abating, we can expect widespread service outages and performance degradations to continue. Knowing this, many organizations go into overdrive as they attempt to improve their resiliency and ensure strong performance levels. But like an auto-immune disease, the addition of processes and technologies can actually have the adverse effect of increasing complexity and risk, by introducing new points of failure.

What is Causing Today's Massive Ripple Effect?

Today, businesses are hosting less and less of what gets delivered on their websites, instead relying on a growing number of externally hosted (third party) web elements to enrich their web properties. These third party internet services are also called web services, and when a major web service goes down, it often takes a portion of the internet with it.

As an example, on August 16, several of Google's websites including email, YouTube and its core search engine suffered a rare four-minute global meltdown. The episode, the cause of which Google has not explained publicly, served to illustrate the staggering volume of global internet traffic served by Google. During the outage, one monitor put the drop in global internet traffic at 40 percent – reinforcing the concept that when major web services fail, they tend to fail spectacularly.

It's certainly true that in recent years, businesses have dramatically increased their use of cloud services. According to Verizon's State of the Enterprise Cloud Report, enterprise use of cloud technology grew by 90 percent between January 2012 and June 2013.

Put another way, to stay competitive, companies have no choice but to provide the best online experience to their online customers and shoppers. That means they must provide the ability to watch a product video, use a coupon, subscribe for a promotion, read customer reviews, share with their friends on Facebook or Twitter, select products and pay for them to be delivered within a defined timeframe. All of this constitutes a lot of minor services. In theory, these functionalities could be developed, hosted and maintained in house, but the cost associated with hardware, software, development, support and maintenance often does not make economic sense.

In the meantime, startups who saw the business opportunity have developed externally hosted packaged solutions for the very functionalities companies wish to offer. The trend is to contract more of these specialized third party service providers, which often results in a company becoming a cloud customer indirectly, without their even knowing it!

Today, a North American website has somewhere between 9 and 13 third party web services contributing to a typical web transaction, according to May 2013 data provided by Compuware's Outage Analyzer. If any one of these third party services slows down or fails, performance for an entire web page, mobile site or application can degrade substantially, wreaking havoc on a company's reputation and revenues.

Third party performance issues occur more frequently than one might think. As an example, Outage Analyzer recently collated data for the six months between March 1 and August 31, 2013 and found 6,217 total outages which included:

- 1,500 full service outages – an average of 125 per month or about four daily – where the entire web service was unavailable in all geographies.

- 4,717 partial service outages – an average of 393 per month or about 13 daily – where only certain geographies or a limited number of user transactions were affected. While full service outages get the most attention, a partial service outage is more likely to occur and affect a limited number of individual web and mobile transactions, while leaving others completely untouched. But all you need is one disgruntled user logging onto Facebook or Twitter to start spreading viral negativity on your brand.

According to Compuware, there are nearly 1,500 distinct third party services available worldwide. Ad servers and social media plug-ins experience the highest number of outages, while online security services and ad verification experience the fewest number. The longest duration for a web service outage in this tracking period was 4,876 minutes (or 3.3 days) for an ad serving firm on March 21, 2013.

Down Goes the Internet (Again) – Part Two: 4 Strategies to Ensure Website Performance

Klaus Enzenhofer is Technology Strategist for Compuware APM’s Center of Excellence.

Hot Topics

The Latest

MEAN TIME TO INSIGHT Podcast - Episode 24: Network Observability Tool Sprawl

May 29, 2026

In MEAN TIME TO INSIGHT Episode 24, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network observability tool sprawl ...

Capacity Isn't a Guess: Observability-Driven Sizing for On-Prem Databases

May 28, 2026

In cloud-native systems, scaling is often as simple as moving a slider. For on-premise databases, the stakes are different. Over-provisioning hardware is expensive. Under-provisioning leads to performance bottlenecks that are difficult to fix once the equipment is in the rack ...

5 Security Principles Every Entrepreneur Should Apply to Leadership

May 27, 2026

When most people think about cybersecurity, they picture firewalls, encryption, and access controls — technical tools designed to protect systems and data. But beneath the technology lies a deeper set of principles about trust, decision-making, and resilience ... The best leaders don't eliminate risk. They manage it intelligently. And in many ways, cybersecurity offers a surprisingly useful playbook for doing exactly that ...

Signs It May Be Time to Reassess Your IT Infrastructure Strategy

May 26, 2026

Many organizations assumed their infrastructure strategy was settled. It had been implemented, optimized and built into long-term plans. Recent changes in technology and vendor consolidation are forcing a second look. Cloud outages and licensing changes have exposed how much dependency exists on a small number of platforms. As a result, organizations are reevaluating whether those decisions still hold up under current conditions ...

Enterprise Edge AI Reaches Inflection Point

May 22, 2026

Edge AI is strategically embedded in core IT and infrastructure spending across industries, according to the 2026 Edge AI Survey from ZEDEDA. The research shows that 83% of C-suite and IT executive respondents say edge AI is important to their core business strategy ...

AI Is Hitting Operational Limits

May 21, 2026

As AI adoption accelerates, operational complexity — not model intelligence — is becoming the primary barrier to reliable AI at scale, according to the State of AI Engineering 2026 from Datadog ... The report highlights a compounding complexity challenge as AI systems scale ... Around 5% of AI model requests fail in production, with nearly 60% of those failures caused by capacity limits ...

Alert Fatigue Is No Longer a Morale Problem, It's a Reliability Risk and a System Failure

May 20, 2026

For years, production operations teams have treated alert fatigue as a quality-of-life problem: something that makes on-call rotations miserable but isn't considered a direct contributor to outages. That framing doesn't capture how these systems fail, and we now have data to show why. More importantly, it's now clear alert fatigue is a symptom of a deeper issue: production systems have outgrown the current operational approaches ...

Most Enterprises Think They Can Switch AI Vendors in a Month ... Most Who've Tried Couldn't

May 19, 2026

I was on a customer call last fall when an enterprise architect said something I haven't been able to shake. Her team had just spent four months trying to swap one AI vendor for another. The original plan said three weeks. "We didn't switch vendors," she told me. "We rebuilt half our integrations and discovered what we'd actually been depending on." Most enterprise leaders don't expect that to be the experience ...

Your Observability Stack Has a Telemetry Pipeline Problem

May 18, 2026

Ask any senior SRE or platform engineer what keeps them up at night, and the answer probably isn't the monitoring tool — it's the data feeding it. The proliferation of APM, observability, and AIOps platforms has created a telemetry sprawl problem that most teams manage reactively rather than architect proactively. Metrics are going to one platform. Traces routed somewhere else. Logs duplicated across multiple backends because nobody wants to be caught without them when something breaks. Every redundant stream costs money ...

Operator to Orchestrator: 80% of IT Pros See Shift in Role as AI Permeates Workflows

May 15, 2026

80% of respondents agree that the IT role is shifting from operators to orchestrators, according to the 2026 IT Trends Report: The Human Side of Autonomous IT from SolarWinds ...

Down Goes the Internet (Again) – Part One: Are You Ready?

October 28, 2013

Learn more about Dynatrace

And it's not all about dollars and cents. Reputation, customer retention, employee satisfaction and overall confidence can be shaken by even a short outage.

For these reasons, media are drawn to outage stories like passers-by of a major roadside accident. Recently, a spate of high-profile outages has once again captured headlines around the world:

- On August 14, the New York Times website experienced a two-hour failure, in which the newspaper had to resort to publishing articles on its Facebook page.

What is Causing Today's Massive Ripple Effect?

- 1,500 full service outages – an average of 125 per month or about four daily – where the entire web service was unavailable in all geographies.

Down Goes the Internet (Again) – Part Two: 4 Strategies to Ensure Website Performance

Klaus Enzenhofer is Technology Strategist for Compuware APM’s Center of Excellence.

Hot Topics

The Latest

MEAN TIME TO INSIGHT Podcast - Episode 24: Network Observability Tool Sprawl

May 29, 2026

In MEAN TIME TO INSIGHT Episode 24, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network observability tool sprawl ...

Capacity Isn't a Guess: Observability-Driven Sizing for On-Prem Databases

May 28, 2026

5 Security Principles Every Entrepreneur Should Apply to Leadership

May 27, 2026

Signs It May Be Time to Reassess Your IT Infrastructure Strategy

May 26, 2026

Enterprise Edge AI Reaches Inflection Point

May 22, 2026

AI Is Hitting Operational Limits

May 21, 2026

Alert Fatigue Is No Longer a Morale Problem, It's a Reliability Risk and a System Failure

May 20, 2026

Most Enterprises Think They Can Switch AI Vendors in a Month ... Most Who've Tried Couldn't

May 19, 2026

Your Observability Stack Has a Telemetry Pipeline Problem

May 18, 2026

Operator to Orchestrator: 80% of IT Pros See Shift in Role as AI Permeates Workflows

May 15, 2026

80% of respondents agree that the IT role is shifting from operators to orchestrators, according to the 2026 IT Trends Report: The Human Side of Autonomous IT from SolarWinds ...

Featured Free Trial

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Free Trial

Featured Webinar

Featured Webinar

Featured eBook

Featured Free Trial

Featured Webinar

Featured Webinar

Featured Webinar

Featured eBook

Featured Webinar

Featured White Paper

Featured Webinar

Featured eBook

Featured Webinar

Featured Webinar

Featured White Paper

Featured Webinar

Featured Free Trial

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Free Trial

Featured eBook

Featured Webinar

Featured Webinar

Featured Free Trial

Featured eBook

Featured Webinar

Featured White Paper

Featured Free Tool

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured Report

Featured Webinar

Featured eBook

Featured Webinar

Featured Webinar

Featured White Paper

Featured Free Trial

Featured Free Trial

Featured Free Tool

Featured Webinar

Featured Free Trial

Featured eBook

Featured Webinar

Featured eBook

Featured Free Trial

Featured Webinar

Featured Webinar

Featured Webinar

Featured Webinar

Featured White Paper

Featured Webinar