The Leading Causes of IT Outages - and How to Prevent Them
November 04, 2019

Mark Banfield
LogicMonitor

Share this

IT outages happen to companies across the globe, regardless of location, annual revenue or size. Even the most mammoth companies are at risk of downtime. Increasingly over the past few years, high-profile IT outages — defined as when the services or systems a business provides suddenly become unavailable — have ended up splashed across national news headlines.

In March 2019, Facebook and Instagram each experienced 14 hours of downtime. A second IT outage struck both — along with WhatsApp — in April 2019, taking all three platforms offline. And in July 2019, all three platforms experienced availability problems that impacted users. British Airways has also faced a series of high-profile IT outages in the past, including one in April that resulted in 100 canceled flights and 200 delayed flights. An outage back in May 2017 also affected more than 1,000 flights, call centers, BA's website and BA's mobile app.

Given all of these recent disruptive and costly outages, LogicMonitor decided to investigate the causes behind downtime, commissioning an independent study investigating the major causes of downtime, the business impact of outages on organizations, and ways to avoid IT outages and brownouts. The IT Outage Impact Study involved surveying 300 IT decision-makers across the United States, Canada, the United Kingdom, Australia and New Zealand.

Outages Lead to Compliance Failures and High Costs

The number one and number two issues were concerns about performance and availability

Among other insights, the survey revealed the top 5 issues keeping IT decision makers up at night. The number one and number two issues were concerns about performance and availability, beating out security and cost-effectiveness worries.

Unfortunately, those self-reported fears about IT teams' ability to maintain availability are well-founded. In fact, 96% of global survey respondents reported that their organizations had suffered at least one IT outage over the past three years. Such outages can have serious implications, including steep costs and low customer satisfaction scores. Heavily regulated industries, such as healthcare and finance, face another dire consequence beyond service disruptions and costs as a result of outages: compliance failure.

"One of our clients is a radiology company, and they need to be up 24/7," said a service desk support engineer for a solution provider. "If they have more than an hour of downtime a year, probably less than that, that's a serious issue. These guys can never go down, for legal reasons."


Human Error is #1 Cause of IT Outages in the US and Canada

The study found that human error was the #1 cause of IT outages in the United States and Canada, and the #3 cause globally. Given this finding, it was no surprise that Network World covered the story of British Airways' May 2017 outage under the headline, "British Airways' outage, like most data center outages, was caused by humans."

The Network World article describes how an engineer working onsite at a data center near the Heathrow airport disconnected a power supply. When the power supply was reconnected, a surge of power caused the outage. The article also cites a 2016 Ponemon Institute study, which found that human error accounted for 11 percent of outages, more than weather (10%), generator failures (6%) or IT equipment malfunction (4%).

Faced with findings like this, it's no wonder that global IT decision makers said 51% of IT outages are avoidable. As a result, more and more teams worldwide are transitioning to monitoring tools that incorporate AIOps and automation to minimize human error and maximize early warning opportunities.

Monitoring Helps Prevent Outages Through Early Warning Systems

Comprehensive monitoring provides visibility into IT infrastructure and can help organizations get ahead of trends that indicate an outage may be rapidly approaching. The top two causes of outages, according to survey respondents, are declining hardware/software performance and IT teams' failure to notice when usage reaches a dangerous level. Artificial intelligence for IT operations (AIOps) and intelligent monitoring offer an effective solution to both of these outage factors.

To minimize your organizations' outage risk, look for monitoring solutions with the following capabilities:

■ A platform that offers a holistic view of your IT systems via a single pane of glass and integrates with all your technologies

■ A tool that builds in a high level of redundancy to eliminate single points of failure

■ A platform that provides early visibility via an early warning system into trends that could indicate future trouble

■ A solution that is able to scale with your business as it grows, making sure your current and future monitoring needs are met.

Mark Banfield is CRO at LogicMonitor
Share this

The Latest

November 14, 2019

A brief introduction to Applications Performance Monitoring (APM), breaking it down to a few key points, followed by a few important lessons which I have learned over the years ...

November 13, 2019

Research conducted by ServiceNow shows that Gen Zs, now entering the workforce, recognize the promise of technology to improve work experiences, are eager to learn from other generations, and believe they can help older generations be more open‑minded ...

November 12, 2019

We're in the middle of a technology and connectivity revolution, giving us access to infinite digital tools and technologies. Is this multitude of technology solutions empowering us to do our best work, or getting in our way? ...

November 07, 2019

Microservices have become the go-to architectural standard in modern distributed systems. While there are plenty of tools and techniques to architect, manage, and automate the deployment of such distributed systems, issues during troubleshooting still happen at the individual service level, thereby prolonging the time taken to resolve an outage ...

November 06, 2019

A recent APMdigest blog by Jean Tunis provided an excellent background on Application Performance Monitoring (APM) and what it does. A further topic that I wanted to touch on though is the need for good quality data. If you are to get the most out of your APM solution possible, you will need to feed it with the best quality data ...

November 05, 2019

Humans and manual processes can no longer keep pace with network innovation, evolution, complexity, and change. That's why we're hearing more about self-driving networks, self-healing networks, intent-based networking, and other concepts. These approaches collectively belong to a growing focus area called AIOps, which aims to apply automation, AI and ML to support modern network operations ...

November 04, 2019

IT outages happen to companies across the globe, regardless of location, annual revenue or size. Even the most mammoth companies are at risk of downtime. Increasingly over the past few years, high-profile IT outages — defined as when the services or systems a business provides suddenly become unavailable — have ended up splashed across national news headlines ...

October 31, 2019

APM tools are ideal for an application owner or a line of business owner to track the performance of their key applications. But these tools have broader applicability to different stakeholders in an organization. In this blog, we will review the teams and functional departments that can make use of an APM tool and how they could put it to work ...

October 30, 2019

Enterprises depending exclusively on legacy monitoring tools are falling behind in business agility and operational efficiency, according to a new study, Prevalence of Legacy Tools Paralyzes Enterprises' Ability to Innovate conducted by Forrester Consulting ...

October 29, 2019

Hyperconverged infrastructure is sometimes referred to as a "data center in a box" because, after the initial cabling and minimal networking configuration, it has all of the features and functionality of the traditional 3-2-1 virtualization architecture (except that single point of failure) ...