The Leading Causes of IT Outages - and How to Prevent Them

November 04, 2019

Mark Banfield

LogicMonitor

IT outages happen to companies across the globe, regardless of location, annual revenue or size. Even the most mammoth companies are at risk of downtime. Increasingly over the past few years, high-profile IT outages — defined as when the services or systems a business provides suddenly become unavailable — have ended up splashed across national news headlines.

In March 2019, Facebook and Instagram each experienced 14 hours of downtime. A second IT outage struck both — along with WhatsApp — in April 2019, taking all three platforms offline. And in July 2019, all three platforms experienced availability problems that impacted users. British Airways has also faced a series of high-profile IT outages in the past, including one in April that resulted in 100 canceled flights and 200 delayed flights. An outage back in May 2017 also affected more than 1,000 flights, call centers, BA's website and BA's mobile app.

Given all of these recent disruptive and costly outages, LogicMonitor decided to investigate the causes behind downtime, commissioning an independent study investigating the major causes of downtime, the business impact of outages on organizations, and ways to avoid IT outages and brownouts. The IT Outage Impact Study involved surveying 300 IT decision-makers across the United States, Canada, the United Kingdom, Australia and New Zealand.

Outages Lead to Compliance Failures and High Costs

The number one and number two issues were concerns about performance and availability

Among other insights, the survey revealed the top 5 issues keeping IT decision makers up at night. The number one and number two issues were concerns about performance and availability, beating out security and cost-effectiveness worries.

Unfortunately, those self-reported fears about IT teams' ability to maintain availability are well-founded. In fact, 96% of global survey respondents reported that their organizations had suffered at least one IT outage over the past three years. Such outages can have serious implications, including steep costs and low customer satisfaction scores. Heavily regulated industries, such as healthcare and finance, face another dire consequence beyond service disruptions and costs as a result of outages: compliance failure.

"One of our clients is a radiology company, and they need to be up 24/7," said a service desk support engineer for a solution provider. "If they have more than an hour of downtime a year, probably less than that, that's a serious issue. These guys can never go down, for legal reasons."

Human Error is #1 Cause of IT Outages in the US and Canada

The study found that human error was the #1 cause of IT outages in the United States and Canada, and the #3 cause globally. Given this finding, it was no surprise that Network World covered the story of British Airways' May 2017 outage under the headline, "British Airways' outage, like most data center outages, was caused by humans."

The Network World article describes how an engineer working onsite at a data center near the Heathrow airport disconnected a power supply. When the power supply was reconnected, a surge of power caused the outage. The article also cites a 2016 Ponemon Institute study, which found that human error accounted for 11 percent of outages, more than weather (10%), generator failures (6%) or IT equipment malfunction (4%).

Faced with findings like this, it's no wonder that global IT decision makers said 51% of IT outages are avoidable. As a result, more and more teams worldwide are transitioning to monitoring tools that incorporate AIOps and automation to minimize human error and maximize early warning opportunities.

Monitoring Helps Prevent Outages Through Early Warning Systems

Comprehensive monitoring provides visibility into IT infrastructure and can help organizations get ahead of trends that indicate an outage may be rapidly approaching. The top two causes of outages, according to survey respondents, are declining hardware/software performance and IT teams' failure to notice when usage reaches a dangerous level. Artificial intelligence for IT operations (AIOps) and intelligent monitoring offer an effective solution to both of these outage factors.

To minimize your organizations' outage risk, look for monitoring solutions with the following capabilities:

■ A platform that offers a holistic view of your IT systems via a single pane of glass and integrates with all your technologies

■ A tool that builds in a high level of redundancy to eliminate single points of failure

■ A platform that provides early visibility via an early warning system into trends that could indicate future trouble

■ A solution that is able to scale with your business as it grows, making sure your current and future monitoring needs are met.

Mark Banfield is CRO at LogicMonitor

Hot Topics

Downtime

Monitoring

AIOps

The Latest

Rising IT Complexity Threatens Modernization - Survey Shows SysAdmins Under Pressure

June 04, 2025

Two in three IT professionals now cite growing complexity as their top challenge — an urgent signal that the modernization curve may be getting too steep, according to the Rising to the Challenge survey from Checkmk ...

State of the Data Center 2025

June 03, 2025

While IT leaders are becoming more comfortable and adept at balancing workloads across on-premises, colocation data centers and the public cloud, there's a key component missing: connectivity, according to the 2025 State of the Data Center Report from CoreSite ...

The Clock Is Ticking: How 47-Day Certificates and Quantum Threats Are Reshaping Cybersecurity

June 02, 2025

A perfect storm is brewing in cybersecurity — certificate lifespans shrinking to just 47 days while quantum computing threatens today's encryption. Organizations must embrace ephemeral trust and crypto-agility to survive this dual challenge ...

MEAN TIME TO INSIGHT Podcast - Episode 14: Hybrid Multi-Cloud Network Observability

May 29, 2025

In MEAN TIME TO INSIGHT Episode 14, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses hybrid multi-cloud network observability...

What's the State of AI Costs in 2025?

May 28, 2025

While companies adopt AI at a record pace, they also face the challenge of finding a smart and scalable way to manage its rapidly growing costs. This requires balancing the massive possibilities inherent in AI with the need to control cloud costs, aim for long-term profitability and optimize spending ...

Bridging the Visibility Gap: A Path to Smarter Telecom Infrastructure

May 27, 2025

Telecommunications is expanding at an unprecedented pace ... But progress brings complexity. As WanAware's 2025 Telecom Observability Benchmark Report reveals, many operators are discovering that modernization requires more than physical build outs and CapEx — it also demands the tools and insights to manage, secure, and optimize this fast-growing infrastructure in real time ...

Redis Monitoring 101: Key Metrics You Need to Watch

May 22, 2025

As businesses increasingly rely on high-performance applications to deliver seamless user experiences, the demand for fast, reliable, and scalable data storage systems has never been greater. Redis — an open-source, in-memory data structure store — has emerged as a popular choice for use cases ranging from caching to real-time analytics. But with great performance comes the need for vigilant monitoring ...

Beyond Traditional Autoscaling: The Future of Kubernetes in AI Infrastructure

May 22, 2025

Kubernetes was not initially designed with AI's vast resource variability in mind, and the rapid rise of AI has exposed Kubernetes limitations, particularly when it comes to cost and resource efficiency. Indeed, AI workloads differ from traditional applications in that they require a staggering amount and variety of compute resources, and their consumption is far less consistent than traditional workloads ... Considering the speed of AI innovation, teams cannot afford to be bogged down by these constant infrastructure concerns. A solution is needed ...

AI Drives Surge in Data Budgets

May 21, 2025

AI is the catalyst for significant investment in data teams as enterprises require higher-quality data to power their AI applications, according to the State of Analytics Engineering Report from dbt Labs ...

Misaligned Architecture Causes Service Disruptions, High Operational Costs and Security Challenges

May 20, 2025

Misaligned architecture can lead to business consequences, with 93% of respondents reporting negative outcomes such as service disruptions, high operational costs and security challenges ...

The Leading Causes of IT Outages - and How to Prevent Them

November 04, 2019

Mark Banfield

LogicMonitor

Learn more about LogicMonitor

Outages Lead to Compliance Failures and High Costs

The number one and number two issues were concerns about performance and availability

Human Error is #1 Cause of IT Outages in the US and Canada

Monitoring Helps Prevent Outages Through Early Warning Systems

To minimize your organizations' outage risk, look for monitoring solutions with the following capabilities:

■ A platform that offers a holistic view of your IT systems via a single pane of glass and integrates with all your technologies

■ A tool that builds in a high level of redundancy to eliminate single points of failure

■ A platform that provides early visibility via an early warning system into trends that could indicate future trouble

■ A solution that is able to scale with your business as it grows, making sure your current and future monitoring needs are met.

Mark Banfield is CRO at LogicMonitor

Hot Topics

Downtime

Monitoring

AIOps

The Latest

Rising IT Complexity Threatens Modernization - Survey Shows SysAdmins Under Pressure

June 04, 2025

State of the Data Center 2025

June 03, 2025

The Clock Is Ticking: How 47-Day Certificates and Quantum Threats Are Reshaping Cybersecurity

June 02, 2025

MEAN TIME TO INSIGHT Podcast - Episode 14: Hybrid Multi-Cloud Network Observability

May 29, 2025

In MEAN TIME TO INSIGHT Episode 14, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses hybrid multi-cloud network observability...

What's the State of AI Costs in 2025?

May 28, 2025

Bridging the Visibility Gap: A Path to Smarter Telecom Infrastructure

May 27, 2025

Redis Monitoring 101: Key Metrics You Need to Watch

May 22, 2025

Beyond Traditional Autoscaling: The Future of Kubernetes in AI Infrastructure

May 22, 2025

AI Drives Surge in Data Budgets

May 21, 2025

Misaligned Architecture Causes Service Disruptions, High Operational Costs and Security Challenges

May 20, 2025

Misaligned architecture can lead to business consequences, with 93% of respondents reporting negative outcomes such as service disruptions, high operational costs and security challenges ...

Featured Webinar

Featured White Paper

Featured eBook

Featured Report

Featured Webinar

Featured Webinar

Featured Free Trial

Featured White Paper

Featured Report

Featured Free Trial

Featured Webinar

Featured Webinar

Featured Webinar

Featured Free Trial

Featured eBook

Featured Free Trial

Featured Free Trial

Featured Webinar

Featured White Paper

Featured eBook

Featured Webinar

Featured White Paper

Featured Webinar

Featured Webinar

Featured eBook

Featured Free Tool

Featured White Paper

Featured White Paper

Featured Webinar

Featured Webinar

Featured White Paper

Featured Webinar

Featured Webinar

Featured Webinar

Featured Report

Featured Free Trial

Featured White Paper

Featured White Paper

Featured eBook

Featured eBook

Featured eBook

Featured Webinar

Featured White Paper

Featured eBook

Featured White Paper

Featured Free Trial

Featured Free Trial

Featured Free Trial

Featured White Paper

Featured eBook

Featured Free Trial

Featured White Paper

Featured Free Trial

Featured White Paper

Featured Report

Featured White Paper

Featured eBook

Featured Webinar

Featured Webinar

Featured White Paper

Featured Webinar

Featured Free Trial

Featured Webinar

Featured Free Trial

Featured Webinar