Skip to main content

FAA Outage: System Downtime Puts an Entire Industry on Hold

Pete Goldin
Editor and Publisher
APMdigest

"The US aviation sector was struggling to return to normal following a nationwide ground stop imposed by Federal Aviation Administration (FAA) early Wednesday over a computer issue that forced a 90-minute halt to all US departing flights," Reuters reported on January 11.


The breakdown showed how much American air travel depends on the computer system that generates alerts called NOTAMs — or Notice to Air Missions, Associated Press reported. The system broke down late Tuesday and was not fixed until midmorning Wednesday. The FAA took the rare step of preventing any planes from taking off for a time, and the cascading chaos led to more than 1,300 flight cancellations and 9,000 delays by early evening on the East Coast, according to flight-tracking website FlightAware.

The FAA said a corrupted file affected both the primary and backup systems.

Speaking of tech problems impacting the aviation industry, this happened only a couple weeks after Southwest Airlines experienced a meltdown during the holidays. National Public Radion reported: "By all accounts Southwest was using badly outdated computer systems to manage that complicated system."

But from the IT Ops perspective, the real take away from this news is not specifically about the FAA or the airline industry. Every organization faces this same concern every day — keeping systems updated and up and running. The alternative can be disastrous.

Many, if not most, companies in the US could not take a hit of this caliber and still maintain business as usual

"Today's FAA outage underscores the great need for modernized infrastructure, especially within organizations that operate on antiquated systems," said Fred Koopmans, BigPanda CPO. "The impact to travelers is obvious in this case, but it's imperative to also consider the internal mechanics the FAA will now have to address to recover from this."

Koopmans continued, "The average cost of a significant IT outage, according to 2022 research, is $6,912/minute or $414,720/hour – that's a $7.4M price tag for the FAA based on reports that issues arose at 3pm ET on Tuesday. Many, if not most, companies in the US could not take a hit of this caliber and still maintain business as usual."

"The outdated SaaS systems that many airlines rely upon are difficult to operate and run using older coding languages that few people still know how to use efficiently," explained Peter Pezaris, SVP of Strategy & User Experience at New Relic. "This means that when issues occur, they can be difficult to locate and fix — especially in a timely manner. Beyond that, they are also susceptible to cascading events, when a system fails and goes on to cause a ripple effect. As companies scale and the average tech stack becomes more complex, the risk of outages only rises. Not only is the IT team trying to get the system back up and running, but they are also fielding what can be a massive influx of requests ranging from internal stakeholders up to the Board level or customer complaints."

"Minimizing the time to understand the issue is critical," Pezaris added. "What makes this difficult is that most companies have observability data scattered everywhere. Observability unifies an organization's data and can provide airlines with a 360-degree view of their entire IT stacks, allowing engineers to detect and resolve issues before they impact flights."

Recently published data from New Relic's 2022 Observability Forecast shows that 45% of respondents experience an outage with a high business impact once per week or more — and 29% of those outages take an hour or more to resolve.

Pete Goldin is Editor and Publisher of APMdigest

Hot Topics

The Latest

From growing reliance on FinOps teams to the increasing attention on artificial intelligence (AI), and software licensing, the Flexera 2025 State of the Cloud Report digs into how organizations are improving cloud spend efficiency, while tackling the complexities of emerging technologies ...

Today, organizations are generating and processing more data than ever before. From training AI models to running complex analytics, massive datasets have become the backbone of innovation. However, as businesses embrace the cloud for its scalability and flexibility, a new challenge arises: managing the soaring costs of storing and processing this data ...

Despite the frustrations, every engineer we spoke with ultimately affirmed the value and power of OpenTelemetry. The "sucks" moments are often the flip side of its greatest strengths ... Part 2 of this blog covers the powerful advantages and breakthroughs — the "OTel Rocks" moments ...

OpenTelemetry (OTel) arrived with a grand promise: a unified, vendor-neutral standard for observability data (traces, metrics, logs) that would free engineers from vendor lock-in and provide deeper insights into complex systems ... No powerful technology comes without its challenges, and OpenTelemetry is no exception. The engineers we spoke with were frank about the friction points they've encountered ...

Enterprises are turning to AI-powered software platforms to make IT management more intelligent and ensure their systems and technology meet business needs for efficiency, lowers costs and innovation, according to new research from Information Services Group ...

The power of Kubernetes lies in its ability to orchestrate containerized applications with unparalleled efficiency. Yet, this power comes at a cost: the dynamic, distributed, and ephemeral nature of its architecture creates a monitoring challenge akin to tracking a constantly shifting, interconnected network of fleeting entities ... Due to the dynamic and complex nature of Kubernetes, monitoring poses a substantial challenge for DevOps and platform engineers. Here are the primary obstacles ...

The perception of IT has undergone a remarkable transformation in recent years. What was once viewed primarily as a cost center has transformed into a pivotal force driving business innovation and market leadership ... As someone who has witnessed and helped drive this evolution, it's become clear to me that the most successful organizations share a common thread: they've mastered the art of leveraging IT advancements to achieve measurable business outcomes ...

More than half (51%) of companies are already leveraging AI agents, according to the PagerDuty Agentic AI Survey. Agentic AI adoption is poised to accelerate faster than generative AI (GenAI) while reshaping automation and decision-making across industries ...

Image
Pagerduty

 

Real privacy protection thanks to technology and processes is often portrayed as too hard and too costly to implement. So the most common strategy is to do as little as possible just to conform to formal requirements of current and incoming regulations. This is a missed opportunity ...

The expanding use of AI is driving enterprise interest in data operations (DataOps) to orchestrate data integration and processing and improve data quality and validity, according to a new report from Information Services Group (ISG) ...

FAA Outage: System Downtime Puts an Entire Industry on Hold

Pete Goldin
Editor and Publisher
APMdigest

"The US aviation sector was struggling to return to normal following a nationwide ground stop imposed by Federal Aviation Administration (FAA) early Wednesday over a computer issue that forced a 90-minute halt to all US departing flights," Reuters reported on January 11.


The breakdown showed how much American air travel depends on the computer system that generates alerts called NOTAMs — or Notice to Air Missions, Associated Press reported. The system broke down late Tuesday and was not fixed until midmorning Wednesday. The FAA took the rare step of preventing any planes from taking off for a time, and the cascading chaos led to more than 1,300 flight cancellations and 9,000 delays by early evening on the East Coast, according to flight-tracking website FlightAware.

The FAA said a corrupted file affected both the primary and backup systems.

Speaking of tech problems impacting the aviation industry, this happened only a couple weeks after Southwest Airlines experienced a meltdown during the holidays. National Public Radion reported: "By all accounts Southwest was using badly outdated computer systems to manage that complicated system."

But from the IT Ops perspective, the real take away from this news is not specifically about the FAA or the airline industry. Every organization faces this same concern every day — keeping systems updated and up and running. The alternative can be disastrous.

Many, if not most, companies in the US could not take a hit of this caliber and still maintain business as usual

"Today's FAA outage underscores the great need for modernized infrastructure, especially within organizations that operate on antiquated systems," said Fred Koopmans, BigPanda CPO. "The impact to travelers is obvious in this case, but it's imperative to also consider the internal mechanics the FAA will now have to address to recover from this."

Koopmans continued, "The average cost of a significant IT outage, according to 2022 research, is $6,912/minute or $414,720/hour – that's a $7.4M price tag for the FAA based on reports that issues arose at 3pm ET on Tuesday. Many, if not most, companies in the US could not take a hit of this caliber and still maintain business as usual."

"The outdated SaaS systems that many airlines rely upon are difficult to operate and run using older coding languages that few people still know how to use efficiently," explained Peter Pezaris, SVP of Strategy & User Experience at New Relic. "This means that when issues occur, they can be difficult to locate and fix — especially in a timely manner. Beyond that, they are also susceptible to cascading events, when a system fails and goes on to cause a ripple effect. As companies scale and the average tech stack becomes more complex, the risk of outages only rises. Not only is the IT team trying to get the system back up and running, but they are also fielding what can be a massive influx of requests ranging from internal stakeholders up to the Board level or customer complaints."

"Minimizing the time to understand the issue is critical," Pezaris added. "What makes this difficult is that most companies have observability data scattered everywhere. Observability unifies an organization's data and can provide airlines with a 360-degree view of their entire IT stacks, allowing engineers to detect and resolve issues before they impact flights."

Recently published data from New Relic's 2022 Observability Forecast shows that 45% of respondents experience an outage with a high business impact once per week or more — and 29% of those outages take an hour or more to resolve.

Pete Goldin is Editor and Publisher of APMdigest

Hot Topics

The Latest

From growing reliance on FinOps teams to the increasing attention on artificial intelligence (AI), and software licensing, the Flexera 2025 State of the Cloud Report digs into how organizations are improving cloud spend efficiency, while tackling the complexities of emerging technologies ...

Today, organizations are generating and processing more data than ever before. From training AI models to running complex analytics, massive datasets have become the backbone of innovation. However, as businesses embrace the cloud for its scalability and flexibility, a new challenge arises: managing the soaring costs of storing and processing this data ...

Despite the frustrations, every engineer we spoke with ultimately affirmed the value and power of OpenTelemetry. The "sucks" moments are often the flip side of its greatest strengths ... Part 2 of this blog covers the powerful advantages and breakthroughs — the "OTel Rocks" moments ...

OpenTelemetry (OTel) arrived with a grand promise: a unified, vendor-neutral standard for observability data (traces, metrics, logs) that would free engineers from vendor lock-in and provide deeper insights into complex systems ... No powerful technology comes without its challenges, and OpenTelemetry is no exception. The engineers we spoke with were frank about the friction points they've encountered ...

Enterprises are turning to AI-powered software platforms to make IT management more intelligent and ensure their systems and technology meet business needs for efficiency, lowers costs and innovation, according to new research from Information Services Group ...

The power of Kubernetes lies in its ability to orchestrate containerized applications with unparalleled efficiency. Yet, this power comes at a cost: the dynamic, distributed, and ephemeral nature of its architecture creates a monitoring challenge akin to tracking a constantly shifting, interconnected network of fleeting entities ... Due to the dynamic and complex nature of Kubernetes, monitoring poses a substantial challenge for DevOps and platform engineers. Here are the primary obstacles ...

The perception of IT has undergone a remarkable transformation in recent years. What was once viewed primarily as a cost center has transformed into a pivotal force driving business innovation and market leadership ... As someone who has witnessed and helped drive this evolution, it's become clear to me that the most successful organizations share a common thread: they've mastered the art of leveraging IT advancements to achieve measurable business outcomes ...

More than half (51%) of companies are already leveraging AI agents, according to the PagerDuty Agentic AI Survey. Agentic AI adoption is poised to accelerate faster than generative AI (GenAI) while reshaping automation and decision-making across industries ...

Image
Pagerduty

 

Real privacy protection thanks to technology and processes is often portrayed as too hard and too costly to implement. So the most common strategy is to do as little as possible just to conform to formal requirements of current and incoming regulations. This is a missed opportunity ...

The expanding use of AI is driving enterprise interest in data operations (DataOps) to orchestrate data integration and processing and improve data quality and validity, according to a new report from Information Services Group (ISG) ...