Skip to main content

Organizations Can Lose $1M+ Per Hour During Unplanned Disruptions

The financial stakes of extended service disruption has made operational resilience a top priority, according to 2026 State of AI-First Operations Report, a report from PagerDuty.

According to survey findings, 95% of respondents believe their leadership understands the competitive advantage that can be gained from reducing incidents and speeding recovery.

The report also shows that organizations are increasingly considering the adoption of AI for digital operations, with 59% indicating they actively incorporate the technology into operations. The AI adopters appear to be experiencing more success than those who may have discussed it but have not yet incorporated it: 75% report improved operational resilience, compared to only 66% of organizations that improved operational resilience but are not yet using AI.

Additional key takeaways from the report include:

Disruptions have become a board-level financial risk

Some organizations (8%) lose more than $1 million per hour, 34% lose at least $500,000 per hour, and more than two thirds (68%) lose more than $300,000 per hour during IT incidents. The cost of disruptions have grown too high for leaders to ignore and the impact extends beyond immediate revenue loss to damaging brand reputation (52%), introducing recovery costs (50%), reducing productivity (48%) and contributing to developer burnout (42%).

Successful organizations prioritize investments in operational resilience

A majority of organizations have made strides from their investments in the past year, with 71% reporting higher resilience and maturity than a year ago. However, progress appears to vary based on two key factors: business performance and investment. While 77% of organizations plan to increase operational resilience budgets over the next 12 months, companies reporting revenue growth are investing at significantly higher rates (82%) than underperformers (62%).

Post-incident learning capabilities gain recognition

Organizations that reported improved resilience most often attributed this progress to tools that combine integration with learning capabilities. Nearly half of organizations (48%) have increased resilience by turning incidents into structured learning opportunities to improve future performance. Successful companies with revenue growth are more likely to see a massive or moderate need for continuous learning (83%) than companies with flat or decreased revenue (77%). This suggests that the most successful platforms will be those that can transform incidents into systematic improvement cycles.

"The 2026 PagerDuty State of AI-First Operations Report further demonstrates how the financial risk of major incidents makes operational resilience a board-level priority," said Katherine Calvert, chief marketing officer at PagerDuty. "AI-first operations enable organizations to accelerate their incident management workflows so they can restore service more quickly during disruption. With PagerDuty, organizations can not only minimize risk, but cut down on teams’ time spent firefighting so they can focus on driving innovation and revenue."

Methodology: The report draws insights based on survey responses from 1,000 business leaders, IT decision makers and senior developers across Australia and New Zealand, France, Germany, Japan, the Nordic countries, the UK and Ireland, and the US.

Hot Topics

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Organizations Can Lose $1M+ Per Hour During Unplanned Disruptions

The financial stakes of extended service disruption has made operational resilience a top priority, according to 2026 State of AI-First Operations Report, a report from PagerDuty.

According to survey findings, 95% of respondents believe their leadership understands the competitive advantage that can be gained from reducing incidents and speeding recovery.

The report also shows that organizations are increasingly considering the adoption of AI for digital operations, with 59% indicating they actively incorporate the technology into operations. The AI adopters appear to be experiencing more success than those who may have discussed it but have not yet incorporated it: 75% report improved operational resilience, compared to only 66% of organizations that improved operational resilience but are not yet using AI.

Additional key takeaways from the report include:

Disruptions have become a board-level financial risk

Some organizations (8%) lose more than $1 million per hour, 34% lose at least $500,000 per hour, and more than two thirds (68%) lose more than $300,000 per hour during IT incidents. The cost of disruptions have grown too high for leaders to ignore and the impact extends beyond immediate revenue loss to damaging brand reputation (52%), introducing recovery costs (50%), reducing productivity (48%) and contributing to developer burnout (42%).

Successful organizations prioritize investments in operational resilience

A majority of organizations have made strides from their investments in the past year, with 71% reporting higher resilience and maturity than a year ago. However, progress appears to vary based on two key factors: business performance and investment. While 77% of organizations plan to increase operational resilience budgets over the next 12 months, companies reporting revenue growth are investing at significantly higher rates (82%) than underperformers (62%).

Post-incident learning capabilities gain recognition

Organizations that reported improved resilience most often attributed this progress to tools that combine integration with learning capabilities. Nearly half of organizations (48%) have increased resilience by turning incidents into structured learning opportunities to improve future performance. Successful companies with revenue growth are more likely to see a massive or moderate need for continuous learning (83%) than companies with flat or decreased revenue (77%). This suggests that the most successful platforms will be those that can transform incidents into systematic improvement cycles.

"The 2026 PagerDuty State of AI-First Operations Report further demonstrates how the financial risk of major incidents makes operational resilience a board-level priority," said Katherine Calvert, chief marketing officer at PagerDuty. "AI-first operations enable organizations to accelerate their incident management workflows so they can restore service more quickly during disruption. With PagerDuty, organizations can not only minimize risk, but cut down on teams’ time spent firefighting so they can focus on driving innovation and revenue."

Methodology: The report draws insights based on survey responses from 1,000 business leaders, IT decision makers and senior developers across Australia and New Zealand, France, Germany, Japan, the Nordic countries, the UK and Ireland, and the US.

Hot Topics

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...