Skip to main content

AI Is Making Incident Management More Complex - Bad Data Makes It Harder

Eric Johnson
PagerDuty

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate.

With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action.

It All Begins with Data

Today, everything is digital, and that makes operations more complex and fragile. As companies add AI tools like customer service bots and coding assistants into the mix, they create even more dependencies and less room for error. That is a significant business risk, particularly when incidents are so costly: PagerDuty found that 68% of organizations lose more than $300,000 per hour during IT incidents, and a third lose at least $500,000.

While AI can contribute to the problem, it can also play a critical role in incident management when applied effectively. IT can enhance triage, run auto-diagnostics, and intelligently route alerts to the right engineers, saving time and reducing alert fatigue. For this to work, however, AI needs a real-time 360-degree view of the environment.

Establishing this level of visibility requires detailed incident origination data, including which systems and services were affected, which rules triggered alerts, and relevant timestamps. It also needs comprehensive resolution data, such as the type of automation used to remediate an issue, status updates, responder communications, and notes that explain why the issue occurred and how it was resolved.

Turning Data Into Value Requires Strong Governance

To be usable, incident datasets must be properly governed, and this governance depends on four key elements:

1. Data integrity

Duplicated records, inconsistent naming, and other issues undermine incident response processes. This leads to slower resolution times, limited ability to continuously improve, and engineers pulled away from their work to deal with false positives. At a time when DevOps talent is in high demand, this could have a direct impact on competitive advantage.

Organizations must work to keep data accurate and standardized, so AI-powered monitoring tools can do their job effectively. This reduces unnecessary noise in incident response workflows, ensures automated runbooks trigger the right actions at the right time, and enables incident post-mortems to generate useful insights.

2. Data security

Beyond the immediate hit to the bottom line, the long-term impact of data breaches on brand reputation and hard-won customer trust can be devastating. Regulators are also increasingly willing to scrutinize security posture and levy financial penalties.

Incident management processes must follow strict data security best practices to protect and encrypt data. These include least-privilege access policies to minimize risk exposure, logging and auditing for compliance, and encryption as a last line of defense. Alerts should also be configured to flag unusual access and usage patterns.

3. Compliance

The compliance landscape is growing more complex by the year. Alongside GDPR and similar data protection laws across US states, the EU AI Act introduces additional obligations for organizations developing, deploying, or selling certain AI systems. With the risk of significant fines, data governance teams must stay informed and continuously evolve their programs in response to changing mandates.

Organizations must meet data sovereignty and residency requirements, optimize data security, and ensure employees are appropriately trained. Supplier vetting is also critical: do data center providers and other partners meet recognized standards such as ISO 27001, SOC 2, and PCI DSS 4.0? Risk does not stop there as fourth-party breaches are an increasing concern.

From an AI perspective, governance teams also need clear visibility into how vendors use AI within their products, including whether these capabilities can be disabled if required. Contracts should define whether customer data is used to train models and under what terms, alongside assurances that human-in-the-loop checks are in place to prevent autonomous high-risk actions.

4. Storage

Storage is often overlooked in incident management, but it plays a critical role in the commercial viability of different approaches. Intelligent alert grouping and suppression reduce the volume of incident records and alerts stored in databases, while runbook automation supports more efficient data lifecycle management once an incident is resolved. This can include automatically deleting unused volumes and snapshots to further reduce costs.

Service graphs can also complement logging by helping teams understand dependencies and relationships across systems, enabling more targeted data collection and storage.

Governance in an Agentic Future

Get these four elements right, and organizations will have a solid foundation for effective incident management. But it's equally important to appreciate how the industry is evolving.

According to our data, three-quarters (75%) of companies have already deployed more than one AI agent, with 25% deploying five or more. These systems promise to automate incident investigation, diagnosis, and remediation to free up human talent to work on higher-value work.

This will only increase the importance of robust data governance. As humans are removed from parts of the decision-making process, data must be as accurate as possible. In a world shaped by autonomous AI, high-quality data is non-negotiable.

Eric Johnson is Chief Information Officer at PagerDuty

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

AI Is Making Incident Management More Complex - Bad Data Makes It Harder

Eric Johnson
PagerDuty

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate.

With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action.

It All Begins with Data

Today, everything is digital, and that makes operations more complex and fragile. As companies add AI tools like customer service bots and coding assistants into the mix, they create even more dependencies and less room for error. That is a significant business risk, particularly when incidents are so costly: PagerDuty found that 68% of organizations lose more than $300,000 per hour during IT incidents, and a third lose at least $500,000.

While AI can contribute to the problem, it can also play a critical role in incident management when applied effectively. IT can enhance triage, run auto-diagnostics, and intelligently route alerts to the right engineers, saving time and reducing alert fatigue. For this to work, however, AI needs a real-time 360-degree view of the environment.

Establishing this level of visibility requires detailed incident origination data, including which systems and services were affected, which rules triggered alerts, and relevant timestamps. It also needs comprehensive resolution data, such as the type of automation used to remediate an issue, status updates, responder communications, and notes that explain why the issue occurred and how it was resolved.

Turning Data Into Value Requires Strong Governance

To be usable, incident datasets must be properly governed, and this governance depends on four key elements:

1. Data integrity

Duplicated records, inconsistent naming, and other issues undermine incident response processes. This leads to slower resolution times, limited ability to continuously improve, and engineers pulled away from their work to deal with false positives. At a time when DevOps talent is in high demand, this could have a direct impact on competitive advantage.

Organizations must work to keep data accurate and standardized, so AI-powered monitoring tools can do their job effectively. This reduces unnecessary noise in incident response workflows, ensures automated runbooks trigger the right actions at the right time, and enables incident post-mortems to generate useful insights.

2. Data security

Beyond the immediate hit to the bottom line, the long-term impact of data breaches on brand reputation and hard-won customer trust can be devastating. Regulators are also increasingly willing to scrutinize security posture and levy financial penalties.

Incident management processes must follow strict data security best practices to protect and encrypt data. These include least-privilege access policies to minimize risk exposure, logging and auditing for compliance, and encryption as a last line of defense. Alerts should also be configured to flag unusual access and usage patterns.

3. Compliance

The compliance landscape is growing more complex by the year. Alongside GDPR and similar data protection laws across US states, the EU AI Act introduces additional obligations for organizations developing, deploying, or selling certain AI systems. With the risk of significant fines, data governance teams must stay informed and continuously evolve their programs in response to changing mandates.

Organizations must meet data sovereignty and residency requirements, optimize data security, and ensure employees are appropriately trained. Supplier vetting is also critical: do data center providers and other partners meet recognized standards such as ISO 27001, SOC 2, and PCI DSS 4.0? Risk does not stop there as fourth-party breaches are an increasing concern.

From an AI perspective, governance teams also need clear visibility into how vendors use AI within their products, including whether these capabilities can be disabled if required. Contracts should define whether customer data is used to train models and under what terms, alongside assurances that human-in-the-loop checks are in place to prevent autonomous high-risk actions.

4. Storage

Storage is often overlooked in incident management, but it plays a critical role in the commercial viability of different approaches. Intelligent alert grouping and suppression reduce the volume of incident records and alerts stored in databases, while runbook automation supports more efficient data lifecycle management once an incident is resolved. This can include automatically deleting unused volumes and snapshots to further reduce costs.

Service graphs can also complement logging by helping teams understand dependencies and relationships across systems, enabling more targeted data collection and storage.

Governance in an Agentic Future

Get these four elements right, and organizations will have a solid foundation for effective incident management. But it's equally important to appreciate how the industry is evolving.

According to our data, three-quarters (75%) of companies have already deployed more than one AI agent, with 25% deploying five or more. These systems promise to automate incident investigation, diagnosis, and remediation to free up human talent to work on higher-value work.

This will only increase the importance of robust data governance. As humans are removed from parts of the decision-making process, data must be as accurate as possible. In a world shaped by autonomous AI, high-quality data is non-negotiable.

Eric Johnson is Chief Information Officer at PagerDuty

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...