Skip to main content

AI Is Making Incident Management More Complex - Bad Data Makes It Harder

Eric Johnson
PagerDuty

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate.

With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action.

It All Begins with Data

Today, everything is digital, and that makes operations more complex and fragile. As companies add AI tools like customer service bots and coding assistants into the mix, they create even more dependencies and less room for error. That is a significant business risk, particularly when incidents are so costly: PagerDuty found that 68% of organizations lose more than $300,000 per hour during IT incidents, and a third lose at least $500,000.

While AI can contribute to the problem, it can also play a critical role in incident management when applied effectively. IT can enhance triage, run auto-diagnostics, and intelligently route alerts to the right engineers, saving time and reducing alert fatigue. For this to work, however, AI needs a real-time 360-degree view of the environment.

Establishing this level of visibility requires detailed incident origination data, including which systems and services were affected, which rules triggered alerts, and relevant timestamps. It also needs comprehensive resolution data, such as the type of automation used to remediate an issue, status updates, responder communications, and notes that explain why the issue occurred and how it was resolved.

Turning Data Into Value Requires Strong Governance

To be usable, incident datasets must be properly governed, and this governance depends on four key elements:

1. Data integrity

Duplicated records, inconsistent naming, and other issues undermine incident response processes. This leads to slower resolution times, limited ability to continuously improve, and engineers pulled away from their work to deal with false positives. At a time when DevOps talent is in high demand, this could have a direct impact on competitive advantage.

Organizations must work to keep data accurate and standardized, so AI-powered monitoring tools can do their job effectively. This reduces unnecessary noise in incident response workflows, ensures automated runbooks trigger the right actions at the right time, and enables incident post-mortems to generate useful insights.

2. Data security

Beyond the immediate hit to the bottom line, the long-term impact of data breaches on brand reputation and hard-won customer trust can be devastating. Regulators are also increasingly willing to scrutinize security posture and levy financial penalties.

Incident management processes must follow strict data security best practices to protect and encrypt data. These include least-privilege access policies to minimize risk exposure, logging and auditing for compliance, and encryption as a last line of defense. Alerts should also be configured to flag unusual access and usage patterns.

3. Compliance

The compliance landscape is growing more complex by the year. Alongside GDPR and similar data protection laws across US states, the EU AI Act introduces additional obligations for organizations developing, deploying, or selling certain AI systems. With the risk of significant fines, data governance teams must stay informed and continuously evolve their programs in response to changing mandates.

Organizations must meet data sovereignty and residency requirements, optimize data security, and ensure employees are appropriately trained. Supplier vetting is also critical: do data center providers and other partners meet recognized standards such as ISO 27001, SOC 2, and PCI DSS 4.0? Risk does not stop there as fourth-party breaches are an increasing concern.

From an AI perspective, governance teams also need clear visibility into how vendors use AI within their products, including whether these capabilities can be disabled if required. Contracts should define whether customer data is used to train models and under what terms, alongside assurances that human-in-the-loop checks are in place to prevent autonomous high-risk actions.

4. Storage

Storage is often overlooked in incident management, but it plays a critical role in the commercial viability of different approaches. Intelligent alert grouping and suppression reduce the volume of incident records and alerts stored in databases, while runbook automation supports more efficient data lifecycle management once an incident is resolved. This can include automatically deleting unused volumes and snapshots to further reduce costs.

Service graphs can also complement logging by helping teams understand dependencies and relationships across systems, enabling more targeted data collection and storage.

Governance in an Agentic Future

Get these four elements right, and organizations will have a solid foundation for effective incident management. But it's equally important to appreciate how the industry is evolving.

According to our data, three-quarters (75%) of companies have already deployed more than one AI agent, with 25% deploying five or more. These systems promise to automate incident investigation, diagnosis, and remediation to free up human talent to work on higher-value work.

This will only increase the importance of robust data governance. As humans are removed from parts of the decision-making process, data must be as accurate as possible. In a world shaped by autonomous AI, high-quality data is non-negotiable.

Eric Johnson is Chief Information Officer at PagerDuty

The Latest

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

AI Is Making Incident Management More Complex - Bad Data Makes It Harder

Eric Johnson
PagerDuty

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate.

With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action.

It All Begins with Data

Today, everything is digital, and that makes operations more complex and fragile. As companies add AI tools like customer service bots and coding assistants into the mix, they create even more dependencies and less room for error. That is a significant business risk, particularly when incidents are so costly: PagerDuty found that 68% of organizations lose more than $300,000 per hour during IT incidents, and a third lose at least $500,000.

While AI can contribute to the problem, it can also play a critical role in incident management when applied effectively. IT can enhance triage, run auto-diagnostics, and intelligently route alerts to the right engineers, saving time and reducing alert fatigue. For this to work, however, AI needs a real-time 360-degree view of the environment.

Establishing this level of visibility requires detailed incident origination data, including which systems and services were affected, which rules triggered alerts, and relevant timestamps. It also needs comprehensive resolution data, such as the type of automation used to remediate an issue, status updates, responder communications, and notes that explain why the issue occurred and how it was resolved.

Turning Data Into Value Requires Strong Governance

To be usable, incident datasets must be properly governed, and this governance depends on four key elements:

1. Data integrity

Duplicated records, inconsistent naming, and other issues undermine incident response processes. This leads to slower resolution times, limited ability to continuously improve, and engineers pulled away from their work to deal with false positives. At a time when DevOps talent is in high demand, this could have a direct impact on competitive advantage.

Organizations must work to keep data accurate and standardized, so AI-powered monitoring tools can do their job effectively. This reduces unnecessary noise in incident response workflows, ensures automated runbooks trigger the right actions at the right time, and enables incident post-mortems to generate useful insights.

2. Data security

Beyond the immediate hit to the bottom line, the long-term impact of data breaches on brand reputation and hard-won customer trust can be devastating. Regulators are also increasingly willing to scrutinize security posture and levy financial penalties.

Incident management processes must follow strict data security best practices to protect and encrypt data. These include least-privilege access policies to minimize risk exposure, logging and auditing for compliance, and encryption as a last line of defense. Alerts should also be configured to flag unusual access and usage patterns.

3. Compliance

The compliance landscape is growing more complex by the year. Alongside GDPR and similar data protection laws across US states, the EU AI Act introduces additional obligations for organizations developing, deploying, or selling certain AI systems. With the risk of significant fines, data governance teams must stay informed and continuously evolve their programs in response to changing mandates.

Organizations must meet data sovereignty and residency requirements, optimize data security, and ensure employees are appropriately trained. Supplier vetting is also critical: do data center providers and other partners meet recognized standards such as ISO 27001, SOC 2, and PCI DSS 4.0? Risk does not stop there as fourth-party breaches are an increasing concern.

From an AI perspective, governance teams also need clear visibility into how vendors use AI within their products, including whether these capabilities can be disabled if required. Contracts should define whether customer data is used to train models and under what terms, alongside assurances that human-in-the-loop checks are in place to prevent autonomous high-risk actions.

4. Storage

Storage is often overlooked in incident management, but it plays a critical role in the commercial viability of different approaches. Intelligent alert grouping and suppression reduce the volume of incident records and alerts stored in databases, while runbook automation supports more efficient data lifecycle management once an incident is resolved. This can include automatically deleting unused volumes and snapshots to further reduce costs.

Service graphs can also complement logging by helping teams understand dependencies and relationships across systems, enabling more targeted data collection and storage.

Governance in an Agentic Future

Get these four elements right, and organizations will have a solid foundation for effective incident management. But it's equally important to appreciate how the industry is evolving.

According to our data, three-quarters (75%) of companies have already deployed more than one AI agent, with 25% deploying five or more. These systems promise to automate incident investigation, diagnosis, and remediation to free up human talent to work on higher-value work.

This will only increase the importance of robust data governance. As humans are removed from parts of the decision-making process, data must be as accurate as possible. In a world shaped by autonomous AI, high-quality data is non-negotiable.

Eric Johnson is Chief Information Officer at PagerDuty

The Latest

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...