Skip to main content

Data Matters More Than Ever in AIOps

New study reveals key hurdles on the road to AIOps
Bhanu Singh

Like any new, potentially disruptive technology, artificial intelligence for IT operations, or AIOps has quickly become a trend, and slowly become a reality. It's only been a few years since Gartner coined the term, and yet, 30% of IT teams in large enterprises will roll out AIOps initiatives by 2023. These IT practitioners are still in experimentation mode with artificial intelligence in many cases, and still have concerns about how credible the technology can be. They have concerns over the results of these implementations, and worry about maintaining service availability and uptime during migration.

Because AIOps is still in its infancy, there hasn't been much reporting on what these concerns specifically are. A recent study from OpsRamp targeted these IT managers who have implemented AIOps, and among other data, reports on the primary concerns of this new approach to operations management.

The Devil is in the Data

The report cites data accuracy as the chief concern for IT pros when it comes to AIOps. Two-thirds (67%) of those surveyed revealed it as their top priority. This could be for a variety of reasons, including:

Data Sources: In a world of distributed, hybrid, multi-cloud infrastructure, it's more difficult than ever to capture data on every level of an organization. Different cloud providers report in different ways. And point tools provide analytics across a host of different metrics. It's next to impossible to compare data sources together for a true contextual view of the organization.

Data Quality: Even when that data is captured, these IT teams aren't necessarily sure that it's accurately reflecting the truth about a system. Modern data can be fragmented, hidden, unparsed or too distributed to make sense.

Data Volume: Today's enterprise infrastructure produces an overwhelming amount of metrics on usage, capacity, performance, availability, security, and more. It's easy to get lost in the noise.

Data Consistency: It's impossible to say, under the crushing weight of data today, that IT teams are seeing consistent reporting and results across the organization. But until data is consistent, it can't be actionable.

Data Culture: This is perhaps the biggest change the world of IT operations will resist as it continues to adopt AIOps. Most organizations today are still process-driven, focusing maniacally on improving, tweaking, and changing the process to get a different result. Tomorrow's AIOps-driven organization will become data-driven, putting that same focus on refining data for better outcomes.

Improving Accuracy by Changing Culture

Becoming a data-driven organization means shifting priorities from process milestones to data-based ones, where data manipulation and governance are critical. It's building an organization where data modeling is as important as product development, and where data drives business outcomes. It's where there's as much focused placed on algorithms as applications. Once this culture is installed, where the focus becomes accuracy, consistency, and context, can an operations team truly trust the data. And this is where AIOps can truly come to life.

Data accuracy isn't the only concern when it comes to AIOps adoption, but it's definitely on the minds of IT managers and infrastructure professionals. Where they once just struggled to find skilled practitioners and leading-edge technology to solve problems, they now must also juggle a focus on data. It's clear that enterprises will need more time to build trust in the relevance and reliability of AIOps recommendations. This also represents an opportunity for AIOps vendors to provide solutions that drive improved accuracy, cleaner data, and greater control. AIOps promises to transform how IT operations is managed and maintained. It's likely to do the same for data.

Hot Topics

The Latest

IT organizations have historically measured success by how quickly they can respond when something goes wrong. The entire discipline of Incident Management has been optimized around mean time to resolution, first-response SLAs and ticket closure rates. But new research suggests that even though this is a well-executed playbook, it's no longer enough to retain customers ...

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Data Matters More Than Ever in AIOps

New study reveals key hurdles on the road to AIOps
Bhanu Singh

Like any new, potentially disruptive technology, artificial intelligence for IT operations, or AIOps has quickly become a trend, and slowly become a reality. It's only been a few years since Gartner coined the term, and yet, 30% of IT teams in large enterprises will roll out AIOps initiatives by 2023. These IT practitioners are still in experimentation mode with artificial intelligence in many cases, and still have concerns about how credible the technology can be. They have concerns over the results of these implementations, and worry about maintaining service availability and uptime during migration.

Because AIOps is still in its infancy, there hasn't been much reporting on what these concerns specifically are. A recent study from OpsRamp targeted these IT managers who have implemented AIOps, and among other data, reports on the primary concerns of this new approach to operations management.

The Devil is in the Data

The report cites data accuracy as the chief concern for IT pros when it comes to AIOps. Two-thirds (67%) of those surveyed revealed it as their top priority. This could be for a variety of reasons, including:

Data Sources: In a world of distributed, hybrid, multi-cloud infrastructure, it's more difficult than ever to capture data on every level of an organization. Different cloud providers report in different ways. And point tools provide analytics across a host of different metrics. It's next to impossible to compare data sources together for a true contextual view of the organization.

Data Quality: Even when that data is captured, these IT teams aren't necessarily sure that it's accurately reflecting the truth about a system. Modern data can be fragmented, hidden, unparsed or too distributed to make sense.

Data Volume: Today's enterprise infrastructure produces an overwhelming amount of metrics on usage, capacity, performance, availability, security, and more. It's easy to get lost in the noise.

Data Consistency: It's impossible to say, under the crushing weight of data today, that IT teams are seeing consistent reporting and results across the organization. But until data is consistent, it can't be actionable.

Data Culture: This is perhaps the biggest change the world of IT operations will resist as it continues to adopt AIOps. Most organizations today are still process-driven, focusing maniacally on improving, tweaking, and changing the process to get a different result. Tomorrow's AIOps-driven organization will become data-driven, putting that same focus on refining data for better outcomes.

Improving Accuracy by Changing Culture

Becoming a data-driven organization means shifting priorities from process milestones to data-based ones, where data manipulation and governance are critical. It's building an organization where data modeling is as important as product development, and where data drives business outcomes. It's where there's as much focused placed on algorithms as applications. Once this culture is installed, where the focus becomes accuracy, consistency, and context, can an operations team truly trust the data. And this is where AIOps can truly come to life.

Data accuracy isn't the only concern when it comes to AIOps adoption, but it's definitely on the minds of IT managers and infrastructure professionals. Where they once just struggled to find skilled practitioners and leading-edge technology to solve problems, they now must also juggle a focus on data. It's clear that enterprises will need more time to build trust in the relevance and reliability of AIOps recommendations. This also represents an opportunity for AIOps vendors to provide solutions that drive improved accuracy, cleaner data, and greater control. AIOps promises to transform how IT operations is managed and maintained. It's likely to do the same for data.

Hot Topics

The Latest

IT organizations have historically measured success by how quickly they can respond when something goes wrong. The entire discipline of Incident Management has been optimized around mean time to resolution, first-response SLAs and ticket closure rates. But new research suggests that even though this is a well-executed playbook, it's no longer enough to retain customers ...

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...