Skip to main content

What Can AIOps Do For IT Ops? - Part 2

APMdigest asked the top minds in the industry what they think AIOps can do for IT Operations. Part 2 covers capabilities supported by AIOps, such as visibility and alerting.

Start with What Can AIOps Do For IT Ops? - Part 1

END-TO-END VISIBILITY

Rapid digitization has made maintaining the performance of key business services across a hybrid cloud environment more imperative than ever. AIOps has quickly become the key technology to provide end-to-end visibility across the technology domains that support business and technology services. AIOps can be used in domain-agnostic settings, applying AI/ML across existing data from monitoring tools, or in a domain-centric way within a specific technology for existing multi-cloud, serverless and container environments. In both instances, the goal is to minimize downtime, speed up code releases and allow developers to get back to writing great software and features rather than spending cycles troubleshooting, and creates opportunities for teams to leverage data-driven decisions to quickly pin-point problems. While many organizations feel that they improved incident response and problem resolution within an area, leveraging proactive AIOps to provide end-to-end visibility across and deep within technology services is a competitive advantage.
Kia Behnia
VP of IT Operations, Splunk

ANOMALY DETECTION

AI (machine learning in particular) is particularly good at recognizing patterns in data — and thus highlights exceptions to those patterns. As a result, AIOps excels at anomaly detection. However, AI drops the ball when the data don't follow clear patterns. The more chaotic the IT environment, the poorer AIOps works.
Jason Bloomberg
President, Intellyx

The anomaly detection from AIOps gives users powerful visualizations comparing the statistics from one period of time compared to another.
Russell Rothstein
Founder and CEO, IT Central Station

PREDICTING INCIDENTS BEFORE THEY BECOME PROBLEMS

Use AI/ML to easily detect evolving incidents across your IT Ops environment as they happen, through the use of cross-domain enrichment
Mohan Kompella, VP Product Marketing,
Adam Blau, Director of Product Marketing,
Anirban Chatterjee, Director of Product Marketing, BigPanda

AIOps is the quintessential analysis of application data coming from the hosting platform, systems logs, Continuous Monitoring tools, CI/CD events, as well as many other sources which can be gradually added as the practice matures (e.g. QA results/metrics output, incidents tracking, etc.) With the help of AI and ML and minimal human intervention, AIOps solutions are able to comb through aggregated data, identify the most relevant pieces of information and correlate them using deterministic analysis. This is the most exciting part of AIOps as it gives way to the true cognitive operations support capable of predicting (and possibly mitigating) incidents before they fully manifested themself in the application's ecosystem.
Oleg Boyko
CTO, Exadel

INTELLIGENT ALERTING

An AIOps tool can predict outcomes and smooth the path for IT operations using intelligent event notification indicating the likely problem within the infrastructure and propose probable solutions. AIOps can reach out to the specific owner or support team providing them with specific information about the issue, suggest remediation, or remediate the problem directly. AIOps tools are transformative for IT operations because they allow operations to concentrate on the business, not on watching screens.
Ron Williams
Analyst, Gigaom

REDUCED ALERT NOISE

AIOps platforms reduce alert noise by grouping interrelated alerts into high-quality incidents that describe to IT Ops teams what the problem is, what's impacted, what the root cause is, what action to take and then kick off automated remediation and resolution workflows.
Mohan Kompella, VP Product Marketing,
Adam Blau, Director of Product Marketing,
Anirban Chatterjee, Director of Product Marketing, BigPanda

AIOps extends value for distributed DevOps teams because of the drastic reduction in alert noise it creates. A majority of alerts currently presented to NOC, IT ops, and SRE teams are deemed irrelevant in mid-to-large sized enterprises. A great deal of toil can be eliminated, freeing up valuable employees to work on the hard, important problems.
Jason English
Principal Analyst, Intellyx

We have seen IT Central Station reviewers using AIOps capabilities for anomaly detection and for troubleshooting, which they say reduces noise when alerts come in.
Russell Rothstein
Founder and CEO, IT Central Station

Now, more than ever, IT teams are overwhelmed by the growing amount of noise from complex infrastructure and applications. Signal noise can impede finding incident root causes in critical moments, lead to longer resolution times, and bring about tedious, manual remediation practices. Teams end up with very little insight into what they need to improve to prevent similar problems from happening in the future. Investing in AIOps helps teams reduce noise from various monitoring tools, find the probable source of significant incidents, and automate as much of the incident response process as possible.
Andrew Marshall
Sr. Director of Product Marketing and Advocacy, PagerDuty

IT Ops teams are drowning in data. The advantage AIOps provides to IT Ops teams is the ability to find a signal through the noise.
Thomas LaRock
Head Geek, SolarWinds

With complicated and multilayered security stacks in play today, companies are finally having to deal with the issue of too many alerts being generated too often, commonly known as "alert fatigue." AIOps technology can easily and effectively reduce alert fatigue by applying a correlation of alerts from several sources and producing actionable alerts that point towards a common root cause. Ultimately, this reduces the noisy background, low priority, or unturned alerts that commonly distracts or hides credible alerts or patterns from security and IT professionals. As a result, there is less chance that they miss a credible alert or threat if AIOps is filtering, correlating, and pattern matching raw alerts into actionable events and alerts.
Chuck Everette
Director of Cybersecurity Advocacy, Deep Instinct

Go to What Can AIOps Do For IT Ops? - Part 3

Hot Topics

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Enterprise IT environments have never been more observable ... Yet many organizations still grapple with outages, lengthy incident resolution cycles, and increasing complexity. Most teams do not suffer from a shortage of data. They struggle to determine what deserves attention and what action to take next ... Enterprise IT operations must move beyond monitoring and visibility. The next stage of maturity is decision operations, an approach that helps teams make faster, better-informed decisions ...

What Can AIOps Do For IT Ops? - Part 2

APMdigest asked the top minds in the industry what they think AIOps can do for IT Operations. Part 2 covers capabilities supported by AIOps, such as visibility and alerting.

Start with What Can AIOps Do For IT Ops? - Part 1

END-TO-END VISIBILITY

Rapid digitization has made maintaining the performance of key business services across a hybrid cloud environment more imperative than ever. AIOps has quickly become the key technology to provide end-to-end visibility across the technology domains that support business and technology services. AIOps can be used in domain-agnostic settings, applying AI/ML across existing data from monitoring tools, or in a domain-centric way within a specific technology for existing multi-cloud, serverless and container environments. In both instances, the goal is to minimize downtime, speed up code releases and allow developers to get back to writing great software and features rather than spending cycles troubleshooting, and creates opportunities for teams to leverage data-driven decisions to quickly pin-point problems. While many organizations feel that they improved incident response and problem resolution within an area, leveraging proactive AIOps to provide end-to-end visibility across and deep within technology services is a competitive advantage.
Kia Behnia
VP of IT Operations, Splunk

ANOMALY DETECTION

AI (machine learning in particular) is particularly good at recognizing patterns in data — and thus highlights exceptions to those patterns. As a result, AIOps excels at anomaly detection. However, AI drops the ball when the data don't follow clear patterns. The more chaotic the IT environment, the poorer AIOps works.
Jason Bloomberg
President, Intellyx

The anomaly detection from AIOps gives users powerful visualizations comparing the statistics from one period of time compared to another.
Russell Rothstein
Founder and CEO, IT Central Station

PREDICTING INCIDENTS BEFORE THEY BECOME PROBLEMS

Use AI/ML to easily detect evolving incidents across your IT Ops environment as they happen, through the use of cross-domain enrichment
Mohan Kompella, VP Product Marketing,
Adam Blau, Director of Product Marketing,
Anirban Chatterjee, Director of Product Marketing, BigPanda

AIOps is the quintessential analysis of application data coming from the hosting platform, systems logs, Continuous Monitoring tools, CI/CD events, as well as many other sources which can be gradually added as the practice matures (e.g. QA results/metrics output, incidents tracking, etc.) With the help of AI and ML and minimal human intervention, AIOps solutions are able to comb through aggregated data, identify the most relevant pieces of information and correlate them using deterministic analysis. This is the most exciting part of AIOps as it gives way to the true cognitive operations support capable of predicting (and possibly mitigating) incidents before they fully manifested themself in the application's ecosystem.
Oleg Boyko
CTO, Exadel

INTELLIGENT ALERTING

An AIOps tool can predict outcomes and smooth the path for IT operations using intelligent event notification indicating the likely problem within the infrastructure and propose probable solutions. AIOps can reach out to the specific owner or support team providing them with specific information about the issue, suggest remediation, or remediate the problem directly. AIOps tools are transformative for IT operations because they allow operations to concentrate on the business, not on watching screens.
Ron Williams
Analyst, Gigaom

REDUCED ALERT NOISE

AIOps platforms reduce alert noise by grouping interrelated alerts into high-quality incidents that describe to IT Ops teams what the problem is, what's impacted, what the root cause is, what action to take and then kick off automated remediation and resolution workflows.
Mohan Kompella, VP Product Marketing,
Adam Blau, Director of Product Marketing,
Anirban Chatterjee, Director of Product Marketing, BigPanda

AIOps extends value for distributed DevOps teams because of the drastic reduction in alert noise it creates. A majority of alerts currently presented to NOC, IT ops, and SRE teams are deemed irrelevant in mid-to-large sized enterprises. A great deal of toil can be eliminated, freeing up valuable employees to work on the hard, important problems.
Jason English
Principal Analyst, Intellyx

We have seen IT Central Station reviewers using AIOps capabilities for anomaly detection and for troubleshooting, which they say reduces noise when alerts come in.
Russell Rothstein
Founder and CEO, IT Central Station

Now, more than ever, IT teams are overwhelmed by the growing amount of noise from complex infrastructure and applications. Signal noise can impede finding incident root causes in critical moments, lead to longer resolution times, and bring about tedious, manual remediation practices. Teams end up with very little insight into what they need to improve to prevent similar problems from happening in the future. Investing in AIOps helps teams reduce noise from various monitoring tools, find the probable source of significant incidents, and automate as much of the incident response process as possible.
Andrew Marshall
Sr. Director of Product Marketing and Advocacy, PagerDuty

IT Ops teams are drowning in data. The advantage AIOps provides to IT Ops teams is the ability to find a signal through the noise.
Thomas LaRock
Head Geek, SolarWinds

With complicated and multilayered security stacks in play today, companies are finally having to deal with the issue of too many alerts being generated too often, commonly known as "alert fatigue." AIOps technology can easily and effectively reduce alert fatigue by applying a correlation of alerts from several sources and producing actionable alerts that point towards a common root cause. Ultimately, this reduces the noisy background, low priority, or unturned alerts that commonly distracts or hides credible alerts or patterns from security and IT professionals. As a result, there is less chance that they miss a credible alert or threat if AIOps is filtering, correlating, and pattern matching raw alerts into actionable events and alerts.
Chuck Everette
Director of Cybersecurity Advocacy, Deep Instinct

Go to What Can AIOps Do For IT Ops? - Part 3

Hot Topics

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Enterprise IT environments have never been more observable ... Yet many organizations still grapple with outages, lengthy incident resolution cycles, and increasing complexity. Most teams do not suffer from a shortage of data. They struggle to determine what deserves attention and what action to take next ... Enterprise IT operations must move beyond monitoring and visibility. The next stage of maturity is decision operations, an approach that helps teams make faster, better-informed decisions ...