The New Normal for IT Ops Deepens Need for AI - Part 1
May 05, 2020

Will Cappelli
Moogsoft

Share this

The global pandemic has radically changed how enterprise IT services are consumed, both in the short and long term. Here's how AIOps can help IT Ops teams.

The current crisis has upended all aspects of our personal and work lives, and IT Ops pros aren't the exception. The abrupt shift to remote work has created unprecedented challenges for IT Ops teams, while increasing pressure on them to prevent outages and provide service assurance.

Specifically, new consumption patterns of enterprise IT services have put stress on systems, architectures and topologies at all stack layers. In response, IT Ops teams must rapidly implement structural and management changes to address both temporary and permanent shifts.

In this turmoil, AIOps has emerged as a lifeline. By streamlining and automating IT operations, AIOps helps IT leaders collaborate remotely and act quickly and precisely to maintain business-critical digital services — during the pandemic and beyond.

Let's look in more detail at these challenges and at how AIOps can help IT Ops teams cope and succeed.

AIOps: A Definition

An AIOps solution must have these five types of algorithms that fully automate and streamline five key dimensions of IT operations monitoring:

■ Data selection: Identifying and surfacing the most relevant information.

■ Pattern discovery: Correlating and finding relationships between events across your tool stack.

■ Inference: Identifying root causes and recurring issues.

■ Collaboration: Notifying appropriate operators, and facilitating collaboration.

■ Automation: Automating remediation

In a real world setting, an AIOps solution ingests heterogeneous data from many different sources. Using entropy algorithms, it removes noise and duplication, and selects only the truly relevant data. It then groups and correlates this relevant information using various criteria, like text, time and topology.

Next, it discovers patterns in the data, and infers which data items signify causes, and which signify events. It then communicates the result of that analysis to a collaborative environment, which will support automated responses to what has been discovered.

As such, an AIOps solution plays the role of organizing and integrating what an organization's domain-specific IT monitoring and management tools do, intelligently integrating the stack's functionalities. AIOps should act as the brain that brings together these tools, and becomes a coordinating, central layer.

Transitioning to the New Normal

As the workforce shifts to remote work, user behaviors will change and different elements of the IT infrastructure, both in-house and publicly sourced, will be stressed. This will result in new, quickly-evolving types of incidents and outages. With AIOps, IT Ops teams can detect and analyze genuinely novel anomalies which can cause incidents and outages rapidly and stealthily.

Cross-regional and intra-regional team collaboration among IT operations and NOC organizations will need to be reinforced virtually as the implicit supports derived from physical co-presence are removed. AIOps can enable and guide virtual collaborative observation, analysis and response efforts, helping IT Ops teams collaborate and communicate despite being physically dispersed.

Sharp and unpredictable levels of staff reduction due to illness and self-isolation will force IT operations and NOC organizations to "do more with less" on both the side of signal observation and the side of signal response. Here again AIOps can help IT Ops teams to respond by both dynamically filtering noisy alert streams, and integrating and automating platforms that support various aspects of incident and problem management.

Go to The New Normal for IT Ops Deepens Need for AI - Part 2

Will Cappelli is Field CTO at Moogsoft
Share this

The Latest

September 25, 2023

A long-running study of DevOps practices ... suggests that any historical gains in MTTR reduction have now plateaued. For years now, the time it takes to restore services has stayed about the same: less than a day for high performers but up to a week for middle-tier teams and up to a month for laggards. The fact that progress is flat despite big investments in people, tools and automation is a cause for concern ...

September 21, 2023

Companies implementing observability benefit from increased operational efficiency, faster innovation, and better business outcomes overall, according to 2023 IT Trends Report: Lessons From Observability Leaders, a report from SolarWinds ...

September 20, 2023

IT leaders are driving an increasing number of automation initiatives as a way to stay competitive, reduce costs and scale as they navigate an unpredictable social and economic environment, according to the 2023 State of Automation in IT survey conducted by Jitterbit ...

September 19, 2023

Customer loyalty is changing as retailers get increasingly competitive. More than 75% of consumers say they would end business with a company after a single bad customer experience. This means that just one price discrepancy, inventory mishap or checkout issue in a physical or digital store, could have customers running out to the next store that can provide them with better service. Retailers must be able to predict business outages in advance, and act proactively before an incident occurs, impacting customer experience ...

September 18, 2023
Digital transformation is key to ensuring companies keep up with the competitive market landscape. Putting digital at the core of a business can significantly reduce operating expenses and inefficiencies. However, this process often means changing the way internal teams work with one another. To help with the transition, this blog offers chief experience officers (CXOs) advice on how to lead a successful digital transformation project ...
September 14, 2023

Earlier this year, New Relic conducted a study on observability ... The 2023 Observability Forecast reveals observability's impact on the lives of technical professionals and businesses' bottom lines. Here are 10 key takeaways from the forecast ...

September 13, 2023
On September 10, MGM Resorts experienced what it called a "cybersecurity issue" that had a major impact on the company's systems, showing how cyberattacks can bring down applications, ultimately causing problems for a company in many ways ...
September 12, 2023

Only 33% of executives are "very confident" in their ability to operate in a public cloud environment, according to the 2023 State of CloudOps report from NetApp. This represents an increase from 2022 when only 21% reported feeling very confident ...

September 11, 2023

The majority of organizations across Australia and New Zealand (A/NZ) breached over the last year had personally identifiable information (PII) compromised, but most have not yet modified their data management policies, according to the Cybersecurity and PII Report from ManageEngine ...

September 07, 2023

A large majority of organizations employ more than one cloud automation solution, and this practice creates significant challenges that are resulting in delays and added costs for businesses, according to Why companies lose efficiency and compliance with cloud automation solutions from Broadcom ...