4 Tips for Dealing with All Those Event Alerts
July 10, 2013

Ariel Gordon

Share this

IT operations handles hundreds, or even thousands, of console messages day in and day out – including weekends. It’s an ongoing 24x7 battle. Data centers keep expanding and increasing in complexity, yet operations is still expected to manage the flood of event alerts pouring in.

Compounding the problem of the sheer volume of events, these alert notifications typically uses technical language that can only be understood by domain experts and come entirely without context.

So, let’s have a look at some tips that will help IT operations personnel deal with all of this by focusing on important events, while understanding their impact on delivery of business services.

1. Add meaning with enrichment rules

Turn cryptic technical messages into meaningful information with text to describe the event including severity prioritization, owner, and if known the service(s) impacted. The illustration below provides an example. This helps to clarify impact of the event alert and provides guidance about the next steps to be taken.

2. Apply correlation rules

Apply correlation rules to help reduce redundant events displayed on the console. Use filtering rules to remove events below a specific impact level – or events that impact less important components such as test servers. It’s also possible to use de-duplication rules to reduce noise related to the same event.

3. Apply tools that define all business service infrastructure components and their interrelationships

Then, you’ll be able to understand the links between IT events and their associated context and impact on business services.

4. Be proactive to understand the impact of changes in the IT infrastructure

It’s a truism in IT that 80 percent of problems originate from changes. Get in front of those event alerts caused by change so you understand “will an upgrade to that problematic switch port take down the customer portal, or does it only affect ordering supplies?” Ensuring safer changes can eliminate many event alerts.

Ariel Gordon is Chief Technology Officer and Co-Founder of Neebula.

Share this

The Latest

July 23, 2024

The rapid rise of generative AI (GenAI) has caught everyone's attention, leaving many to wonder if the technology's impact will live up to the immense hype. A recent survey by Alteryx provides valuable insights into the current state of GenAI adoption, revealing a shift from inflated expectations to tangible value realization across enterprises ... Here are five key takeaways that underscore GenAI's progression from hype to real-world impact ...

July 22, 2024
A defective software update caused what some experts are calling the largest IT outage in history on Friday, July 19. The impact reverberated through multiple industries around the world ...
July 18, 2024

As software development grows more intricate, the challenge for observability engineers tasked with ensuring optimal system performance becomes more daunting. Current methodologies are struggling to keep pace, with the annual Observability Pulse surveys indicating a rise in Mean Time to Remediation (MTTR). According to this survey, only a small fraction of organizations, around 10%, achieve full observability today. Generative AI, however, promises to significantly move the needle ...

July 17, 2024

While nearly all data leaders surveyed are building generative AI applications, most don't believe their data estate is actually prepared to support them, according to the State of Reliable AI report from Monte Carlo Data ...

July 16, 2024

Enterprises are putting a lot of effort into improving the digital employee experience (DEX), which has become essential to both improving organizational performance and attracting and retaining talented workers. But to date, most efforts to deliver outstanding DEX have focused on people working with laptops, PCs, or thin clients. Employees on the frontlines, using mobile devices to handle logistics ... have been largely overlooked ...

July 15, 2024

The average customer-facing incident takes nearly three hours to resolve (175 minutes) while the estimated cost of downtime is $4,537 per minute, meaning each incident can cost nearly $794,000, according to new research from PagerDuty ...

July 12, 2024

In MEAN TIME TO INSIGHT Episode 8, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses AutoCon with the conference founders Scott Robohn and Chris Grundemann ...

July 11, 2024

Numerous vendors and service providers have recently embraced the NaaS concept, yet there is still no industry consensus on its definition or the types of networks it involves. Furthermore, providers have varied in how they define the NaaS service delivery model. I conducted research for a new report, Network as a Service: Understanding the Cloud Consumption Model in Networking, to refine the concept of NaaS and reduce buyer confusion over what it is and how it can offer value ...

July 10, 2024

Containers are a common theme of wasted spend among organizations, according to the State of Cloud Costs 2024 report from Datadog. In fact, 83% of container costs were associated with idle resources ...

July 10, 2024

Companies prefer a mix of on-prem and cloud environments, according to the 2024 Global State of IT Automation Report from Stonebranch. In only one year, hybrid IT usage has doubled from 34% to 68% ...