Datadog On-Call Introduced
June 26, 2024
Share this

Datadog announced Datadog On-Call, an on-call experience with observability-enriched paging and seamless incident management workflows.

Datadog On-Call instantly coordinates teams with relevant context for faster issue resolution, better incident control and improved collaboration.

By unifying observability and paging into one seamless platform, Datadog On-Call solves these issues and eliminates the inefficiencies of multiple disjointed tools, allowing engineers to focus on resolving incidents quickly and effectively without the added stress of switching contexts or missing critical information.

“Being on-call is one of the most challenging aspects of an engineer’s job, where redundant service configurations between various tools can lead to brittle, error-prone setups. The general overhead of maintaining on-call schedules and the ambiguity around service and team ownership make it a grueling ordeal, especially during critical times,” said Michael Whetten, VP of Product at Datadog. “Datadog On-Call addresses these pain points with a team-centric design that clarifies ownership, reduces redundancy and minimizes errors. This approach ensures that every team member knows their role and responsibilities, leading to quicker and more effective incident response.”

Datadog On-Call helps DevOps, SRE, Security and IT Operations teams:

- Act Quickly and Stay Informed: Paging with integrated observability and seamless incident management ensures critical insights and data are readily available within a single platform, eliminating the need for context switching.

- Connect with the Tools They Use Every Day: On-Call integrates with a rich ecosystem of third-party monitoring, alerting and service management tools so teams don’t have to learn new workflows or spend resources on training.

- Ensure Clear Service and Team Ownership: Break down knowledge silos and avoid confusion by associating teams with their respective services to simplify configuration, address ownership gaps and ensure the right responders are paged during an alert. Instantly trace upstream and downstream services affected by an outage or issue.

- Implement Intuitive Scheduling and Notifications: Automate scheduling and escalation policies to ensure continuous coverage and timely responses, reducing the burden on individual team members and enhancing overall team coordination.

- Measure On-Call Performance: Rich and customizable analytics measure on-call performance to help ensure system reliability, improve mean-time-to-resolution and optimize the well-being of on-call teams.

Datadog On-Call is in beta now.

Share this

The Latest

July 23, 2024

The rapid rise of generative AI (GenAI) has caught everyone's attention, leaving many to wonder if the technology's impact will live up to the immense hype. A recent survey by Alteryx provides valuable insights into the current state of GenAI adoption, revealing a shift from inflated expectations to tangible value realization across enterprises ... Here are five key takeaways that underscore GenAI's progression from hype to real-world impact ...

July 22, 2024
A defective software update caused what some experts are calling the largest IT outage in history on Friday, July 19. The impact reverberated through multiple industries around the world ...
July 18, 2024

As software development grows more intricate, the challenge for observability engineers tasked with ensuring optimal system performance becomes more daunting. Current methodologies are struggling to keep pace, with the annual Observability Pulse surveys indicating a rise in Mean Time to Remediation (MTTR). According to this survey, only a small fraction of organizations, around 10%, achieve full observability today. Generative AI, however, promises to significantly move the needle ...

July 17, 2024

While nearly all data leaders surveyed are building generative AI applications, most don't believe their data estate is actually prepared to support them, according to the State of Reliable AI report from Monte Carlo Data ...

July 16, 2024

Enterprises are putting a lot of effort into improving the digital employee experience (DEX), which has become essential to both improving organizational performance and attracting and retaining talented workers. But to date, most efforts to deliver outstanding DEX have focused on people working with laptops, PCs, or thin clients. Employees on the frontlines, using mobile devices to handle logistics ... have been largely overlooked ...

July 15, 2024

The average customer-facing incident takes nearly three hours to resolve (175 minutes) while the estimated cost of downtime is $4,537 per minute, meaning each incident can cost nearly $794,000, according to new research from PagerDuty ...

July 12, 2024

In MEAN TIME TO INSIGHT Episode 8, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses AutoCon with the conference founders Scott Robohn and Chris Grundemann ...

July 11, 2024

Numerous vendors and service providers have recently embraced the NaaS concept, yet there is still no industry consensus on its definition or the types of networks it involves. Furthermore, providers have varied in how they define the NaaS service delivery model. I conducted research for a new report, Network as a Service: Understanding the Cloud Consumption Model in Networking, to refine the concept of NaaS and reduce buyer confusion over what it is and how it can offer value ...

July 10, 2024

Containers are a common theme of wasted spend among organizations, according to the State of Cloud Costs 2024 report from Datadog. In fact, 83% of container costs were associated with idle resources ...

July 10, 2024

Companies prefer a mix of on-prem and cloud environments, according to the 2024 Global State of IT Automation Report from Stonebranch. In only one year, hybrid IT usage has doubled from 34% to 68% ...