
PagerDuty launched new capabilities and upgrades for the PagerDuty Operations Cloud.
The new capabilities are critical to enterprises that are modernizing their operations centers, standardizing automation practices, transforming incident management, and automating their remote-location operations. Now teams can take advantage of AI and automation and more powerful end-to-end incident management capabilities to anticipate, identify and resolve operational issues more quickly than ever.
The PagerDuty Operations Cloud combines Incident Management, AIOps, Automation, Customer Service Operations and PagerDuty Copilot (early access) into a flexible, easy-to-use platform designed for mission-critical, time-sensitive, high-impact work across IT, DevOps, security and business teams. The platform is enhanced by APIs that allow organizations to integrate with multiple technology stacks, delivering reliable availability for operational transformation.
“To remain competitive, companies must innovate rapidly and deliver an always-on, immediate digital experience that consumers expect. With the massive amount of noise coming in across teams and tools, it’s challenging to act quickly when your teams are mired in antiquated systems and manual processes, especially at scale,” said Jeffrey Hausman, Chief Product Development Officer at PagerDuty. “The PagerDuty Operations Cloud makes it easy for business and IT leaders to cross the operational chasm by giving them advanced AI and automation capabilities to address some of the most complex, cross-functional processes of enterprise operations, which frees up time and resources to focus on the most mission-critical work to drive their business.”
PagerDuty Copilot (early access) — the generative AI assistant embedded in the PagerDuty Operations Cloud — augments and scales operations teams with AI and automation to manage mission-critical work faster and more effectively. By interpreting the results of automated diagnostics, providing responders with helpful incident context, drafting status updates and generating drafts of post-incident reviews with the click of a button, PagerDuty Copilot allows teams to eliminate time-consuming and repetitive tasks so they can focus on high-priority needs. If the user asks PagerDuty Copilot to generate a post-incident review, it can generate a draft in seconds — reduced from the hours it typically takes.
In addition, PagerDuty Copilot can provide a quick synopsis of the incident through simple prompts, which creates a summary view of the incident. Responders coming into incidents can leverage PagerDuty generative AI to rapidly summarize incident details, Slack notes and customer impact in a moment. PagerDuty Copilot saves responders precious time because they no longer need to spend time collecting dispersed data points and important details.
PagerDuty Operations Console (early access), the latest AIOps offering, serves as a single source of truth on newly created incidents, providing a live, shared view of operational health. Flexible filters ensure issues are immediately discoverable so that network operations center (NOC) and ITOps teams can triage and take action on issues quickly to minimize business impact and protect customer experience. Teams can accelerate triage and resolution using valuable context surfaced directly in a single view, including the impact of an issue, key insights, the ability to run automated diagnostics, recommended actions, and the ability to predict the next likely incident.
PagerDuty Automation helps organizations standardize operations, improve efficiency, and enhance customer experiences by connecting and automating critical work across teams, systems and environments. PagerDuty Workflow Automation allows both developers and non-developers to fully automate complex and manual operations processes — including human steps such as gathering approvals, making decisions or providing updates — and can leverage runbooks in PagerDuty Runbook Automation as part of the process. As a result, teams reduce risks associated with human error and see dramatic improvements in operational efficiencies.
New capabilities for PagerDuty Runbook Automation enable organizations to build, deploy, run and manage automation jobs at scale to standardize automation across the business. Project-based runner management (early access) helps organizations increase the adoption of automation while allowing each team to operate efficiently within their particular technical requirements and dependencies.
PagerDuty Incident Management - an enterprise-grade solution that unites PagerDuty’s industry-leading incident management product with the power of Jeli’s innovative post-incident review capabilities into a single end-to-end offering. PagerDuty empowers organizations to standardize processes with guided remediation and automated workflows directly from Slack, turning every incident into an opportunity to learn and improve. The dynamic narrative builder can drag content from Slack directly into post-incident review. Incident analysis is critical to identify patterns for what happened and why so that teams can adjust processes and avoid repeat issues. This sets organizations up with a more proactive approach to managing incidents that can deliver more resilient operations over time.
The PagerDuty Operations Console is currently in Early Access and will be generally available in Q3 of 2024.
Workflow Automation is generally available.
Runbook Automation’s project-based runner management is currently in Early Access and will be generally available in Q3 of 2024.
PagerDuty Copilot is currently in Early Access and will be generally available in Q3 of 2024.
Enterprise plan for PagerDuty Incident Management, including Jeli Post-Incident Reviews, is generally available.
The Latest
Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...
AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...
Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...
Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...
Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...
In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ...
Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...
Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...
This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...
There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...