Skip to main content

4 Ways Agentic AI Could Transform IT Operations

Eric Johnson
PagerDuty

The next generation of AI is already here. It may have been mere months since organizations adopted generative AI (GenAI), but now there's a new kid on the block and it promises to offer even greater benefits to businesses and IT operations teams in particular. In fact, research reveals that more than half of companies in the US, UK, Australia and Japan have already adopted agentic AI, with most expecting an ROI of over 100%.

The key to success will be to avoid repeating the adoption mistakes of the past and to start small with manageable projects.

A New Era of Productivity

Agentic AI promises another great leap forward.

The machine-based intelligence is capable of working autonomously to achieve pre-determined goals. In the process, it is capable of adjusting to any unseen bumps in the road through reasoning, iterative planning and adaptive problem-solving. While GenAI focuses on creating content and requires human input in the form of prompts, its agentic cousin acts independently with little or no intervention, making decisions based on data and objectives. Unlike GenAI, it continuously learns and adapts.

It's not difficult to see the huge potential here for enhancing business performance, upskilling workers and relieving them from manual toil. Gartner believes that at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028. Last year, the figure was zero.

Use cases are almost limitless in scope.

Take financial services. An AI agent could be set to work continuously, monitoring transactions for anomalies indicative of fraud. It would leverage adaptive learning to tell the difference between legitimate and criminal activity and take actions to remediate the problem, such as blocking a transaction and notifying the customer. All of this minimizes the delays and customer frustration that come with traditional manual reviews, not to mention the risk of fraudulent transactions sneaking through. From the bank's perspective, it frees staff to work on more satisfying and higher-value tasks, mitigating fraud losses and keeping customers happy.

Companies Are Bullish

There are many more examples like this. In healthcare, agentic AI could automate patient scheduling based on doctor availability, patient history and urgency, sending appointment reminders and even predicting potential cancellations. In local government, it could automate the time-consuming, paperwork-heavy process of granting business or construction permits by analyzing submitted documents and referencing regulatory requirements. The impact on citizens, businesses and under-staffed, under-funded local authorities could be immense.

It's no surprise that organizations are so optimistic about the technology. By 2027, 86% of companies expect to have deployed AI agents operationally, with early GenAI adopters leading the way.

4 Ways to Change ITOps

On a more granular level, agentic AI promises to transform IT operations (ITOps). We've already seen how GenAI has helped by intelligently automating alert correlation, root cause analysis, ticketing and reporting, as well as how it can empower and upskill incident responders as they struggle to deal with manual toil and alert overload. Agentic AI offers more, working through problems even if there are multiple steps and disparate tools involved.

Here are four specific use cases:

1. Site Reliability Engineering (SRE): SRE expertise is in short supply, but it's still much needed as application complexity and customer expectations increase. Agentic AI could help engineers fix problems faster by identifying and classifying operational issues, flagging important historical context and suggesting recommended actions, allowing engineers to focus on innovation.

2. Operations insight: Complexity is the enemy of effective IT operations. ITOps teams sometimes struggle to make sense of their environment given the number of tools they have to manage across distributed, hybrid cloud and on-premises systems. An AI agent could analyze data from across this potentially large ecosystem of tools, uncover trends, surface insights and recommend actions for improved decision making.

3. Scheduling: Today's customers expect seamless digital experiences, and if they don't get them, they are more likely than ever to move to a competitor. That makes it especially critical to ensure seamless responder coverage. Agentic AI can help by taking on a painful manual process, pre-empting scheduling and availability conflicts by dynamically adjusting on-call shifts. Not only will this help drive faster incident resolution, it could also reduce operational costs.

4. Incident response: Incidents are a fact of life in a modern, digital-first organization. But as the business impact of such issues increases, the pressure mounts on ITOps to get ahead of system failures. Agentic AI can reduce response times and human error by stepping in to help and proactively identifying anomalies, taking action to resolve them and continuously learning from past incidents. As with all of the above examples, this is about freeing up human talent to work on higher-value tasks, which also means happier employees.

Some Lessons Learned

As much as organizations are keen to harness the benefits of agentic AI, they're also aware of repeating the mistakes of the past. Many know they didn't train employees enough on GenAI to truly optimize their use of the technology. That's why nearly two-thirds (61%) are planning organization-wide seminars or training initiatives. Others cite lessons learned, such as insufficient planning, not having well-defined ROI expectations and a failure to put in place the right data infrastructure first.

Operational guidelines and guardrails will also be critically important as organizations rush to embrace a technology that operates autonomously.

Agentic AI promises to help ITOps do critical work better, faster and smarter, but success requires careful planning.

Eric Johnson is Chief Information Officer at PagerDuty

Hot Topics

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

4 Ways Agentic AI Could Transform IT Operations

Eric Johnson
PagerDuty

The next generation of AI is already here. It may have been mere months since organizations adopted generative AI (GenAI), but now there's a new kid on the block and it promises to offer even greater benefits to businesses and IT operations teams in particular. In fact, research reveals that more than half of companies in the US, UK, Australia and Japan have already adopted agentic AI, with most expecting an ROI of over 100%.

The key to success will be to avoid repeating the adoption mistakes of the past and to start small with manageable projects.

A New Era of Productivity

Agentic AI promises another great leap forward.

The machine-based intelligence is capable of working autonomously to achieve pre-determined goals. In the process, it is capable of adjusting to any unseen bumps in the road through reasoning, iterative planning and adaptive problem-solving. While GenAI focuses on creating content and requires human input in the form of prompts, its agentic cousin acts independently with little or no intervention, making decisions based on data and objectives. Unlike GenAI, it continuously learns and adapts.

It's not difficult to see the huge potential here for enhancing business performance, upskilling workers and relieving them from manual toil. Gartner believes that at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028. Last year, the figure was zero.

Use cases are almost limitless in scope.

Take financial services. An AI agent could be set to work continuously, monitoring transactions for anomalies indicative of fraud. It would leverage adaptive learning to tell the difference between legitimate and criminal activity and take actions to remediate the problem, such as blocking a transaction and notifying the customer. All of this minimizes the delays and customer frustration that come with traditional manual reviews, not to mention the risk of fraudulent transactions sneaking through. From the bank's perspective, it frees staff to work on more satisfying and higher-value tasks, mitigating fraud losses and keeping customers happy.

Companies Are Bullish

There are many more examples like this. In healthcare, agentic AI could automate patient scheduling based on doctor availability, patient history and urgency, sending appointment reminders and even predicting potential cancellations. In local government, it could automate the time-consuming, paperwork-heavy process of granting business or construction permits by analyzing submitted documents and referencing regulatory requirements. The impact on citizens, businesses and under-staffed, under-funded local authorities could be immense.

It's no surprise that organizations are so optimistic about the technology. By 2027, 86% of companies expect to have deployed AI agents operationally, with early GenAI adopters leading the way.

4 Ways to Change ITOps

On a more granular level, agentic AI promises to transform IT operations (ITOps). We've already seen how GenAI has helped by intelligently automating alert correlation, root cause analysis, ticketing and reporting, as well as how it can empower and upskill incident responders as they struggle to deal with manual toil and alert overload. Agentic AI offers more, working through problems even if there are multiple steps and disparate tools involved.

Here are four specific use cases:

1. Site Reliability Engineering (SRE): SRE expertise is in short supply, but it's still much needed as application complexity and customer expectations increase. Agentic AI could help engineers fix problems faster by identifying and classifying operational issues, flagging important historical context and suggesting recommended actions, allowing engineers to focus on innovation.

2. Operations insight: Complexity is the enemy of effective IT operations. ITOps teams sometimes struggle to make sense of their environment given the number of tools they have to manage across distributed, hybrid cloud and on-premises systems. An AI agent could analyze data from across this potentially large ecosystem of tools, uncover trends, surface insights and recommend actions for improved decision making.

3. Scheduling: Today's customers expect seamless digital experiences, and if they don't get them, they are more likely than ever to move to a competitor. That makes it especially critical to ensure seamless responder coverage. Agentic AI can help by taking on a painful manual process, pre-empting scheduling and availability conflicts by dynamically adjusting on-call shifts. Not only will this help drive faster incident resolution, it could also reduce operational costs.

4. Incident response: Incidents are a fact of life in a modern, digital-first organization. But as the business impact of such issues increases, the pressure mounts on ITOps to get ahead of system failures. Agentic AI can reduce response times and human error by stepping in to help and proactively identifying anomalies, taking action to resolve them and continuously learning from past incidents. As with all of the above examples, this is about freeing up human talent to work on higher-value tasks, which also means happier employees.

Some Lessons Learned

As much as organizations are keen to harness the benefits of agentic AI, they're also aware of repeating the mistakes of the past. Many know they didn't train employees enough on GenAI to truly optimize their use of the technology. That's why nearly two-thirds (61%) are planning organization-wide seminars or training initiatives. Others cite lessons learned, such as insufficient planning, not having well-defined ROI expectations and a failure to put in place the right data infrastructure first.

Operational guidelines and guardrails will also be critically important as organizations rush to embrace a technology that operates autonomously.

Agentic AI promises to help ITOps do critical work better, faster and smarter, but success requires careful planning.

Eric Johnson is Chief Information Officer at PagerDuty

Hot Topics

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...