Skip to main content

4 Ways Agentic AI Could Transform IT Operations

Eric Johnson
PagerDuty

The next generation of AI is already here. It may have been mere months since organizations adopted generative AI (GenAI), but now there's a new kid on the block and it promises to offer even greater benefits to businesses and IT operations teams in particular. In fact, research reveals that more than half of companies in the US, UK, Australia and Japan have already adopted agentic AI, with most expecting an ROI of over 100%.

The key to success will be to avoid repeating the adoption mistakes of the past and to start small with manageable projects.

A New Era of Productivity

Agentic AI promises another great leap forward.

The machine-based intelligence is capable of working autonomously to achieve pre-determined goals. In the process, it is capable of adjusting to any unseen bumps in the road through reasoning, iterative planning and adaptive problem-solving. While GenAI focuses on creating content and requires human input in the form of prompts, its agentic cousin acts independently with little or no intervention, making decisions based on data and objectives. Unlike GenAI, it continuously learns and adapts.

It's not difficult to see the huge potential here for enhancing business performance, upskilling workers and relieving them from manual toil. Gartner believes that at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028. Last year, the figure was zero.

Use cases are almost limitless in scope.

Take financial services. An AI agent could be set to work continuously, monitoring transactions for anomalies indicative of fraud. It would leverage adaptive learning to tell the difference between legitimate and criminal activity and take actions to remediate the problem, such as blocking a transaction and notifying the customer. All of this minimizes the delays and customer frustration that come with traditional manual reviews, not to mention the risk of fraudulent transactions sneaking through. From the bank's perspective, it frees staff to work on more satisfying and higher-value tasks, mitigating fraud losses and keeping customers happy.

Companies Are Bullish

There are many more examples like this. In healthcare, agentic AI could automate patient scheduling based on doctor availability, patient history and urgency, sending appointment reminders and even predicting potential cancellations. In local government, it could automate the time-consuming, paperwork-heavy process of granting business or construction permits by analyzing submitted documents and referencing regulatory requirements. The impact on citizens, businesses and under-staffed, under-funded local authorities could be immense.

It's no surprise that organizations are so optimistic about the technology. By 2027, 86% of companies expect to have deployed AI agents operationally, with early GenAI adopters leading the way.

4 Ways to Change ITOps

On a more granular level, agentic AI promises to transform IT operations (ITOps). We've already seen how GenAI has helped by intelligently automating alert correlation, root cause analysis, ticketing and reporting, as well as how it can empower and upskill incident responders as they struggle to deal with manual toil and alert overload. Agentic AI offers more, working through problems even if there are multiple steps and disparate tools involved.

Here are four specific use cases:

1. Site Reliability Engineering (SRE): SRE expertise is in short supply, but it's still much needed as application complexity and customer expectations increase. Agentic AI could help engineers fix problems faster by identifying and classifying operational issues, flagging important historical context and suggesting recommended actions, allowing engineers to focus on innovation.

2. Operations insight: Complexity is the enemy of effective IT operations. ITOps teams sometimes struggle to make sense of their environment given the number of tools they have to manage across distributed, hybrid cloud and on-premises systems. An AI agent could analyze data from across this potentially large ecosystem of tools, uncover trends, surface insights and recommend actions for improved decision making.

3. Scheduling: Today's customers expect seamless digital experiences, and if they don't get them, they are more likely than ever to move to a competitor. That makes it especially critical to ensure seamless responder coverage. Agentic AI can help by taking on a painful manual process, pre-empting scheduling and availability conflicts by dynamically adjusting on-call shifts. Not only will this help drive faster incident resolution, it could also reduce operational costs.

4. Incident response: Incidents are a fact of life in a modern, digital-first organization. But as the business impact of such issues increases, the pressure mounts on ITOps to get ahead of system failures. Agentic AI can reduce response times and human error by stepping in to help and proactively identifying anomalies, taking action to resolve them and continuously learning from past incidents. As with all of the above examples, this is about freeing up human talent to work on higher-value tasks, which also means happier employees.

Some Lessons Learned

As much as organizations are keen to harness the benefits of agentic AI, they're also aware of repeating the mistakes of the past. Many know they didn't train employees enough on GenAI to truly optimize their use of the technology. That's why nearly two-thirds (61%) are planning organization-wide seminars or training initiatives. Others cite lessons learned, such as insufficient planning, not having well-defined ROI expectations and a failure to put in place the right data infrastructure first.

Operational guidelines and guardrails will also be critically important as organizations rush to embrace a technology that operates autonomously.

Agentic AI promises to help ITOps do critical work better, faster and smarter, but success requires careful planning.

Eric Johnson is Chief Information Officer at PagerDuty

Hot Topics

The Latest

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

4 Ways Agentic AI Could Transform IT Operations

Eric Johnson
PagerDuty

The next generation of AI is already here. It may have been mere months since organizations adopted generative AI (GenAI), but now there's a new kid on the block and it promises to offer even greater benefits to businesses and IT operations teams in particular. In fact, research reveals that more than half of companies in the US, UK, Australia and Japan have already adopted agentic AI, with most expecting an ROI of over 100%.

The key to success will be to avoid repeating the adoption mistakes of the past and to start small with manageable projects.

A New Era of Productivity

Agentic AI promises another great leap forward.

The machine-based intelligence is capable of working autonomously to achieve pre-determined goals. In the process, it is capable of adjusting to any unseen bumps in the road through reasoning, iterative planning and adaptive problem-solving. While GenAI focuses on creating content and requires human input in the form of prompts, its agentic cousin acts independently with little or no intervention, making decisions based on data and objectives. Unlike GenAI, it continuously learns and adapts.

It's not difficult to see the huge potential here for enhancing business performance, upskilling workers and relieving them from manual toil. Gartner believes that at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028. Last year, the figure was zero.

Use cases are almost limitless in scope.

Take financial services. An AI agent could be set to work continuously, monitoring transactions for anomalies indicative of fraud. It would leverage adaptive learning to tell the difference between legitimate and criminal activity and take actions to remediate the problem, such as blocking a transaction and notifying the customer. All of this minimizes the delays and customer frustration that come with traditional manual reviews, not to mention the risk of fraudulent transactions sneaking through. From the bank's perspective, it frees staff to work on more satisfying and higher-value tasks, mitigating fraud losses and keeping customers happy.

Companies Are Bullish

There are many more examples like this. In healthcare, agentic AI could automate patient scheduling based on doctor availability, patient history and urgency, sending appointment reminders and even predicting potential cancellations. In local government, it could automate the time-consuming, paperwork-heavy process of granting business or construction permits by analyzing submitted documents and referencing regulatory requirements. The impact on citizens, businesses and under-staffed, under-funded local authorities could be immense.

It's no surprise that organizations are so optimistic about the technology. By 2027, 86% of companies expect to have deployed AI agents operationally, with early GenAI adopters leading the way.

4 Ways to Change ITOps

On a more granular level, agentic AI promises to transform IT operations (ITOps). We've already seen how GenAI has helped by intelligently automating alert correlation, root cause analysis, ticketing and reporting, as well as how it can empower and upskill incident responders as they struggle to deal with manual toil and alert overload. Agentic AI offers more, working through problems even if there are multiple steps and disparate tools involved.

Here are four specific use cases:

1. Site Reliability Engineering (SRE): SRE expertise is in short supply, but it's still much needed as application complexity and customer expectations increase. Agentic AI could help engineers fix problems faster by identifying and classifying operational issues, flagging important historical context and suggesting recommended actions, allowing engineers to focus on innovation.

2. Operations insight: Complexity is the enemy of effective IT operations. ITOps teams sometimes struggle to make sense of their environment given the number of tools they have to manage across distributed, hybrid cloud and on-premises systems. An AI agent could analyze data from across this potentially large ecosystem of tools, uncover trends, surface insights and recommend actions for improved decision making.

3. Scheduling: Today's customers expect seamless digital experiences, and if they don't get them, they are more likely than ever to move to a competitor. That makes it especially critical to ensure seamless responder coverage. Agentic AI can help by taking on a painful manual process, pre-empting scheduling and availability conflicts by dynamically adjusting on-call shifts. Not only will this help drive faster incident resolution, it could also reduce operational costs.

4. Incident response: Incidents are a fact of life in a modern, digital-first organization. But as the business impact of such issues increases, the pressure mounts on ITOps to get ahead of system failures. Agentic AI can reduce response times and human error by stepping in to help and proactively identifying anomalies, taking action to resolve them and continuously learning from past incidents. As with all of the above examples, this is about freeing up human talent to work on higher-value tasks, which also means happier employees.

Some Lessons Learned

As much as organizations are keen to harness the benefits of agentic AI, they're also aware of repeating the mistakes of the past. Many know they didn't train employees enough on GenAI to truly optimize their use of the technology. That's why nearly two-thirds (61%) are planning organization-wide seminars or training initiatives. Others cite lessons learned, such as insufficient planning, not having well-defined ROI expectations and a failure to put in place the right data infrastructure first.

Operational guidelines and guardrails will also be critically important as organizations rush to embrace a technology that operates autonomously.

Agentic AI promises to help ITOps do critical work better, faster and smarter, but success requires careful planning.

Eric Johnson is Chief Information Officer at PagerDuty

Hot Topics

The Latest

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ...