Skip to main content

Agentic AI, Realistic Expectations and the Future of IT Operations in 2026

Phil Christianson
Xurrent

Agentic AI is a major buzzword for 2026. Many tech companies are making bold promises about this technology, but many aren't grounded in reality, at least not yet. This coming year will likely be shaped by reality checks for IT teams, and progress will only come from a focus on strong foundations and disciplined execution.

Agentic AI Will Be Restricted to Basic IT Tasks in 2026

In 2026, there's going to be a significant gap between what vendors market and what IT leaders actually allow AI agents to do. IT teams are rightfully cautious about this technology because the risk profile of enterprise infrastructure is fundamentally different from other AI use cases.

The blast radius of missteps is very large. Autonomous infrastructure actions, like adding memory or scaling resources, can trigger outages, security issues or cascading failures. These expensive mistakes can actually set implementation back by months.

Plus, many "simple" tasks aren't actually simple. For example, adding memory to a virtual machine may require a restart that introduces downtime, increases cloud costs or violates internal policies tied to licensing or compliance. In many cases, the alert triggering the action isn't even caused by a resource shortage, but by an application issue or upstream dependency. Experienced IT professionals know when not to follow the playbook — context that AI agents don't yet consistently understand.

I had a leader of a mid-size engineering team say to me, "Why would I automate site reliability and infrastructure work? It's what we do better than our competition and the reason we hired some of the best talent available to do it." In other words, he believed that IT operations were part of their competitive advantage and thus should be invested in, not outsourced. While AI may introduce new avenues for automation, it doesn't change basic business principles.

That mindset also highlights the importance of scale. In smaller or mid-sized organizations, an SRE team may spend 5-10 hours a month handling routine code release issues. While AI tools exist that could automate some of that work, they come with real tradeoffs — cost, setup and ongoing oversight. In those cases, eliminating a small amount of effort may not justify the investment, especially if operational excellence is part of the company's differentiation.

However, this equation looks different for Fortune 100 companies with tens of thousands of engineers; the same recurring issues can quickly multiply into hundreds or even thousands of hours. At this scale, automation can become less about optimization at the margins and more about operational necessity. The real challenge for IT leaders is knowing where their teams spend time and applying automation selectively — where it meaningfully supports the business, rather than undermines what makes it competitive.

Most AI Investments in Service Management Will Underperform

Many IT organizations are going to be disappointed with their AI investments in 2026. The disappointment stems from skipping the foundational work.

AI can't clean up a messy knowledge base or fix poorly documented processes. These issues will limit AI's effectiveness while also exposing — or even amplifying — process gaps. The result is friction rather than efficiency.

Some organizations also adopted tools without knowing where their real bottlenecks are (see above). In these scenarios, AI becomes a solution in search of a problem. For example, an organization may deploy an AI chatbot to reduce ticket volume when the real issue is outdated knowledge articles and unclear request workflows. The tool won't solve the underlying problem, making it difficult to demonstrate ROI.

It's also hard to quantify a tool's value when you don't have a baseline. IT leaders must specifically define the problem AI will solve, then measure outcomes, such as ticket deflection rates and mean time to resolution, before and after AI implementation. If you don't know what you're measuring, you can't prove improvements.

Infrastructure Monitoring Will Be a Strategic IT Priority

In 2026, the limits of the traditional IT service desk model will be hard to ignore. The strategy doesn't support modern IT infrastructure. Teams aren't just dealing with laptops, printers and mobile devices that generate tickets when something goes wrong.

Most organizations now run critical systems in the cloud, rely on complex integrations and support applications that sit outside established reporting workflows. A broken server doesn't call the service desk; it just stops working, often in the middle of the night, and the consequences are far more severe than a single employee issue.

This shift to distributed systems forces IT teams to rethink workflows, tooling and escalation models. Adapting is not as simple as implementing continuous monitoring. Poor monitoring creates human exhaustion. When alerts lack context or urgency, on-call teams are forced to respond blindly, often waking up for issues that aren't critical. IT departments need systems that can detect failures, correlate signals across multiple tools and route issues to the right people with the right level of urgency.

Proactive IT support strategies also make the department a strategic partner rather than a cost center. Leaders can use monitoring data to clearly explain system health, risk exposure and potential downstream impact. This perspective informs business decisions grounded in operational reality.

Looking Ahead

IT teams must modernize their systems without losing control. There are plenty of promising AI capabilities emerging, but that doesn't mean the technology is ready to run the show.

In 2026, IT leaders must clean up their processes and workflow, find the bottlenecks and adopt tools that solve a defined problem, not an assumption. This discipline sets teams up for successful implementation. Let's see how agentic AI does on the basic tasks in the year ahead, and maybe 2027 will bring closer alignment between agentic AI marketing promises and operational realities.

Phil Christianson is Chief Product Officer at Xurrent

Hot Topics

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Agentic AI, Realistic Expectations and the Future of IT Operations in 2026

Phil Christianson
Xurrent

Agentic AI is a major buzzword for 2026. Many tech companies are making bold promises about this technology, but many aren't grounded in reality, at least not yet. This coming year will likely be shaped by reality checks for IT teams, and progress will only come from a focus on strong foundations and disciplined execution.

Agentic AI Will Be Restricted to Basic IT Tasks in 2026

In 2026, there's going to be a significant gap between what vendors market and what IT leaders actually allow AI agents to do. IT teams are rightfully cautious about this technology because the risk profile of enterprise infrastructure is fundamentally different from other AI use cases.

The blast radius of missteps is very large. Autonomous infrastructure actions, like adding memory or scaling resources, can trigger outages, security issues or cascading failures. These expensive mistakes can actually set implementation back by months.

Plus, many "simple" tasks aren't actually simple. For example, adding memory to a virtual machine may require a restart that introduces downtime, increases cloud costs or violates internal policies tied to licensing or compliance. In many cases, the alert triggering the action isn't even caused by a resource shortage, but by an application issue or upstream dependency. Experienced IT professionals know when not to follow the playbook — context that AI agents don't yet consistently understand.

I had a leader of a mid-size engineering team say to me, "Why would I automate site reliability and infrastructure work? It's what we do better than our competition and the reason we hired some of the best talent available to do it." In other words, he believed that IT operations were part of their competitive advantage and thus should be invested in, not outsourced. While AI may introduce new avenues for automation, it doesn't change basic business principles.

That mindset also highlights the importance of scale. In smaller or mid-sized organizations, an SRE team may spend 5-10 hours a month handling routine code release issues. While AI tools exist that could automate some of that work, they come with real tradeoffs — cost, setup and ongoing oversight. In those cases, eliminating a small amount of effort may not justify the investment, especially if operational excellence is part of the company's differentiation.

However, this equation looks different for Fortune 100 companies with tens of thousands of engineers; the same recurring issues can quickly multiply into hundreds or even thousands of hours. At this scale, automation can become less about optimization at the margins and more about operational necessity. The real challenge for IT leaders is knowing where their teams spend time and applying automation selectively — where it meaningfully supports the business, rather than undermines what makes it competitive.

Most AI Investments in Service Management Will Underperform

Many IT organizations are going to be disappointed with their AI investments in 2026. The disappointment stems from skipping the foundational work.

AI can't clean up a messy knowledge base or fix poorly documented processes. These issues will limit AI's effectiveness while also exposing — or even amplifying — process gaps. The result is friction rather than efficiency.

Some organizations also adopted tools without knowing where their real bottlenecks are (see above). In these scenarios, AI becomes a solution in search of a problem. For example, an organization may deploy an AI chatbot to reduce ticket volume when the real issue is outdated knowledge articles and unclear request workflows. The tool won't solve the underlying problem, making it difficult to demonstrate ROI.

It's also hard to quantify a tool's value when you don't have a baseline. IT leaders must specifically define the problem AI will solve, then measure outcomes, such as ticket deflection rates and mean time to resolution, before and after AI implementation. If you don't know what you're measuring, you can't prove improvements.

Infrastructure Monitoring Will Be a Strategic IT Priority

In 2026, the limits of the traditional IT service desk model will be hard to ignore. The strategy doesn't support modern IT infrastructure. Teams aren't just dealing with laptops, printers and mobile devices that generate tickets when something goes wrong.

Most organizations now run critical systems in the cloud, rely on complex integrations and support applications that sit outside established reporting workflows. A broken server doesn't call the service desk; it just stops working, often in the middle of the night, and the consequences are far more severe than a single employee issue.

This shift to distributed systems forces IT teams to rethink workflows, tooling and escalation models. Adapting is not as simple as implementing continuous monitoring. Poor monitoring creates human exhaustion. When alerts lack context or urgency, on-call teams are forced to respond blindly, often waking up for issues that aren't critical. IT departments need systems that can detect failures, correlate signals across multiple tools and route issues to the right people with the right level of urgency.

Proactive IT support strategies also make the department a strategic partner rather than a cost center. Leaders can use monitoring data to clearly explain system health, risk exposure and potential downstream impact. This perspective informs business decisions grounded in operational reality.

Looking Ahead

IT teams must modernize their systems without losing control. There are plenty of promising AI capabilities emerging, but that doesn't mean the technology is ready to run the show.

In 2026, IT leaders must clean up their processes and workflow, find the bottlenecks and adopt tools that solve a defined problem, not an assumption. This discipline sets teams up for successful implementation. Let's see how agentic AI does on the basic tasks in the year ahead, and maybe 2027 will bring closer alignment between agentic AI marketing promises and operational realities.

Phil Christianson is Chief Product Officer at Xurrent

Hot Topics

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...