Skip to main content

When the Observability Bill Arrives for Agentic AI

Andi Mann
Apica

I have been building enterprise software for more than 20 years, at BMC Software, CA Technologies, Splunk, now Apica. One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now.

New research from Omdia, surveying 300+ enterprise IT decision-makers across North America and Western Europe, puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here.

The Volume Numbers Are Not Projections

54% of enterprises saw their telemetry data volume triple in the past 12 months. Not projected. Not modeled. Running right now, in the environments they operate today. AI and machine learning workloads account for roughly 43% of that growth, making AI the single largest driver of telemetry volume in enterprise environments.

And this isn't the peak. Keir Walker, Senior Market Research Analyst at Omdia, said it directly: "This is just the beginning of the growth curve. We are at the beginning of the hockey stick."

Survey respondents expect an average 9.5x increase in telemetry from agentic AI workloads within two years. I'd pay less attention to that average than to the spread: 44% of respondents expect growth somewhere between 6x and 100x. When your most experienced enterprise IT leaders can't put a ceiling on expected data volume, the architecture has to be built for extremes, not averages.

The Cost Surprise Isn't Compute

I have been in technology long enough to know that cost surprises don't come from where you expect them. For agentic AI, the surprise isn't compute. It isn't talent. It's observability.

In 69% of agentic AI projects, observability costs already exceed compute and infrastructure costs combined. The average enterprise spends $3.17M annually on observability. Nearly 20% spend more than $5M. Budgets are growing 28% year over year on average, and for more than a third of enterprises that growth exceeds 52% annually.

Those numbers have started forcing real decisions. 59% of organizations have terminated or delayed at least one agentic AI deployment because monitoring costs were too high. And the agents getting shelved aren't experimental ones: Cybersecurity, legal and compliance, and fraud detection lead the list. The cost of leaving those agents unmonitored isn't measured in IT budget. It's measured in business risk.

Torsten Volk, Principal Analyst at Omdia, put it plainly: "Scalability is the main reason why people can't have agentic projects. They have no way of deploying them without exposing themselves to operational, legal, and security risk."

That's not an IT problem. That's a C-suite problem.

35% Deployment Is Not What It Sounds Like

The research shows 35% of enterprises claiming widespread agentic AI deployment, which is striking given the category has only existed since 2024. The Omdia analysts were direct about what's behind that number: Competitive pressure, not infrastructure readiness. Every CEO is saying the company has to be agentic. The pipelines underneath those agents weren't built for what the agents actually generate.

The gap is measurable. Organizations unfamiliar with agentic AI are 4.5x less likely to be prepared for the data volumes it produces. The organizations most exposed are the ones least likely to know it.

What the Prepared Organizations Did Differently

The research is consistent here. 97% of enterprises have implemented or are actively evaluating a telemetry pipeline solution, routing, filtering, governing, and contextualizing telemetry before it reaches any downstream observability or analytics platform. That's not an emerging concept, that's the market's answer to the problem.

Pipeline adopters are 50% more likely to be prepared for agentic AI data scale within 24 months. They're 80% more likely to have avoided the operational cost problems that are stalling everyone else. That second number matters. Pipeline adoption isn't correlated with readiness by coincidence. The organizations that scaled successfully built the pipeline layer first.

68% of enterprises plan to evaluate changes to their observability solutions within six months. 70% plan to evaluate pipeline solutions in the same window. Most organizations reading this are already in or entering that cycle.

In 20+ years of enterprise infrastructure, the pattern holds. The organizations that come out ahead in a technology shift are the ones that fix the infrastructure problem before it becomes a budget crisis.

The data from this study is not a warning about a future problem. It's a measurement of a present one.

Andi Mann is Chief Product & Technology Officer at Apica

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

When the Observability Bill Arrives for Agentic AI

Andi Mann
Apica

I have been building enterprise software for more than 20 years, at BMC Software, CA Technologies, Splunk, now Apica. One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now.

New research from Omdia, surveying 300+ enterprise IT decision-makers across North America and Western Europe, puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here.

The Volume Numbers Are Not Projections

54% of enterprises saw their telemetry data volume triple in the past 12 months. Not projected. Not modeled. Running right now, in the environments they operate today. AI and machine learning workloads account for roughly 43% of that growth, making AI the single largest driver of telemetry volume in enterprise environments.

And this isn't the peak. Keir Walker, Senior Market Research Analyst at Omdia, said it directly: "This is just the beginning of the growth curve. We are at the beginning of the hockey stick."

Survey respondents expect an average 9.5x increase in telemetry from agentic AI workloads within two years. I'd pay less attention to that average than to the spread: 44% of respondents expect growth somewhere between 6x and 100x. When your most experienced enterprise IT leaders can't put a ceiling on expected data volume, the architecture has to be built for extremes, not averages.

The Cost Surprise Isn't Compute

I have been in technology long enough to know that cost surprises don't come from where you expect them. For agentic AI, the surprise isn't compute. It isn't talent. It's observability.

In 69% of agentic AI projects, observability costs already exceed compute and infrastructure costs combined. The average enterprise spends $3.17M annually on observability. Nearly 20% spend more than $5M. Budgets are growing 28% year over year on average, and for more than a third of enterprises that growth exceeds 52% annually.

Those numbers have started forcing real decisions. 59% of organizations have terminated or delayed at least one agentic AI deployment because monitoring costs were too high. And the agents getting shelved aren't experimental ones: Cybersecurity, legal and compliance, and fraud detection lead the list. The cost of leaving those agents unmonitored isn't measured in IT budget. It's measured in business risk.

Torsten Volk, Principal Analyst at Omdia, put it plainly: "Scalability is the main reason why people can't have agentic projects. They have no way of deploying them without exposing themselves to operational, legal, and security risk."

That's not an IT problem. That's a C-suite problem.

35% Deployment Is Not What It Sounds Like

The research shows 35% of enterprises claiming widespread agentic AI deployment, which is striking given the category has only existed since 2024. The Omdia analysts were direct about what's behind that number: Competitive pressure, not infrastructure readiness. Every CEO is saying the company has to be agentic. The pipelines underneath those agents weren't built for what the agents actually generate.

The gap is measurable. Organizations unfamiliar with agentic AI are 4.5x less likely to be prepared for the data volumes it produces. The organizations most exposed are the ones least likely to know it.

What the Prepared Organizations Did Differently

The research is consistent here. 97% of enterprises have implemented or are actively evaluating a telemetry pipeline solution, routing, filtering, governing, and contextualizing telemetry before it reaches any downstream observability or analytics platform. That's not an emerging concept, that's the market's answer to the problem.

Pipeline adopters are 50% more likely to be prepared for agentic AI data scale within 24 months. They're 80% more likely to have avoided the operational cost problems that are stalling everyone else. That second number matters. Pipeline adoption isn't correlated with readiness by coincidence. The organizations that scaled successfully built the pipeline layer first.

68% of enterprises plan to evaluate changes to their observability solutions within six months. 70% plan to evaluate pipeline solutions in the same window. Most organizations reading this are already in or entering that cycle.

In 20+ years of enterprise infrastructure, the pattern holds. The organizations that come out ahead in a technology shift are the ones that fix the infrastructure problem before it becomes a budget crisis.

The data from this study is not a warning about a future problem. It's a measurement of a present one.

Andi Mann is Chief Product & Technology Officer at Apica

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...