Skip to main content

Observability Gaps Are Costing You. Here's How to Fix Them - Fast

Mehdi Daoudi
Catchpoint

It's 7 in the morning. You get an alert from your team. A critical service is down. Yet, your monitoring systems show no critical alerts. Where is the problem? You are considering calling for a war room. It will be a massive distraction for the best people on your team, but what option do you have?

Your next thought may be: why was this not caught with our APM tools? We spend a fortune on them. These incidents are happening in places you cannot see. From global SaaS disruptions to regional ISP failures, to APIs your systems rely on and cloud services in multiple availability zones, the Internet has become a critical extension of enterprise infrastructure. And yet, many teams are still relying on legacy observability strategies that were never built for internet-centric dependencies. The result? Ongoing blind spots, impatient users, and increasing operational costs.

And then there are the micro-outages your team might not even be aware of. Regional incidents, small hiccups, and process failures that are likely happening, often undetected, often unreported and like paper cuts, they cut into user satisfaction, damage the business and erode trust. In fact, our annual Internet Resilience Report found that, in 2025, one in eight businesses now lose over $10 million a month to disruptions and half lose over a million plus a month.

Thus, it's clear a new approach is needed — one that complements APM's detailed view into code, infrastructure, and events with a broad view of the internet stack and more user-centric monitoring. Let's break down how organizations are taking this approach to close the gaps and why the cost of ignoring them is only getting higher.

Why Observability Is Falling Short

Observability tools' main focus is to monitor internal systems: servers, containers, microservices — code traces, metrics, logs and events (MELT). But the modern enterprise is no longer built on custom applications that run on the infrastructure they manage. Cloud apps, SaaS platforms, APIs, and third-party services are now integral to delivering digital experiences. And all of them rely on the health of the Internet: DNS, SSL, BGP, routing, ISPs, etc.

That's where APM alone starts to fail. They were not built to monitor a massively distributed service-oriented multi-party applications. They offer insufficient insight into the routes, external services, internet protocols, and the regional performance that determine whether users can access your app at all.

In fact, what really matters isn't backend-system health; instead, it's real-world user experience. The customer waiting at the rental car counter doesn't care that your servers are humming along at 72% CPU utilization. They care that they need to get to a meeting and the person at the other side of the counter says "Sorry, my computer is slow today". And if you can't tell whether the root cause is your code, your cloud provider, the local internet, DNS resolution times, latency for an API, or a BGP routing issue in some part of the world, you're in trouble without the visibility you need.

APM + IPM = End-to-End Visibility

To solve this, forward-looking enterprises are covering their visibility gap by enhancing the visibility they get from APM tools with Internet Performance Monitoring (IPM). On one side, APM delivers the inside-out view, including instrumentation, tracing, and system health. On the other, IPM offers the outside-in perspective, including real user experience, the health of the global Internet, and proactive testing of everything that may impact a user including first and third party dependencies — from APIs to cloud services to VPNs to database timeouts.

Together, they provide true end-to-end observability, a model is already proving invaluable for global enterprises like SAP, IKEA, and Akamai. APM tools paired with IPM are delivering the unified view of performance that teams need, from the application code to the end user's screen, wherever in the world they are.

With this approach, teams are moving way faster and resolving issues more rapidly and in this way are aligning themselves better to meet business outcomes by making customer experience KPIs the primary objective of observability teams. For instance, they can measure the impact of outages on customer satisfaction and revenue, not just uptime and latency.

The Role of OpenTelemetry

If APM and IPM are the two sides of the observability coin, OpenTelemetry is the glue that binds them. OTel has emerged as the de facto standard for integrating monitoring data, including traces, logs, and metrics, from multiple components of an ecosystem. Its adoption is accelerating because it helps teams break vendor lock-in, standardize data collection, and reduce the cost of managing multiple tools.

In fact, most enterprises now require OTel support as a prerequisite for any observability solution. The best outcomes happen when OpenTelemetry is part of a broader strategy that includes governance, platform selection, and integration with both APM and IPM tools.

As an example, an OTel SDK on a native mobile application could feed telemetry to both APM and IPM systems and both of these could feed a central system with a unified dashboard and/or an alerting or AIOps system. What is possible with OTel is growing and becoming more practical over time.

Centralized Observability Is on the Rise

With greater complexity and greater stakes, enterprises are shifting observability decisions to centralized teams. These groups, sometimes part of architecture, sometimes under operations, are tasked with standardizing vendors, enforcing best practices, and ensuring observability aligns with business needs.

Trend-wise, this is a direct response to tool sprawl and rising costs. According to a recent Elastic survey, many organizations are actively consolidating their observability stacks to improve collaboration and reduce licensing and training expenses.

Centralized observability teams are also the ones most likely to invest in IPM, recognizing that the user's path through the Internet is as important as the path through the code. EMA research recently confirmed this, noting that "Internet Performance Monitoring tools have become just as important as application performance management, if not more so."

Real Results from Modern Observability

Enterprises that embrace this model APM + IPM + OTel, led by a centralized team are already seeing results. They include:

  • Faster time to resolution: By monitoring beyond the firewall, teams spot and diagnose issues quicker.
  • Cost savings: Fewer tools, better data, less duplication.
  • Improved user experience: Outages that used to take hours to triage now take minutes to fix.
  • Greater alignment with business goals: IT teams can tie observability metrics to user impact and revenue risk.

By integrating Internet Performance Monitoring alongside APM, adopting OpenTelemetry for data consistency, and empowering centralized observability teams to lead the way, enterprises can close their performance blind spots and deliver better digital experiences faster and more reliably.

In 2025, observability isn't just about keeping the lights on. It's about creating resilience, reducing cost, and proving the value of IT across the business. And that starts with seeing the whole picture, inside and out.

Mehdi Daoudi is CEO and Co-Founder of Catchpoint

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Observability Gaps Are Costing You. Here's How to Fix Them - Fast

Mehdi Daoudi
Catchpoint

It's 7 in the morning. You get an alert from your team. A critical service is down. Yet, your monitoring systems show no critical alerts. Where is the problem? You are considering calling for a war room. It will be a massive distraction for the best people on your team, but what option do you have?

Your next thought may be: why was this not caught with our APM tools? We spend a fortune on them. These incidents are happening in places you cannot see. From global SaaS disruptions to regional ISP failures, to APIs your systems rely on and cloud services in multiple availability zones, the Internet has become a critical extension of enterprise infrastructure. And yet, many teams are still relying on legacy observability strategies that were never built for internet-centric dependencies. The result? Ongoing blind spots, impatient users, and increasing operational costs.

And then there are the micro-outages your team might not even be aware of. Regional incidents, small hiccups, and process failures that are likely happening, often undetected, often unreported and like paper cuts, they cut into user satisfaction, damage the business and erode trust. In fact, our annual Internet Resilience Report found that, in 2025, one in eight businesses now lose over $10 million a month to disruptions and half lose over a million plus a month.

Thus, it's clear a new approach is needed — one that complements APM's detailed view into code, infrastructure, and events with a broad view of the internet stack and more user-centric monitoring. Let's break down how organizations are taking this approach to close the gaps and why the cost of ignoring them is only getting higher.

Why Observability Is Falling Short

Observability tools' main focus is to monitor internal systems: servers, containers, microservices — code traces, metrics, logs and events (MELT). But the modern enterprise is no longer built on custom applications that run on the infrastructure they manage. Cloud apps, SaaS platforms, APIs, and third-party services are now integral to delivering digital experiences. And all of them rely on the health of the Internet: DNS, SSL, BGP, routing, ISPs, etc.

That's where APM alone starts to fail. They were not built to monitor a massively distributed service-oriented multi-party applications. They offer insufficient insight into the routes, external services, internet protocols, and the regional performance that determine whether users can access your app at all.

In fact, what really matters isn't backend-system health; instead, it's real-world user experience. The customer waiting at the rental car counter doesn't care that your servers are humming along at 72% CPU utilization. They care that they need to get to a meeting and the person at the other side of the counter says "Sorry, my computer is slow today". And if you can't tell whether the root cause is your code, your cloud provider, the local internet, DNS resolution times, latency for an API, or a BGP routing issue in some part of the world, you're in trouble without the visibility you need.

APM + IPM = End-to-End Visibility

To solve this, forward-looking enterprises are covering their visibility gap by enhancing the visibility they get from APM tools with Internet Performance Monitoring (IPM). On one side, APM delivers the inside-out view, including instrumentation, tracing, and system health. On the other, IPM offers the outside-in perspective, including real user experience, the health of the global Internet, and proactive testing of everything that may impact a user including first and third party dependencies — from APIs to cloud services to VPNs to database timeouts.

Together, they provide true end-to-end observability, a model is already proving invaluable for global enterprises like SAP, IKEA, and Akamai. APM tools paired with IPM are delivering the unified view of performance that teams need, from the application code to the end user's screen, wherever in the world they are.

With this approach, teams are moving way faster and resolving issues more rapidly and in this way are aligning themselves better to meet business outcomes by making customer experience KPIs the primary objective of observability teams. For instance, they can measure the impact of outages on customer satisfaction and revenue, not just uptime and latency.

The Role of OpenTelemetry

If APM and IPM are the two sides of the observability coin, OpenTelemetry is the glue that binds them. OTel has emerged as the de facto standard for integrating monitoring data, including traces, logs, and metrics, from multiple components of an ecosystem. Its adoption is accelerating because it helps teams break vendor lock-in, standardize data collection, and reduce the cost of managing multiple tools.

In fact, most enterprises now require OTel support as a prerequisite for any observability solution. The best outcomes happen when OpenTelemetry is part of a broader strategy that includes governance, platform selection, and integration with both APM and IPM tools.

As an example, an OTel SDK on a native mobile application could feed telemetry to both APM and IPM systems and both of these could feed a central system with a unified dashboard and/or an alerting or AIOps system. What is possible with OTel is growing and becoming more practical over time.

Centralized Observability Is on the Rise

With greater complexity and greater stakes, enterprises are shifting observability decisions to centralized teams. These groups, sometimes part of architecture, sometimes under operations, are tasked with standardizing vendors, enforcing best practices, and ensuring observability aligns with business needs.

Trend-wise, this is a direct response to tool sprawl and rising costs. According to a recent Elastic survey, many organizations are actively consolidating their observability stacks to improve collaboration and reduce licensing and training expenses.

Centralized observability teams are also the ones most likely to invest in IPM, recognizing that the user's path through the Internet is as important as the path through the code. EMA research recently confirmed this, noting that "Internet Performance Monitoring tools have become just as important as application performance management, if not more so."

Real Results from Modern Observability

Enterprises that embrace this model APM + IPM + OTel, led by a centralized team are already seeing results. They include:

  • Faster time to resolution: By monitoring beyond the firewall, teams spot and diagnose issues quicker.
  • Cost savings: Fewer tools, better data, less duplication.
  • Improved user experience: Outages that used to take hours to triage now take minutes to fix.
  • Greater alignment with business goals: IT teams can tie observability metrics to user impact and revenue risk.

By integrating Internet Performance Monitoring alongside APM, adopting OpenTelemetry for data consistency, and empowering centralized observability teams to lead the way, enterprises can close their performance blind spots and deliver better digital experiences faster and more reliably.

In 2025, observability isn't just about keeping the lights on. It's about creating resilience, reducing cost, and proving the value of IT across the business. And that starts with seeing the whole picture, inside and out.

Mehdi Daoudi is CEO and Co-Founder of Catchpoint

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...