Skip to main content

APM and Observability: Cutting Through the Confusion — Part 5

Pete Goldin
APMdigest

Many of the experts see Observability as an evolution of APM, providing even greater visibility.

Start with: APM and Observability - Cutting Through the Confusion - Part 4

"APM remains a cornerstone in the toolkit for application performance management, crucial for pinpointing and resolving application-specific issues. Observability, however, is the evolution of this concept, expanding the scope to encompass distributed systems and cloud environments," explains Gab Menachem, VP ITOM at ServiceNow.

"Based on the current needs for complex, cloud-native systems, many APM tools have evolved into observability platforms," Ajay Khanna, CMO at Yugabyte, agrees. "This is evident in how the Gartner Magic Quadrant name has evolved over the years — from APM to APM & Observability to the current iteration of Observability Platforms."

Observability encompasses a broader spectrum of monitoring and diagnostic capabilities compared to traditional APM tools, Khanna observes. As systems grow in scale and complexity, the value of observability as something broader and more adaptable than APM is becoming clearer. Ultimately, while APM is useful for maintaining performance baselines and triggering alerts, observability provides the depth, flexibility, and adaptability required to manage modern, dynamic systems. It empowers teams not only to detect that something is wrong, but also to understand why it's happening — even in cases where the problem was previously unknown or poorly understood.

Sven Delmas, VP of Research at Mezmo adds, "The boundaries are fluid and changing in these kinds of dynamic systems — APM use morphs into observability; observability implementations draw from the more predefined capabilities/solutions of APM."

In fact, some experts believe that the confusion between APM and Observability is rooted in this evolution. Severin Neumann, Head of Community & Developer Relations at Causely, says, "There is confusion in the market about APM vs. observability, largely because the shift has been more evolutionary than revolutionary. Many observability concepts build on capabilities that APM tools have offered for years behind vendors' closed gardens, like code-level tracing and analytics, just with more flexibility, scale for cloud native systems and with open standards. This overlap blurs the lines, especially as both types of tools adopt similar language and features."

Bringing Everything Together

Some experts see Observability's role as bringing a range of capabilities, including APM, together. Mimi Shalash, Observability Advisor at Splunk, a Cisco Company, explains, "Observability brings once-separate monitoring domains (like application performance monitoring, infrastructure monitoring, digital experience monitoring, AIOps and log analytics) together in order to enable unified visibility and eliminate blind spots. This is especially critical with the shift to cloud native as technology environments become more distributed. A comprehensive observability practice should include all of these components mentioned above and leverage artificial intelligence (AI) and machine learning (ML) to drive earlier detection and faster investigation of business impacting incidents."

Andreas Grabner, Fellow DevRel and CNCF Ambassador, Dynatrace, adds, "Observability provides a broader, real-time view of system health by integrating signals from across the entire stack — not just the application layer. This includes infrastructure telemetry, cloud services, security events, and user behavior data. It enables proactive problem detection, faster root-cause isolation, and more effective collaboration across DevOps, SRE, and business teams."

Another way Observability has evolved from APM is the addition of standardized open source elements. Neumann, from Causely says, "Observability's shift from proprietary APM vendor agents to open, standardized data has significantly expanded the surface of applications that can be instrumented."

Observability Is Essential

Many of the experts see this evolution making Observability essential in today's dynamic distributed enterprise.

Carlos Casanova, Principal Analyst at Forrester, explains, "The features/functions of APM do not go away and are still very much needed. They are just added onto with the investigative capabilities of observability. I personally don't see why organizations would settle for just APM these days when there are so many options to do much more with observability. The tracing, alerting, analytics are all vital elements that observability needs in order to dig deeper and explore without the pre-instrumentation. Observability provides the high cardinality and multi-dimensional support for a system that you don't get with just APM."

"For modern, distributed architectures, observability has become essential," says Brian Douglas, Head of Ecosystem, Cloud Native Computing Foundation (CNCF). "It allows teams to understand relationships between components, identify emergent issues, and trace performance degradations across microservices and infrastructure layers."

CNCF's 2025 Tech Radar validates this shift, with OpenTelemetry and Cortex positioned in the "adopt" tier, reflecting their growing role in powering flexible, telemetry-driven operations.

"Observability is like having a superpower for IT operations," asserts Varma Kunaparaju, SVP and GM for Cloud Platform and OpsRamp Software at HPE. "It dives into logs, metrics, and traces, revealing hidden issues across distributed systems. Designed for microservices and cloud-native architectures, it provides end-to-end tracing and correlates data from multiple sources, offering a holistic view of system behavior."

Observability Limitations

Although Observability is considered essential by the majority of experts, this does not mean the technology always lives up to this endorsement. Some experts outline the issues here:

Lack of True Insight

APM or application performance monitoring is agent-based application instrumentation that measures standard stats such as memory utilization and transaction latency, and is focused on known failure modes. Very little has changed with the rise of observability — the same concepts apply with modest efforts to meet the definition of observability. The fundamental focus of observability is finding insights into your data, centered around the idea of unknowns, which enable the discovery of unknown failure modes. Current observability platforms struggle to deliver the definition of observability and instead are primarily traditional APM with a better UI, that dabble in observability.

Observability is supposed to provide teams with insight into known failure states, but in practice, the ability to provide true insight is limited. Observability has become more about the upsell than delivering actual value.
Ed Bailey
Field CISO, Cribl

Costly Data Volumes

Many vendors claim to offer "observability" but still force customers into the same costly tradeoffs — sampling traces, limiting log retention, or splitting data across multiple tools. Whether you call it APM or observability is irrelevant if the platform can't actually handle modern data volumes economically.
Rakesh Gupta
Head of Product Management, Observe

Go to: APM and Observability: Cutting Through the Confusion - Part 6, covering the differing use cases of APM and Observability.

Pete Goldin is Editor and Publisher of APMdigest

Hot Topics

The Latest

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

APM and Observability: Cutting Through the Confusion — Part 5

Pete Goldin
APMdigest

Many of the experts see Observability as an evolution of APM, providing even greater visibility.

Start with: APM and Observability - Cutting Through the Confusion - Part 4

"APM remains a cornerstone in the toolkit for application performance management, crucial for pinpointing and resolving application-specific issues. Observability, however, is the evolution of this concept, expanding the scope to encompass distributed systems and cloud environments," explains Gab Menachem, VP ITOM at ServiceNow.

"Based on the current needs for complex, cloud-native systems, many APM tools have evolved into observability platforms," Ajay Khanna, CMO at Yugabyte, agrees. "This is evident in how the Gartner Magic Quadrant name has evolved over the years — from APM to APM & Observability to the current iteration of Observability Platforms."

Observability encompasses a broader spectrum of monitoring and diagnostic capabilities compared to traditional APM tools, Khanna observes. As systems grow in scale and complexity, the value of observability as something broader and more adaptable than APM is becoming clearer. Ultimately, while APM is useful for maintaining performance baselines and triggering alerts, observability provides the depth, flexibility, and adaptability required to manage modern, dynamic systems. It empowers teams not only to detect that something is wrong, but also to understand why it's happening — even in cases where the problem was previously unknown or poorly understood.

Sven Delmas, VP of Research at Mezmo adds, "The boundaries are fluid and changing in these kinds of dynamic systems — APM use morphs into observability; observability implementations draw from the more predefined capabilities/solutions of APM."

In fact, some experts believe that the confusion between APM and Observability is rooted in this evolution. Severin Neumann, Head of Community & Developer Relations at Causely, says, "There is confusion in the market about APM vs. observability, largely because the shift has been more evolutionary than revolutionary. Many observability concepts build on capabilities that APM tools have offered for years behind vendors' closed gardens, like code-level tracing and analytics, just with more flexibility, scale for cloud native systems and with open standards. This overlap blurs the lines, especially as both types of tools adopt similar language and features."

Bringing Everything Together

Some experts see Observability's role as bringing a range of capabilities, including APM, together. Mimi Shalash, Observability Advisor at Splunk, a Cisco Company, explains, "Observability brings once-separate monitoring domains (like application performance monitoring, infrastructure monitoring, digital experience monitoring, AIOps and log analytics) together in order to enable unified visibility and eliminate blind spots. This is especially critical with the shift to cloud native as technology environments become more distributed. A comprehensive observability practice should include all of these components mentioned above and leverage artificial intelligence (AI) and machine learning (ML) to drive earlier detection and faster investigation of business impacting incidents."

Andreas Grabner, Fellow DevRel and CNCF Ambassador, Dynatrace, adds, "Observability provides a broader, real-time view of system health by integrating signals from across the entire stack — not just the application layer. This includes infrastructure telemetry, cloud services, security events, and user behavior data. It enables proactive problem detection, faster root-cause isolation, and more effective collaboration across DevOps, SRE, and business teams."

Another way Observability has evolved from APM is the addition of standardized open source elements. Neumann, from Causely says, "Observability's shift from proprietary APM vendor agents to open, standardized data has significantly expanded the surface of applications that can be instrumented."

Observability Is Essential

Many of the experts see this evolution making Observability essential in today's dynamic distributed enterprise.

Carlos Casanova, Principal Analyst at Forrester, explains, "The features/functions of APM do not go away and are still very much needed. They are just added onto with the investigative capabilities of observability. I personally don't see why organizations would settle for just APM these days when there are so many options to do much more with observability. The tracing, alerting, analytics are all vital elements that observability needs in order to dig deeper and explore without the pre-instrumentation. Observability provides the high cardinality and multi-dimensional support for a system that you don't get with just APM."

"For modern, distributed architectures, observability has become essential," says Brian Douglas, Head of Ecosystem, Cloud Native Computing Foundation (CNCF). "It allows teams to understand relationships between components, identify emergent issues, and trace performance degradations across microservices and infrastructure layers."

CNCF's 2025 Tech Radar validates this shift, with OpenTelemetry and Cortex positioned in the "adopt" tier, reflecting their growing role in powering flexible, telemetry-driven operations.

"Observability is like having a superpower for IT operations," asserts Varma Kunaparaju, SVP and GM for Cloud Platform and OpsRamp Software at HPE. "It dives into logs, metrics, and traces, revealing hidden issues across distributed systems. Designed for microservices and cloud-native architectures, it provides end-to-end tracing and correlates data from multiple sources, offering a holistic view of system behavior."

Observability Limitations

Although Observability is considered essential by the majority of experts, this does not mean the technology always lives up to this endorsement. Some experts outline the issues here:

Lack of True Insight

APM or application performance monitoring is agent-based application instrumentation that measures standard stats such as memory utilization and transaction latency, and is focused on known failure modes. Very little has changed with the rise of observability — the same concepts apply with modest efforts to meet the definition of observability. The fundamental focus of observability is finding insights into your data, centered around the idea of unknowns, which enable the discovery of unknown failure modes. Current observability platforms struggle to deliver the definition of observability and instead are primarily traditional APM with a better UI, that dabble in observability.

Observability is supposed to provide teams with insight into known failure states, but in practice, the ability to provide true insight is limited. Observability has become more about the upsell than delivering actual value.
Ed Bailey
Field CISO, Cribl

Costly Data Volumes

Many vendors claim to offer "observability" but still force customers into the same costly tradeoffs — sampling traces, limiting log retention, or splitting data across multiple tools. Whether you call it APM or observability is irrelevant if the platform can't actually handle modern data volumes economically.
Rakesh Gupta
Head of Product Management, Observe

Go to: APM and Observability: Cutting Through the Confusion - Part 6, covering the differing use cases of APM and Observability.

Pete Goldin is Editor and Publisher of APMdigest

Hot Topics

The Latest

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ...