Skip to main content

APM and Observability: Cutting Through the Confusion — Part 4

Pete Goldin
APMdigest

Observability truly offers a wealth of capabilities that reach far beyond what we traditionally expect from APM, according to Arun Balachandran, Senior Product Marketing Manager, ManageEngine APM Solutions.

While APM excels at meticulously tracking application metrics and promptly alerting us when things go awry, observability empowers our teams to delve much deeper, he continues. It's about enabling us to truly get to the bottom of things, correlating data across logs, traces, and metrics to more easily investigate those unexpected or intricate issues that inevitably arise. This more exploratory approach helps us uncover hidden dependencies, pinpoint subtle performance bottlenecks, and identify potential failure points that might never surface through standard monitoring.

Start with: APM and Observability - Cutting Through the Confusion - Part 3

"APM is still important but it's not enough," adds Gurjeet Arora, CEO and Co-Founder of Observo AI.

The following is a list of several ways experts see Observability surpassing APM or offering capabilities beyond APM:

Broader Coverage

Observability platforms typically provide a much wider lens, encompassing infrastructure, network layers, and even third-party services. This gives us a richer context and a far more complete understanding of how our systems are truly behaving.
Arun Balachandran
Senior Product Marketing Manager, ManageEngine APM Solutions

Observability goes far beyond APM to cover other domains of computing and networking, from routing and switching systems, to compute and storage server infrastructure, to edge systems and endpoints, to orchestration layers and operating systems, to microservices, API gateways, hypervisors and microkernels, to databases and even the data itself.
Peter Corless
Director, Product Marketing, StarTree

Observability should provide a much deeper understanding of system performance within the context of business services and the broader IT environment. Ideally, it should see everything, from cloud-native microservices, containers and VMs, to physical devices and hardware like network devices and laptops.
Douglas James
VP, Solutions & Ecosystem, ScienceLogic

Richer Data Stack

APM generally focuses on specific or predefined metrics for the application layer. It is great at tracking predefined thresholds for known problem areas to ensure specific SLAs and health. Observability can do all that; however, by using a much richer dataset from all aspects of the tech stack, it can paint a larger picture and support much richer root cause analysis.
Sven Delmas
VP of Research, Mezmo

Observability also provides system-wide context by aggregating data from a variety of sources, including servers, containers, databases, and cloud services. This breadth of visibility extends beyond the application runtime, giving teams a holistic view of system behavior and performance.
Gurjeet Arora
CEO and Co-Founder, Observo AI

Comprehensive View of a System's Internal State

While often used interchangeably, observability goes further than APM. It provides a comprehensive view of a system's internal state through extensive data collection — metrics, logs, and traces — across distributed systems.
Varma Kunaparaju
SVP and GM for Cloud Platform and OpsRamp Software, HPE

Application Performance Monitoring (APM), has a narrower focus, concentrating on user experience (UX) by tracking metrics like latency, error rates, and throughput. It's valuable for surfacing known issues through dashboards and alerts, particularly in traditional environments. Observability, on the other hand, is about the practice of understanding a system's internal state based on its external outputs. This broader approach enables teams to investigate, analyze, and resolve issues even when the root cause is not immediately clear. Observability includes multiple tools and technologies such as application monitoring, log management, tracing, telemetry pipelines, and anomaly detection. Where APM might indicate that latency has spiked, observability helps explain why, offering a deeper view into the system's health and behaviors across the stack.
Ajay Khanna
CMO, Yugabyte

Deeper Insights into Complex Distributed Systems

APM has traditionally been rooted in performance monitoring, with a clear focus on application-level metrics and diagnostics, while observability represents a more expansive and adaptive approach that is aimed at providing deeper insights into complex, distributed systems.
Arun Balachandran
Senior Product Marketing Manager, ManageEngine APM Solutions

The killer innovation of the earliest APM tools was really in design. APM tools were the first tools to bring opinionated workflows to working with production monitoring data. But as systems become more complex, it's necessary to expand beyond the capabilities of traditional APM tools. The core APM workflows are still an incredible starting point for system understanding, as long as we don't mistake them for full observability, which can address a breadth of use cases where traditional monitoring falls short. We may even be able to re-use or re-implement these core workflows in observability tools in many cases.
Emily Nakashima
VP of Engineering, Honeycomb

Understanding How Everything Connects

Observability provides a holistic view across the entire infrastructure stack, enabling teams to understand complex interactions between components that might not be visible through an APM-only lens.
Paul Appleby
CEO, Virtana

APM is a subset of observability. It provides valuable insights into application behavior, typically through metrics and transaction traces. But observability not only encompasses performance, it includes understanding availability, dependencies, infrastructure health, security signals, and system-wide behavior. It's not just about measuring how fast something is; it's about making sense of how everything connects.
Gurjeet Arora
CEO and Co-Founder, Observo AI

Identifying Unknown Unknowns

Observability isn't just APM with more data types, it's about enabling teams to ask arbitrary questions about system behavior, not just track predefined KPIs. It's a shift from monitoring known issues to exploring unknown unknowns. True observability requires a mindset (and architecture) focused on correlation, context, and flexibility, rather than dashboards alone.
Brian Douglas
Head of Ecosystem, Cloud Native Computing Foundation (CNCF)

While both disciplines rely on telemetry harvested from systems, they address different needs; APM typically focuses on answering pre-defined questions about application health using curated views, whereas observability equips teams to explore the unknown and ask questions they hadn't anticipated. Think of it as knowing the difference between tending to a specific plant based on its known needs versus analyzing the entire garden's soil health to understand unexpected issues.

Observability extends beyond APM by empowering sophisticated users to investigate unpredicted scenarios within complex systems. Where APM often provides answers to questions we already know to ask (like "What is the latency of this service?" and "What are my slowest DB queries?" observability provides the tools — often rich query languages and flexible visualization platforms — to explore unanticipated behavior, debug novel failure modes, and understand emergent properties of the system as a whole. It's about having the capability to diagnose why a specific section of your garden is behaving unexpectedly, rather than just monitoring the known growth patterns of individual plants.
Juraci Paixão Kröhling
Software Engineer, OllyGarden

The key element that observability should provide is insights into data that help teams identify the unknown unknowns. This is what "should" make observability an upgrade over APM. These concepts matter because unknown failure states are the hardest element of traditional monitoring. For example, traditional monitoring asks how I define failure so as a first step, I will create alerts to look for how I define failure. The alert will fire, and teams will respond. The big issue that observability helps with is how to find failure states I do not know about. When an unknown failure state occurs, a traditional monitor will miss the problem and teams will be late to respond, if at all. The damage from the outage is magnified because of the delay, which centers around the inability to detect unknown failure states.
Ed Bailey
Field CISO, Cribl

High-Cardinality Filtering

Modern observability practices support high-cardinality filtering, which allows teams to minimize noise from unqueried metrics or overly granular tagging. This level of filtering goes beyond the capabilities of most traditional APM tools, enabling more efficient and meaningful analysis.
Gurjeet Arora
CEO and Co-Founder, Observo AI

Getting to the Root Cause of Performance Issues

Application-level performance is a critical lens, especially for user-facing services, but it's only one part of the larger story. Observability enables teams to go beyond metrics and get to the root cause of performance issues by correlating logs, traces, and other machine data across systems. In modern, distributed environments, observability is essential to understanding the context of what's going wrong, not just that something is.
Gurjeet Arora
CEO and Co-Founder, Observo AI

APM and observability serve distinct but complementary roles. APM focuses on ensuring application performance by monitoring key metrics to quickly detect and resolve issues, ensuring a smooth UX. Observability also supports deeper capabilities, like performance troubleshooting, where root cause analysis may rely on fine-grained metrics or trace data unavailable in a traditional APM tool. It supports fine-grained monitoring, such as tracking specific components or behaviors, and adds additional context, which is especially useful for resolving complex, systemic issues or ensuring stability during critical changes like upgrades or migrations. While APM is focused on performance at the application level, observability offers broader, more detailed visibility across the entire system.
Ajay Khanna
CMO, Yugabyte

Delivering Actionable Insights

Observability helps to answer system-wide questions such as, "What's the impact of any state/change/issue on the business in terms of availability, performance, costs, security, and compliance?". Observability answers these questions by turning telemetry into something users can act on, ultimately improving overall system behavior, operational awareness and excellence across teams.
Hugo Kaczmarek
Director of Product, APM Suite, Datadog

Observability transcends mere detection, offering a deeper understanding of the 'why' behind system behaviors. By integrating data from diverse sources, observability facilitates rapid root cause analysis, enhances team collaboration, and empowers organizations to foresee and address potential issues before they affect the customer experience. It's about transforming data into actionable insights. With the addition of GenAI, observability can dynamically orchestrate the collection, analysis, and action phases of IT operations, further enhancing its capabilities.
Gab Menachem
VP ITOM, ServiceNow

Go to: APM and Observability: Cutting Through the Confusion — Part 5, for more insight in how Observability evolved from APM.

Pete Goldin is Editor and Publisher of APMdigest

The Latest

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

APM and Observability: Cutting Through the Confusion — Part 4

Pete Goldin
APMdigest

Observability truly offers a wealth of capabilities that reach far beyond what we traditionally expect from APM, according to Arun Balachandran, Senior Product Marketing Manager, ManageEngine APM Solutions.

While APM excels at meticulously tracking application metrics and promptly alerting us when things go awry, observability empowers our teams to delve much deeper, he continues. It's about enabling us to truly get to the bottom of things, correlating data across logs, traces, and metrics to more easily investigate those unexpected or intricate issues that inevitably arise. This more exploratory approach helps us uncover hidden dependencies, pinpoint subtle performance bottlenecks, and identify potential failure points that might never surface through standard monitoring.

Start with: APM and Observability - Cutting Through the Confusion - Part 3

"APM is still important but it's not enough," adds Gurjeet Arora, CEO and Co-Founder of Observo AI.

The following is a list of several ways experts see Observability surpassing APM or offering capabilities beyond APM:

Broader Coverage

Observability platforms typically provide a much wider lens, encompassing infrastructure, network layers, and even third-party services. This gives us a richer context and a far more complete understanding of how our systems are truly behaving.
Arun Balachandran
Senior Product Marketing Manager, ManageEngine APM Solutions

Observability goes far beyond APM to cover other domains of computing and networking, from routing and switching systems, to compute and storage server infrastructure, to edge systems and endpoints, to orchestration layers and operating systems, to microservices, API gateways, hypervisors and microkernels, to databases and even the data itself.
Peter Corless
Director, Product Marketing, StarTree

Observability should provide a much deeper understanding of system performance within the context of business services and the broader IT environment. Ideally, it should see everything, from cloud-native microservices, containers and VMs, to physical devices and hardware like network devices and laptops.
Douglas James
VP, Solutions & Ecosystem, ScienceLogic

Richer Data Stack

APM generally focuses on specific or predefined metrics for the application layer. It is great at tracking predefined thresholds for known problem areas to ensure specific SLAs and health. Observability can do all that; however, by using a much richer dataset from all aspects of the tech stack, it can paint a larger picture and support much richer root cause analysis.
Sven Delmas
VP of Research, Mezmo

Observability also provides system-wide context by aggregating data from a variety of sources, including servers, containers, databases, and cloud services. This breadth of visibility extends beyond the application runtime, giving teams a holistic view of system behavior and performance.
Gurjeet Arora
CEO and Co-Founder, Observo AI

Comprehensive View of a System's Internal State

While often used interchangeably, observability goes further than APM. It provides a comprehensive view of a system's internal state through extensive data collection — metrics, logs, and traces — across distributed systems.
Varma Kunaparaju
SVP and GM for Cloud Platform and OpsRamp Software, HPE

Application Performance Monitoring (APM), has a narrower focus, concentrating on user experience (UX) by tracking metrics like latency, error rates, and throughput. It's valuable for surfacing known issues through dashboards and alerts, particularly in traditional environments. Observability, on the other hand, is about the practice of understanding a system's internal state based on its external outputs. This broader approach enables teams to investigate, analyze, and resolve issues even when the root cause is not immediately clear. Observability includes multiple tools and technologies such as application monitoring, log management, tracing, telemetry pipelines, and anomaly detection. Where APM might indicate that latency has spiked, observability helps explain why, offering a deeper view into the system's health and behaviors across the stack.
Ajay Khanna
CMO, Yugabyte

Deeper Insights into Complex Distributed Systems

APM has traditionally been rooted in performance monitoring, with a clear focus on application-level metrics and diagnostics, while observability represents a more expansive and adaptive approach that is aimed at providing deeper insights into complex, distributed systems.
Arun Balachandran
Senior Product Marketing Manager, ManageEngine APM Solutions

The killer innovation of the earliest APM tools was really in design. APM tools were the first tools to bring opinionated workflows to working with production monitoring data. But as systems become more complex, it's necessary to expand beyond the capabilities of traditional APM tools. The core APM workflows are still an incredible starting point for system understanding, as long as we don't mistake them for full observability, which can address a breadth of use cases where traditional monitoring falls short. We may even be able to re-use or re-implement these core workflows in observability tools in many cases.
Emily Nakashima
VP of Engineering, Honeycomb

Understanding How Everything Connects

Observability provides a holistic view across the entire infrastructure stack, enabling teams to understand complex interactions between components that might not be visible through an APM-only lens.
Paul Appleby
CEO, Virtana

APM is a subset of observability. It provides valuable insights into application behavior, typically through metrics and transaction traces. But observability not only encompasses performance, it includes understanding availability, dependencies, infrastructure health, security signals, and system-wide behavior. It's not just about measuring how fast something is; it's about making sense of how everything connects.
Gurjeet Arora
CEO and Co-Founder, Observo AI

Identifying Unknown Unknowns

Observability isn't just APM with more data types, it's about enabling teams to ask arbitrary questions about system behavior, not just track predefined KPIs. It's a shift from monitoring known issues to exploring unknown unknowns. True observability requires a mindset (and architecture) focused on correlation, context, and flexibility, rather than dashboards alone.
Brian Douglas
Head of Ecosystem, Cloud Native Computing Foundation (CNCF)

While both disciplines rely on telemetry harvested from systems, they address different needs; APM typically focuses on answering pre-defined questions about application health using curated views, whereas observability equips teams to explore the unknown and ask questions they hadn't anticipated. Think of it as knowing the difference between tending to a specific plant based on its known needs versus analyzing the entire garden's soil health to understand unexpected issues.

Observability extends beyond APM by empowering sophisticated users to investigate unpredicted scenarios within complex systems. Where APM often provides answers to questions we already know to ask (like "What is the latency of this service?" and "What are my slowest DB queries?" observability provides the tools — often rich query languages and flexible visualization platforms — to explore unanticipated behavior, debug novel failure modes, and understand emergent properties of the system as a whole. It's about having the capability to diagnose why a specific section of your garden is behaving unexpectedly, rather than just monitoring the known growth patterns of individual plants.
Juraci Paixão Kröhling
Software Engineer, OllyGarden

The key element that observability should provide is insights into data that help teams identify the unknown unknowns. This is what "should" make observability an upgrade over APM. These concepts matter because unknown failure states are the hardest element of traditional monitoring. For example, traditional monitoring asks how I define failure so as a first step, I will create alerts to look for how I define failure. The alert will fire, and teams will respond. The big issue that observability helps with is how to find failure states I do not know about. When an unknown failure state occurs, a traditional monitor will miss the problem and teams will be late to respond, if at all. The damage from the outage is magnified because of the delay, which centers around the inability to detect unknown failure states.
Ed Bailey
Field CISO, Cribl

High-Cardinality Filtering

Modern observability practices support high-cardinality filtering, which allows teams to minimize noise from unqueried metrics or overly granular tagging. This level of filtering goes beyond the capabilities of most traditional APM tools, enabling more efficient and meaningful analysis.
Gurjeet Arora
CEO and Co-Founder, Observo AI

Getting to the Root Cause of Performance Issues

Application-level performance is a critical lens, especially for user-facing services, but it's only one part of the larger story. Observability enables teams to go beyond metrics and get to the root cause of performance issues by correlating logs, traces, and other machine data across systems. In modern, distributed environments, observability is essential to understanding the context of what's going wrong, not just that something is.
Gurjeet Arora
CEO and Co-Founder, Observo AI

APM and observability serve distinct but complementary roles. APM focuses on ensuring application performance by monitoring key metrics to quickly detect and resolve issues, ensuring a smooth UX. Observability also supports deeper capabilities, like performance troubleshooting, where root cause analysis may rely on fine-grained metrics or trace data unavailable in a traditional APM tool. It supports fine-grained monitoring, such as tracking specific components or behaviors, and adds additional context, which is especially useful for resolving complex, systemic issues or ensuring stability during critical changes like upgrades or migrations. While APM is focused on performance at the application level, observability offers broader, more detailed visibility across the entire system.
Ajay Khanna
CMO, Yugabyte

Delivering Actionable Insights

Observability helps to answer system-wide questions such as, "What's the impact of any state/change/issue on the business in terms of availability, performance, costs, security, and compliance?". Observability answers these questions by turning telemetry into something users can act on, ultimately improving overall system behavior, operational awareness and excellence across teams.
Hugo Kaczmarek
Director of Product, APM Suite, Datadog

Observability transcends mere detection, offering a deeper understanding of the 'why' behind system behaviors. By integrating data from diverse sources, observability facilitates rapid root cause analysis, enhances team collaboration, and empowers organizations to foresee and address potential issues before they affect the customer experience. It's about transforming data into actionable insights. With the addition of GenAI, observability can dynamically orchestrate the collection, analysis, and action phases of IT operations, further enhancing its capabilities.
Gab Menachem
VP ITOM, ServiceNow

Go to: APM and Observability: Cutting Through the Confusion — Part 5, for more insight in how Observability evolved from APM.

Pete Goldin is Editor and Publisher of APMdigest

The Latest

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ...