Skip to main content

APM and Observability: Cutting Through the Confusion — Part 9

Pete Goldin
APMdigest

The story of the evolution of Observability to encompass APM and other IT performance management capabilities would not be complete without discussing the monumental impact of open source.

Start with: APM and Observability - Cutting Through the Confusion - Part 8

Open source is transforming how organizations approach APM and observability by providing vendor neutral standards for collecting and exporting telemetry types, says Mimi Shalash, Observability Advisor at Splunk, a Cisco Company.

Solutions like OpenTelemetry simplify integration across platforms, reduce vendor lock-in, and improve interoperability in complex environments, Shalash continues. Prometheus enhances this approach with robust metrics and alerting, especially systems like Kubernetes. And together these tools enable flexible, cost-effective stacks designed to scale and evolve with modern infrastructure.

“Open source tools like OpenTelemetry and Prometheus are becoming essential building blocks for observability in modern, cloud-native environments,” explains Andreas Grabner, Fellow DevRel and CNCF Ambassador, Dynatrace. “They empower organizations with greater flexibility and standardization in how telemetry data is collected. The broader industry trend is moving toward interoperability and data unification—using open standards for collection while relying on more advanced platforms to contextualize, analyze and act on that data at scale. This hybrid model allows teams to preserve their existing investments in open source while benefiting from automation, AI and enterprise grade observability.”

“The observability space is a prime target for OSS,” Sven Delmas, VP of Research at Mezmo, agrees. “Between dealing with a tech-savvy and curious audience, constant pressure on cost control, and the need for transparency and avoiding vendor lock-in, there has been — and will be — an ever-increasing push to OSS.”

Driving Observability's Evolution

Open source is changing the center of gravity in observability from tools to telemetry, according to Brian Douglas, Head of Ecosystem, Cloud Native Computing Foundation (CNCF). Developers are adopting Prometheus, OpenTelemetry, and Fluent Bit not just because they're free or flexible, but because they represent an open, portable foundation. These tools make it easier to switch vendors, build internal platforms, and innovate on top of shared standards. They're not just part of the observability conversation; they're shaping the future of how observability is defined.

APM is one specific implementation of observability, not its full scope, Douglas continues. It answers questions like, 'Is this app performing within expected parameters?' Observability, in contrast, supports deeper exploration: 'Why did latency spike in a downstream service for certain regions?' Projects like Prometheus and OpenTelemetry enable this broader context by collecting high-dimensional metrics, distributed traces, and logs which gives teams the raw, interoperable data needed to connect the dots.”

Observability supports cross-signal correlation and open-ended investigation, Douglas adds. Rather than focusing solely on applications, it lets teams visualize the full stack, from container runtimes and infrastructure to network topology and business-level SLIs.

  • Prometheus provides robust, flexible metrics, while Cortex scales them across environments.
  • Fluent Bit and Fluentd handle log aggregation and routing across edge and core environments.
  • OpenTelemetry standardizes telemetry collection and enriches it with context, making it easier for tools and teams to interoperate without reinventing the wheel.

“What's important is interoperability,” Douglas of CNCF explains. “With standards like OpenTelemetry and protocols like the Prometheus exposition format, teams can adopt a modular approach: instrument once, analyze anywhere. This lets them use best-in-class components rather than be locked into a monolithic solution. Observability isn't a single tool, it's a strategy backed by open, composable tooling.”

With OpenTelemetry, users can build a composable observability stack where each tool plays to its strengths: one might excel at exploratory debugging, another at automated root cause analysis, and a third at cost-effective long-term storage, Severin Neumann, Head of Community & Developer Relations at Causely, elaborates. This flexibility lets teams get the best outcomes for their specific needs without duplicating instrumentation or locking themselves into a one-size-fits-all solution.

OpenTelemetry: Reshaping APM and Observability

OpenTelemetry, in particular, has already profoundly reshaped the APM market and the broader observability field, explains Juraci Paixão Kröhling, Software Engineer at OllyGarden. “Initially, some established players might have overlooked it, but strong customer demand has made OpenTelemetry support almost table stakes now; it's rare to find a vendor unable to ingest the standard OTLP format.”

OpenTelemetry is an open source standard, framework and suite of tools facilitating the generation, collection, and exporting of telemetry data.

“OpenTelemetry is having a huge impact on the industry with studies showing that nearly half of organizations polled are using OpenTelemetry with another 25-percent-plus looking to adopt in the near term,” says Harald Burose, Director, Product Management, Research & Development – Engineering, OpenText.

Download the EMA Report: Taking Observability to the Next Level - OpenTelemetry’s Emerging Role in IT Performance and Reliability

Kröhling from OllyGarden continues, “I expect vendors will increasingly embrace OpenTelemetry more natively, treating its semantic conventions not just as data points but as first-class citizens for richer understanding. The era of requiring proprietary agents for basic data collection is closing; customers now expect tools not only to handle open formats but to do so meaningfully, respecting the common language defined by standards like OpenTelemetry. This shared foundation allows everyone to cultivate better systems.”

OpenTelemetry provides teams with greater flexibility, standardization, and control over their telemetry data and has become the de facto standard for data ingestion, according to Bahubali Shetti, Senior Director, Product Marketing, Elastic. Whether deployments use standard OTel SDKs, auto-instrumentation, OTel Collectors, or a combination of these, users can avoid vendor lock-in and reduce the need for future retooling.

“OpenTelemetry isn't just shaping the future of observability, it's quickly becoming the standard that modern, scalable systems are built on,” concludes Shalash from Splunk.

Go to: APM and Observability: Cutting Through the Confusion — Part 10, discussing AI's impact on APM and Observability.

Pete Goldin is Editor and Publisher of APMdigest

The Latest

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...

77% of leaders say their teams need AI skills urgently. 64% say their organization plans to train current employees rather than hire new ones. So far, so reasonable. The part that surprised me is who's been put in charge: 34% of those leaders say IT and engineering own the AI skills mandate. Learning and Development or HR own it at 7% of organizations. That's roughly five-to-one in favor of the people who understand the tools, over the people whose actual job is teaching adults how to learn new ones ...

In the ever-evolving digital landscape, enterprises are increasingly focused on enhancing their observability stacks to gain deeper insights into their IT environments. Observability has become a cornerstone of modern IT operations, enabling organizations to monitor, diagnose, and optimize their systems with unprecedented precision. However, a critical piece of the puzzle often goes unnoticed in this transformation: IBM i ...

We just surveyed 300 frontend and mobile engineers across 16 countries, and the finding that keeps sticking with me isn't the one about AI. It's this: 74% of engineering teams rate themselves in the "middle" of the observability maturity scale. Not reactive, not strategic. Stuck in the middle. They have dashboards, they have tracing, they have alerts. And yet when something goes wrong, they still can't tell you why ...

In MEAN TIME TO INSIGHT Episode 25, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses  AI's impact on the Wide Area Network (WAN) ... 

Application performance monitoring (APM) dashboards are only as useful as what they are configured to measure. The default setup covers obvious failure modes such as downtime, error spikes, and latency breaches, but it does not cover everything. Some failures produce no alerts or anomalies. The dashboard stays green while users experience a broken product. Here are six signs that is happening ...

The race to deploy AI is largely over. Most enterprises have entered it. The question now is not whether artificial intelligence is running inside the organization. The question is whether anyone is genuinely responsible for what it does. That is not a technical question. It is a leadership one. And most organizations are not yet structured to answer it honestly ...

A new analysis of 250 real-world queries across common retail tasks, such as product pricing, availability, ratings, shipping and specifications, reveals systemic inefficiency at the heart of web-based AI agents. On average, 97.9% of the data retrieved by agents from live web pages is irrelevant to the query being answered. Specifically, the average page ingested ran nearly 9,000 characters, while the average answer was just 32 characters, resulting in a noise-to-signal ratio of 278:1. Price queries were the most extreme outlier, with noise rates approaching 99.5%. That's not a rounding error. That's a structural problem ...

The enterprises that will define the next decade are not the ones that deployed the most technology. They are the ones who understood what their technology was actually doing. That distinction is not a philosophical point. It is the central operational challenge facing every organization that has spent the last five years modernizing at speed ...

AI is becoming the operating system of the enterprise. It acts as an invisible coordination layer that understands intent, connects systems, and executes work across complex SaaS environments. Previously, employees had to click through multiple systems — CRM, ERP, support tools, collaboration platforms — to complete a single task. Now, instead of navigating each application manually, they can simply state what they need to accomplish ...

APM and Observability: Cutting Through the Confusion — Part 9

Pete Goldin
APMdigest

The story of the evolution of Observability to encompass APM and other IT performance management capabilities would not be complete without discussing the monumental impact of open source.

Start with: APM and Observability - Cutting Through the Confusion - Part 8

Open source is transforming how organizations approach APM and observability by providing vendor neutral standards for collecting and exporting telemetry types, says Mimi Shalash, Observability Advisor at Splunk, a Cisco Company.

Solutions like OpenTelemetry simplify integration across platforms, reduce vendor lock-in, and improve interoperability in complex environments, Shalash continues. Prometheus enhances this approach with robust metrics and alerting, especially systems like Kubernetes. And together these tools enable flexible, cost-effective stacks designed to scale and evolve with modern infrastructure.

“Open source tools like OpenTelemetry and Prometheus are becoming essential building blocks for observability in modern, cloud-native environments,” explains Andreas Grabner, Fellow DevRel and CNCF Ambassador, Dynatrace. “They empower organizations with greater flexibility and standardization in how telemetry data is collected. The broader industry trend is moving toward interoperability and data unification—using open standards for collection while relying on more advanced platforms to contextualize, analyze and act on that data at scale. This hybrid model allows teams to preserve their existing investments in open source while benefiting from automation, AI and enterprise grade observability.”

“The observability space is a prime target for OSS,” Sven Delmas, VP of Research at Mezmo, agrees. “Between dealing with a tech-savvy and curious audience, constant pressure on cost control, and the need for transparency and avoiding vendor lock-in, there has been — and will be — an ever-increasing push to OSS.”

Driving Observability's Evolution

Open source is changing the center of gravity in observability from tools to telemetry, according to Brian Douglas, Head of Ecosystem, Cloud Native Computing Foundation (CNCF). Developers are adopting Prometheus, OpenTelemetry, and Fluent Bit not just because they're free or flexible, but because they represent an open, portable foundation. These tools make it easier to switch vendors, build internal platforms, and innovate on top of shared standards. They're not just part of the observability conversation; they're shaping the future of how observability is defined.

APM is one specific implementation of observability, not its full scope, Douglas continues. It answers questions like, 'Is this app performing within expected parameters?' Observability, in contrast, supports deeper exploration: 'Why did latency spike in a downstream service for certain regions?' Projects like Prometheus and OpenTelemetry enable this broader context by collecting high-dimensional metrics, distributed traces, and logs which gives teams the raw, interoperable data needed to connect the dots.”

Observability supports cross-signal correlation and open-ended investigation, Douglas adds. Rather than focusing solely on applications, it lets teams visualize the full stack, from container runtimes and infrastructure to network topology and business-level SLIs.

  • Prometheus provides robust, flexible metrics, while Cortex scales them across environments.
  • Fluent Bit and Fluentd handle log aggregation and routing across edge and core environments.
  • OpenTelemetry standardizes telemetry collection and enriches it with context, making it easier for tools and teams to interoperate without reinventing the wheel.

“What's important is interoperability,” Douglas of CNCF explains. “With standards like OpenTelemetry and protocols like the Prometheus exposition format, teams can adopt a modular approach: instrument once, analyze anywhere. This lets them use best-in-class components rather than be locked into a monolithic solution. Observability isn't a single tool, it's a strategy backed by open, composable tooling.”

With OpenTelemetry, users can build a composable observability stack where each tool plays to its strengths: one might excel at exploratory debugging, another at automated root cause analysis, and a third at cost-effective long-term storage, Severin Neumann, Head of Community & Developer Relations at Causely, elaborates. This flexibility lets teams get the best outcomes for their specific needs without duplicating instrumentation or locking themselves into a one-size-fits-all solution.

OpenTelemetry: Reshaping APM and Observability

OpenTelemetry, in particular, has already profoundly reshaped the APM market and the broader observability field, explains Juraci Paixão Kröhling, Software Engineer at OllyGarden. “Initially, some established players might have overlooked it, but strong customer demand has made OpenTelemetry support almost table stakes now; it's rare to find a vendor unable to ingest the standard OTLP format.”

OpenTelemetry is an open source standard, framework and suite of tools facilitating the generation, collection, and exporting of telemetry data.

“OpenTelemetry is having a huge impact on the industry with studies showing that nearly half of organizations polled are using OpenTelemetry with another 25-percent-plus looking to adopt in the near term,” says Harald Burose, Director, Product Management, Research & Development – Engineering, OpenText.

Download the EMA Report: Taking Observability to the Next Level - OpenTelemetry’s Emerging Role in IT Performance and Reliability

Kröhling from OllyGarden continues, “I expect vendors will increasingly embrace OpenTelemetry more natively, treating its semantic conventions not just as data points but as first-class citizens for richer understanding. The era of requiring proprietary agents for basic data collection is closing; customers now expect tools not only to handle open formats but to do so meaningfully, respecting the common language defined by standards like OpenTelemetry. This shared foundation allows everyone to cultivate better systems.”

OpenTelemetry provides teams with greater flexibility, standardization, and control over their telemetry data and has become the de facto standard for data ingestion, according to Bahubali Shetti, Senior Director, Product Marketing, Elastic. Whether deployments use standard OTel SDKs, auto-instrumentation, OTel Collectors, or a combination of these, users can avoid vendor lock-in and reduce the need for future retooling.

“OpenTelemetry isn't just shaping the future of observability, it's quickly becoming the standard that modern, scalable systems are built on,” concludes Shalash from Splunk.

Go to: APM and Observability: Cutting Through the Confusion — Part 10, discussing AI's impact on APM and Observability.

Pete Goldin is Editor and Publisher of APMdigest

The Latest

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...

77% of leaders say their teams need AI skills urgently. 64% say their organization plans to train current employees rather than hire new ones. So far, so reasonable. The part that surprised me is who's been put in charge: 34% of those leaders say IT and engineering own the AI skills mandate. Learning and Development or HR own it at 7% of organizations. That's roughly five-to-one in favor of the people who understand the tools, over the people whose actual job is teaching adults how to learn new ones ...

In the ever-evolving digital landscape, enterprises are increasingly focused on enhancing their observability stacks to gain deeper insights into their IT environments. Observability has become a cornerstone of modern IT operations, enabling organizations to monitor, diagnose, and optimize their systems with unprecedented precision. However, a critical piece of the puzzle often goes unnoticed in this transformation: IBM i ...

We just surveyed 300 frontend and mobile engineers across 16 countries, and the finding that keeps sticking with me isn't the one about AI. It's this: 74% of engineering teams rate themselves in the "middle" of the observability maturity scale. Not reactive, not strategic. Stuck in the middle. They have dashboards, they have tracing, they have alerts. And yet when something goes wrong, they still can't tell you why ...

In MEAN TIME TO INSIGHT Episode 25, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses  AI's impact on the Wide Area Network (WAN) ... 

Application performance monitoring (APM) dashboards are only as useful as what they are configured to measure. The default setup covers obvious failure modes such as downtime, error spikes, and latency breaches, but it does not cover everything. Some failures produce no alerts or anomalies. The dashboard stays green while users experience a broken product. Here are six signs that is happening ...

The race to deploy AI is largely over. Most enterprises have entered it. The question now is not whether artificial intelligence is running inside the organization. The question is whether anyone is genuinely responsible for what it does. That is not a technical question. It is a leadership one. And most organizations are not yet structured to answer it honestly ...

A new analysis of 250 real-world queries across common retail tasks, such as product pricing, availability, ratings, shipping and specifications, reveals systemic inefficiency at the heart of web-based AI agents. On average, 97.9% of the data retrieved by agents from live web pages is irrelevant to the query being answered. Specifically, the average page ingested ran nearly 9,000 characters, while the average answer was just 32 characters, resulting in a noise-to-signal ratio of 278:1. Price queries were the most extreme outlier, with noise rates approaching 99.5%. That's not a rounding error. That's a structural problem ...

The enterprises that will define the next decade are not the ones that deployed the most technology. They are the ones who understood what their technology was actually doing. That distinction is not a philosophical point. It is the central operational challenge facing every organization that has spent the last five years modernizing at speed ...

AI is becoming the operating system of the enterprise. It acts as an invisible coordination layer that understands intent, connects systems, and executes work across complex SaaS environments. Previously, employees had to click through multiple systems — CRM, ERP, support tools, collaboration platforms — to complete a single task. Now, instead of navigating each application manually, they can simply state what they need to accomplish ...