Skip to main content

APM, Observability and AIOps - a Way Forward

Ron Williams
Gigaom

What's coming in operations management tooling? In a nutshell, a shift from observability to intelligent operations and the longer-term move towards AI-enabled operations in support of the business, but application performance management (APM) still has a place.

Let's break these pieces down. First, APM could be perceived as becoming passé, in tooling terms. All larger companies use it, and tools vendors pull it into their observability suites. Companies still need APM as a starting point if they are unready for the integration heavy lifting, coordination between multiple departments, and political capital that more advanced solutions require.

Many vendors recognize this, selling APM at a reasonable cost with bundled access to other features — but there's a catch. Historically, APM licensing has been based on users, rather than data consumed. But now, vendors are using data as the driving factor for cost. The focus now is on data consumption models: If you're consuming a certain volume of logs, telemetry, and traces, these will drive your cost.

This means less predictability. If someone is temporarily consuming a lot of data, even legitimately (for example, for a new project), they'll have a blip in their billing. In addition, a user can say, "Oh, I can use this feature too," meaning they consume more data, which makes more money for vendors. APM is almost the gateway drug to observability, feature by feature.

Some companies make it easier for you to add another of their little tools because it's convenient. One company has 26 products — if you use one, you can access the others. Suddenly, finance goes, "Wait a minute, why do we suddenly have this big cost increase?" And you have to go back and look and realize, "Oh, George added this one, Sarah used that one, and Sam used the other one, and wow, our bill just quadrupled."

We're also seeing the rise of generative AI in Ops. Predictive AI and machine learning have long been in the mix, but this is the first year that genAI will appear in products. I expect every vendor will offer something related, but the offerings will almost universally be bad. It's not the vendors' fault, but nobody knows what we can, or should be doing with this capability. So vendors will include the feature, whether or not it's useful or really answers the questions businesses have.

For this reason, I'm updating one of my models. Historically, I have shown the evolution from monitoring to observability to awareness. This year, I'll change from monitoring to observability to intelligence. Under "intelligence" I have questions such as:

Is the business OK?

What was the result of last month's marketing campaign?

Sales has a new initiative; what will impact our services and support?

Unless you're in the business of IT, your real questions are not about IT but the business. If you fly people from point A to point B, you want to ask questions about that, not whether the revenue management system is working.

Observability didn't look to answer these questions, but now that we have more intelligence in tools, we must address them. You want to ask your chat interface that connects to your AIOps that question, rather than going over to revenue management and then going over to this group, that group, or the other group, for the answers.

These tools still have the same problems with AI: choosing the right algorithm at the right time, explainable AI, and AI bias — these are not going away. Let's say I train my AI on all my data … stop there, I don't have all my data because, for example, the guys over in desktop support didn't want to give me their data, but the guys over in networking did. I've trained the models on network data, and the AI now knows networking. So, what is every problem going to be? You guessed it, a networking problem.

Being able to train the AI and getting beyond its biases are going to be challenging. Additionally, generative AIs can hallucinate, presenting nonsense data as fact. Trusting AI as we train it to learn our businesses and help us run more efficiently is part of the new paradigm in business operations.

That'll set the scene for 2024: I expect them to have something, but it won't really help. It may be a little more focused in 2025, but by year three and on — that's when I really believe the AI they're putting into some of these tools will be truly useful. That is, it can answer questions about the condition of the enterprise, not the condition of IT.

That's the direction I see the industry taking, and I'm pushing to see how vendors will impact how the entire business operates. In three years, we should see the hype turn into real changes. For now, the nascent large language models show promise; but with planning and focus, generative AI won't be another promise broken.

Ron Williams is an Analyst at Gigaom

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...

APM, Observability and AIOps - a Way Forward

Ron Williams
Gigaom

What's coming in operations management tooling? In a nutshell, a shift from observability to intelligent operations and the longer-term move towards AI-enabled operations in support of the business, but application performance management (APM) still has a place.

Let's break these pieces down. First, APM could be perceived as becoming passé, in tooling terms. All larger companies use it, and tools vendors pull it into their observability suites. Companies still need APM as a starting point if they are unready for the integration heavy lifting, coordination between multiple departments, and political capital that more advanced solutions require.

Many vendors recognize this, selling APM at a reasonable cost with bundled access to other features — but there's a catch. Historically, APM licensing has been based on users, rather than data consumed. But now, vendors are using data as the driving factor for cost. The focus now is on data consumption models: If you're consuming a certain volume of logs, telemetry, and traces, these will drive your cost.

This means less predictability. If someone is temporarily consuming a lot of data, even legitimately (for example, for a new project), they'll have a blip in their billing. In addition, a user can say, "Oh, I can use this feature too," meaning they consume more data, which makes more money for vendors. APM is almost the gateway drug to observability, feature by feature.

Some companies make it easier for you to add another of their little tools because it's convenient. One company has 26 products — if you use one, you can access the others. Suddenly, finance goes, "Wait a minute, why do we suddenly have this big cost increase?" And you have to go back and look and realize, "Oh, George added this one, Sarah used that one, and Sam used the other one, and wow, our bill just quadrupled."

We're also seeing the rise of generative AI in Ops. Predictive AI and machine learning have long been in the mix, but this is the first year that genAI will appear in products. I expect every vendor will offer something related, but the offerings will almost universally be bad. It's not the vendors' fault, but nobody knows what we can, or should be doing with this capability. So vendors will include the feature, whether or not it's useful or really answers the questions businesses have.

For this reason, I'm updating one of my models. Historically, I have shown the evolution from monitoring to observability to awareness. This year, I'll change from monitoring to observability to intelligence. Under "intelligence" I have questions such as:

Is the business OK?

What was the result of last month's marketing campaign?

Sales has a new initiative; what will impact our services and support?

Unless you're in the business of IT, your real questions are not about IT but the business. If you fly people from point A to point B, you want to ask questions about that, not whether the revenue management system is working.

Observability didn't look to answer these questions, but now that we have more intelligence in tools, we must address them. You want to ask your chat interface that connects to your AIOps that question, rather than going over to revenue management and then going over to this group, that group, or the other group, for the answers.

These tools still have the same problems with AI: choosing the right algorithm at the right time, explainable AI, and AI bias — these are not going away. Let's say I train my AI on all my data … stop there, I don't have all my data because, for example, the guys over in desktop support didn't want to give me their data, but the guys over in networking did. I've trained the models on network data, and the AI now knows networking. So, what is every problem going to be? You guessed it, a networking problem.

Being able to train the AI and getting beyond its biases are going to be challenging. Additionally, generative AIs can hallucinate, presenting nonsense data as fact. Trusting AI as we train it to learn our businesses and help us run more efficiently is part of the new paradigm in business operations.

That'll set the scene for 2024: I expect them to have something, but it won't really help. It may be a little more focused in 2025, but by year three and on — that's when I really believe the AI they're putting into some of these tools will be truly useful. That is, it can answer questions about the condition of the enterprise, not the condition of IT.

That's the direction I see the industry taking, and I'm pushing to see how vendors will impact how the entire business operates. In three years, we should see the hype turn into real changes. For now, the nascent large language models show promise; but with planning and focus, generative AI won't be another promise broken.

Ron Williams is an Analyst at Gigaom

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...